Image matching cross-scene anti-interference target detection method based on sift-bow

By introducing SIFT-BOW technology into the deep learning object detection algorithm, dictionary construction and image vectorization representation are carried out. Combined with the YOLO algorithm, the problem of performance degradation of deep learning algorithms when detecting unsampled interfering targets is solved, and anti-interference recognition and real-time detection of invisible categories are achieved.

CN119963802APending Publication Date: 2025-05-09SOUTH WEST INST OF TECHN PHYSICS
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411898672.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing deep learning object detection algorithms have significantly reduced their performance when detecting unsampled interfering targets. Especially when the target domain data is not available in the training data, and labeling is scarce or even no labeling, it is difficult to effectively identify invisible interference categories.

Method used

Using the cross-scene anti-interference object detection method of image matching based on SIFT-BOW, the dictionary is constructed and vectorized image representation is performed through technologies such as SIFT feature extraction, K-means clustering, KNN and Soft-VQ. Combined with the deep learning algorithm YOLO, the overall architecture of the detection and filtering algorithm is designed to achieve anti-interference recognition of invisible categories.

Benefits of technology

The combination of traditional SIFT feature extraction algorithm and deep learning algorithm YOLO is realized, overcoming the limitations of deep learning's high requirements for spatial consistency of input image categories, improving the detection performance on unsampled interference targets, and meeting the real-time detection requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963802A_ABST
    Figure CN119963802A_ABST
Patent Text Reader

Abstract

The invention relates to an image matching cross-scene anti-interference target detection method based on sift-bow, and belongs to the technical field of image processing. According to the method, a traditional image feature extraction algorithm SIFT and a deep learning algorithm YOLO are combined, the limitation that deep learning has a high requirement for category space consistency of input images is overcome for the situation that target domain data in training data cannot be obtained, so that annotations are scarce and even no annotations exist, and the method is better in performance in an actual test scene; starting from the K-Means algorithm, the SIFT descriptor soft quantization method combined with the KNN algorithm is provided, local errors caused by hard quantization are overcome, and in the test process, it is proved that the soft quantization method provides more stable distance measurement of the target frame filtering algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and in particular relates to an image matching cross-scene anti-interference target detection method based on SIFT-Bow. Background Art

[0002] In real-world scenarios, it is impossible to obtain target data of complete categories. For example, in the military field, there are interference target images that cannot be obtained, such as new target drones and vehicle models; in the civilian field, data constraints make it impossible to collect user data of specific categories. However, the target detection algorithm based on deep learning has limitations on the spatial consistency of the category of the input image, and the performance of the model in detecting unsampled interference targets has dropped significantly. Summary of the invention

[0003] 1. Technical issues to be resolved

[0004] The technical problem to be solved by the present invention is: to solve the problem of target recognition when the target domain data in the training data is unavailable, resulting in scarce or even no annotations, a target detection method is designed to filter and detect invisible interference categories based on the visible category training model.

[0005] (II) Technical solution

[0006] 1. In order to solve the above technical problems, the present invention provides an image matching cross-scene anti-interference target detection method based on sift-bow, which is characterized by comprising the following steps:

[0007] Step 1: Extract SIFT features

[0008] First, construct a SIFT feature training set of the target, which contains a variety of images. Then, apply the SIFT feature extraction algorithm to perform feature detection on all images in the training set, and combine the extracted N SIFT descriptors together to form a SIFT

[0009] Feature training set, SIFT feature is SIFT descriptor;

[0010] Step 2: Construct a dictionary using the K-means clustering algorithm

[0011] The length of the constructed dictionary is K, that is, SIFT feature training is clustered into K classes; K-means

[0012] The process of clustering algorithm is as follows:

[0013] (1) In the SIFT feature training set {α1,α2,...,α N} randomly extract K 128-dimensional SIFT descriptors as the initial cluster centers {β1,β2,...,β K}, next

[0014] Use formula 1 to calculate the SIFT feature training set and all other SIFT descriptors and K

[0015] The similarity between cluster centers is d ij , and other SIFT descriptors are assigned to

[0016] The clusters are classified into the K cluster centers with the largest similarity, thus completing an iteration.

[0017] The generation process, at the same time, obtains the initial clustering results;

[0018] d ij =‖α i -β j ‖ Formula 1

[0019] (2) Recalculate the cluster center based on the clustering results output in the previous iteration.

[0020] Calculate the similarity between all other SIFT descriptors and these K new cluster centers.

[0021] And classify them separately, and get new clustering results at the same time; the cluster center is:

[0022]

[0023] W j ={α k |d kj =min(α k ,β j ),k=1,...,N}

[0024] Where |W j | represents the clustering result W j The number of SIFT descriptors in;

[0025] (3) Repeat (2) until the cluster centers calculated in the two iterations are exactly the same.

[0026] Or the preset number of iterations is reached;

[0027] After the K-means clustering algorithm, K cluster centers are obtained. Each cluster center is a 128-dimensional SIFT descriptor. These SIFT descriptors are connected together as visual vocabulary to form the final trained K-dimensional dictionary.

[0028] Step 3: Use KNN and Soft-VQ to vectorize the image

[0029] After the dictionary is trained, the next step is to vectorize the image, using the visual vocabulary in the dictionary to vectorize all local features of the image into a K-dimensional global vector;

[0030] Step 4: Calculate the minimum value of the distance between the global vector and each vector in the preset vector template library. If the minimum value is greater than the preset threshold, the target in the image represented by the vector is regarded as an interference target; otherwise, it is reported normally and regarded as a real target.

[0031] The present invention also provides a system for implementing the method.

[0032] The invention also provides an image processing method implemented by the method.

[0033] The invention also provides an image processing system implemented based on the system.

[0034] (III) Beneficial effects

[0035] Compared with the prior art, the cross-scene anti-interference target detection method based on SIFT-Bow image matching in the present invention has the following beneficial effects:

[0036] (1) It combines the traditional image feature extraction algorithm SIFT with the deep learning algorithm YOLO. This overcomes the limitation of deep learning that requires high spatial consistency of the input image category, and performs better in actual test scenarios.

[0037] (2) Starting from the K-Means algorithm, a SIFT descriptor soft quantization method combined with the KNN algorithm is proposed to overcome the local errors caused by hard quantization. In the experimental process, it is confirmed that the soft quantization method provides a more stable distance measurement for the target box filtering algorithm;

[0038] (3) In order to meet the real-time requirements of cross-scenario anti-interference target detection, the present invention, with the support of SIFT and YOLO algorithms, takes real-time detection as the core point, designs the overall architecture of the detection and filtering algorithm, and forms a relatively complete detection and evaluation system for unknown domain invisible categories. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a flow chart of a cross-scene anti-interference target detection method based on SIFT-Bow image matching in the present invention;

[0040] Figure 2 This is a flowchart of converting SIFT features of the training set into BOW dictionary based on K-means;

[0041] Figure 3 It is an image query set that extracts positive sample data based on the training data set, which is suitable for the feature matching stage in the second step;

[0042] Figure 4 This is an example of the detection result of positive sample 1;

[0043] Figure 5 This is an example of the detection result of positive sample 2;

[0044] Figure 6 This is an example of the detection result of interference sample 1;

[0045] Figure 7 This is an example of the detection result of interference sample 2;

[0046] Figure 8 This is an example of the detection result of interference sample 3. DETAILED DESCRIPTION

[0047] In order to make the purpose, content and advantages of the present invention more clear, the specific implementation methods of the present invention are further described in detail below in conjunction with the drawings and examples.

[0048] The present invention aims at the problem of target recognition in the case of scarce or even no annotation due to the unobtainability of target domain data in training data. By constraining the output result of the detection model, the invisible interference category is detected by filtering the visible category training model, and specifically involves the image anti-interference technology of scale-invariant feature transform algorithm (SIFT), K-means clustering and soft quantization (Soft-VQ), k-nearest neighbor algorithm (KNN) feature matching, so as to realize the anti-interference recognition of invisible categories and meet the real-time detection requirements. Specifically, in order to achieve the effective anti-interference ability of cross-scene interference targets, the present invention mainly adopts a method combining target detection and image matching. In the first step, all candidate targets in the current frame are identified by using the target detection algorithm; in the second step, the real targets and interference targets among all candidate targets are screened out by using feature extraction and the matching algorithm set by the present invention, and the real-time requirements are met. Through the soft quantization of features and the distance definition of the matching mechanism and the organic combination of the two types of algorithms, the false alarm rate of the model to the unsampled interference targets is eliminated, thereby improving the model performance.

[0049] In the image matching process, the most important link is feature extraction. Traditional feature extraction algorithms have good effects in manually defined and designed image tasks in certain specific scenarios, and are also interpretable. Therefore, they still have wide application value in the era of deep learning algorithms. The SIFT (Scale-Invariant Feature Transform) algorithm with scale-invariant features was first proposed by the famous British scholar David Lowe in 1999. After further summary and method improvement, it was published in IJCV2004, a top journal in the field of computer vision. The algorithm constructs a scale space, simulates the multi-scale nature of the image, detects extreme points in it, and retains its scale, position and direction information. It also has excellent robustness under various conditions such as lighting, affine transformation, noise, etc., so it is widely used in the field of image matching.

[0050] The present invention designs the overall architecture of the image matching algorithm with real-time as the core. When the Euclidean distance is used to measure the similarity between images in the image matching stage, high requirements are placed on the computational efficiency and storage. However, since the features extracted by SIFT are high-dimensional vectors and the number is different, the architecture requirements of image matching cannot be directly met. The bag of words model BOW (Bag of Words) model can generate a K-dimensional vector using the positional relationship of SIFT feature points to solve the above problems. BOW was first applied in text information retrieval with significant results. It can also be widely used in various directions such as image retrieval, target recognition, and classification. In text retrieval, the BOW model quantifies each document in the document library into an N-dimensional vector, and in subsequent document classification, retrieval and other applications, these vectors are operated, which not only avoids the tedious processing of the entire document information, but also improves time efficiency and ensures accuracy. At the same time, converting documents into vectors greatly saves storage space, making it possible for computers to process massive document information. Applying the BOW model in the field of computer vision also applies the above basic principles. First, let's compare a few concepts. We can think of an image as a document and imagine that it is composed of many "small pieces". These "small pieces" are named "visual vocabulary", which is equivalent to the words in the document. Of course, visual vocabulary is the same as words. They are independent of each other and have no connection. However, for images, visual vocabulary is not ready-made, which is different from words in documents. But we can use computer vision related methods to find visual vocabulary. Analogous to words, visual vocabulary refers to the representative "small pieces" that can represent image information, and the representation of image information can naturally think of image features. Here, we take the local feature SIFT of the image as an example. The visual vocabulary is each 128-dimensional SIFT descriptor.

[0051] The research of the present invention is attributed to reducing the total number of SIFT features. The core idea of ​​BOW is to find similar attributes between local feature points of the image and quantify them into visual words, which effectively reduces the dimension of image feature description, reduces the amount of calculation, and improves the calculation speed. In the field of computer vision, it constructs an effective middle-level semantic expression method for the object to be identified, and has good generalization ability and adaptability. In the present invention, we also use BOW to process the extracted SIFT features. BOW can be specifically established through the following steps: first, extract the local feature points of the training sample image to form a local feature descriptor; then, use a clustering algorithm to analyze and calculate the feature descriptor set to establish a visual dictionary; then, according to the visual dictionary, quantify the local feature vector of the image to be described; finally, count the frequency of each visual word in the image in the visual dictionary to generate the image's visual dictionary vector.

[0052] In view of the situation that the interference target data in the training data is unavailable, resulting in scarce or even no annotations, the present invention mainly adopts a method combining target detection and feature matching to achieve the real-time demand of cross-scene anti-interference, specifically involving the image anti-interference technology of scale-invariant feature transform algorithm (SIFT), K-means clustering and soft quantization (Soft-VQ), k-nearest neighbor algorithm (KNN) feature matching, so as to achieve anti-interference recognition of invisible categories and meet the real-time detection requirements. The first step is to use the target detection algorithm to identify all candidate targets in the current frame; the second step is to use feature matching to select the real target among all candidate targets. Although the research idea of ​​the present invention is simple, it realizes the combination of traditional image algorithms and deep learning, and performs well on the test data set.

[0053] The present invention comprises the following steps:

[0054] Step 1: Extract SIFT features

[0055] First, construct the target's SIFT feature training set, which should contain as many different images as possible to make it representative enough. Then apply the SIFT feature extraction algorithm to perform feature detection on all images in the training set, and combine the extracted N SIFT descriptors together to form a SIFT feature training set.

[0056] Step 2: Construct a dictionary using the K-means clustering algorithm

[0057] The most commonly used clustering algorithm is K-means, which is an indirect clustering method that uses the similarity between sample points. The length of the constructed dictionary is K, that is, the SIFT feature training is clustered into K classes. The algorithm process is as follows:

[0058] (4) In the SIFT feature training set {α1,α2,...,α N} randomly extract K 128-dimensional SIFT descriptors as the initial cluster centers {β1,β2,...,β K}, then use Formula 1 to calculate the similarity d between all other SIFT descriptors (SIFT features) in the SIFT feature training set and the K cluster centers ij , and other

[0059] SIFT descriptors are classified into K clusters with the largest similarity (closest distance)

[0060] In this way, an iterative process is completed and the initial clustering result is obtained.

[0061] d ij =||α i -β j || Formula 1

[0062] (5) Recalculate the cluster center of the clustering result output in the previous iteration, and then calculate the similarity (distance) between all other SIFT descriptors and the K new cluster centers, and classify them respectively, and obtain new clustering results at the same time; the cluster center is:

[0063]

[0064] W j ={α k |d kj =min(α k ,β j ),k=1,...,N}

[0065] Where |W j | represents the clustering result W j The number of SIFT descriptors in .

[0066] (6) Repeat (2) until the cluster centers calculated in the previous and subsequent iterations are exactly the same, or the preset number of iterations is reached.

[0067] After the above K-means algorithm, K cluster centers are obtained. Each cluster center is a 128-dimensional SIFT descriptor. These SIFT descriptors are connected together as visual vocabulary to form the final trained K-dimensional dictionary.

[0068] Step 3: Use KNN and Soft-VQ to vectorize the image

[0069] After training the dictionary according to the above steps, the next step is to vectorize the image and use the visual vocabulary in the dictionary to vectorize all local features of the image into a K-dimensional global vector. The main steps are as follows:

[0070] (1) Input an image and extract the SIFT of the local image within the deep learning detection box

[0071] Descriptor;

[0072] (2) Map each local SIFT feature of the image to a visual word in the dictionary and find the visual word with the closest distance (maximum similarity); finally, count the number of times each visual word in the dictionary appears in the image and represent the image as a global vector of size K.

[0073] After the above steps, each image in the data set is represented by a K-dimensional vector. In subsequent applications such as image retrieval, the unit for comparing images is no longer a large number of local feature descriptors, but a quantized global vector. The processed features are applied in retrieval to ensure accuracy and significantly reduce detection time. For each image, its features can be described by a set of visual vocabulary frequencies. By counting the frequency of each visual word in the visual dictionary in the image, a high-dimensional vector in the form of a histogram is formed in the order of the visual vocabulary index as the final representation of the image features.

[0074] The traditional visual dictionary representation adopts the hard quantization method (Hard-VQ), which uses the nearest neighbor search algorithm to directly divide the feature descriptor into the category of the nearest visual vocabulary. However, this is not the best choice. Imagine if there are two or more visual vocabulary closest to a descriptor, or the distance between the descriptor and the nearest visual vocabulary and the second nearest visual vocabulary is very small, the hard quantization method will still only select the nearest visual vocabulary, which will make the quantization error larger and cannot well reflect the probability distribution of the image feature descriptor in the visual dictionary, which will eventually affect the detection effect of subsequent images. Based on the above considerations, the present invention adopts an improved visual vocabulary histogram representation method based on soft quantization (Soft-VQ). The core idea of ​​this method is to use the k-nearest neighbor algorithm (KNN) to find the k visual vocabulary closest to the feature descriptor in the visual dictionary, and then set the corresponding weights for the k visual vocabulary according to the different distances between the feature and the k visual vocabulary, that is, the different contributions to the visual vocabulary, and finally sum these weights as the histogram representation of the image based on the visual dictionary.

[0075] The specific process of step 3 is as follows:

[0076] Assume that the SIFT feature descriptor of the image to be detected I is α i (i=1,2,...,N), N is the total number of combined features detected in I, and the visual dictionary V consists of K visual words β j (j=1,2,...,K) composition:

[0077]

[0078] It is only necessary to assign corresponding weights to the k nearest neighbor visual words. The weight distribution in the k nearest neighbor visual words is the same as α i and visual vocabulary β j The distance is negatively correlated, that is, the closer the distance, the greater the weight assigned to the visual vocabulary, which can be expressed by the following formula:

[0079]

[0080] Where w(α i ,β j ) is the weight function, α i represents the i-th feature descriptor, β j represents the jth visual word, N k (α i ) is α i The set of k nearest neighbor words of . σ is a smoothing factor, and σ is 10 in the present invention.

[0081] After obtaining the weight function, normalize its weight:

[0082]

[0083] Then use the weight function w(α i ,β j ) respectively calculate the feature descriptor α i The weight distribution p(α i ), quantize the SIFT descriptor to the entire visual dictionary:

[0084] p(α i )=[w(α i ,β1),w(α i ,β2),...,w(α i ,β k )] T ,p(α i )∈R K×1 Formula 6

[0085] Similarly, by soft-quantizing the N SIFT descriptors (features) of the detection image I and assigning them to the visual dictionary, we can get the weight matrix P(I) of the entire image:

[0086] P(I)=[p(α1),p(α2),...,p(α N )],P(I)∈R K×N Formula 7

[0087] Finally, the sum of the weights of each row of the weight matrix P(I) is calculated, the frequency of occurrence of K visual words is counted, and standardized calculations are performed to obtain the visual dictionary statistical vector S(I) with a dimension of K for the entire image:

[0088]

[0089] where Q = [1,1,...1] T ,Q∈R N×1 .

[0090] The minimum value of the distance between the vector S(I) and each vector in the preset vector template library is calculated respectively. If the minimum value is greater than the preset threshold, the target in the image represented by the vector is regarded as an interference target, and the reporting information is shielded on the candidate target box corresponding to the yolov5 detector; otherwise, it is reported normally and regarded as a real target.

[0091] A brief summary of one or more aspects is given below to provide a basic understanding of these aspects. This summary is not an exhaustive overview of all conceived aspects, and is neither intended to identify the key or critical elements of all aspects nor to define the scope of any or all aspects. Its only purpose is to give some concepts of one or more aspects in a simplified form as a prelude to a more detailed description that will be given later.

[0092] like Figure 1 The process shown is used for a cross-scene anti-interference target detection method based on sift-bow image matching. It is a method that uses clustering and bag-of-words models to construct a vectorized representation of the image and overcome the problems of high SIFT feature dimensions, inconsistent numbers, and inability to meet real-time requirements. The training and detection process of this method includes the following steps:

[0093] 1. Training phase:

[0094] The collected images are divided into training set Data_train and test set Data_test, and are annotated and segmented into training set Data_train_cut and test set Data_test_cut of the real target according to the annotated information. While using Data_train and Data_test to train the yolov5 target detector, use Data_train_cut to train the K-means clustering center julei and obtain the representation vector zhenshi_train of the real target through soft quantization (based on steps 1 to 3), and store it in the vector template library for subsequent query. Use the Data_test_cut data set and the julei vector to obtain its representation vector zhenshi_test, and query the L2 closest distance with zhenshi_train (traverse the dictionary vector array once to find the nearest neighbor vector), calculate its average value, and obtain the distance threshold Threshold.

[0095]

[0096] Where α∈zhenshi_test, β∈zhenshi_train.

[0097]

[0098] 2. Testing phase:

[0099] The yolov5 target detector in the training phase identifies the candidate target box in the detection area. The target box image is soft-quantized according to the julei template library to obtain the representation vector houxuan (based on steps 1 to 3). The zhenshi_train vector template library is traversed to calculate the minimum L2 distance between houxuan and all vectors in the library zhenshi_train.

[0100]

[0101] Where β∈zhenshi_train.

[0102] If dis_houxuan>Threshold, it is regarded as an interference target, and the reporting information is blocked on the corresponding candidate target frame of the yolov5 detector; otherwise, it is reported normally and regarded as a real target.

[0103] Using the vehicle target, put in 3 interference target vehicles, and test the confidence and distance as shown in the following table:

[0104] Table 1 Yolov5 confidence and distance results of test instances

[0105]

[0106] It can be seen that if only YOLOv5 training is used, in the test phase, there is no obvious difference between the true target and the interference target in the confidence dimension, and the interference target cannot be effectively removed by threshold control, and only the true target is reported. However, there is an obvious difference in the distance designed by the present invention, and the interference target can be removed by post-processing, and only the true target is reported. This shows the effectiveness of this method.

[0107] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A cross-scene anti-interference target detection method based on sift-bow image matching, characterized in that: The following steps are involved: Step 1: Extract SIFT features First, construct a SIFT feature training set of the target, which contains a variety of images. Then, apply the SIFT feature extraction algorithm to perform feature detection on all images in the training set, and combine the extracted N SIFT descriptors together to form a SIFT feature training set. The SIFT feature is the SIFT descriptor. Step 2: Construct a dictionary using the K-means clustering algorithm The length of the constructed dictionary is K, that is, the SIFT feature training is clustered into K classes; the process of the K-means clustering algorithm is as follows: (1) In the SIFT feature training set {α1,α2,...,α N } randomly extract K 128-dimensional SIFT descriptors as the initial cluster centers {β1,β2,...,β K }, then use Formula 1 to calculate the similarity d between all other SIFT descriptors in the SIFT feature training set and the K cluster centers ij , and classify other SIFT descriptors into the K cluster centers with the largest similarity, thus completing an iterative process and obtaining the initial clustering result; d ij = ||α i - β j || Equation 1 (2) Recalculate the cluster center of the clustering result output in the previous iteration, and then calculate the similarity between all other SIFT descriptors and the K new cluster centers, and classify them respectively, and obtain new clustering results at the same time; the cluster center is: W j ={a k |d kj =min(a k ,b j ),k=1,...,N} Where |W j | represents the clustering result W j The number of SIFT descriptors in; (3) Repeat (2) until the cluster centers calculated in the previous and next iterations are exactly the same, or the preset number of iterations is reached; After the K-means clustering algorithm, K cluster centers are obtained. Each cluster center is a 128-dimensional SIFT descriptor. These SIFT descriptors are connected together as visual vocabulary to form the final trained K-dimensional dictionary. Step 3: Use KNN and Soft-VQ to vectorize the image After the dictionary is trained, the next step is to vectorize the image, using the visual vocabulary in the dictionary to vectorize all local features of the image into a K-dimensional global vector; Step 4: Calculate the minimum value of the distance between the global vector and each vector in the preset vector template library respectively. If the minimum value is greater than a preset threshold, the target in the image represented by the vector is regarded as an interference target. Otherwise, it will be reported normally and regarded as the real target.

2. The method according to claim 1, characterized in that Step 3 is as follows: (1) Input an image and extract the SIFT descriptor of the local image; (2) Map each local SIFT feature of the image to a visual word in the dictionary and find the visual word that is closest to it; finally, count the number of times each visual word in the dictionary appears in the image and represent the image as a global vector of size K.

3. The method according to claim 1, characterized in that In step 3, the k-nearest neighbor algorithm is used to find the k visual words that are the nearest neighbors to the feature descriptor in the visual dictionary. Then, according to the different distances between the feature and the k visual words, that is, the different contributions to the visual vocabulary, corresponding weights are set for them. Finally, the sum of these weights is used as the histogram representation of the image based on the visual dictionary.

4. The method according to claim 1, characterized in that The specific process of step 3 is as follows: Assume that the SIFT feature descriptor of the image to be detected I is α i (i=1,2,...,N), N is the total number of combined features detected in I, and the visual dictionary V consists of K visual words β j (j=1,2,...,K) composition: Assign corresponding weights to the k nearest neighbor visual words; the weight distribution in the k nearest neighbor visual words is similar to α i and visual vocabulary β j The distance is negatively correlated, that is, the closer the distance, the greater the weight assigned to the visual vocabulary, which can be expressed by the following formula: Where w(α i ,β j ) is the weight function, α i represents the i-th feature descriptor, β j represents the jth visual word, N k (α i ) is α i The set of k nearest neighbor words, σ is the smoothing factor; After obtaining the weight function, normalize its weight: Then use the weight function w(α i ,β j ) respectively calculate the feature descriptor α i The weight distribution p(α i ), quantize the SIFT descriptor to the entire visual dictionary: p(α i ) = [w(α i , β1), w(α i , β2),..., w(α i , β k )] T , p(α i ) ∈ R K×1 Formula 6 Similarly, the N SIFT descriptors of the detection image I are soft-quantized and assigned to the visual dictionary to obtain the weight matrix P(I) of the entire image: P(I) = [p(α1), p(α2),..., p(α N )], P(I) ∈ R K×N Formula 7 Finally, the sum of the weights of each row of the weight matrix P(I) is calculated, the frequency of occurrence of K visual words is counted, and a standardized calculation is performed to obtain the visual dictionary statistical vector S(I) of the entire image dimension K as the global vector: where Q = [1,1,...1] T ,Q∈R N×1 .

5. The method according to claim 1, characterized in that When an interfering target is detected, the reporting information is blocked on the candidate target frame corresponding to the yolov5 detector.

6. The method according to claim 4, characterized in that σ is taken as 10.

7. Application of this method in image processing.

8. A system for implementing the method according to any one of claims 1 to 6.

9. An image processing method implemented by the method according to any one of claims 1 to 6.

10. An image processing system implemented based on the system according to claim 8.

Citation Information

Patent Citations

  • Method for detecting landslip from remotely sensed image by adopting image classification technology

    CN102542295A

  • Image retrieval method

    CN103488664A

  • Automobile wheel hub classification method based on word bag model and support vector machine

    CN106570514A

  • Remote image target recognizing method

    CN106951873A

  • Method of generating feature vector, generating histogram, and learning classifier for recognition of behavior

    KR1020150088157A