Machine vision-oriented image global feature representation method

CN117351024BActive Publication Date: 2026-09-25BEIJING UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211006393.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-22
Publication Date
2026-09-25
Estimated Expiration
2042-08-22

AI Technical Summary

Technical Problem

该方法的缺点是并没有考虑区域块之间的空间关联信息

Benefits of technology

[0123]1.本发明方法概念简单,没有增加空间金字塔模型SPM特征的维数。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117351024B_ABST
    Figure CN117351024B_ABST
Patent Text Reader

Abstract

The application provides a kind of machine vision-oriented image global feature representation method, comprising extracting image, further comprising the following steps: the image is divided by space pyramid, and the feature representation of whole image is obtained;The spatial correlation information between the region blocks is calculated for each time of the region division of SPM, and a spatial correlation matrix is obtained, and then the classical SPM feature is transformed by matrix to obtain a new feature.The application provides a kind of machine vision-oriented image global feature representation method, by matrix transformation, the spatial pyramid model SPM is integrated into spatial correlation information, without increasing the feature dimension of the spatial pyramid model SPM, but the new image-level feature representation has stronger discrimination ability and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method for representing global features of images for machine vision. Background Technology

[0002] Image features are descriptions of the properties or attributes of an image. The extraction and representation of image features are fundamental to image processing, and feature extraction is crucial when representing images. The BoF (Browser-of-Flight) model is one of the most widely used feature types in computer vision in recent years, applied to image classification, object recognition, image retrieval, robot localization, and texture recognition. Numerous studies have demonstrated the excellent performance of BoF features in computer vision. The steps for constructing BoF features include feature extraction, dictionary generation, feature encoding, and feature pooling. First, local feature descriptors are extracted from the image; a visual word dictionary is obtained through clustering; local features are encoded using the visual words in the dictionary; finally, feature pooling is performed on each visual word within a specific image region. Feature encoding and feature pooling operations result in a more compact final image representation, better adapting to rotation and translation invariance. When performing feature pooling, it is necessary to select the pooling region, i.e., within which the feature pooling operation is performed. The BoF model uses the entire image for pooling, a classic method for selecting pooling regions. Spatial Pyramid Merging (SPM) [S. Lazebnik, C. Schmid, and J. Ponce. Beyond bags of features: spatial pyramid matching for recognizing natural scene categories[C]. Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, 2012. 2169–2178.] is also a classic feature pooling method. It involves constructing a three-layer pyramid: Given three different visual words in the image, represented by diamonds, crosses, and dots; the image region is continuously subdivided: the first subdivision results in the entire image, the second divides it into four sub-regions, and the third divides it into sixteen sub-regions; statistical features are calculated for each sub-region, and a frequency histogram of each visual word is plotted within each region; finally, the statistical histograms of all sub-regions are concatenated to obtain the feature representation of the entire image. The process of SPM successively subdividing the image into sub-regions is the process of feature region selection. The Spatial Pyramid Model (SPM) incorporates more spatial information than the BoF model, which illustrates the importance of spatial pooling operations in image space. However, there is still room for improvement in the Spatial Pyramid Model (SPM).

[0003] The article "Research on Image Representation Methods Based on BoF Model" by Liang Ye, Yu Jian, and Liu Hongzhe was published in the March 2014 issue of *Computer Science*. This article points out that designing suitable image representations is one of the most important problems in computer vision. BoF feature representation is very popular and has been widely applied in image classification, object recognition, image retrieval, robot localization, and texture recognition. BoF features represent an image as an unordered set of features. Although this method lacks structural and spatial information, it is conceptually simple, computationally straightforward, and in some applications, its performance can even rival the best current methods. The article carefully studies the BoF model, focusing on the typical techniques involved in its three stages: local feature extraction, feature quantization and encoding, and feature aggregation. Finally, based on an analysis of various research methods, it summarizes the current research problems and possible development directions. A drawback of this method is that it does not consider the spatial correlation information between regions. Summary of the Invention

[0004] To address the aforementioned technical issues, this invention proposes a global image feature representation method for machine vision. Through matrix transformation, spatial correlation information is incorporated into the Spatial Pyramid Model (SPM) without increasing the feature dimension of the SPM. However, the new image-level feature representation has stronger discriminative power and robustness.

[0005] This invention provides a global feature representation method for images for machine vision, including image extraction, and further including the following steps:

[0006] Step 1: Perform spatial pyramid partitioning on the image to obtain the feature representation of the entire image;

[0007] Step 2: For each SPM region division, the spatial association information between region blocks must be calculated to obtain the spatial association matrix. Then, the classic SPM features are transformed into new features.

[0008] Preferably, step 1 includes the following sub-steps:

[0009] Step 11: Divide the image into three parts;

[0010] Step 12: Calculate the frequency histogram of each visual word in each region block;

[0011] Step 13: Connect the statistical histogram features of all sub-regions to obtain the feature representation of the entire image.

[0012] In any of the above schemes, step 11 preferably includes a first division result of the entire image, a second division of the image into 4 sub-region blocks, and a third division of the image into 16 sub-region blocks.

[0013] In any of the above solutions, step 12 preferably includes the following sub-steps:

[0014] Step 121: Assume the number of visual words in the visual word dictionary is M, and L represents the total number of times the image is divided into regions. In the l-th division, the image is divided into 2... l-1 There are 3 regions, where 1 ≤ l ≤ L, and L = 3;

[0015] Step 122: Perform feature pooling on the encoded features of region k in layer l to obtain the pooled features. Among them, R M It is an M-dimensional vector;

[0016] Step 123: After the pooling operation of the l-th layer, the resulting feature is f. l ,

[0017] Step 124: After performing pooling operations on all layers of the image L, we obtain the feature f, f∈R. s , Features are represented as

[0018] In any of the above solutions, step 2 preferably includes the following sub-steps:

[0019] Step 21: The first segmentation result is still the original image, with feature f. 1 ;

[0020] Step 22: Perform feature calculations on the second partitioning result;

[0021] Step 23: Perform feature calculations on the third partitioning result;

[0022] Step 24: Obtain new image-level feature representations.

[0023] In any of the above embodiments, step 22 preferably further includes the following sub-steps:

[0024] Step 221: When performing the second division, there are a total of 4 regions. Label the positions of the 4 regions.

[0025] Step 222: Calculate the distance between two positions (i,j) and (p,q), where 1≤i≤2, 1≤j≤2, 1≤p≤2, and 1≤q≤2;

[0026] Step 223: Calculate the distance between every two of the four positions to obtain a 4×4 distance matrix C. 2 ;

[0027] Step 224: From the distance matrix C 2 Calculate the correlation matrix M 2 The correlation matrix M 2 It is a 4×4 square matrix;

[0028] Step 225: For M 2 Perform Cholesky decomposition, which yields matrix A. 2 A 2 It is a 4×4 diagonal matrix;

[0029] Step 226: For f 2 After transformation, we obtain a 4×128 matrix B. 2 , for B 2 Perform matrix transformations to transform B 2 Transform it again into a 4×128 dimensional matrix

[0030] In any of the above schemes, it is preferable that the distance between the two positions (i,j) and (p,q) is

[0031] (ip) 2 +(jq) 2 .

[0032] In any of the above schemes, the preferred embodiment is the correlation matrix M. 2 elements The calculation formula is:

[0033]

[0034] Where i∈1...4, j∈1...4, and a is a parameter. Let C be the distance matrix. 2 The elements in.

[0035] In any of the above solutions, the preferred option is the one for B. 2 The formula for performing matrix transformations is B. 2= A 2 *B 2 .

[0036] In any of the above embodiments, step 23 preferably includes the following sub-steps:

[0037] Step 231: When performing the third division, there are a total of 16 regions. Label the positions of the 16 regions.

[0038] Step 232: Calculate the distance between two positions (i,j) and (p,q), where 1≤i≤4, 1≤j≤4, 1≤p≤4, and 1≤q≤4;

[0039] Step 233: Calculate the distance between every two of the 16 positions to obtain a 16×16 distance matrix C. 3 ;

[0040] Step 234: From the distance matrix C 3 Calculate the correlation matrix M 3 The correlation matrix M 3 It is a 16×16 square matrix;

[0041] Step 235: For M 3 Perform Cholesky decomposition, which yields matrix A. 3 A 3 It is a 16×16 diagonal matrix;

[0042] Step 236: For f 3 After transformation, we obtain a 16×128 matrix B. 3 , for B 3 Perform matrix transformations to transform B 3 Transform it again into a 16×128 dimensional matrix

[0043] In any of the above schemes, it is preferable that the distance between the two positions (i,j) and (p,q) is

[0044] (ip) 2 +(jq) 2 .

[0045] In any of the above schemes, the preferred embodiment is the correlation matrix M. 3 elements The calculation formula is:

[0046]

[0047] Where i∈1...16, j∈1...16, Let C be the distance matrix. 3 The elements in.

[0048] In any of the above solutions, the preferred option is the one for B. 3 The formula for performing matrix transformations is B. 3= A 3 *B 3 .

[0049] Preferably, in any of the above schemes, the new image-level feature representation is as follows:

[0050]

[0051] This invention proposes a global image feature representation method for machine vision. The concept is simple, does not increase the dimension of the Spatial Pyramid Model (SPM) features, and considers the spatial correlation between image regions. The new global image features have stronger discriminative power and robustness. Attached Figure Description

[0052] Figure 1 This is a flowchart of a preferred embodiment of the image global feature representation method for machine vision according to the present invention.

[0053] Figure 2 This is a schematic diagram illustrating an embodiment of the spatial pyramid feature (SPM) construction process according to the machine vision-oriented image global feature representation method of the present invention. Detailed Implementation

[0054] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0055] Example 1

[0056] like Figure 1 As shown, perform step 1000 to extract the image.

[0057] Step 1100 involves performing spatial pyramid partitioning on the image to obtain the feature representation of the entire image, including the following sub-steps:

[0058] Step 1110: Divide the image into three parts. The first part is the whole image. The second part is to divide the image into 4 sub-region blocks. The third part is to divide the image into 16 sub-region blocks.

[0059] Step 1120: Calculate the frequency histogram of each visual word in each region block, including the following sub-steps:

[0060] Step 1121: Assume the number of visual words in the visual word dictionary is M, and L represents the total number of times the image is divided into regions. In the l-th division, the image is divided into 2... l-1 There are 3 regions, where 1 ≤ l ≤ L, and L = 3;

[0061] Step 1122: Perform feature pooling on the encoded features of region k in layer l to obtain the pooled features. Among them, R M It is an M-dimensional vector;

[0062] Step 1123, after the pooling operation of the l-th layer, the feature obtained is f l ,

[0063] Step 1124: After performing pooling operations on all layers of the image L, we obtain the feature f, f∈R. s , Features are represented as

[0064] Step 1130: Connect the statistical histogram features of all sub-regions to obtain the feature representation of the entire image.

[0065] Step 1200 involves calculating the spatial correlation information between regions for each SPM region division to obtain a spatial correlation matrix. Then, a matrix transformation is performed on the classic SPM features to obtain new features, including the following sub-steps:

[0066] Execute step 1210, where the first segmentation result is still the original image, with feature f 1 .

[0067] Step 1220 involves performing feature calculations on the second partitioning result, including the following sub-steps:

[0068] When performing step 1221 and dividing the area for the second time, there are a total of 4 regions. Label the positions of the 4 regions.

[0069] Perform step 1222 to calculate the distance between two positions (i,j) and (p,q). The distance between two positions (i,j) and (p,q) is (ip). 2 +(jq) 2 Where 1≤i≤2, 1≤j≤2, 1≤p≤2, 1≤q≤2;

[0070] Performing step 1223, the distance is calculated for every two of the four positions, resulting in a 4×4 distance matrix C. 2 ;

[0071] Execute step 1224, using the distance matrix C 2 Calculate the correlation matrix M 2 The correlation matrix M 2 It is a 4×4 square matrix, and the correlation matrix M 2 elements The calculation formula is:

[0072]

[0073] Where i∈1...4, j∈1...4, and a is a parameter. Let C be the distance matrix. 2 The elements in.

[0074] Execute step 1225, for M 2 Perform Cholesky decomposition, which yields matrix A. 2 A 2 It is a 4×4 diagonal matrix;

[0075] Execute step 1226, for f 2 After transformation, we obtain a 4×128 matrix B. 2 , for B 2 Perform matrix transformations, the formula is B 2 = A 2 *B 2 B 2 Transform it again into a 4×128 dimensional matrix

[0076] Step 1230, which involves feature calculation based on the third partitioning result, further includes the following sub-steps:

[0077] When performing step 1231 and making the third division, there are a total of 16 region blocks. Label the positions of the 16 regions.

[0078] Perform steps 1, 2, 3, and 2 to calculate the distance between two positions (i,j) and (p,q). The distance between two positions (i,j) and (p,q) is (ip). 2 +(jq) 2 Where 1≤i≤4, 1≤j≤4, 1≤p≤4, 1≤q≤4;

[0079] Perform steps 1233 to calculate the distance between every two of the 16 positions, resulting in a 16×16 distance matrix C. 3 ;

[0080] Execute step 1234, using the distance matrix C 3 Calculate the correlation matrix M 3 The correlation matrix M 3 It is a 16×16 square matrix, and the correlation matrix M 3 elements The calculation formula is:

[0081]

[0082] Where i∈1...16, j∈1...16, Let C be the distance matrix. 3 The elements in.

[0083] Execute steps 1235 for M 3 Perform Cholesky decomposition, which yields matrix A.3 A 3 It is a 16×16 diagonal matrix;

[0084] Execute steps 1236 to process f 3 After transformation, we obtain a 16×128 matrix B. 3 , for B 3 Perform matrix transformations, the formula is B 3= A 3 *B 3 B 3 Transform it again into a 16×128 dimensional matrix

[0085] Executing step 1240 yields a new image-level feature representation.

[0086] Example 2

[0087] While the Spatial Pyramid Model (SPM) achieves good performance, it does not consider the spatial relationships between regions when performing image region segmentation. This invention incorporates spatial relationship information into the SPM through matrix transformation, without increasing the feature dimension of the SPM. The new image-level feature representation exhibits stronger discriminative power and robustness.

[0088] The implementation method of this invention is as follows:

[0089] 1. Perform spatial pyramid partitioning on the image. The first partitioning results in the entire image. The second partitioning divides the image into 4 sub-regions. The third partitioning divides the image into 16 sub-regions. Calculate the statistical features of each sub-region and generate a frequency histogram for each visual word in each region. Finally, connect the statistical histogram features of all sub-regions to obtain the feature representation of the entire image.

[0090] Assume the visual word dictionary size is M, and L represents the total number of times the image is divided into regions, where L = 3. In the l-th division, the image is divided into 2... l-1 There are regions, where 1 ≤ l ≤ L.

[0091] Perform feature pooling on the encoded features of region k in layer l to obtain the pooled features. M represents the number of visual words.

[0092] After performing the pooling operation on the l-th layer, the resulting feature is f. l ,

[0093] After performing pooling operations on all layers of the image L, we obtain features f, f∈R. s , Features are represented as

[0094]

[0095] 2. For each SPM region division, the spatial correlation information between region blocks needs to be calculated to obtain a spatial correlation matrix. Then, the classic SPM features are transformed into new features, as shown in Table 1.

[0096] (2,1) (2,2) (2,3) (2,4) (3,1) (3,2) (3,3) (3,4) (4,1) (4,2) (4,3) (4,4)

[0097] Table 1. Labels of the regions in the third division

[0098] Given two positions (i,j) and (p,q), where 1≤i≤4, 1≤j≤4, 1≤p≤4, and 1≤q≤4, the distance between the two positions is:

[0099] (ip) 2 +(jq) 2

[0100] The distance between every two of the 16 positions is calculated, resulting in a 16×16 distance matrix C. 3 .

[0101] From the distance matrix C 3 Calculate the correlation matrix M 3 Correlation matrix M 3 It is a 16×16 square matrix with elements The calculation formula is as follows:

[0102]

[0103] Where i∈1...16, j∈1...16, and a is a parameter that needs to be determined. The calculated M... 3 Afterwards, regarding M 3 Perform Cholesky decomposition, which yields matrix A. 3 A 3 It is a 16x16 diagonal matrix.

[0104] In SPM, the local feature descriptor is a 128-dimensional SIFT feature. The third region division consists of 16 blocks, so the feature f... 3 It is a 16×128 dimensional matrix. Let f 3 After transformation, we obtain a 16×128 matrix B. 3 , for B 3 Perform the following matrix transformations,

[0105] B 3 =A 3 *B3

[0106] B 3 Transform it again into a 16×128 dimensional matrix

[0107] (2) When performing the second division, there are a total of 4 regions. The locations of the 4 regions are labeled as shown in Table 2.

[0108] (2,1) (2,2)

[0109] Table 2. Labels of the regions in the second division

[0110] Given two positions (i,j) and (p,q) such that 1≤i≤2, 1≤j≤2, 1≤p≤2, and 1≤q≤2, the distance between the two positions is:

[0111] (ip) 2 +(jq) 2

[0112] The distance between every two of the four positions is calculated, resulting in a 4×4 distance matrix C. 2 .

[0113] From the distance matrix C 2 Calculate the correlation matrix M 2 Correlation matrix M 2 It is a 4×4 square matrix with elements The calculation formula is as follows:

[0114]

[0115] Where i∈1...4, j∈1...4, and a is a parameter that needs to be determined. The result of calculating M is... 2 Afterwards, regarding M 2 Perform Cholesky decomposition, which yields matrix A. 2 A 2 It is a 4×4 diagonal matrix.

[0116] In SPM, the local feature descriptor is a 128-dimensional SIFT feature. The second region division consists of 4 blocks, therefore the feature f... 2 It is a 4×128 dimensional matrix. Let f 2 After transformation, we obtain a 4*128 matrix B. 2 , for B 2 Perform the following matrix transformations,

[0117] B 2 =A 2 *B 2

[0118] B 2 Transform it again into a 4×128 dimensional matrix

[0119] (3) The first segmentation result is still the original image, so no matrix transformation is needed, and the feature is still f. 1 .

[0120] (4) After the above three steps, the new image-level feature representation is obtained as follows:

[0121]

[0122] The present invention has the following advantages and beneficial effects:

[0123] 1. The method of this invention is conceptually simple and does not increase the dimensionality of the features of the Spatial Pyramid Model (SPM).

[0124] 2. The method of the present invention takes into account the spatial correlation between image region blocks, and the new global image features have stronger discriminative ability and robustness.

[0125] Example 3

[0126] The construction process of the spatial pyramid feature SPM is as follows: Figure 2 As shown.

[0127] The image contains 3 visual words. The image is divided into 3 layers: the first layer contains 1 sub-region, the second layer contains 4 sub-regions, and the third layer contains 16 sub-regions. For each sub-region, a frequency histogram of all visual words is calculated.

[0128] Example 4

[0129] When the image is divided into 16 regions, the positions of the 16 regions are labeled as shown in Table 3.

[0130] (2,1) (2,2) (2,3) (2,4) (3,1) (3,2) (3,3) (3,4) (4,1) (4,2) (4,3) (4,4)

[0131] Table 3. Labels of the 16 area blocks

[0132] Given two positions (i,j) and (p,q), where 1≤i≤4, 1≤j≤4, 1≤p≤4, and 1≤q≤4, the distance between the two positions is:

[0133] (ip) 2 +(jq) 2

[0134] The distance between every two of the 16 positions is calculated, resulting in a 16x16 distance matrix C. 3 C 3 The elements are:

[0135]

[0136]

[0137] From the distance matrix C 3 Calculate the correlation matrix M 3 Correlation matrix M 3 It is a 16x16 square matrix, with elements... as follows:

[0138] 1.388888889 0.603608623 0.049547213 0.000768173 0.6036086230.262327226 0.02153313 0.000333846 0.049547213 0.02153313 0.001767547 2.74E-05 0.000768173 0.000333846 2.74E-05 4.25E-07

[0139] 0.603608623 1.388888889 0.603608623 0.049547213 0.2623272260.603608623 0.262327226 0.02153313 0.02153313 0.049547213 0.021533130.001767547 0.000333846 0.000768173 0.000333846 2.74E-05

[0140] 0.049547213 0.603608623 1.388888889 0.603608623 0.021533130.262327226 0.603608623 0.262327226 0.001767547 0.02153313 0.0495472130.02153313 2.74E-05 0.000333846 0.000768173 0.000333846

[0141] 0.000768173 0.049547213 0.603608623 1.388888889 0.0003338460.02153313 0.262327226 0.603608623 2.74E-05 0.001767547 0.021533130.049547213 4.25E-07 2.74E-05 0.000333846 0.000768173

[0142] 0.603608623 0.262327226 0.02153313 0.000333846 1.3888888890.603608623 0.049547213 0.000768173 0.603608623 0.262327226 0.021533130.000333846 0.049547213 0.02153313 0.001767547 2.74E-05 0.262327226 0.603608623 0.262327226 0.02153313 0.6036086231.388888889 0.603608623 0.049547213 0.262327226 0.603608623 0.2623272260.02153313 0.02153313 0.049547213 0.02153313 0.001767547 0.02153313 0.262327226 0.603608623 0.262327226 0.0495472130.603608623 1.388888889 0.603608623 0.02153313 0.262327226 0.6036086230.262327226 0.001767547 0.02153313 0.049547213 0.02153313

[0145] 0.000333846 0.02153313 0.262327226 0.603608623 0.0007681730.049547213 0.603608623 1.388888889 0.000333846 0.02153313 0.2623272260.603608623 2.74E-05 0.001767547 0.02153313 0.049547213

[0146] 0.049547213 0.02153313 0.001767547 2.74E-05 0.603608623 0.2623272260.02153313 0.000333846 1.388888889 0.603608623 0.049547213 0.0007681730.603608623 0.262327226 0.02153313 0.000333846 0.02153313 0.049547213 0.02153313 0.001767547 0.262327226 0.6036086230.262327226 0.02153313 0.603608623 1.388888889 0.603608623 0.0495472130.262327226 0.603608623 0.262327226 0.02153313 0.001767547 0.02153313 0.049547213 0.02153313 0.02153313 0.2623272260.603608623 0.262327226 0.049547213 0.603608623 1.388888889 0.6036086230.02153313 0.262327226 0.603608623 0.262327226

[0149] 2.74E-05 0.001767547 0.02153313 0.049547213 0.000333846 0.021533130.262327226 0.603608623 0.000768173 0.049547213 0.603608623 1.3888888890.000333846 0.02153313 0.262327226 0.603608623

[0150] 0.000768173 0.000333846 2.74E-05 4.25E-07 0.049547213 0.021533130.001767547 2.74E-05 0.603608623 0.262327226 0.02153313 0.0003338461.388888889 0.603608623 0.049547213 0.000768173

[0151] 0.000333846 0.000768173 0.000333846 2.74E-05 0.02153313 0.0495472130.02153313 0.001767547 0.262327226 0.603608623 0.262327226 0.021533130.603608623 1.388888889 0.603608623 0.049547213

[0152] 2.74E-05 0.000333846 0.000768173 0.000333846 0.001767547 0.021533130.049547213 0.02153313 0.02153313 0.262327226 0.603608623 0.2623272260.049547213 0.603608623 1.388888889 0.603608623 4.25E-07 2.74E-05 0.0003338460.000768173 2.74E-05 0.001767547 0.02153313 0.049547213 0.000333846 0.02153313 0.262327226 0.603608623 0.000768173 0.049547213 0.603608623 1.388888889

[0153] Example 5

[0154] When the image is divided into 4 regions, the positions of the 4 regions are labeled as shown in Table 4.

[0155] (2,1) (2,2)

[0156] Table 4. Labels of the four area blocks

[0157] Given two positions (i,j) and (p,q), where 1≤i≤2, 1≤j≤2, 1≤p≤2, and 1≤q≤2, the distance between the two positions is:

[0158] (ip) 2 +(jq) 2

[0159] The distance between every two of the four positions is calculated, resulting in a 4x4 distance matrix C. 2 C 2 as follows:

[0160]

[0161]

[0162] From the distance matrix C 2 Calculate the correlation matrix M 2 Correlation matrix M 2 It is a 4x4 square matrix with elements as follows: 1.38888888888889 0.603608622926498 0.603608622926498 0.262327226163280 0.603608622926498 1.38888888888889 0.262327226163280 0.603608622926498 0.603608622926498 0.262327226163280 1.388888888888890.603608622926498 0.262327226163280 0.603608622926498 0.603608622926498 1.38888888888889

[0167] To better understand this invention, specific embodiments have been described in detail above, but these are not intended to limit the invention. Any simple modifications made to the above embodiments based on the technical essence of this invention still fall within the scope of this invention. Each embodiment in this specification focuses on its differences from other embodiments; similar or identical parts between embodiments can be referred to mutually. For system embodiments, since they basically correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

Claims

1. A global feature representation method for images for machine vision, comprising image extraction, characterized in that, It also includes the following steps: Step 1: Perform spatial pyramid partitioning on the image to obtain the feature representation of the entire image, including the following sub-steps: Step 11: Divide the image into three parts; Step 12: Calculate the frequency histogram of each visual word in each region block; Step 13: Connect the statistical histogram features of all sub-regions to obtain the feature representation of the entire image; Step 2: For each SPM region division, the spatial correlation information between region blocks needs to be calculated to obtain the spatial correlation matrix. Then, the classic SPM features are transformed into new features, including the following sub-steps: Step 21: For the first segmentation result, which is still the original image, the feature is f 1 ; Step 22: Perform feature calculations on the second partitioning result, which also includes the following sub-steps: Step 221: When performing the second division, there are a total of 4 regions. Label the positions of the 4 regions. Step 222: Calculate the distance between two positions (i,j) and (p,q), where 1≤i≤2, 1≤j≤2, 1≤p≤2, and 1≤q≤2; Step 223: Calculate the distance between every two of the four positions to obtain a 4×4 distance matrix C. 2 ; Step 224: From the distance matrix C 2 Calculate the correlation matrix M 2 The correlation matrix M 2 It is a 4×4 square matrix; Step 225: For M 2 Performing Cholesky decomposition yields matrix A. 2 A 2 It is a 4×4 diagonal matrix; Step 226: For f 2 After transformation, we obtain a 4×128 matrix B. 2 , for B 2 Perform matrix transformation B 2 = A 2 *B 2 B 2 Transform it again into a 4×128 dimensional matrix ; Step 23: Perform feature calculations based on the third partitioning result; Step 24: Obtain new image-level feature representations.

2. The image global feature representation method for machine vision as described in claim 1, characterized in that, Step 11 includes the first division result being the entire image, the second division dividing the image into 4 sub-region blocks, and the third division dividing the image into 16 sub-region blocks.

3. The image global feature representation method for machine vision as described in claim 2, characterized in that, Step 12 includes the following sub-steps: Step 121: Assume the number of visual words in the visual word dictionary is M, and L represents the total number of times the image is divided into regions. In the l-th division, the image is divided into 2... l-1 There are 3 regions, where 1 ≤ l ≤ L, and L = 3; Step 122: For the first The encoded features of region k in the layer are subjected to feature pooling to obtain the pooled features. , 1≤k≤2 l-1 , , where R M It is an M-dimensional vector; Step 123: For the first After layer pooling, the resulting features are: , ; Step 124: After performing pooling operations on all layers of the image L, we obtain the feature f, f∈R. s , The features are represented as .

4. The image global feature representation method for machine vision as described in claim 3, characterized in that, The distance between two positions (i,j) and (p,q) is 5. The image global feature representation method for machine vision as described in claim 4, characterized in that, The correlation matrix M 2 elements The calculation formula is: ; in, , , It is a parameter. Let C be the distance matrix. 2 The elements in.

6. The image global feature representation method for machine vision as described in claim 5, characterized in that, Step 23 also includes the following sub-steps: Step 231: When performing the third division, there are a total of 16 regions. Label the positions of the 16 regions. Step 232: Calculate the distance between two positions (i,j) and (p,q). The distance between two positions (i,j) and (p,q) is... Where 1≤i≤4, 1≤j≤4, 1≤p≤4, 1≤q≤4; Step 233: Calculate the distance between every two of the 16 positions to obtain a 16×16 distance matrix C. 3 ; Step 234: From the distance matrix C 3 Calculate the correlation matrix M 3 The correlation matrix M 3 It is a 16×16 square matrix, and the correlation matrix M 3 elements The calculation formula is: ; in, , , Let C be the distance matrix. 3 Elements in; Step 235: For M 3 Performing Cholesky decomposition yields matrix A. 3 A 3 It is a 16×16 diagonal matrix; Step 236: For f 3 After transformation, we obtain a 16×128 matrix B. 3 , for B 3 Perform matrix transformations to transform B 3 Transform it again into a 16×128 dimensional matrix .

Citation Information

Patent Citations

  • Linear spatial pyramid matching using sparse coding

    US20100124377A1