An intelligent knowledge base retrieval system

Through the intelligent knowledge base search system of image preprocessing and dynamic adjustment of feature points, the problem of low accuracy in crop image retrieval is solved, and efficient and accurate crop feature matching is achieved in complex scenarios, improving the accuracy of crop pest diagnosis and variety identification.

CN120179850BActive Publication Date: 2025-07-22HANGZHOU JINYUAN BIAOJU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510665455.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-22
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The existing SIFT algorithm is difficult to distinguish complex crop subjects and backgrounds in crop image retrieval, resulting in low retrieval accuracy and cannot dynamically adjust according to image characteristics.

Method used

Through image preprocessing, a binarized mask matrix and RGB feature map are generated, combined with the modular entropy value and dynamic scale adjustment, a multi-scale feature point set is extracted, and feature point density adjustment and quality evaluation is carried out, and a comprehensive similarity threshold and dynamic weight sorting mechanism is constructed to achieve the screening and matching of high-quality feature points.

Benefits of technology

It significantly improves the accuracy and efficiency of crop image retrieval, can accurately match crop characteristics in complex scenarios, reduces the impact of scale mismatch and feature distortion, and improves the accuracy of crop pest diagnosis and variety identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179850B_ABST
    Figure CN120179850B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of agricultural knowledge base retrieval, and discloses an intelligent knowledge base retrieval system, including: an image preprocessing unit, an image adjustment unit, an image matching unit, and a retrieval judgment unit. By image enhancement and pixel clustering, a binary mask matrix is generated, and the modulus entropy value is calculated in combination with parameters to provide a basis for feature extraction. Dynamic scale adjustment realizes multi-scale feature adaptive extraction, avoiding feature distortion. Through density dynamic adjustment and quality evaluation, feature points are evenly distributed, and high-quality feature points are screened. By constructing a comprehensive evaluation method, double screening of matching pairs is realized, and then by dynamically adjusting the similarity threshold and sorting, the deficiencies of the SIFT algorithm are solved, and the accuracy of image retrieval in complex scenes is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of agricultural knowledge base retrieval, and particularly relates to an intelligent retrieval system for a knowledge base. Background Art

[0002] In the field of image retrieval, there are currently various technical methods for retrieving images from a knowledge base. In the agricultural field, crop image retrieval is of great significance. For example, in scenarios such as crop pest and disease diagnosis and variety identification, it is necessary to quickly and accurately find images similar to the input image from a large number of crop knowledge base images to assist in decision-making. Staff take crop images in the fields and use an intelligent retrieval system to find similar images in the knowledge base to obtain corresponding information.

[0003] However, currently, the knowledge base often uses the SIFT algorithm to achieve retrieval. By detecting local feature points in the image and describing their features for matching and retrieval, in practical applications, due to the influence of factors such as the complex shooting environment of crop images and the morphology of crops, and the SIFT algorithm only focuses on the extraction and description of local feature points during retrieval, it is difficult to distinguish complex crop subjects and backgrounds, and at the same time, it cannot be dynamically adjusted according to the image characteristics, resulting in low accuracy of the output images after retrieval. Summary of the Invention

[0004] In view of the deficiencies of the prior art, the present invention provides an intelligent retrieval system for a knowledge base, which solves the above problems.

[0005] The above technical object of the present invention is achieved through the following technical solutions:

[0006] An intelligent retrieval system for a knowledge base, comprising:

[0007] An image preprocessing unit, configured to obtain an input image taken by a user in a farmland, unify the input image and the knowledge base image to a standard size, separate the standard size image, generate a binary mask matrix and a corresponding RGB feature map, and at the same time calculate the binary mask matrix to generate a modulus entropy value;

[0008] An image adjustment unit, configured to extract the RGB feature map to obtain a multi-scale set of feature points; generate a feature point distribution density map based on the binary mask matrix, dynamically adjust the feature point distribution density map to obtain a feature point density coefficient, calculate a dual index of the RGB feature map, and generate a feature point quality distribution heat map based on the dual index;

[0009] An image matching unit for bidirectionally matching the input image feature point set and the knowledge base image feature point set in the feature point set to obtain matching feature points, calculating the matching feature points to obtain the comprehensive similarity score between the input image feature point set and the knowledge base image feature point set, calculating the modulus entropy value, the heat map of feature point quality distribution, and the feature point density coefficient to obtain the similarity threshold, and at the same time marking the images higher than the similarity threshold as candidate matches and establishing a collection of the marked images;

[0010] A retrieval judgment unit for calculating the dynamic weight of the knowledge base image according to the comprehensive similarity score; performing a three-level sorting on the collection according to the dynamic weight.

[0011] Furthermore, separating the standard-size image to generate a binary mask matrix and a corresponding RGB feature map, including:

[0012] Enhancing the standard-size image to obtain an enhanced image;

[0013] According to the enhanced image, clustering the image pixels to divide the image into two parts, foreground and background, and generating an initial binary mask matrix;

[0014] Optimizing the initial binary mask matrix to obtain a final binary mask matrix;

[0015] Extracting the RGB features of each pixel of the enhanced image according to the final binary mask matrix to form an RGB feature map;

[0016] Statistical analysis of the ratio of the number of pixels in the target area to the total number of pixels in the binary mask matrix to obtain the pixel ratio parameter;

[0017] Calculating the gray-scale change frequency between adjacent pixels in the binary mask matrix to obtain the neighborhood change rate parameter;

[0018] Fusing the pixel ratio parameter and the neighborhood change rate parameter to obtain the modulus entropy value, specifically as follows: amplifying the original neighborhood change rate to obtain an amplified neighborhood change rate; adding the amplified neighborhood change rate to the base unit value and performing a non-linear mapping through the natural logarithm function to obtain a mapping result; multiplying the foreground pixel ratio by the mapping result to obtain a fusion result; calculating a penalty term according to the absolute difference between the foreground pixel ratio and the neighborhood change rate parameter; using the fusion result as the numerator term, adding the penalty term to the unit reference value as the denominator term, and performing the ratio operation of the two to obtain the modulus entropy value.

[0019] Furthermore, extracting the RGB feature map to obtain a multi-scale feature point set, including:

[0020] Calculate the texture complexity distribution in different regions of the RGB feature map to generate dynamic scale adjustment parameters;

[0021] Based on the dynamic scale adjustment parameters, perform hierarchical processing on the RGB feature map, and generate a set of local feature vectors with scale labels for each layer, specifically including: making a three-level division according to the range of the dynamic scale adjustment parameter values:

[0022] Divide the region with the normalized parameter value > 0.7 into a high texture complexity region and match it with a 5×5 small-scale window. Divide the interval where 0.3 ≤ normalized parameter value < 0.7 into a medium complexity region and match it with a 7×7 medium-scale window. Divide the interval with the normalized parameter value < 0.3 into a low complexity region and match it with a 9×9 large-scale window;

[0023] Calculate the weight coefficients of the feature vectors in each layer of the set of local feature vectors, and generate a set of feature points containing multi-scale spatial associations through weighted fusion.

[0024] Furthermore, based on the binary mask matrix, generate a feature point distribution density map, and perform dynamic adjustment on the feature point distribution density map to obtain the feature point density coefficient, including:

[0025] Evenly divide the image in the binary mask matrix into several small grids, calculate the number of feature points in each grid, and generate a two-dimensional density distribution map;

[0026] Calculate the average value of the distribution density in the two-dimensional density distribution map;

[0027] In the grids with a density higher than the average value, remove some feature points at a ratio of 50%;

[0028] In the grids with a density lower than the average value, generate new feature points by copying adjacent points to fill the sparse regions;

[0029] Calculate the variance of the feature point density in each grid after the removal and filling operations, and through normalization and mapping it to the 0-1 range, obtain the feature point density coefficient, specifically as follows: Calculate the ratio of the mean value to the standard deviation of the grid density, and then square this ratio to obtain a measure of the density dispersion. Use this as the numerator, and the denominator is 1 plus the measure of the density dispersion. Divide the numerator by the denominator to obtain the influence coefficient of the dispersion degree. At the same time, calculate the ratio of the adjusted total number of feature points to the original total number of feature points, and then substitute it into the hyperbolic tangent function to obtain the feature point number adjustment coefficient. Multiply the influence coefficient of the dispersion degree by the feature point number adjustment coefficient to obtain the feature point density coefficient.

[0030] Furthermore, the dual metrics of the RGB feature map include the local contrast value and the clustering degree value. Calculate the dual metrics of the RGB feature map, and generate a heat map of the feature point quality distribution based on the dual metrics, including:

[0031] In the RGB feature map, a circular area with a fixed radius is taken centered on each feature point, and the pixel gradient magnitude within the circular area is calculated to obtain the local contrast value;

[0032] The feature points are grouped and counted to obtain the clustering degree value;

[0033] The local contrast value and the clustering degree value are linearly combined to obtain the quality score of each feature point;

[0034] The RGB feature map is divided into several blocks of a fixed size, and the average value of the quality scores of the remaining feature points within each block is statistically calculated, and it is converted into a two-dimensional heat map through continuous color mapping, and finally a heat map of the feature point quality distribution is generated.

[0035] Furthermore, the input image feature point set and the knowledge base image feature point set in the feature point set are bidirectionally matched to obtain the matching feature points, including:

[0036] The feature point set includes an input image feature point set and a knowledge base image feature point set;

[0037] The cosine similarity of the feature point descriptors between the input image feature point set and the knowledge base image feature point set is calculated to obtain a rough matching result, and the feature point pairs that do not meet the geometric consistency in the rough matching result are excluded;

[0038] The deep features of the local area around the feature point pair are extracted, and the semantic similarity of the feature point pair is calculated. If the semantic similarity of the feature point pair is less than or equal to 30%, it is excluded;

[0039] Based on the rough matching result and the semantic similarity, double screening is performed to obtain the matching feature points that are both geometrically and semantically consistent.

[0040] Furthermore, the matching feature points are calculated to obtain the comprehensive similarity score of the input image feature point set and the knowledge base image feature point set, including:

[0041] Select and calculate the minimum sample set of the feature point pairs in the matching feature points to generate an initial affine transformation matrix;

[0042] The coordinates of the knowledge base image feature points are transformed to the input image coordinate system using the initial affine transformation matrix to obtain the predicted coordinates;

[0043] Calculate the distance residual between the actual coordinates and the predicted coordinates of the input image feature points to obtain the corresponding inliers;

[0044] Fit all the inliers of the initial affine transformation matrix to generate an affine transformation matrix;

[0045] Count the total number of feature point pairs in the matching feature points to generate image quantity indicators;

[0046] Calculate the input image feature point set and the knowledge base image feature point set in the matching feature points to generate the image quality mean index;

[0047] Calculate the internal points and matching feature points of the affine transformation matrix to generate the geometric consistency index of the image;

[0048] Calculate the average semantic similarity of feature point pairs to generate the semantic similarity index of the image;

[0049] The image quantity index, quality mean index, geometric consistency index and semantic similarity index are normalized and weighted summed to obtain a comprehensive similarity score.

[0050] Furthermore, the model entropy value, the heat map of the feature point mass distribution, and the feature point density coefficient are calculated to obtain the similarity threshold. At the same time, the images above the similarity threshold are marked as candidate matches, and the marked images are collected, including:

[0051] The normalized modulus entropy value, the characteristic point mass distribution heat map, and the characteristic point density coefficient are adjusted to obtain a first adjustment coefficient, a second adjustment coefficient, and a third adjustment coefficient, respectively;

[0052] Set the initial similarity threshold;

[0053] The three adjustment coefficients are calculated with the initial similarity threshold to obtain the similarity threshold;

[0054] Compare the composite similarity score of the input image with the similarity threshold:

[0055] If the comprehensive similarity score of the knowledge base image is greater than or equal to the similarity threshold, it is marked as a candidate match;

[0056] If the comprehensive similarity score of the knowledge base image is less than the similarity threshold, it will not be marked;

[0057] All candidate matching images are aggregated into a collection.

[0058] Furthermore, the dynamic weight of the knowledge base image is calculated based on the comprehensive similarity score, including:

[0059] Calculate the comprehensive similarity scores of all candidate images in the collection and generate an index result;

[0060] Calculate the sum of all index results as the denominator;

[0061] When the denominator is less than or equal to 0, all candidate images are assigned equal weights, which is the dynamic weight, and the weight value is the inverse of the number of candidate images;

[0062] When the denominator is greater than 0, divide the exponential result of each candidate image by the denominator to generate normalized weights that satisfy the sum of all weights being 1, which are the dynamic weights.

[0063] Calculate the dynamic weights of all images in the collection to obtain the standard deviation.

[0064] Take the natural logarithm of the number of candidate images to obtain the empirical coefficient.

[0065] Multiply the standard deviation by the empirical coefficient to obtain the tolerance threshold.

[0066] Furthermore, perform a three-level sorting on the collection according to the dynamic weights, including:

[0067] The three-level sorting is specifically as follows:

[0068] The first-level sorting is: sort the candidate images in the collection in descending order according to the dynamic weights from high to low.

[0069] The second-level sorting is: when there are groups of images with the same dynamic weights in the first-level sorting, perform a secondary sorting within the group in descending order according to the comprehensive similarity scores from high to low.

[0070] The third-level sorting is: when there are still images with the same comprehensive similarity scores in the second-level sorting, sort them in ascending order according to the image storage timestamp, and select the image with the earliest timestamp as the output image; if the storage timestamps are the same, then sort them in ascending order according to the image unique ID and select the image with the smallest ID as the output image.

[0071] Among the candidate images after the three-level sorting, obtain the dynamic weights of the first and second ranked images, calculate and obtain the difference.

[0072] If the difference after the three-level sorting is greater than the tolerance threshold, directly output the first ranked image after the three-level sorting.

[0073] If the difference after the three-level sorting is less than or equal to the tolerance threshold, trigger the local feature comparison mechanism.

[0074] In summary, the present invention mainly has the following beneficial effects:

[0075] Through image enhancement and pixel clustering techniques, the standard-sized image is segmented into foreground (main crop body) and background, generating an optimized binary mask matrix that can accurately depict the contour and distribution of the target area. Combining the modulus entropy value calculated from the pixel ratio parameter and the neighborhood change rate parameter, the structural complexity of the target area can be quantified, providing a priori mask guidance for subsequent feature extraction, effectively filtering out the interference of background noise. In the feature extraction stage, through dynamic scale adjustment, the image is divided into high, medium, and low complexity regions according to texture complexity, and sliding windows of scales 5×5, 7×7, and 9×9 are respectively matched to achieve the adaptive extraction of multi-scale local feature vectors. Compared with the fixed 16×16 scale feature description method of the SIFT algorithm, this application can dynamically adjust the receptive field for different morphological features such as crop leaf veins (high complexity regions) and large areas of solid-color leaf surfaces (low complexity regions), avoiding feature distortion caused by scale mismatch. By weighted fusion of multi-scale feature point sets, a feature representation containing spatial correlation information is constructed, significantly enhancing the feature characterization ability for complex-shaped crops, thereby making the matching ability during retrieval more accurate.

[0076] Through the dual optimization of density dynamic adjustment and quality assessment, in the density adjustment stage, a feature point distribution density map is generated through grid division. For regions with a density higher than the average, redundant feature points are removed at a ratio of 50%, and for low-density regions, they are filled by replicating neighboring points, making the feature points evenly distributed within the target area, avoiding matching ambiguities caused by local feature aggregation (such as the interference of feature point clusters in leaf overlapping regions). The feature point density coefficient obtained through normalization can quantitatively describe the degree of balance of feature distribution, providing a structural constraint for similarity calculation. In terms of feature point quality assessment, the local contrast value measures the significance of the feature point by calculating the average pixel gradient of the feature point neighborhood, and the clustering degree value evaluates the structural relevance of the feature point by statistically eliminating isolated points based on the cluster scale. The quality score obtained by fusing the two can effectively screen out high-quality feature points with both significance and structural stability. The further generated feature point quality distribution heat map can intuitively reflect the quality differences of feature points in different regions. Compared with the method of treating all feature points equally in the SIFT algorithm, it significantly reduces the interference of low-quality feature points on the matching result, improves the robustness of feature matching in complex lighting and occlusion scenarios, enhances the retrieval ability, and reduces the retrieval interference.

[0077] By constructing a comprehensive evaluation method that includes geometric transformation, semantic similarity, feature quality, and distribution, during the bidirectional matching process, first, rough matching is performed through cosine similarity to eliminate feature point pairs that do not meet geometric consistency. Further, deep semantic features are extracted to calculate semantic similarity. The double screening ensures the consistency of the matching pairs in terms of spatial position and semantic category, effectively solving the problem of cross-category mis-matching caused by the lack of semantic constraints in the SIFT algorithm (such as the misjudgment between pest and disease spots and natural damage patches). The comprehensive similarity score integrates four major indicators: the number of feature point pairs, the mean quality, the geometric consistency ratio, and the mean semantic similarity. Through normalization and weighting, multi-dimensional measurement of the overall similarity of the image is achieved. Based on the similarity threshold dynamically adjusted by the modulus entropy value, the feature point density coefficient, and the quality heat map, the screening criteria for candidate matches can be adaptively set to avoid retrieval bias of fixed thresholds in different scenarios. During the sorting stage, through dynamic weight calculation and a three-level sorting mechanism (weight first, score second best, timestamp as a fallback), it is ensured that the retrieval results not only meet the similarity priority but also conform to reality, and a tolerance threshold mechanism is added to intelligently judge whether to trigger local feature comparison, avoiding the omission of the optimal solution image due to minor score differences while ensuring retrieval efficiency. Finally, the collaborative improvement of retrieval accuracy and recall rate in complex farmland scenarios is realized. Compared with the single feature matching and fixed sorting strategy of the traditional SIFT algorithm, in scenarios with extremely high precision requirements such as crop pest and disease diagnosis and variety identification, the retrieval efficiency and matching accuracy are improved, making the retrieved output images more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 is a block diagram of the knowledge base intelligent retrieval system of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0079] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present invention.

[0080] Refer to Figure 1 , a knowledge base intelligent retrieval system, including:

[0081] An image preprocessing unit for obtaining an input image taken by a user in a farmland, unifying the input image and the knowledge base image to a standard size, separating the standard size image to generate a binary mask matrix and a corresponding RGB feature map, and at the same time calculating the binary mask matrix to generate a modulus entropy value;

[0082] An image adjustment unit for extracting RGB feature maps to obtain a multi-scale set of feature points; generating a feature point distribution density map based on a binary mask matrix, dynamically adjusting the feature point distribution density map to obtain a feature point density coefficient, calculating a dual index of the RGB feature map, and generating a feature point quality distribution heat map based on the dual index;

[0083] An image matching unit for bidirectionally matching the input image feature point set and the knowledge base image feature point set in the set of feature points to obtain matching feature points, calculating the comprehensive similarity score of the input image feature point set and the knowledge base image feature point set, calculating the modulus entropy value, the feature point quality distribution heat map, and the feature point density coefficient to obtain a similarity threshold, and at the same time marking the images higher than the similarity threshold as candidate matches and establishing a collection of the marked images;

[0084] A retrieval judgment unit for calculating the dynamic weight of the knowledge base image according to the comprehensive similarity score; performing a three-level sorting on the collection according to the dynamic weight.

[0085] The input image and the knowledge base image are standardized through the image preprocessing unit to generate a binary mask matrix and an RGB feature map, and at the same time the modulus entropy value is calculated, providing standardized and diverse basic data for subsequent image analysis and ensuring the integrity and analyzability of image information. The image adjustment unit extracts a multi-scale set of feature points, combines the feature point distribution density map and the feature point density coefficient obtained by dynamic adjustment, and the feature point quality distribution heat map generated based on the dual index, comprehensively characterizing the image features from multiple dimensions such as scale, density, and quality, and improving the depth and accuracy of feature analysis. The image matching unit obtains matching feature points through bidirectional matching, calculates the comprehensive similarity score, and combines the modulus entropy value, etc. to obtain a similarity threshold to screen candidate matching images, avoiding the limitations of a single index and improving the accuracy and reliability of image matching. The retrieval judgment unit calculates the dynamic weight according to the comprehensive similarity score and performs a three-level sorting on the candidate images, making the retrieval results more in line with actual needs and significantly improving the retrieval efficiency and quality.

[0086] In one case of this embodiment, separating the standard-size image to generate a binary mask matrix and the corresponding RGB feature map includes:

[0087] Enhancing the standard-size image to obtain an enhanced image, specifically including: dividing the standard-size image into non-overlapping sub-blocks of 8×8 pixels, calculating the cumulative distribution function for each sub-block and performing gray mapping, limiting the contrast enhancement factor to 4.0, then using a 5×5 median filter for denoising, scanning the image pixel by pixel, sorting the pixel values in the neighborhood of this point and taking the median to replace the original pixel value, and then performing edge sharpening using the Laplace operator, using the template Perform a convolution operation to enhance the high-frequency edge information in the image, normalize the gray values of the processed image to the range of 0-255 through linear transformation, and generate the enhanced image;

[0088] According to the enhanced image, cluster the image pixels to divide the image into two parts: foreground (object) and background, and generate an initial binary mask matrix. Specifically, it includes: converting the image from the RGB space to the HSV space, extracting the three-channel features of hue, saturation, and brightness, using the K-means clustering algorithm to cluster the pixels, setting K = 2, using the three-channel features of hue, saturation, and brightness as feature vectors, stopping when iterating 100 times, calculating the Euclidean distance from each pixel to the two cluster centers, marking the category with the smaller distance as the foreground (object), and the other category as the background to obtain the clustering result. Perform morphological processing on the clustering result, use a 3×3 structuring element for opening operation to remove small noise points, and then perform closing operation to fill internal holes, finally generating the initial binary mask matrix, where the foreground pixel value is 255 and the background is 0;

[0089] Optimize the initial binary mask matrix to obtain the final binary mask matrix. Specifically, it includes: using the Sobel operator to calculate the gradients in the horizontal and vertical directions of the initial binary mask matrix, taking the square root of the sum of the squares of the gradient magnitudes in the two directions to obtain the total gradient magnitude, setting the gradient threshold to 80, marking the pixels with gradient magnitudes greater than this threshold as edge candidate points, and applying the Canny edge detection algorithm to these candidate points: first perform non-maximum suppression to retain the maximum points in the local gradient direction, and then use double-threshold processing to connect the strong edges detected by the low threshold of 40 and the high threshold of 120 to form a continuous and closed edge contour;

[0090] Perform a morphological dilation operation on the edge, use a 3×3 cross-shaped structuring element to expand the edge width by 1 pixel, perform a logical AND operation on the dilated edge and the initial mask to correct the edge offset caused by clustering errors, and use the foreground area in the initial mask as the seed point to perform region growing on the enhanced RGB feature map: calculate the Euclidean distance between the neighborhood pixels and the average RGB value of the seed region, set the threshold to 15, and if the Euclidean distance is less than the threshold of 15, then incorporate the pixel into the growing region;

[0091] Calculate the number of pixels in each connected growing region, retain the 3 regions with the largest area, remove the noise regions with too small area, smooth the boundaries of the retained regions, and use Gaussian filtering to eliminate jagged edges to generate the final binary mask matrix;

[0092] Perform per-pixel RGB feature extraction on the enhanced image according to the final binary mask matrix to form an RGB feature map, specifically including: traversing the binary mask matrix pixel by pixel. For foreground pixels with a value of 255, directly extract the RGB triple values at the corresponding positions from the enhanced image. For background pixels with a value of 0, use the KNN interpolation algorithm for processing: search for foreground pixels within a 3×3 neighborhood. If there are at least 2 foreground pixels, take the weighted average of their RGB values (where the weight is inversely proportional to the distance), otherwise retain the RGB values of the original enhanced image. After extraction, perform median filtering (3×3 window) on the RGB feature map to remove isolated noise points, and finally generate the RGB feature map;

[0093] Statistically calculate the ratio of the number of pixels in the target area to the total number of pixels in the binary mask matrix to obtain the pixel ratio parameter, specifically including: obtain the size of the mask matrix. Assuming the image resolution is M×N pixels, the total number of pixels is M×N. Traverse each pixel in the mask matrix row by row and column by column. For pixels with a pixel value of 255 (target area), accumulate the count to obtain the number of pixels in the target area. Divide the number of target pixels by the total number of pixels to obtain the pixel ratio parameter;

[0094] Calculate the gray-level change frequency between adjacent pixels in the binary mask matrix to obtain the neighborhood change rate parameter, specifically including: determine the four-neighborhood range of each pixel in the binary mask matrix, that is, the four directly adjacent pixel positions above, below, left, and right (where boundary pixels do not participate in the calculation due to lack of a complete neighborhood), and traverse each pixel in the binary mask matrix row by row and column by column (except for the rows and columns at the image edges). For the current pixel value (0 represents the background, 255 represents the foreground), check the gray-level values of its four adjacent pixels above, below, left, and right respectively. If the gray-level value of the current pixel is different from that of any adjacent pixel, it is regarded as a valid gray-level change and counted. After traversing the entire image, count the total number of all gray-level changes, and at the same time calculate the total number of adjacent pixel pairs in the horizontal and vertical directions of the image, specifically:

[0095] In the horizontal direction, each row has "image width minus 1" pairs of adjacent pixels (left and right adjacent), and in the vertical direction, each column has "image height minus 1" pairs of adjacent pixels (up and down adjacent). Add the two to get the total number of adjacent pixel pairs in the entire image. Divide the total number of gray-level changes by the total number of adjacent pixel pairs to obtain the neighborhood change rate parameter;

[0096] Fuse the pixel ratio parameter and the neighborhood change rate parameter to obtain the modulus entropy value, including: amplify the original neighborhood change rate to obtain the amplified neighborhood change rate ; Add the amplified neighborhood change rate to the base unit value and perform a non-linear mapping through the natural logarithm function to obtain the mapping result ; Multiply the foreground pixel ratio with the mapping result to obtain a fusion result ; Calculate the penalty term based on the absolute difference between the pixel ratio parameter and the neighborhood change rate parameter ; Use the fusion result as the numerator term and the penalty term plus the unit reference value as the denominator term , and perform the ratio operation on the two to obtain the modulus entropy value . When specifically applied, it can be achieved through the following calculation formula, for example:

[0097] ;

[0098] In the formula, is the modulus entropy value, represents the foreground pixel ratio (number of target area pixels / total number of pixels), represents the neighborhood change rate, represents the pixel ratio weight coefficient (default is 1.2), represents the neighborhood change rate amplification factor (default is 0.8), represents the difference penalty factor (default is 0.5), represents the constant to prevent a zero denominator, represents the natural logarithm function.

[0099] By dividing the image into non - overlapping 8×8 pixel sub - blocks for cumulative distribution function calculation and gray - level mapping, and restricting the contrast enhancement factor to 4.0, not only the local contrast of the image is improved, but also the noise amplification caused by over - enhancement is avoided. The denoising process of the 5×5 median filter effectively eliminates random noises such as salt - and - pepper noise in the image, while retaining the edge information of the image. The edge sharpening of the Laplacian operator further enhances the high - frequency edge information in the image, making the details of the image clearer and significantly improving the quality of the image.

[0100] By converting the image from the RGB space to the HSV space for K-means clustering, the foreground and background can be more effectively distinguished. The opening operation and closing operation in morphological processing remove small noise points and fill internal holes respectively, improving the quality of the mask matrix. Through the accurate extraction and optimization of edges by the Sobel operator and the Canny edge detection algorithm, and the correction of edges by the region growing algorithm, the accuracy of the mask matrix is further improved. Finally, the 3 regions with the largest areas are retained and the boundaries are smoothed, effectively removing the noise regions and making the generated binary mask matrix more accurate. Based on the RGB feature map extracted from this mask matrix, the background pixels are processed by the KNN interpolation algorithm, which not only retains the overall information of the image but also improves the quality of the feature map. The calculation of the pixel ratio parameter and the neighborhood change rate parameter depicts the features of the image from different angles, providing rich information for the calculation of the modulus entropy value and enabling the modulus entropy value to more comprehensively reflect the features of the image.

[0101] In one case of this embodiment, the RGB feature map is extracted to obtain a multi-scale feature point set, including:

[0102] Calculate the texture complexity distribution of different regions in the RGB feature map to generate a dynamic scale adjustment parameter, specifically including: processing each pixel point in the RGB feature map with a multi-scale sliding window (using three odd sizes of 5×5, 7×7, and 9×9). Taking the pixel at the center of the window as a reference, count the number of occurrences of the gray value combinations of adjacent pixel pairs in four main directions of 0°, 45°, 90°, and 135° to form a gray-level co-occurrence matrix. For any two pixels within each window, record the frequency of the gray value combinations at these two positions at a specified direction and a 1-pixel distance to generate a co-occurrence probability matrix. The entropy value is obtained by accumulating the product of each probability value in the co-occurrence probability matrix and the logarithm of this probability value. The entropy value is used to reflect the degree of chaos of the gray distribution within the region. The higher the entropy value, the more complex the texture (such as edge and noise regions), and the lower the entropy value, the simpler the texture (such as smooth regions). After calculating pixel by pixel and window by window, statistically calculate the maximum and minimum values of the global entropy value, and map the entropy values of each region to the [0, 1] interval through linear transformation to obtain the dynamic scale adjustment parameter;

[0103] Based on the dynamic scale adjustment parameter, the RGB feature map is hierarchically processed, and a set of local feature vectors with scale labels is generated for each layer, specifically including: performing three-level division according to the range of the dynamic scale adjustment parameter values:

[0104] Divide the region with the normalized parameter value > 0.7 into a high texture complexity region and match it with a 5×5 small-scale window;

[0105] Divide the interval of 0.3 ≤ normalized parameter value < 0.7 into a medium complexity region and match it with a 7×7 medium-scale window;

[0106] Divide the interval with a normalized parameter value less than 0.3 into a low-complexity region and match a 9×9 large-scale window;

[0107] During division, mirror padding is performed on the boundary pixels to ensure that all pixels can fit into a complete window. For each level of pixels, with the current pixel as the center, extract the RGB three-channel pixel values within the local region according to the window scale. For each channel, calculate the mean and variance of the pixel values within the window. At the same time, calculate the covariance between the R-G, G-B, and R-B channels. Concatenate the original three-channel values, three-channel means, three-channel variances, and three pairs of covariances, a total of 9 eigenvalue vectors, to form a local feature vector. When generating the feature vector, add a clear scale identifier to each feature vector. The identifier content directly corresponds to the window size specification used, and the label content is directly associated with the window size used. When processing the image boundary, if the window exceeds the image range, fill the missing area by mirroring and copying the edge pixels to ensure the consistency of feature extraction. After completing the above operations pixel by pixel, aggregate the feature vectors by scale level to form a set of local feature vectors containing small, medium, and large scale labels;

[0108] Calculate the weight coefficients of the feature vectors at each level in the set of local feature vectors and generate a set of feature points with multi-scale spatial associations through weighted fusion. Specifically, for each pixel point, use its corresponding normalized parameter value as the weight for the small-scale layer (5×5), the weight for the medium-scale layer (7×7) is "1 minus the parameter value", and the weight for the large-scale layer (9×9) is "1 - small-scale layer weight - large-scale layer weight", ensuring that the sum of the weights of the three layers is 1. Multiply the small, medium, and large scale feature vectors of the same pixel point by their corresponding weights and then accumulate them to generate a fusion vector containing 9-dimensional statistical features. Retain the scale labels of each scale as context association parameters during fusion to generate a set of feature points with multi-scale spatial associations.

[0109] Through dynamic scale adjustment parameters and hierarchical processing, precise adaptation to the complexity of image texture and multi-dimensional feature extraction are achieved. When calculating the dynamic scale adjustment parameters, multi-scale sliding windows and gray-level co-occurrence matrices are used to statistically calculate the frequency of gray-level value combinations in four main directions, obtaining the entropy value that can reflect the degree of texture chaos. Then, through linear transformation, it is mapped to the interval [0, 1], so that each region has a corresponding parameter value. Based on this parameter value, a three-level division is carried out, enabling high, medium, and low texture complexity regions to respectively match different scale windows. Small windows can capture fine features in regions with rich details such as edges and noises, and large windows can retain the overall structure of smooth regions, avoiding the limitations of a single scale window in complex texture scenarios. At the same time, for each channel, the mean, variance, and covariance between channels are calculated, and nine eigenvalue vectors are concatenated to form a local feature vector, comprehensively covering the intensity distribution, fluctuation situation, and inter-channel correlation of pixels within the region, effectively enhancing the richness and accuracy of image feature expression.

[0110] By calculating the weight coefficients of each layer, using the normalized parameter value as the weight of the small-scale layer, the weight of the medium-scale layer is "1 minus the parameter value", and the weight of the large-scale layer is obtained by balancing the weights of the first two layers, so that the sum of the three-layer weights is always 1. This weight allocation strategy based on texture complexity enables the feature vectors in high texture complexity regions to fuse more detailed information extracted by small-scale windows, while low texture complexity regions focus on the overall structural features captured by large-scale windows, meeting the feature expression requirements of different texture regions, avoiding the blindness of weight allocation. At the same time, when generating the fusion vector, the scale labels of each scale are retained as context correlation parameters, so that the final feature point set not only contains 9-dimensional statistical features but also carries clear scale identifiers, providing rich scale context information for subsequent image feature analysis. In image recognition, it can more accurately capture the multi-scale features of objects and improve the image recognition ability.

[0111] In one case of this embodiment, based on the binary mask matrix, a feature point distribution density map is generated, and the feature point distribution density map is dynamically adjusted to obtain the feature point density coefficient, including:

[0112] The image in the binarized mask matrix is evenly divided into several small grids, the number of feature points in each grid is calculated, and a two-dimensional density distribution map is generated, specifically including: the image in the binarized mask matrix is horizontally divided into aq grids and vertically divided into bq grids, so that the entire image forms a grid array of aq×bq. In a point-by-point traversal manner, for each feature point in the multi-scale feature point set, according to its coordinate position in the image, the horizontal grid serial number and vertical grid serial number to which the point belongs are determined, and the feature points are sequentially assigned to the corresponding grids. Each feature point is traversed one by one, and the count of the grid cell to which it belongs is incremented by 1. Finally, the number of feature points in each grid is obtained, and these numerical values are arranged in the row and column order of the grids to form a two-dimensional matrix, which is the feature point distribution density map;

[0113] Calculate the average distribution density in the two-dimensional density distribution map, specifically including: traversing the feature point distribution density map row by row and column by column, accumulating and summing the number of feature points in each grid to obtain the total number of feature points in all grids, then calculating the total number of grids, that is, the product of the number of horizontal grids aq and the number of vertical grids bq, and finally dividing the total by the total number of grids to obtain the average number of feature points in each grid, which is the average distribution density;

[0114] In the grids with a density higher than the average value, part of the feature points are removed at a ratio of 50%, specifically including: traversing the two-dimensional density distribution map row by row and column by column, comparing the number of feature points in each grid with the average distribution density, screening out the grids with the number of feature points greater than the average value. For each such grid, multiply the current number of feature points in the grid by 50% (round down to ensure an integer), and remove these numbers of feature points;

[0115] In the grids with a density lower than the average value, new feature points are generated by copying adjacent points to fill the sparse areas, specifically including: for the grids with a density lower than the average value, first screen out the grids with the number of feature points less than the average distribution density, calculate the number of feature points to be supplemented (that is, the average value minus the current number, rounded up), determine the four neighboring grids of the grid (for boundary grids, only the existing neighboring grids are taken), extract all feature points from the neighboring grids, sort them in ascending order according to the Euclidean distance from the coordinate of the feature point in the image to the center of the current grid, preferentially select the feature points with a closer distance, and according to the number of feature points to be supplemented, equally spaced select adjacent feature points, copy their multi-scale feature vectors and coordinate information to generate new feature points, and insert the new feature points into the current grid equally divided by the grid rows and columns, so that the number of feature points in the adjusted grid reaches the average distribution density, and the sparse area filling is completed;

[0116] Calculate the variance of the feature point density in each grid after removal and filling operations. Through normalization and mapping it to the range of 0-1, the feature point density coefficient is obtained. When the density coefficient approaches 0, it indicates a highly uniform distribution, and when it approaches 1, it indicates a highly dispersed distribution. The specific calculation formula of the density coefficient is as follows:

[0117] ;

[0118] In the formula, represents the feature point density coefficient, represents the standard deviation of the adjusted grid density, represents the mean of the adjusted grid density, represents the total number of adjusted feature points, is the original total number of feature points, represents the hyperbolic tangent function.

[0119] Through the method of grid division and dynamic adjustment, the uniformity and rationality of the feature point distribution are significantly improved. The binary mask image is evenly divided into a grid array, and feature points are accurately assigned to the corresponding grids based on coordinates to form a two-dimensional matrix reflecting the regional feature density. Remove feature points in the high-density area at a ratio of 50%, which can not only avoid the computational redundancy caused by excessive local features but also prevent the matching ambiguity caused by the aggregation of feature points, ensuring that the feature points in the key area maintain sufficient distinctiveness. For the low-density area, through the Euclidean distance sorting and equally spaced replication of the feature points in adjacent grids, the sparse area can be filled directionally. On the premise of retaining the original feature vector and coordinate information, the number of grid feature points can be accurately matched with the average distribution density. This two-way adjustment strategy effectively balances the global feature density and avoids the problems of feature distortion or information loss in traditional retrieval, providing a basis of feature points with uniform distribution and moderate density for subsequent image matching, and significantly improving the robustness of the algorithm in complex scenarios.

[0120] By mapping the feature distribution uniformity to the range of 0-1, the accurate characterization of the distribution state is realized. Among them, the non-linear mapping of the tanh function to the change of the total number of feature points before and after adjustment effectively reflects the dynamic balance of the feature point scale, and the normalization process converts the relative relationship between the mean and the standard deviation into a comparable quantitative index, making the density coefficient not only sensitive to capture the distribution difference but also applicable to cross-scale scenarios. When the density coefficient approaches 0, it indicates that the feature points are highly uniformly distributed in the global grid, which can minimize the interference of local feature imbalance on the algorithm performance. When it approaches 1, it warns of a significant distribution deviation, providing a clear direction for subsequent targeted optimization. It not only provides a traceable evaluation basis for the adjustment effect of feature points but also ensures the effectiveness and authenticity of the supplementary feature points through the replication strategy of retaining multi-scale feature vectors, facilitating the subsequent image matching.

[0121] In one case of this embodiment, the dual metrics of the RGB feature map include the local contrast value and the clustering degree value. Calculating the dual metrics of the RGB feature map and generating a heat map of the feature point quality distribution based on the dual metrics. Calculating the dual metrics of the RGB feature map and generating a heat map of the feature point quality distribution includes:

[0122] In the RGB feature map, take a circular region with a fixed radius centered on each feature point, calculate the average value of the pixel gradient magnitudes within the circular region, and map this average value to a local contrast value between 0 and 1. Specifically, it includes: In the RGB feature map, for each feature point, construct a circular region centered on its coordinates with a set radius. By judging whether the Euclidean distance from the pixel to the center is less than or equal to the radius, filter out all pixels within the region where the Euclidean distance is less than or equal to the radius. Use the Sobel operator to calculate the horizontal and vertical gradients of these pixels respectively. Obtain the gradient magnitude of each pixel by taking the square root of the sum of squares. Calculate the mean value of the gradient magnitudes of all pixels within the region to get the mean value of the gradient magnitudes of this region. Then compare this mean value of the gradient magnitudes with the minimum and maximum values of the mean values of the gradient magnitudes of all circular regions in the whole map. Use the global normalization method of (current mean - minimum value) divided by (maximum value - minimum value + a very small amount) to map the mean value to a local contrast value in the range of 0 to 1;

[0123] Group the feature points to eliminate isolated points, count the cluster size (number of feature points) of each point's belonging cluster, and map the cluster size to a clustering degree value between 0 and 1. Specifically, it includes: First, set the neighborhood radius and the minimum number of included points. Traverse each feature point, calculate the number of feature points within the radius range of its neighborhood. If the number ≥ the minimum number of included points, it is determined as a core point. If the number < the minimum number of included points but within the neighborhood of a core point, it is a boundary point. Otherwise, it is regarded as a noise point (isolated point) and eliminated. Divide the core points and their reachable boundary points into the same cluster to form a connected region. Count the total number of feature points (i.e., the cluster size) of each effective feature point's belonging cluster. Through global normalization, subtract the minimum size of all non - isolated point clusters in the whole map from each cluster size, and then divide by (the maximum cluster size in the whole map - the minimum cluster size + a very small amount) to linearly map the cluster size to a clustering degree value between 0 and 1;

[0124] Linearly combine the local contrast value and the clustering degree value according to the weight ratio of 60% and 40% to obtain the quality score of each feature point. Among them, sort the feature points from high to low according to the quality score, retain the top 70% of the high - quality feature points, and delete the remaining feature points. The specific calculation formula of the quality score is as follows:

[0125] ;

[0126] In the formula, is the quality score, Represents local contrast, Represents clustering degree, And Respectively represent the first exponential slope parameter and the second exponential slope parameter (default values are 0.5 and 0.3 respectively), Represents the natural constant, Represents the normalized gradient magnitude;

[0127] The RGB feature map is divided into several blocks of fixed size, and the average value of the quality scores of the remaining feature points in each block is statistically calculated. It is converted into a two-dimensional heat map through continuous color mapping, and finally a heat map of the feature point quality distribution is generated. Specifically, in the RGB feature map, the image is divided into non-overlapping rectangular blocks according to a fixed size (M×N pixels) to cover the entire image. For each block, the top 70% of the high-quality feature points previously retained are screened out. If there are feature points in the block, the average value of their quality scores is calculated. If it can be regarded as the background, the average value of the quality scores of all blocks is linearly normalized and mapped to the color space (0-255 interval), and pixel coloring is performed using a continuous color gradient (from cold color to warm color representing low value to high value), and bilinear interpolation is used to smooth the boundaries of adjacent blocks, and finally a heat map of the feature point quality distribution is generated.

[0128] Through the dual-index design of the local contrast value and the clustering degree value, a multi-dimensional feature point quality evaluation is constructed. The local contrast value is based on the mean value of the regional gradient magnitude extracted by the Sobel operator, which can accurately depict the texture complexity and edge significance of the feature point neighborhood, enabling feature points in high-contrast regions (such as object contours and rich texture details) to obtain higher weights and avoiding interference from invalid features in low-texture regions. The clustering degree value uses the DBSCAN clustering idea to eliminate isolated points and reflects the spatial aggregation degree of feature points through cluster scale normalization, ensuring that the retained feature points have good spatial correlation and structural stability. Especially in complex scenes, it can effectively filter out noise points (such as random noise and outlier pixels), improving the overall reliability of the feature point set. The two are linearly combined with weights to form a quality score, which not only pays attention to the significance of local detail features but also takes into account the relevance of the global structure, making the feature point quality evaluation have both local sensitivity and global structural properties.

[0129] The fixed-size block division strategy ensures the consistency of the evaluation units, which is convenient for subsequent quantitative analysis and cross-regional comparison. The mechanism of retaining the top 70% high-quality feature points avoids excessive information loss while filtering noise, and balances the quantity and quality of feature points. The visualization scheme based on continuous color gradient (cold colors represent low values and warm colors represent high values) can clearly present the spatial differences in the quality of feature points. In the image semantic segmentation task, it can quickly locate high-value feature-enriched areas (such as the boundaries of the target body). The auxiliary algorithm optimizes the feature extraction strategy, and the bilinear interpolation smoothing technology eliminates the jagged effect of the block boundary, making the transition of the heat map natural, more in line with the laws of human visual perception, and ensuring the recognition ability of the later image.

[0130] In one case of this embodiment, bidirectional matching is performed on the input image feature point set and the knowledge base image feature point set in the feature point set to obtain matching feature points, including:

[0131] The feature point set (the first 70% high-quality feature point set after arrangement) includes the input image feature point set and the knowledge base image feature point set;

[0132] The cosine similarity of the feature point descriptors between the input image feature point set P1 and the knowledge base image feature point set P2 is calculated to obtain a rough matching result, and the feature point pairs that do not meet the geometric consistency in the rough matching result are eliminated, specifically including: for the input image feature point set and the knowledge base image feature point set, each feature point descriptor (ORB vector) is first extracted, and the cosine similarity algorithm is used to calculate the cosine value of each feature point descriptor in the input image feature point set and all feature point descriptors in the knowledge base image feature point set. The two nearest feature points are found through cosine similarity calculation. Only when the ratio of the similarity of the nearest neighbor descriptor to the similarity of the next nearest neighbor descriptor is less than the empirical threshold (0.8), the matching pair is retained to obtain a rough matching result, and the projection error of the matching pair in the rough matching result is calculated through the RANSAC algorithm, and the non-consistent pairs with errors exceeding the threshold (3 pixels) are eliminated, and the geometrically consistent matching point pairs are retained;

[0133] Extract the deep features of the local area around the feature point pairs, calculate the semantic similarity of the feature point pairs. If the semantic similarity of the feature point pairs is less than or equal to 30%, then eliminate them, specifically including: for the geometrically consistent feature point pairs, respectively intercept local areas with a fixed size (32×32 pixels) centered on the feature points in the input image and the knowledge base image. After size adjustment (adjusted to 224×224 pixels) and normalization preprocessing, extract high-dimensional feature vectors (2048 dimensions) as deep semantic representations through a multi-level feature extraction method. For the high-dimensional feature vectors of each pair of feature points, use the cosine similarity algorithm to calculate their similarity. If the similarity value ≤ 30%, it is determined that the semantics do not match, eliminate the feature point pair, and retain the effective matching pairs with similarity higher than the threshold;

[0134] Based on the rough matching results and semantic similarity, perform double screening to obtain matching feature points that are both geometrically and semantically consistent.

[0135] Construct a matching set through the top 70% high-quality feature points, effectively ensuring the retrieval reliability of agricultural image feature points, avoiding the influence of low-quality features introduced by complex field environments (such as weed interference, leaf overlap, etc.) on the matching results, and improving the retrieval quality of basic agricultural image data. In the rough matching stage, combining the cosine similarity and the nearest neighbor and next-nearest neighbor ratio threshold screening strategy can specifically filter out irrelevant feature point pairs in agricultural images caused by natural light changes, background noise, etc., while retaining the effective matching pairs, eliminating about 20% of the false matches, laying a data foundation for agricultural image retrieval.

[0136] Dynamically calculate the projection error through the RANSAC algorithm, which can robustly exclude inconsistent matching pairs caused by factors such as differences in crop growth forms (such as changes in plant perspectives at different growth stages) and field occlusions, making the matching pairs strictly meet the spatial geometric constraints in the agricultural scenario, significantly improving the spatial structure consistency of similar images in the retrieval results. The deep semantic feature screening effectively captures agricultural semantic associations that are difficult to describe by traditional geometric features (such as different disease symptoms with approximate textures, varietal differences of crops with similar morphologies) by extracting deep features such as crop leaf textures, lesion morphologies, and ear structures in agricultural images and calculating the cosine similarity, and precisely eliminates feature pairs that are only geometrically matched but semantically irrelevant, breaking through the retrieval difficulties of similar textures and complex structures of the same type of objects in agricultural images. The double screening mechanism forms a collaborative optimization of geometric constraints and semantic constraints, ensuring both the spatial position consistency of the retrieved images in complex field lighting, multi-angle shooting, partial occlusion and other scenarios, and strengthening the relevance of key semantic information such as crop varieties and pests and diseases, so that the retrieval results of agricultural knowledge base images maintain high accuracy in practical applications with variable natural environments and subtle differences in target features.

[0137] In a case of this embodiment, the comprehensive similarity score of the input image feature point set and the knowledge base image feature point set is calculated by calculating the matching feature points, including:

[0138] Select the minimum sample set of the feature point pairs in the matching feature points, calculate the minimum sample set, and generate an initial affine transformation matrix, specifically including: randomly select 3 pairs of matching points (6 independent parameters are required for affine transformation, and each pair of two-dimensional points provides 2 coordinate constraints), ensure that the 3 pairs of matching points are not collinear in the image plane, subtract the coordinate means of all matching points in their respective images from the horizontal and vertical coordinates of each pair of points respectively, and then scale them to a unified scale. Based on the 3 pairs of preprocessed coordinates, analyze the transformation relationships of translation, rotation, scaling, and shear between the input image and the knowledge base image. By minimizing the coordinate conversion error (that is, finding a set of transformation parameters to minimize the position difference between the input points after transformation and the corresponding points in the knowledge base), calculate the preliminary transformation parameter combination including the translation amount, rotation angle, scaling factor, and shear coefficient. For each feature point coordinate in the input image, calculate its predicted coordinate in the knowledge base image, and then calculate the Euclidean distance (i.e., pixel-level position deviation) between the predicted coordinate and the actual matching point coordinate. If the deviation of a pair of feature points is less than 3 pixels, it is determined as an inlier (a valid match with geometric consistency), otherwise it is regarded as an outlier (an invalid point affected by noise or incorrect matching). Count the number and proportion of inliers. If the number of inliers is sufficient (more than 50% of the total number of matching points), the current initial transformation matrix is considered reliable. If the number of inliers is insufficient (not more than 50% of the total number of matching points), randomly select 3 new non-collinear sample points from the matching point set again and continue the calculation until a set of initial transformation matrices that meet the requirements of the number of inliers is obtained, and the initial affine transformation matrix is obtained;

[0139] Apply the initial affine transformation matrix to all matching feature points, and transform the knowledge base image feature point coordinates to the input image coordinate system through the initial affine transformation matrix to obtain the predicted coordinates, specifically including: expand the two-dimensional coordinates (x, y) of each feature point in the knowledge base image into homogeneous coordinates (x, y, 1), and then perform the translation, rotation, scaling, and shear operations in the initial affine transformation matrix on each homogeneous coordinate vector in turn (first update the coordinates through linear transformation, and then add the translation parameters), and then convert the transformed homogeneous coordinates (X, Y, W) into two-dimensional coordinates (X / W, Y / W) to obtain the predicted coordinates;

[0140] Calculate the distance residuals between the actual coordinates and the predicted coordinates of the feature points in the input image. Take the median of the distance residuals of all feature points as the residual threshold. Mark the feature points with distance residuals less than the residual threshold as inliers and record the number of inliers. Specifically, for each pair of matched feature points, calculate the Euclidean distance between the actual coordinates of the feature points in the input image and the predicted coordinates obtained by mapping the feature points of the knowledge base image to the input image coordinate system through the initial affine transformation matrix as the distance residual of this pair of feature points. After collecting the distance residual values of all pairs of matched feature points, sort them in ascending order and take the middle value of the sorted residual sequence as the residual threshold. Traverse all pairs of matched feature points again, compare the distance residual of each pair of feature points with this threshold, mark the feature point pairs with distance residuals less than the threshold as inliers, and count the number of all feature point pairs marked as inliers;

[0141] Fit the number of inliers of all initial affine transformation matrices and select the transformation matrix with the largest number of inliers to generate the affine transformation matrix. Specifically, for each initial affine transformation matrix randomly sampled each time, count the corresponding number of inliers, and then record the transformation matrix and its number of inliers obtained in each iteration. When the number of iterations reaches 100 times, select the transformation matrix with the largest number of inliers. Based on this matrix, use all the corresponding inlier coordinates to re-estimate the transformation parameters by the least squares method to generate the final optimized affine transformation matrix;

[0142] Count the total number of feature point pairs in the matched feature points to generate the image quantity index. Specifically, count the valid matched feature point pairs after double verification of geometric consistency and semantic similarity to obtain the total number of pairs. Calculate the ratios of the total number of pairs to the total number of feature points in the input image and the total number of feature points in the knowledge base image respectively, and denote them as the first ratio and the second ratio. Take the geometric mean of these two ratios as the image quantity index;

[0143] Calculate the average of the quality scores of the set of feature points in the input image and the set of feature points in the knowledge base image among the matched feature points to generate the quality mean index of the image (the average of the quality scores of the retained high-quality feature points). Specifically, extract the quality scores of the feature points in the input image and the quality scores of the corresponding feature points in the knowledge base image among all valid matched feature point pairs, calculate the average quality score of the matched feature points in the input image and the average quality score of the matched feature points in the knowledge base image respectively, and divide the sum of these two averages by 2 to obtain the quality mean index of the image;

[0144] Calculate the ratio of the number of inliers to the total number of feature point pairs in the matching feature points, and generate a geometric consistency index for the image, which specifically includes: obtaining the number of inliers after screening and the total number of matching feature point pairs verified by both geometric consistency and semantic similarity, dividing the number of inliers by the total number of matching feature point pairs to obtain the ratio value, which is the geometric consistency index;

[0145] Calculate the average semantic similarity of the feature point pairs in the matching feature points, and generate a semantic similarity index for the image, which specifically includes: extracting the semantic similarity values corresponding to all valid matching feature point pairs retained after geometric consistency and semantic screening, summing up the semantic similarity values of all valid point pairs, and dividing by the total number of valid point pairs to obtain the average value, which is the semantic similarity index for the image;

[0146] Normalize the image quantity index, quality mean index, geometric consistency index, and semantic similarity index respectively, and then perform weighted summation to obtain the comprehensive similarity score.

[0147] ;

[0148] In the formula, represents the comprehensive similarity score, represents the image quantity index, represents the geometric consistency index, represents the semantic similarity index, represents the number of inliers, represents the quality mean index, , , and are all weights (default ), and the sum of the four weights is equal to 1, represents the maximum value of the image quantity index, represents the maximum value of the quality mean index, represents the maximum value of the geometric consistency index, represents the maximum value of the semantic similarity index.

[0149] Through geometric transformation estimation and inlier screening, the reliability and anti-interference ability of feature matching in agricultural image retrieval are effectively improved. In agricultural scenarios, images often undergo geometric deformations (such as scale and rotation changes caused by leaf pitching and fruit occlusion) due to differences in shooting angles, lighting conditions, and plant growth stages. Traditional matching methods are easily affected by noise and mismatched points. In this application, a random sampling minimum sample set is used to generate an initial affine transformation matrix, and inliers are iteratively screened to dynamically exclude invalid matches caused by perspective changes and noise interference. When retrieving crop leaf images at different growth stages, this mechanism can accurately identify rotation or scaling differences caused by different degrees of leaf unfolding, retain truly corresponding feature point pairs. At the same time, through 100 iterations of optimization using the RANSAC idea, the transformation matrix with the largest number of inliers is selected and combined with the least squares method to refine the parameters, further enhancing the adaptability to complex agricultural image deformations.

[0150] Through multi-dimensional index fusion and weighted evaluation strategies, the comprehensive matching quality of agricultural images is comprehensively quantified, significantly improving the semantic relevance and domain adaptability of retrieval results. Agricultural image retrieval not only requires geometric matching of visual features but also needs to combine semantic information (such as fruit maturity and disease feature regions) and feature point quality (such as the stability of key point detection). Through the constructed image quantity index (proportion of matching point pairs), quality mean index (average score of high-quality feature points), geometric consistency index (proportion of inliers), and semantic similarity index (average semantic score of point pairs), a three-dimensional evaluation is formed from four dimensions of "matching scale - feature quality - geometric accuracy - semantic association". For example, in pest and disease image retrieval, the quality mean index can preferentially retain high-stability feature points describing the edges of disease spots, while the semantic similarity index focuses on the semantic label matching of disease feature regions, avoiding false retrievals caused by background interference. Through normalized weighted summation (default weights take into account the importance of each index), different retrieval requirements can be balanced. For example, variety identification focuses on geometric consistency, and disease diagnosis relies on semantic similarity. The finally generated comprehensive similarity score not only conforms to the complex feature distribution of agricultural images but also can be adapted to different retrieval scenarios through weight adjustment, providing a scientific and flexible quantitative basis for the accurate retrieval of large-scale images in agricultural knowledge bases.

[0151] In one case of this embodiment, the modulus entropy value, the heat map of feature point quality distribution, and the feature point density coefficient are calculated to obtain a similarity threshold. At the same time, images with a similarity higher than the threshold are marked as candidate matches, and a collection of the marked images is established, including:

[0152] The normalized modulus entropy value, the feature point mass distribution heat map and the feature point density coefficient are adjusted to obtain a first adjustment coefficient, a second adjustment coefficient and a third adjustment coefficient respectively, specifically including: for the normalized modulus entropy value, the global minimum and maximum values of the modulus entropy values of all images are counted, and the modulus entropy value of each image is calculated by (current value-global minimum value) / (global maximum value-global minimum value), and the modulus entropy value is mapped to the interval [0, 1] to obtain the first adjustment coefficient;

[0153] For the feature point quality distribution heat map, first calculate the average quality score of each block (that is, the average quality score of the feature points retained in each fixed-size rectangular block), then perform Gaussian smoothing on the heat map, and map the pixel values to [0, 1] through global normalization to obtain the second adjustment coefficient;

[0154] For the density coefficient of the feature points, the average value of all the image density coefficients is calculated as the reference value. For the density coefficient higher than the reference value, it is enhanced according to (current value - reference value) / (1 - reference value). For the density coefficient lower than the reference value, it is attenuated according to (current value / reference value) to obtain the third adjustment coefficient.

[0155] Set the initial similarity threshold;

[0156] The three adjustment coefficients are calculated with the initial similarity threshold to obtain the similarity threshold, specifically including: assigning weights (0.4, 0.3 and 0.3) to the first adjustment coefficient, the second adjustment coefficient and the third adjustment coefficient respectively, calculating the comprehensive adjustment factor by linear combination (i.e., summing up the coefficients after multiplying the corresponding weights), and then multiplying the initial threshold with the comprehensive adjustment factor to obtain the dynamically corrected similarity threshold;

[0157] Compare the composite similarity score of the input image with the similarity threshold:

[0158] If the comprehensive similarity score of the knowledge base image is greater than or equal to the similarity threshold, it is marked as a candidate match;

[0159] If the comprehensive similarity score of the knowledge base image is less than the similarity threshold, it will not be marked;

[0160] All candidate matching images are aggregated into a collection.

[0161] The similarity threshold adaptive technology of dynamic fusion of multi-dimensional features effectively solves the problem of threshold fixation caused by scene complexity and feature distribution differences in agricultural image retrieval, and significantly improves the accuracy and robustness of candidate matching image screening. In the agricultural field, the feature point distribution of different types of images (such as fruit phenotypes, pest and disease leaves, and field weeds) is significantly different. The image may cause local feature point quality fluctuations due to light reflection. Pest and disease images are often accompanied by abnormally high density of feature points in the diseased area, and field panoramic images may have significant changes in feature point modulus entropy values due to background clutter. In this application, the difference in image structure complexity is quantified by the modulus entropy adjustment coefficient, and healthy leaves with uniform texture and diseased leaves with mottled spots are distinguished at the structural level. Gaussian smoothing and normalization of feature point quality heat maps are performed. The method can suppress noise interference and retain the quality distribution trend of reliable feature points (such as giving priority to high-stability feature areas such as the edge of lesions). The nonlinear enhancement and attenuation strategy of the density coefficient effectively balances the retrieval requirements of sparse features (such as single seedlings) and dense features (such as weeds). Through linear combination through weight allocation, the dynamically corrected similarity threshold can adapt to the characteristic characteristics of different images. For example, it automatically increases the threshold strictness for pest and disease images with high feature point density and uneven quality distribution, and appropriately reduces the threshold looseness for crop root images with sparse features but stable structure, avoiding the problem of missed detection or false detection caused by fixed thresholds. The final collection of candidate matching images not only filters out invalid matches caused by abnormal feature fluctuations, but also retains effective retrieval results that conform to the semantics of the agricultural field.

[0162] In one case of this embodiment, the dynamic weight of the knowledge base image is calculated based on the comprehensive similarity score, including:

[0163] Calculating the comprehensive similarity scores of all candidate images in the collection to generate an index result, specifically including: performing an exponential operation on the comprehensive similarity score of each candidate image in the collection, then summing the exponential operation results of all candidate images, and finally dividing the exponential value of each candidate image by the sum to obtain a normalized index result;

[0164] Calculate the sum of all index results as the denominator;

[0165] When the denominator is less than or equal to 0, all candidate images are assigned equal weights, which is the dynamic weight, and the weight value is the inverse of the number of candidate images;

[0166] When the denominator is greater than 0, divide the exponential result of each candidate image by the denominator to generate normalized weights that satisfy the sum of all weights being 1, which are the dynamic weights. Specifically, for example, if there are 3 candidate images now, with comprehensive similarity scores of 0.6, 0.8, and 0.4 respectively, the exponential result value of image 1 is 1.822, the exponential result value of image 2 is 2.225, and the exponential result value of image 3 is 1.492. 1.822 + 2.225 + 1.492 = 5.539. At this time, 5.539 is greater than 0. The normalized weight of image 1 is 1.822 / 5.539 ≈ 0.329, the normalized weight of image 2 is 2.225 / 5.539 ≈ 0.402, and the normalized weight of image 3 is 1.492 / 5.539 ≈ 0.269;

[0167] Calculate the dynamic weights of all images in the collection to obtain the standard deviation;

[0168] Perform the natural logarithm operation on the number of candidate images to obtain the empirical coefficient;

[0169] Multiply the standard deviation by the empirical coefficient to obtain the tolerance threshold.

[0170] Through the dynamic weight allocation and tolerance threshold calculation mechanism, the problems of unreasonable weight allocation and uneven similarity distribution of candidate matching images in agricultural image retrieval are effectively solved, and the accuracy of retrieval results is significantly improved. In the agricultural knowledge base, candidate images often show significant differentiation in comprehensive similarity scores due to differences in shooting angles, feature integrity, etc. For example, typical symptom images of the same disease may have extremely high similarity scores, while non-target images with approximate symptoms have lower scores. The dynamic weight calculation amplifies the score differences through exponential operations, making high-similarity images dominant in weight allocation, ensuring that the retrieval results focus on the most relevant core images, avoiding interference from low-quality matches to decision-making, and effectively coping with the weight failure problem in extreme cases (such as extremely low similarity of all candidate images) through the robust design of denominator discrimination, thus ensuring the stability of retrieval.

[0171] In one case of this embodiment, the collection is sorted in three levels according to the dynamic weights, including:

[0172] The three-level sorting is specifically as follows:

[0173] The first-level sorting is: sort the candidate images in the collection in descending order according to the dynamic weights from high to low;

[0174] The second-level sorting is: when there are groups of images with the same dynamic weights in the first-level sorting, perform a secondary sorting within the group in descending order according to the comprehensive similarity scores.

[0175] The third - level sorting is as follows: when there are still images with the same comprehensive similarity score in the second - level sorting, sort them in ascending order of the image storage timestamp, and select the image with the earliest timestamp as the output image. If the storage timestamps are the same, sort them in ascending order of the image unique ID and select the image with the smallest ID as the output image;

[0176] Among the candidate images after three - level sorting, obtain the dynamic weights of the first - ranked and second - ranked images, calculate and get the difference;

[0177] If the difference after three - level sorting is greater than the tolerance threshold, directly output the first - ranked image after three - level sorting;

[0178] If the difference after three - level sorting is less than or equal to the tolerance threshold, trigger the local feature comparison mechanism. Specifically:

[0179] Respectively retrieve the binarized mask matrices of the first two images, compare pixel by pixel and count the number of different pixels, and select the image with fewer different pixels as the final output. If the number of different pixels is the same, directly output the first - ranked image after three - level sorting.

[0180] Through the three - level hierarchical sorting and dynamic tolerance decision - making mechanism, the problems of fuzzy priority of candidate matching images and difficult distinction of similar features in agricultural image retrieval are effectively solved, and the accuracy and practicality of the retrieval results are significantly improved. In the agricultural knowledge base, candidate images often show similar dynamic weights due to close varieties, similar growth stages, or slight differences in disease symptoms (such as different varieties with similar fruit colors, different disease levels with similar lesion areas). Traditional single - sorting rules are difficult to accurately distinguish the core matching images. The three - level sorting strategy, through different hierarchical logics, not only ensures the priority output of highly relevant images (such as typical disease images with high dynamic weights are presented first), but also ensures the accuracy of image retrieval through the ordered processing of timestamps and unique IDs.

[0181] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent knowledge base retrieval system, characterized in that, Including: An image preprocessing unit, which is used to obtain an input image taken by a user in a farmland, unify the input image and the knowledge base image to a standard size, separate the standard size image, generate a binary mask matrix and a corresponding RGB feature map, and at the same time calculate the binary mask matrix to generate a modulus entropy value; An image adjustment unit, which is used to extract the RGB feature map to obtain a multi-scale set of feature points; Based on the binary mask matrix, generate a feature point distribution density map, dynamically adjust the feature point distribution density map to obtain a feature point density coefficient, calculate the dual index of the RGB feature map, and generate a feature point quality distribution heat map based on the dual index; An image matching unit, which is used to bidirectionally match the input image feature point set and the knowledge base image feature point set in the set of feature points to obtain matching feature points, calculate the matching feature points to obtain the comprehensive similarity score of the input image feature point set and the knowledge base image feature point set, calculate the modulus entropy value, the feature point quality distribution heat map and the feature point density coefficient to obtain a similarity threshold, and at the same time mark the images higher than the similarity threshold as candidate matches and establish a collection of the marked images; A retrieval judgment unit, which is used to calculate the dynamic weight of the knowledge base image according to the comprehensive similarity score; sort the collection in three levels according to the dynamic weight.

2. An intelligent retrieval system for a knowledge base according to claim 1, characterized in that, Separating the standard size image to generate a binary mask matrix and a corresponding RGB feature map, including: Enhancing the standard size image to obtain an enhanced image; According to the enhanced image, clustering the image pixels, dividing the image into two parts: foreground and background, and generating an initial binary mask matrix; Optimizing the initial binary mask matrix to obtain a final binary mask matrix; Performing pixel-by-pixel RGB feature extraction on the enhanced image according to the final binary mask matrix to form an RGB feature map; Statistical analysis of the ratio of the number of pixels in the target area to the total number of pixels in the binary mask matrix to obtain a pixel occupancy parameter; Calculating the gray level change frequency between adjacent pixels in the binary mask matrix to obtain a neighborhood change rate parameter; Fusing the pixel occupancy parameter and the neighborhood change rate parameter to obtain a modulus entropy value.

3. An intelligent retrieval system for a knowledge base according to claim 2, characterized in that, Extracting the RGB feature map to obtain a multi-scale set of feature points, including: Calculating the texture complexity distribution of different regions in the RGB feature map to generate a dynamic scale adjustment parameter; Based on the dynamic scale adjustment parameter, performing hierarchical processing on the RGB feature map, and generating a set of local feature vectors with scale labels in each layer, specifically including: performing three-level division according to the range of the dynamic scale adjustment parameter value: Dividing the area where the normalized parameter value > 0.7 into a high texture complexity area and matching a 5×5 small scale window, dividing the interval where 0.3 ≤ normalized parameter value < 0.7 into a medium complexity area and matching a 7×7 medium scale window, and dividing the interval where the normalized parameter value < 0.3 into a low complexity area and matching a 9×9 large scale window; Calculating the weight coefficients of the feature vectors in each layer in the set of local feature vectors, and generating a set of feature points containing multi-scale spatial associations through weighted fusion.

4. An intelligent retrieval system for a knowledge base according to claim 2, characterized in that, Based on the binary mask matrix, a feature point distribution density map is generated, and the feature point distribution density map is dynamically adjusted to obtain the feature point density coefficient, including: The image in the binary mask matrix is evenly divided into several small grids, the number of feature points in each grid is calculated, and a two-dimensional density distribution map is generated; Calculate the average value of the distribution density in the two-dimensional density distribution map; In the grids with density higher than the average value, some feature points are removed at a ratio of 50%; In grids with density lower than the average, new feature points are generated by copying neighboring points to fill in sparse areas; Calculate the density variance of the feature points in each grid after removal and filling operations, normalize it and map it to the range of 0-1 to obtain the feature point density coefficient, as follows: calculate the ratio of the mean and standard deviation of the grid density, square the ratio to obtain a measure of the density dispersion, use it as the numerator and the denominator as 1 plus the measure of the density dispersion, divide the numerator by the denominator to obtain the dispersion influence coefficient, and calculate the ratio of the adjusted total number of feature points to the original total number of feature points, then substitute it into the hyperbolic tangent function to obtain the feature point quantity adjustment coefficient, multiply the dispersion influence coefficient by the feature point quantity adjustment coefficient to obtain the feature point density coefficient.

5. An intelligent retrieval system for a knowledge base according to claim 2, characterized in that, The dual indicators of the RGB feature map include the local contrast value and the clustering value. The dual indicators of the RGB feature map are calculated, and the feature point quality distribution heat map is generated based on the dual indicators, including: In the RGB feature map, a circular area with a fixed radius is taken with each feature point as the center, and the pixel gradient amplitude in the circular area is calculated to obtain the local contrast value; Group and count the feature points to get the clustering value; The local contrast value and the clustering value are linearly combined to obtain the quality score of each feature point; The RGB feature map is divided into several blocks of fixed size, and the average quality score of the retained feature points in each block is calculated. It is converted into a two-dimensional thermal distribution map through continuous color mapping, and finally a feature point quality distribution thermal map is generated.

6. An intelligent retrieval system for a knowledge base according to claim 3, wherein Perform bidirectional matching on the input image feature point set and the knowledge base image feature point set in the feature point set to obtain matching feature points, including: The feature point set includes an input image feature point set and a knowledge base image feature point set; The cosine similarity of the feature point descriptors between the input image feature point set and the knowledge base image feature point set is calculated to obtain a rough matching result, and the feature point pairs that do not meet the geometric consistency in the rough matching result are eliminated; Extract the deep features of the local area around the feature point pair, calculate the semantic similarity of the feature point pair, and remove it if the semantic similarity of the feature point pair is less than or equal to 30%; Double screening is performed based on the rough matching results and semantic similarity to obtain matching feature points that are consistent in both geometry and semantics.

7. An intelligent retrieval system for a knowledge base according to claim 6, wherein, Calculate the matching feature points to obtain the comprehensive similarity score of the input image feature point set and the knowledge base image feature point set, including: Select and calculate the minimum sample set of feature point pairs in the matching feature points to generate an initial affine transformation matrix; Use the initial affine transformation matrix to transform the coordinates of the feature points of the knowledge base image to the input image coordinate system to obtain the predicted coordinates; Calculate the distance residual between the actual coordinates of the input image feature points and the predicted coordinates to obtain the corresponding inliers; Fit all the internal points of the initial affine transformation matrix to generate an affine transformation matrix; Count the total number of feature point pairs in the matching feature points to generate image quantity indicators; Calculate the input image feature point set and the knowledge base image feature point set in the matching feature points to generate the image quality mean index; Calculate the internal points and matching feature points of the affine transformation matrix to generate the geometric consistency index of the image; Calculate the average semantic similarity of feature point pairs to generate the semantic similarity index of the image; The image quantity index, quality mean index, geometric consistency index and semantic similarity index are normalized and weighted summed to obtain a comprehensive similarity score.

8. An intelligent retrieval system for a knowledge base according to claim 5, characterized in that, The model entropy value, feature point mass distribution heat map and feature point density coefficient are calculated to obtain the similarity threshold. At the same time, images above the similarity threshold are marked as candidate matches, and the marked images are collected, including: The normalized modulus entropy value, the characteristic point mass distribution heat map, and the characteristic point density coefficient are adjusted to obtain a first adjustment coefficient, a second adjustment coefficient, and a third adjustment coefficient, respectively; Set the initial similarity threshold; The three adjustment coefficients are calculated with the initial similarity threshold to obtain the similarity threshold; Compare the composite similarity score of the input image with the similarity threshold: If the comprehensive similarity score of the knowledge base image is greater than or equal to the similarity threshold, it is marked as a candidate match; If the comprehensive similarity score of the knowledge base image is less than the similarity threshold, it will not be marked; All candidate matching images are aggregated into a collection.

9. An intelligent retrieval system for a knowledge base according to claim 8, characterized in that, The dynamic weight of the knowledge base image is calculated based on the comprehensive similarity score, including: Calculate the comprehensive similarity scores of all candidate images in the collection and generate an index result; Calculate the sum of all index results as the denominator; When the denominator is less than or equal to 0, all candidate images are assigned equal weights, which is the dynamic weight, and the weight value is the inverse of the number of candidate images; When the denominator is greater than 0, the exponential result of each candidate image is divided by the denominator to generate a normalized weight that satisfies the sum of all weights to be 1, which is the dynamic weight; Calculate the dynamic weights of all images in the collection and get the standard deviation; Perform natural logarithm operation on the number of candidate images to obtain empirical coefficients; Multiply the standard deviation by the empirical coefficient to obtain the tolerance threshold.

10. A knowledge base intelligent retrieval system according to claim 9, characterized in that, The collection is sorted in three levels according to dynamic weights, including: The three-level sorting is as follows: The first level sorting is: sorting the candidate images in the collection in descending order from high to low according to the dynamic weight; The second level sorting is: when there are image groups with the same dynamic weight in the first level sorting, the image groups are sorted again from high to low according to the comprehensive similarity scores; The third-level sorting is as follows: when there are still images with the same comprehensive similarity score in the second-level sorting, sort them in ascending order of the image storage timestamp, and select the image with the earliest timestamp as the output image; if the storage timestamps are the same, sort them in ascending order of the image unique ID and select the image with the smallest ID as the output image; Among the candidate images after the three-level sorting, obtain the dynamic weights of the first and second ranked images, calculate and obtain the difference; If the difference after the three-level sorting is greater than the tolerance threshold, directly output the first image after the three-level sorting; If the difference after the three-level sorting is less than or equal to the tolerance threshold, trigger the local feature comparison mechanism.

Citation Information

Patent Citations

  • Multilayer image feature extraction and model training method based on matching task

    CN117809051A

  • Feature extraction with keypoint resampling and fusion (KRF)

    US20210241022A1