A fuzzy hybrid instance search method combining image and instance features

By combining the blurred hybrid instance search method of image and instance features, the instance search problem with motion blur images is solved, and the target tracking and object recognition performance of autonomous driving and intelligent robots in complex environments is improved.

CN115757845BActive Publication Date: 2025-07-04EAST CHINA NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210839198.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-18
Publication Date
2025-07-04
Estimated Expiration
2042-07-18

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently search for instances from images with motion blur, especially in autonomous driving and intelligent robots, which affects the accuracy of environmental recognition and navigation.

Method used

The fuzzy mixed instance search method combining image features and instance features improves the accuracy of search through fuzzy image enhancement, mixing processing, feature extraction based on deep convolutional neural networks and reordering of regional suggestions networks.

Benefits of technology

More efficient instance search is achieved in images with motion blur, improving the target tracking and object recognition capabilities of autonomous driving and intelligent robots in complex contexts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115757845B_ABST
    Figure CN115757845B_ABST
Patent Text Reader

Abstract

The present invention discloses a fuzzy hybrid instance search method combining image and instance features. This method is used for instance search in images / videos with motion blur, and specifically includes: query fuzzy hybridization, sorting based on image features, and re - sorting based on object features. The present invention will enhance the search effect of instance search for search objects of motion - blurred picture data, and can obtain better search effects in image searches with similar backgrounds and images containing the same entity. It provides better retrieval support for autonomous driving and mobile robots to process vision - related retrieval tasks, such as object tracking, object re - identification, etc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision object recognition and fuzzy information retrieval, and in particular to a search method based on a convolutional neural network that combines complete image feature information and recognition target object feature information. Background Art

[0002] Autonomous driving and intelligent robots usually use cameras to collect videos and continuous images to enhance tasks such as environmental recognition and navigation. When moving vehicles and robots collect images / videos, images / videos with motion blur frequently appear. These blurred images / videos may become the targets of information retrieval. Instance Search has been studied in the field of computer vision for many years. As a basic task, instance search directly retrieves a query object from the entire scene image, and is applicable to multimedia and video image semantic analysis in environmental recognition and navigation scenarios even if the instance only occupies a small part of the image. The goal of the instance search task is to retrieve those images from the database that contain the query instance (for example, provided as a bounding box in the query image). Therefore, compared with Content Based Information Retrieval (CBIR), instance search is a more advanced and challenging task due to the diversity of queries and the timeliness of responses. In the current research field, there is a lack of instance search for the above-mentioned blurred images.

[0003] In real scenarios, due to the calibration of cameras mounted on mobile intelligent vehicles and robots, a large number of blurred frames will appear in the collected videos and images, which will affect the performance of instance search. Currently, many algorithms have been proposed to solve the problem of the decline in object detection accuracy due to blur in images. Generally, these algorithms use some image argumentation methods to generate a dataset containing blurred images and train a Convolutional Neural Network (CNN) for object detection. However, there is currently no literature addressing the problem of directly searching for instances from blurred images. Summary of the Invention

[0004] The object of the present invention is to provide a fuzzy hybrid query method that combines image features and instance features to perform instance search on blurred images. And a dataset similar to TRECvid INS was also designed and collected for performance experiments.

[0005] The specific technical solution for achieving the object of the present invention is as follows:

[0006] A fuzzy hybrid instance search method combining image and instance features, the method comprising the following specific steps:

[0007] Step 1: Query fuzzy hybrid

[0008] Perform an instance search query in the image library to obtain a query image I. Perform fuzzy enhancement on the query image. Use simulated fuzzy kernels with different degrees of blurriness and adopt a fuzzification algorithm to convert the query image into a fuzzy image, obtaining n fuzzy images I with different degrees of blurriness b1 ,I b2 ,...I bn ; Then, mix the fuzzy image with the original image. The mixing specifically includes:

[0009] 1) Pixel-level mixing

[0010] Directly use the RGB values of the image to add the query image and its fuzzy image at each pixel, then take the average to obtain a mixed image; input the mixed image as a new query image into the deep convolutional neural network CNN structure for feature extraction; the RGB value of the query image I at pixel (x,y) is V I x,y , then the RGB value of the mixed image I M is calculated by the following formula:

[0011]

[0012] 2) Feature-level mixing

[0013] First, extract the CNN features of the query image and its fuzzy images I b1 ,I b2 ,...I bn , and then perform mixing; and use the new mixed features as the actual query features for subsequent queries; represent the features extracted from the query image I through CNN with the vector F I , where each dimension z is f z I ; The mixing of fuzzy features is divided into three: addition average and maximum The calculation formulas are as follows respectively:

[0014]

[0015]

[0016]

[0017] Step 2: Sorting based on image features

[0018] The result of mixing will enter the sorting as image features, and the entire image library will be filtered to obtain a rough sorting; the following similarity function is used for sorting, where F a is the feature vector of image a, |·| is for calculating the inner product, and · is vector multiplication;

[0019]

[0020] For each picture F in the image library c and the query picture I after mixing M vectors perform similarity calculation, and sort according to the similarity to obtain the sorting result;

[0021] Step 3: Re-ranking based on instance features

[0022] First, perform query expansion QE; QE will select the top K results in the sorting result based on image features as the expanded object images; the global features of the query image after fuzzy mixing in Step 1 and the region proposal network RPN features of the query in the CNN are used as query features to perform a new round of search; the features provided by RPN include the instance features and ROI region features in the image, denoted as R1, R2,..., R m , and the final feature with the ROI region features added is ε; apply the operator to calculate ε, normalized maximum value, sum or average can be used, and the calculation formula is:

[0023]

[0024] Then calculate the re-ranking result based on the final feature ε, and use the top results as the query results.

[0025] The present invention uses image features to complete the first sorting, and then uses the re-sorting of instance features to enhance the query.

[0026] In the present invention, the RCNN network with faster calculation speed is used to construct the CNN network. Both the image features and the instance features use the ResNet backbone network in the CNN network. Therefore, the calculation of ε is obtained and calculated in the same network; in particular, during the process of obtaining the instance features of the region proposal network RPN, the Conv5_3 version of RCNN is used, and its ROI pooling layer is separated for feature extraction.

[0027] The example search with blurred images in the present invention is implemented as follows: It consists of three parts: query processing for blurring, CNN-based image features for the first sorting, and instance features based on the Region Proposal Network (RPN) for re-sorting. As mentioned before, there is still no dataset suitable for physical and real scenarios in instance search. Existing instance search datasets focus on clear images and rarely contain blurred images. Based on COCO-C [9] and Blurred Video Tracking (abbreviated as BVT), an instance search benchmark is proposed in the present invention, which is specifically used to access the task robustness of images with motion blur. It is named the Blur Instance Search (BlurINS) benchmark for the IS task. Blur instance search satisfies both the above BlurINS image and video scenarios.

[0028] For instance search, the present invention relies on the task of re-identifying objects from videos or images with blur. First, a target tracking algorithm is used to identify specific types of targets, and the targets in these images are marked with blur in the real dataset used as a benchmark. Second, a pre-trained convolutional network is used for instance search to obtain frames that may contain the same instance in the query. To verify the effectiveness of the solution, the degree of video / image blur is defined. Therefore, experiments can be conducted on image queries with different blur degrees. The mean average precision is also used as a test metric, which is defined by the position of the correctly retrieved items in the ranked retrieval results. The key GOPRO data is used as a blurred video or image library for instance search.

[0029] The present invention will enhance the search effect of instance search for motion-blurred picture data and can obtain better search effects in image searches with similar backgrounds and images containing the same entity. It provides better retrieval support for vision-related retrieval tasks such as target tracking and object re-identification in autonomous driving and mobile robots. Description of the Drawings

[0030] Figure 1 It is the flowchart of the present invention;

[0031] Figure 2 It is a synthetic motion blur map with different blur kernels and random blur angles;

[0032] Figure 3 It is a schematic diagram of the distribution of the number of positive example samples and test query retrievals in the Blur-Ins dataset;

[0033] Figure 4 It is a schematic diagram of sampling the confusion legend in the dataset. Detailed Implementation Manner

[0034] The present invention will be further described in detail below with reference to the accompanying drawings.

[0035] Referring to Figure 1 , the present invention includes the following specific steps:

[0036] Step 1: Query Blurring and Mixing

[0037] Submit an instance search query to the search system. It will be used for blurring enhancement, that is, using simulated blurring kernels with different blurring degrees (shown in the dashed box on the far left) to blur the query image. Then, there are two branches in the method of the present invention. The query picture I contains n blurred images I with different blurring degrees b1 , I b2 ,... I bn . Different mixing strategies are used in the mixing stage:

[0038] 1) Pixel-level mixing

[0039] Directly add the query image and its per-pixel blurred image in terms of RGB values, and then take the average. The mixed image is input as a new query image into a typical deep convolutional neural network structure, such as ResNet50, VGG16, etc. These features can express global information by activating these bottleneck network blocks. If the RGB value of the image I at the pixel (x, y) is V I x,y , then the RGB value of the mixed image I M is calculated by the following formula:

[0040]

[0041] 2) Feature-level mixing

[0042] First, extract the CNN features of the query image and its blurred image. Feature-level blurring and mixing are divided into three steps: addition, averaging, and merging. And the new mixed features are processed as the actual query features for subsequent use. The features extracted from the image I by CNN are represented by the vector F I , where each dimension z is f z I . The mixing of the blurred features and the blurred features is divided into three steps: addition, averaging, and merging, and the calculation formulas are as follows respectively:

[0043]

[0044]

[0045]

[0046] Step 2: Ranking Based on Image Features

[0047] Then, the result of the mixing stage will be passed as image features to the ranking stage, which is beneficial for the first filtering of the entire image library (see Figure 1 ) to find a rough ranking, thus providing the most relevant query results. Only the top-ranked images will be passed to the next re-ranking process. In the actual process, some conventional retrieval techniques will be applied to enhance performance, such as Query Expansion (QE) before the re-ranking stage. QE means that the image features of the top results in the ranking stage are averaged with the query image features to perform a new search. The ranking uses the following similarity function:

[0048]

[0049] Step 3: Re-ranking based on instance features

[0050] The RPN features are used to describe the instance information. RPN has received extensive attention in object detection and other computer vision tasks. It provides local and subtle clues indicating the instance features and locations in the image. With the help of the previous process, only a few images are re-ranked, so the computational amount is greatly reduced. In addition, RPN can also predict the position of the query instance in the retrieved images. Among these three steps, Step 1 can be achieved by balancing the weights of each ambiguous query instance. The other two steps are introduced by constructing a network such as faster RCNN, and the image features and instance features can be easily separated from the ResNet backbone network and the ROI pooling layer, and the latter is applied by RPN to the proposed instance boxes. The features of the ROI regions on each image are R1, R2,..., R m , note that the number of ROIs m for each image is different, and the global image feature is The final feature ε with the candidate region features added, and different operators can be taking the maximum, taking the sum, or taking the average for normalization, and the calculation formula is:

[0051]

[0052] Then, calculate the re-ranking result with the final feature ε, and use the top results as the query results.

[0053] Dataset construction

[0054] Among all the existing datasets for instance search, TRECvid INS is a subset of the TREC Video Retrieval Evaluation. It contains more than 20,000 key frames from TRECvid. Some other classic image retrieval datasets, such as Oxford5k and Paris6k, usually define the queries first and then collect the image galleries. It makes the target object appear clearly and prominently in the images. While TRECvid INS first determines the data and then selects specific objects or people, which may not be the main part of the images. Therefore, it is not only closer to the actual application but also more challenging.

[0055] However, TRECvid is only open to participants and there are not enough specific images to evaluate the motion blur robustness of instance search in other datasets. For these reasons and inspired by TRECvid, a new dataset named Blur-INS was constructed in the same mode as TRECvid-INS. 33 sets of sequential images and their simulated motion blur were borrowed from the deblurring dataset GOPRO. Image enhancement techniques were also applied to generate synthetic motion blur with different blur kernels and random blur angles, see Figure 2 Example. The ground truth bounding boxes were generated on the clear images using the SOT algorithm, and then these annotations were directly used as the bounding boxes for the corresponding blurred images. For each image sequence, an instance in the first frame was used as the query and the other frames were used to create the image gallery. However, only the frames where the instance is shown as a true positive for the query were considered. Figure 3 Briefly shows the distribution of true positives in Blur-INS with blur.

[0056] To verify the solution, ablation experiments were conducted to ensure that the motion blur does affect the results of instance search and to verify the technical correctness and feasibility of the method of the present invention. The data was divided into two categories, one is the images with mild blur and the other is the images with severe blur. Here, a blur degree judgment algorithm was used, which utilizes some blur kernels to adjust or measure the blur degree of a specific image. It uses a parameter to set the blur degree, and this parameter is larger when the image is more blurred. The images were set to be mildly blurred with K = 10, moderately blurred with K = 30, and severely blurred with K = 60 (see Figure 2 ).

[0057] A large number of imperceptible negative examples and interfering images were also added to the dataset to better evaluate the robustness. They have some functions similar to the queries, such as having the same background or scene. Figure 4 There is an example of a cluttered image in

[0058] The mean average precision (mAP) metric is used to evaluate the performance of the retrieval task. The results are summarized in Table 1. The image gallery is divided into two categories: the concise set (without confusing images) and the confusing set. For each type of set, they are further divided into different blurriness parts. For authenticity and fairness, the method used to generate the blurred database images is different from the method used in the mixing stage. For each partition, different mixing or QE strategies are tried, and the average mAP values in Table 1 are calculated. It can be seen that the effects in the concise set and the lighter blurriness parts are better than other parts. In addition, this shows that both mixing strategies are beneficial to the search results. The best map values are shown in bold. The research shows that the higher the difficulty, the weaker the QE effect, or even ineffective, such as in the heavy concise and medium-confusing parts. When the difficulty is high, the pixel-level mixing strategy is far inferior to the feature-level, and its performance in the medium-confusing case is even worse than the sharp query without any processing. The situation where the pixel-level mixed fuzzy transform seems to have noise is analyzed, otherwise it will have a negative impact on the results. When the difficulty is reduced, the effects of the two mixing strategies are almost the same.

[0059] Table 1 Effects of example search of the present invention under different difficulties and blurriness

[0060]

[0061] The above is only a further description of the present invention and is not intended to limit the present invention. All equivalent implementations of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A fuzzy hybrid instance search method combining image and instance features, characterized in that, The method includes the following specific steps: Step 1: Query fuzzy mixing Perform an instance search query in the image library to obtain a query image I. Blur and enhance the query image using simulated blur kernels with different blur degrees. Apply a blurring algorithm to convert the query image into a blurred image, obtaining n blurred images I b1 , I b2 ,... I bn ; Then, blend the blurred images with the original images. The blending specifically includes: 1) Pixel-level mixing Directly add the query image and its blurred image based on the RGB values of the images. Add the blurred image for each pixel, then take the average to obtain the mixed image. Input the mixed image as a new query image into the deep convolutional neural network (CNN) structure for feature extraction. The RGB value of the query image I at pixel (x, y) is V I x,y , then the RGB value of the mixed image I M is calculated by the following formula: 2) Feature-level mixing First, extract the CNN features of the query image and its blurred images I b1 , I b2 ,... I bn , and then perform mixing; and use the new mixed features as the actual query features for subsequent queries; represent the features extracted from the query image I by the CNN with the vector F I , where each dimension z is f z I ; The mixing of the blurred features is divided into three: addition average and maximum The calculation formulas are as follows respectively: Step 2: Sorting based on image features The result of the mixture will enter the sorting as an image feature, and the entire image library will be filtered to obtain a rough sorting; the following similarity function is used for sorting, where F a is the feature vector of image a, |·| is for taking the inner product, and · is vector multiplication; According to each picture F in the image library c and the mixed query picture I M vectors calculate the similarity, and sort according to the similarity to obtain the sorting result; Step 3: Re-sorting based on instance features First, perform query expansion QE; QE will select the top K results in the ranking results based on image features as the expanded object images; the global features of the query image after the blurring and mixing in step 1 and the region proposal network RPN features of the query in the CNN are used as query features to perform a new round of search; the features provided by the RPN include the instance features in the image and the features of the ROI regions, denoted as R1, R2,..., R m , and the final feature with the ROI region features added is ε; apply the operator to calculate ε, the normalization of taking the maximum value, taking the sum or taking the average can be adopted, and the calculation formula is: Then, calculate the re-sorting result with the final feature ε, and take the top results as the query results.

2. The fuzzy hybrid instance search method combining image and instance features according to claim 1, characterized in that, Use image features to complete the first sorting, and then use the re-sorting of instance features to enhance the query improvement.

3. A fuzzy hybrid instance search method combining image and instance features according to claim 1, characterized in that, The CNN network is constructed using the faster-computing RCNN network. Among them, both image features and instance features use the ResNet backbone network in the CNN network. Therefore, the calculation of ε is obtained and calculated in the same network. In particular, during the process of obtaining the instance features of the Region Proposal Network (RPN), the Conv5_3 version of RCNN is adopted, and its ROI pooling layer is separated for feature extraction.

Citation Information

Patent Citations

  • Searchable encrypted image retrieval method based on deep convolutional network features

    CN110659379A

  • Image retrieval method based on multi-feature fusion

    CN111125416A