A method for multi-object detection in small images based on super-resolution

By generating high-resolution feature maps and performing multi-target detection, the problem of blurry small targets in low-resolution images is solved, and accurate identification and stable detection of multi-target obstacles in guide devices for the visually impaired are achieved.

CN121392797BActive Publication Date: 2026-04-03QINGDAO ZENDINGXUAN INNOVATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-19
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Camera-based visual assistance systems struggle to accurately identify small and distant targets in low-resolution images, resulting in low obstacle recognition accuracy and a high false alarm rate. This impacts the real-time performance and safety of guide devices, especially in dynamic environments where detection performance is insufficient.

Method used

By acquiring low-resolution images, restoring edge texture information to generate high-resolution feature maps, extracting target response regions at different scales, constructing obstacle feature sets, and performing multi-target detection and boundary deduplication, the results are finally back-projected onto the original low-resolution image coordinate system to output multi-target obstacle detection results for use in guide devices.

Benefits of technology

It significantly enhances the identifiability of small targets, achieves accurate identification of multiple scales and multiple targets, reduces duplicate detection and false identification, ensures the reliability and accuracy of obstacle detection results, and improves the obstacle recognition capability and detection stability of guide devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121392797B_ABST
    Figure CN121392797B_ABST
Patent Text Reader

Abstract

This invention relates to the field of multi-object detection technology, and more particularly to a method for multi-object detection of small images based on super-resolution. The method includes the following steps: acquiring a low-resolution image generated by a guide camera unit, and obtaining small image data from the low-resolution image; recovering edge texture information from the small image data, generating a high-resolution feature map, extracting target response regions of different scales from the high-resolution feature map, and constructing a feature set of obstacles; performing multi-object detection operations on the feature set to form obstacle detection results, and performing boundary deduplication on the obstacle detection results to obtain small image multi-object detection results; back-projecting the small image multi-object detection results to the original low-resolution image coordinate system, and outputting the multi-object obstacle detection results used by the guide device; this invention improves the long-distance obstacle recognition capability and detection reliability of guide devices in complex environments through small image multi-object detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of multi-target detection technology, and in particular to a method for multi-target detection of small images based on super-resolution. Background Technology

[0002] Camera-based visual assistance systems typically rely on low-resolution images for obstacle detection. Due to limited image detail, small and distant targets appear blurry or almost indistinguishable, leading to low obstacle recognition accuracy and high false alarm rates, impacting the real-time performance and safety of guide devices. Existing methods often employ single-scale features or traditional image enhancement techniques, but their detection performance remains insufficient for practical applications in scenarios with multiple targets and mixed obstacles at varying distances. The significant loss of target edge and texture information in low-resolution images hinders high-precision fusion of multi-target segmentation, classification, and localization. This is especially problematic in dynamic environments where rapidly moving targets or complex backgrounds can cause detection delays and misidentifications. Summary of the Invention

[0003] Therefore, it is necessary to provide a multi-target detection method for small images based on super-resolution to solve at least one of the above-mentioned technical problems.

[0004] To achieve the above objectives, a multi-object detection method for small images based on super-resolution includes the following steps:

[0005] Step S1: Acquire low-resolution images generated by the guide camera unit and obtain small image data from the low-resolution images;

[0006] Step S2: Recover edge texture information from small image data, generate high-resolution feature maps, extract target response regions at different scales from the high-resolution feature maps, and construct a feature set of obstacles;

[0007] Step S3: Perform multi-target detection on the feature set to generate obstacle detection results, and perform boundary deduplication on the obstacle detection results to obtain multi-target detection results for small images;

[0008] Step S4: Backproject the multi-target detection results of the small image to the coordinate system of the original low-resolution image, and output the multi-target obstacle detection results used by the guide device.

[0009] The beneficial effects of this invention are as follows: By extracting small image data and restoring edge texture information from low-resolution images acquired by the guide camera unit, a high-resolution feature map is generated, significantly enhancing the identifiability of small targets in the image. Utilizing the high-resolution feature map to extract target response regions at different scales and constructing an obstacle feature set, this method can effectively cover obstacles at both long and short distances, achieving accurate multi-scale and multi-target identification. During multi-target detection, by matching the feature set one by one with a preset obstacle category benchmark, and combining the center position and confidence level of the response region for filtering and boundary deduplication, duplicate detection and false identification can be effectively reduced, ensuring the reliability and accuracy of obstacle detection results. Simultaneously, by back-projecting the multi-target detection results of the small image to the coordinate system of the original low-resolution image, the guide device can obtain the precise location and category information of each target within the original image reference frame, facilitating subsequent path planning, voice prompts, or navigation guidance. This method combines deep super-resolution processing, feature fusion, multi-target detection, and back-projection techniques to form a complete closed loop, solving the problems of blurred small targets and indistinct features in traditional low-resolution images, thereby improving the ability to identify obstacles at a distance and the detection stability in multi-target environments. Attached Figure Description

[0010] Figure 1 This is a flowchart illustrating the steps of a multi-object detection method for small images based on super-resolution.

[0011] Figure 2 This is a diagram showing the results of multi-target detection in a small image.

[0012] Figure 3 To identify high-resolution feature maps in an image;

[0013] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0014] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0015] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.

[0016] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated items listed.

[0017] To achieve the above objectives, please refer to Figures 1 to 3 A method for multi-object detection in small images based on super-resolution includes the following steps:

[0018] Step S1: Acquire low-resolution images generated by the guide camera unit and obtain small image data from the low-resolution images;

[0019] Step S2: Recover edge texture information from small image data, generate high-resolution feature maps, extract target response regions at different scales from the high-resolution feature maps, and construct a feature set of obstacles;

[0020] Step S3: Perform multi-target detection on the feature set to generate obstacle detection results, and perform boundary deduplication on the obstacle detection results to obtain multi-target detection results for small images;

[0021] Step S4: Backproject the multi-target detection results of the small image to the coordinate system of the original low-resolution image, and output the multi-target obstacle detection results used by the guide device.

[0022] In one embodiment, an environmental image sequence is acquired using a guide camera unit. The camera unit includes a low-resolution wide-angle camera (640×480 resolution, 30fps frame rate) and an auxiliary infrared illumination device. During acquisition, the camera is mounted at a height of 1.2m and a horizontal viewing angle of 85° at the front of the guide device, continuously acquiring scene images within a range of approximately 5m in front of the user. After the acquired raw images are filtered by a time synchronization module, keyframes are selected in groups of 5 frames to obtain a low-resolution keyframe image set. For each frame, a block segmentation algorithm based on local variance is used to divide the image into several small image blocks. The gradient magnitude and contrast index of each small image block are calculated. If its local information content is lower than a set threshold, the background block is discarded, and only small image data containing obstacle edges or significant texture features are retained.

[0023] In one embodiment, edge texture restoration is performed on the small image data obtained above. A two-stage super-resolution reconstruction network is used to upscale each small image patch from its original size to a higher resolution. The first stage uses an upsampling module based on a residual structure to reconstruct preliminary high-frequency textures; the second stage uses edge-enhancing convolutional layers to finely repair edge contour regions. Multi-scale feature extraction is then performed on the reconstructed high-resolution image, extracting response features at different scales through three-layer scale convolutional kernels (sizes 3×3, 5×5, and 7×7), forming corresponding target response regions in the feature space. Each response region is constrained by feature intensity value, edge direction consistency, and texture continuity to obtain a set of suspected obstacle regions.

[0024] For the obtained obstacle feature set, a multi-object detection network is used to perform recognition and boundary simplification operations. The feature set is input to the convolutional detection module, which generates a target confidence score for each candidate region; target boxes are filtered based on a confidence threshold (0.5); and non-maximum suppression (NMS) is performed to eliminate overlapping detection boxes. After the above processing, preliminary obstacle detection results are obtained. During the training phase, the detection network uses a road scene dataset as samples and optimizes multi-class differentiation through cross-entropy loss.

[0025] The multi-target detection results of the small image are back-projected onto the coordinate system of the original low-resolution image. Based on the upsampling ratio and block segmentation index relationship during super-resolution reconstruction, the corresponding coordinates of each detection box in the low-resolution image are calculated. Secondly, the pixel offset caused by distortion is corrected by the affine matrix so that the boundary position is aligned with the original image.

[0026] In another embodiment, to further improve detection stability, low-resolution images are buffered using a dual-channel buffer during acquisition to ensure that the time difference between two consecutive frames is less than 50ms, thereby ensuring the time consistency of the subsequent feature recovery process. To reduce feature redundancy, principal component analysis is performed on the aforementioned response region set, retaining the first 20 main feature components to form an obstacle feature set. Statistically, a single frame image can generate approximately 15–30 salient feature regions, mainly distributed in the center of the road and at the edges of objects on both sides. The boundary coordinates of the detection results are recalculated using a center point offset method, merging overlapping detection results of the same object in different small image blocks into a single target, forming a small image multi-target detection result. This detection result is output in the form of target category, confidence score, and boundary coordinates.

[0027] During the back-projection process, the spatial pose of the detection results is fine-tuned using attitude angle information provided by the inertial measurement unit (IMU), thereby ensuring stable display of the detection frame position as the guide device moves. The output result data packet includes obstacle category, distance estimate, and image coordinate position.

[0028] Of particular importance, step S1 includes:

[0029] The guide camera unit is activated to capture raw image frames of the external scene in real time;

[0030] Resolution sampling is performed on the captured raw image frames, preserving the low-resolution image of the central region;

[0031] Record the number of pixel rows and columns and the grayscale value matrix of the low-resolution image;

[0032] The grayscale matrix is ​​copied to the processing buffer to generate small image data.

[0033] In one embodiment, the guide camera unit is activated to enter real-time acquisition mode. The guide camera unit employs a low-resolution wide-angle camera, mounted at the front of the guide device, with the acquisition direction aligned with the user's forward direction. Upon activation, it continuously captures raw image frames of the external scene and performs resolution sampling on each frame, retaining only the central region of the image as a low-resolution image. During sampling, the retained region is set to the center pixel block of the original frame to reduce the impact of edge distortion. Subsequently, the pixel row and column count and grayscale value matrix of the retained low-resolution image are calculated, and this matrix is ​​written to the buffer in row-major order. To ensure real-time performance, the buffer uses a dual-channel structure to alternately store the grayscale data of two consecutive frames, thereby avoiding data overwriting. After buffering is complete, the grayscale value matrix is ​​copied from the acquisition buffer to the processing buffer, generating small image data that can be used for subsequent image enhancement. Each small image data frame occupies approximately 110KB of memory, and a complete acquisition and buffering operation can be completed within 200ms.

[0034] In another embodiment, the guide camera unit employs a binocular RGB camera group with a center-to-center distance of 65mm between the left and right cameras. Upon startup, a synchronization trigger signal controls the left and right cameras to simultaneously capture raw frames, and the central area of ​​the left camera is used as the low-resolution image. To adapt to different lighting conditions, the exposure time and gain value are automatically adjusted during sampling to ensure that the dynamic range of the grayscale matrix is ​​between [20, 235]. In low-light scenes, the image grayscale matrix is ​​smoothed by linear interpolation before being written to the buffer to suppress the influence of random noise. The resulting small image data not only contains grayscale matrix information but also includes a timestamp and frame sequence number identifier for subsequent temporal registration and feature recovery operations.

[0035] Preferably, step S2 includes:

[0036] Recover the edge texture and detail information of small targets in small image data to generate high-resolution feature maps;

[0037] Align feature layers of different scales on a high-resolution feature map to generate aligned multi-scale features;

[0038] In the aligned multi-scale features, the target response region is extracted, and the feature intensity and edge continuity of each target response region are calculated.

[0039] Determine the distribution range of the target response area and distinguish between distant and near obstacles based on the scale difference of the distribution range;

[0040] By integrating the feature information of distant and near obstacles, a feature set containing both distant and near obstacles is constructed.

[0041] In one embodiment, super-resolution reconstruction is performed on small image data. Convolutional interpolation algorithms are used to spatially upsample the small image data, increasing the original resolution to a higher resolution, such as a positive integer multiple of the original resolution. Bicubic interpolation is then used to smooth gray-level transitions. Edge regions of small targets are detected based on edge gradient change rates, and edge texture and detail information are recovered using local contrast enhancement methods to generate a high-resolution feature map. After obtaining the high-resolution feature map, its... , and The feature layers are scaled and aligned using affine transformations to ensure consistent spatial coordinates across all scales, generating aligned multi-scale features. Gradient response analysis is then used to calculate the target response region at each scale, statistically analyzing its feature intensity and edge continuity. Regions with feature intensity exceeding a preset threshold and exhibiting good edge continuity are identified as effective response regions.

[0042] Based on the distribution range of the effective response region in the multi-scale features, if the response region is large and has clear edges on the low-scale layer, it is identified as a near obstacle; if the response region only appears on the high-scale layer and has a small shape, it is identified as a far obstacle. The feature information of far and near obstacles is integrated to construct a complete set of obstacle features.

[0043] In another embodiment, super-resolution restoration based on deep residual learning is performed on small image data. This embodiment employs a pre-trained lightweight residual network structure to compensate for details in low-resolution input images. During the restoration process, local feature blocks are first extracted using 3×3 convolutional kernels, edge texture residual information of each feature block is calculated, and high-resolution feature reconstruction is achieved through layer-by-layer iterative stacking. The resulting high-resolution feature map contains 64 feature channels in its structure, which can more completely preserve the detailed information of small targets. In this feature map, a pyramid pooling strategy is further used to extract multi-scale feature layers and align them to a uniform resolution. When performing target response analysis on features at each scale, the feature intensity of the region is measured by calculating channel attention weights, while edge direction consistency is used as an indicator of edge continuity. When the detected target response region shows a significant shrinkage trend in scale changes, it is determined to be a distant obstacle; conversely, targets that maintain a stable contour in all multi-scale layers are marked as near obstacles. The spatial distribution information of the two types of obstacles is merged with the feature vectors to form a unified obstacle feature set for subsequent target recognition.

[0044] Please refer to [link / reference needed] for further information. Figure 3 On a test road with clearly defined boundaries, the road surface uses different paving materials to simulate real road conditions and is marked with a grid of dashed lines. A navigation robot is moving on the sidewalk covered by the grid coordinate system and collecting its perception data relative to the grid marks and objects at near and far distances. The performance of its multi-scale target detection in distinguishing near-end road features from far-end obstacles can be quantitatively evaluated.

[0045] Preferably, restoring the edge texture and detail information of small targets in small image data to generate corresponding high-resolution feature maps includes:

[0046] Calculate the gray-level difference between adjacent pixels for each pixel of the small image data to generate an initial feature map;

[0047] The initial feature map is enlarged layer by layer into a pixel grid, up to four times its original size;

[0048] Calculate the grayscale change amplitude of each pixel in the magnified feature map, and retain pixels whose grayscale change amplitude is greater than a preset change threshold.

[0049] The retained pixels are filled with the grayscale transition values ​​of adjacent pixels, which are used as the filled feature map;

[0050] The filled feature map is then merged with the initial feature map pixel by pixel using a weighted merging process to generate a high-resolution feature map.

[0051] In one embodiment, the grayscale difference between adjacent pixels is calculated pixel by pixel for the obtained small image data. The absolute value of the grayscale difference between each pixel and its right and bottom neighboring pixels is added to generate an initial feature map. The initial feature map is then enlarged twice, layer by layer, in 2×2 pixel blocks, to expand the overall pixel grid to four times its original size. For the enlarged feature map, the grayscale change amplitude is calculated for each pixel. Pixels with amplitudes greater than a preset threshold (e.g., 1.5 times the average grayscale difference) are retained, while the rest are ignored. For the retained pixels, the average grayscale value is calculated for the four neighboring pixels (top, bottom, left, and right). This average value is then weighted and synthesized with the original pixel grayscale value at a ratio of 70%:30% to obtain an updated grayscale value, resulting in a filled feature map. The filled feature map and the initial feature map are then weighted and merged pixel by pixel, with the initial feature map having a weight of 0.6 and the filled feature map having a weight of 0.6. This generates a high-resolution feature map. This feature map displays clear contours in edge regions and can provide input for multi-scale feature extraction.

[0052] In another embodiment, the initial feature map is generated by calculating the gray-level differences in three directions: horizontal, vertical, and diagonal. The gray-level change amplitude of each pixel is taken as the square root of the sum of the squares of the gray-level differences in the three directions, enhancing edge response. Subsequently, the initial feature map is enlarged layer by layer to four times its size, and the local gray-level change amplitude is calculated using a neighborhood sliding window. The retained pixels are determined by adding a fixed threshold to the local average. For the retained pixels, the average gray-level of its 3×3 neighborhood pixels is selected as the fill value, and then compared with the original pixel gray-level... : A proportionally weighted synthesis process is used to generate a padded feature map. This padded feature map is then fused pixel-by-pixel with the initial feature map. The fusion ratio is dynamically adjusted by calculating weight coefficients based on pixel gradients to obtain a high-resolution feature map. This feature map can restore details and textures in a balanced manner, making it suitable for subsequent multi-scale target response extraction and obstacle recognition.

[0053] Preferably, the grayscale change amplitude of each pixel in the magnified feature map is calculated, and pixels with grayscale change amplitudes greater than a preset change threshold are retained, including:

[0054] The magnified feature map is traversed pixel by pixel starting from the top left corner to obtain the absolute value of the grayscale difference between the current pixel and its right-side neighboring pixel.

[0055] Get the absolute value of the grayscale difference between the current pixel and its adjacent pixel below;

[0056] Add the absolute value of the gray level difference on the right to the absolute value of the gray level difference below to obtain the gray level change range of the current pixel;

[0057] The average gray level difference is obtained by summing the gray level changes of all pixels in the magnified feature map and dividing by the total number of pixels.

[0058] Determine the preset change threshold using the average grayscale difference;

[0059] If the grayscale change of the current pixel is greater than the preset change threshold, the grayscale value and position coordinates of the pixel are retained.

[0060] In one embodiment, starting from the top-left pixel, the magnified feature map is traversed pixel by pixel. For the current pixel, the absolute difference between its grayscale value and that of its right-hand neighboring pixel is calculated first, and then the absolute difference between its grayscale value and that of its bottom neighboring pixel is calculated. The two are added together to obtain the grayscale change amplitude of the pixel. The grayscale change amplitudes of all pixels are summed and divided by the total number of pixels to obtain the average grayscale difference. A preset change threshold is set at 1.2 times the average grayscale difference. When the grayscale change amplitude of a pixel is greater than the preset threshold, the grayscale value of that pixel and its row and column coordinates are recorded as reserved pixels for subsequent feature filling and high-resolution feature map generation.

[0061] In another embodiment, the magnified feature map is processed using a block-by-block traversal method, processing 8×8 pixel blocks at a time. For each pixel within a block, the sum of the absolute values ​​of its grayscale differences with its right, bottom, and bottom-right diagonal neighboring pixels is calculated as the pixel's grayscale change amplitude. The average grayscale change amplitudes of all pixels are averaged to obtain the average grayscale difference of the block, and then the global average of all block averages is calculated as a preset change threshold. If the grayscale change amplitude of a pixel is greater than the global threshold, the grayscale value and row and column coordinates of that pixel are retained to form the input dataset for the subsequent high-resolution feature map.

[0062] Preferably, filling the retained pixels with grayscale transition values ​​of adjacent pixels includes:

[0063] Centered on each retained pixel, select its four adjacent pixels in the four directions above, below, left, and right;

[0064] When the position of an adjacent pixel exceeds the image boundary, the gray value of the boundary pixel is used as the replacement value.

[0065] The grayscale values ​​of adjacent pixels in four directions are summed, and the average value is used as the transition fill value.

[0066] The updated grayscale value is obtained by summing 70% of the original grayscale value of the retained pixel with 30% of the transition fill value.

[0067] The gray values ​​of the original retained pixels are replaced with the updated gray values ​​to form the filled feature map.

[0068] In one embodiment, for each retained pixel, neighboring pixels in the four directions (up, down, left, and right) are selected with that pixel as the center. If a neighboring pixel's position exceeds the image boundary, the grayscale value of the boundary pixel is directly used as a replacement. The grayscale values ​​in the four directions are summed and averaged to obtain the transition fill value for the retained pixel. Subsequently, 70% of the original grayscale value of the retained pixel is weighted and summed with 30% of the transition fill value to obtain the updated grayscale value, which replaces the original grayscale value, forming the filled feature map. This operation processes all retained pixels sequentially to ensure that edge details are smoothly filled.

[0069] In another embodiment, for each retained pixel, its four neighboring pixels (top, bottom, left, and right) are also selected as references. If a neighboring pixel exceeds the image boundary, the gray value of the nearest boundary pixel is used instead. The gray values ​​in the four directions are summed, and the average is taken as an initial transition value. This value is then weighted and fused with the original gray value of the retained pixel at a ratio of 60%:40% to generate an updated gray value. This updated gray value is then applied to the retained pixel position, and this operation is repeated for all retained pixels to obtain the filled feature map. This method balances edge information recovery and pixel gray-level smoothing, providing a foundation for high-resolution feature map generation.

[0070] Preferably, the process of weighted merging of the padded feature map and the initial feature map pixel by pixel to generate a high-resolution feature map includes:

[0071] The filled feature map is enlarged to have the same number of rows and columns as the initial feature map;

[0072] The weight allocation coefficients are calculated based on the pixel gradient difference between the magnified feature map and the initial feature map.

[0073] For each pixel location, determine the weight allocation coefficients of the initial feature map and the weight allocation coefficients of the padded feature map, and make the sum of the two equal to 1;

[0074] Starting from the top left corner, traverse the corresponding positions of the two feature maps pixel by pixel, multiply the gray value of the corresponding pixel in the initial feature map by its weight allocation coefficient, multiply the gray value of the corresponding pixel in the filled feature map by its weight allocation coefficient, and add the results to obtain the gray value of the fused pixel.

[0075] The grayscale values ​​of the merged pixels are rearranged according to their original pixel positions to generate a high-resolution feature map.

[0076] In one embodiment, the padded feature map is enlarged row-and-column to the same size as the initial feature map, ensuring a one-to-one correspondence between each pixel. For each pixel location, the pixel gradient difference between the grayscale values ​​of the initial feature map and the padded feature map is calculated. The reciprocal of the gradient difference is used as the weighting benchmark to obtain the weighting coefficients for the initial and padded feature maps, and these coefficients are adjusted so that their sum is 1. Subsequently, starting from the top left corner, the corresponding positions of the two feature maps are traversed pixel by pixel. The grayscale value of the corresponding pixel in the initial feature map is multiplied by its weighting coefficient, and the grayscale value of the corresponding pixel in the padded feature map is multiplied by its weighting coefficient. The two are then added together to obtain the grayscale value of the fused pixel. After all pixel calculations are completed, the fused grayscale values ​​are rearranged according to their original pixel positions to generate a high-resolution feature map. This feature map exhibits enhanced performance in both edge and detail regions and can be directly used for multi-scale target response extraction.

[0077] In another embodiment, the filled feature map is enlarged to the same number of rows and columns as the initial feature map using bilinear interpolation. For each pixel location, its grayscale gradient and the average gradient of its surrounding 3×3 neighborhood are calculated. This gradient value is normalized and used as the weight coefficient of the initial feature map. The weight coefficient of the filled feature map is obtained by subtracting 1 from this weight, ensuring that the sum of the two weights is 1. Subsequently, the process is repeated pixel by pixel from the top left corner. The grayscale value of the initial feature map is multiplied by the initial weight, and the grayscale value of the filled feature map is multiplied by the filling weight. The two are then added together to obtain the grayscale value of the fused pixel. After the traversal is complete, all fused pixels are rearranged according to their original coordinate positions to obtain a high-resolution feature map. This method can dynamically adjust the local pixel weights during the fusion process, enhance the preservation of high-frequency textures, and smooth low-frequency regions, achieving full recovery of small target edges and detail information.

[0078] Preferably, in the aligned multi-scale features, extracting the target response region and calculating the feature intensity and edge continuity of each response region includes:

[0079] The aligned multi-scale features are divided into fixed-size grid cells layer by layer;

[0080] Calculate the sum of all pixel grayscale values ​​within each grid cell, and use this as the grid intensity value;

[0081] Traverse all grid cells. If the grid intensity value is greater than the preset intensity threshold, mark the grid cell as a candidate response region.

[0082] For each candidate response region, the proportion of adjacent pixels belonging to the same category in the boundary pixels is counted to obtain the edge continuity value;

[0083] Regions in the candidate response region whose edge continuity value is greater than a preset continuity threshold are identified as target response regions.

[0084] In one embodiment, the aligned multi-scale feature map is divided layer by layer into fixed-size grid cells, each containing 8×8 pixels. The grayscale values ​​of the pixels within each grid cell are summed to obtain the grid intensity value. All grid cells are traversed, and when the grid intensity value is greater than a preset intensity threshold, the grid is marked as a candidate response region. The boundary pixels of each candidate response region are statistically analyzed, and the proportion of adjacent pixels belonging to the same region is calculated to obtain the edge continuity value. Adjacent pixels are pixels that are adjacent on the boundary of a candidate region and belong to the same region (same category). Candidate regions with edge continuity values ​​greater than a preset continuity threshold are identified as target response regions, and their positions and sizes are recorded for subsequent multi-target detection. This process can filter out high-intensity, well-continuous response regions while suppressing noise and isolated pixel interference.

[0085] In another embodiment, the aligned multi-scale feature map is divided layer by layer into 10×10 pixel grid cells. The grayscale values ​​of all pixels within each grid cell are summed to obtain the grid intensity value. All grid cells are traversed, and grids with intensity values ​​greater than the global average intensity plus 10% are marked as candidate response regions. For each candidate region, the proportion of pixels in its boundary pixels with a grayscale difference of less than 5 from their neighboring pixels is counted to obtain the edge continuity value. Candidate regions with edge continuity values ​​higher than 80% are confirmed as target response regions, and their boundary coordinates and pixel distribution are recorded. In this way, target response regions at different scales can be effectively extracted, providing an accurate spatial basis for constructing feature sets for both near and far-range obstacles.

[0086] Preferably, step S3 includes:

[0087] For each target response region in the feature set, a preset obstacle category benchmark is matched one by one;

[0088] Calculate the similarity score between the target response region and the matching preset obstacle category benchmark, and record the obstacle category corresponding to the highest score;

[0089] The center pixel coordinates of the target response region are used as the target location. The obstacle category is combined with the target location to generate the initial obstacle detection result.

[0090] The initial obstacle detection results are iterated through, and the results with a confidence score greater than a preset confidence threshold are retained as the detection retention results.

[0091] Iterate through all detection and retention results. If the center distance between any two results is less than a preset distance threshold, merge them into one result as the multi-target detection result for the small image.

[0092] In one embodiment, each target response region in the feature set is sequentially matched with a preset obstacle category benchmark. For each matching operation, the similarity score between the grayscale, texture, and edge features of the target response region and the matching benchmark is calculated, and the obstacle category corresponding to the highest similarity score for each target response region is recorded. Subsequently, the center pixel coordinates of each target response region are calculated, and the center coordinates are combined with the highest similarity category to form an initial obstacle detection result. The initial detection results are iterated through, retaining results with confidence scores greater than a preset threshold, while deleting low-confidence results. The retained detection results are then compared pairwise. If the sum of the squared row and column distances of the center coordinates of any two results is less than the squared preset distance threshold, the two results are merged into one, and the one with the higher confidence score is retained as the final position. All merged detection results are arranged in order of center coordinates to obtain the small image multi-target detection result, which is used as input for subsequent backprojection.

[0093] In another embodiment, for each target response region in the feature set, its similarity score with the obstacle category benchmark in the grayscale histogram and edge gradient direction histogram is calculated sequentially. For each target response region, the obstacle category with the highest similarity is recorded, and the row and column center coordinates of that region are used as the target location. After forming the initial obstacle detection results, the results are first filtered according to the similarity threshold, retaining those with a similarity greater than 0.7. Then, the retained results are compared one by one, and the Euclidean distance between the center coordinates of the two target response regions in the row and column directions is calculated. If the distance is less than 10 pixels, they are considered to belong to the same obstacle, and the two results are merged, with the higher of the two confidence values ​​being taken as the merged result. The above operation is repeated until all overlapping or adjacent detection results are merged. All merged detection results are then sorted in ascending order by the row value of the center coordinates to generate small image multi-target detection results, providing complete input for multi-target recognition in guide vision devices.

[0094] Please refer to [link / reference needed] for further information. Figure 2 The image is scanned at different scales to identify all regions that are obstacles, i.e., the feature set. Four sub-images contain red boxes of varying sizes, representing the responses to the same target at different scales or viewing distances. The center point of each red box is recorded as the target location, and its corresponding highest similarity category (e.g., "stones") is determined. These two are combined to form the detection result. The most accurate red box is retained at the stone location.

[0095] Preferably, the center pixel coordinates of the target response region are used as the target location, and the obstacle category is combined with the target location to generate the initial obstacle detection result, including:

[0096] During the detection phase, the average row and column coordinates of the top-left and bottom-right corner coordinates of the target response region are calculated, and the average value is rounded to determine the row and column coordinates of the center pixel.

[0097] Record the determined row and column coordinates of the center pixel as the target position;

[0098] Based on the obstacle category information in the detection results, the recorded target location is matched with the obstacle category;

[0099] The paired obstacle categories and their corresponding target location information are stored together in the results list to generate preliminary obstacle detection results.

[0100] In one embodiment, for each target response region, the row and column values ​​of its top-left and bottom-right corner coordinates are obtained. The average of the row coordinates is taken and rounded down to obtain the row coordinates of the center pixel; the average of the column coordinates is taken and rounded down to obtain the column coordinates of the center pixel. The calculated center pixel coordinates are recorded as the target position of the target response region. Subsequently, the corresponding obstacle category information is obtained and paired with the target position. All paired obstacle categories and target positions are sequentially stored in the result list to form a preliminary obstacle detection result, which is used for subsequent confidence filtering and boundary deduplication processing.

[0101] In another embodiment, for each target response region, the row and column coordinates of its top-left and bottom-right corners are obtained, and the average values ​​of the row and column coordinates are calculated. The average values ​​are then rounded to obtain the row and column coordinates of the center pixel. These center coordinates are recorded as the target location. The corresponding obstacle category information is matched one-to-one with the recorded target location, and the pairing information is written into the target detection result list. Pixel area and boundary width and height information are appended to each target response region for distance determination and merging operations in the subsequent multi-target detection stage. After completing the operations for all target response regions, a complete initial obstacle detection result list is obtained, providing an accurate basis for the next step of filtering and merging.

[0102] Preferably, the detection results are grouped according to their location coordinates. If the center distance between any two results is less than a preset distance threshold, they are merged into one result, which is used as the multi-target detection result for the small image.

[0103] Iterate through all the retained detection results and compare the center coordinates of any two detection results in turn.

[0104] Calculate the squared difference between the row coordinates and the squared difference between the column coordinates of the two center coordinates, and add the squared difference between the row coordinates and the squared difference between the column coordinates to generate the sum of squares of the coordinate interpolation;

[0105] When the sum of squares is less than the square of a preset distance threshold, the two detection results are determined to belong to the same target area, and the detection result with higher confidence is retained.

[0106] Delete detection results with low confidence and merge the detection results;

[0107] The merged detection results are sorted in ascending order according to the row coordinate values ​​of the center coordinate to generate multi-target detection results.

[0108] In one embodiment, the retained list of detection results is traversed, and any two detection results are selected sequentially to obtain the row and column coordinates of their center pixels. The squared difference between the row and column coordinates of the two center coordinates is calculated, and the squared differences are added to generate the squared sum of coordinate interpolation, which is the squared Euclidean distance. When the squared Euclidean distance is less than the square of a preset distance threshold, the two detection results are determined to belong to the same target region. After determining that they belong to the same target region, the confidence scores of the two detection results are compared, the detection result with higher confidence is retained, the result with lower confidence is deleted, and the retained result is updated to the final detection result list. The above steps are repeated until all detection results have been compared pairwise and all overlapping or adjacent results have been merged. The merged detection results are sorted in ascending order according to the row coordinate values ​​of the center coordinates to generate a small image multi-target detection result, which is used as the output result for the guide device.

[0109] In another embodiment, a temporary two-dimensional index table is established for the retained detection results based on their center pixel coordinates. Each detection result is traversed, and its row and column Euclidean distances to the center coordinates of other detection results are calculated sequentially. If the center distance between any two detection results is less than a preset distance threshold (e.g., 10 pixels), the two results are determined to correspond to the same obstacle. For detection results of the same obstacle, their confidence scores are compared, and the higher score is used as the final position and category of the merged target; the lower score is removed from the list. After the merging operation is completed, all detection results are sorted in ascending row coordinate order and column coordinate order to generate a final small image multi-target detection result list. Simultaneously, the center coordinates, obstacle category, and confidence information of each target are recorded, providing input for subsequent backprojection.

[0110] Of particular importance, step S4 includes:

[0111] Record the center row and column coordinates of each target in the multi-target detection results of a small image;

[0112] Divide the center row coordinates by the preset magnification factor to obtain the original row coordinates; divide the center column coordinates by the preset magnification factor to obtain the original column coordinates.

[0113] The back-projection location is combined with the corresponding obstacle category to generate back-projection detection results;

[0114] The back-projection detection results are sorted in ascending order according to the original row coordinates, and the multi-target obstacle detection results used by the guide equipment are output.

[0115] In one embodiment, the multi-target detection result list of the small image is traversed, and the row and column coordinates of the center pixel of each target are recorded. Then, the center row coordinates are divided by a preset magnification factor (e.g., 4 times) to obtain the row coordinates in the original image coordinate system; the center column coordinates are divided by the same magnification factor to obtain the original column coordinates. The original row and column coordinates obtained from backprojection are combined with the corresponding obstacle category information to generate a backprojection detection result list. After backprojection of all targets is completed, the backprojection detection results are sorted in ascending order by the original row coordinates. If there are identical row coordinates, they are then sorted in ascending order by column coordinates to form the final multi-target obstacle detection result list used by the guide device, so that the device can provide target prompts and navigation references in the original image coordinate system.

[0116] In another embodiment, for each target in the multi-target detection results of a small image, its center row and column coordinates and obstacle category information are obtained. The center row and column coordinates are divided by a preset magnification factor to obtain the corresponding original row and column coordinates, and the results are rounded down to adapt to the pixel index. The original coordinates of the backprojection are paired with the obstacle category and stored in a temporary list of backprojection results. After processing all targets, the backprojection result list is sorted in ascending order by row coordinates, and multi-targets with the same row coordinates are sorted in ascending order by column coordinates to generate a complete multi-target obstacle detection result. This result contains the original row and column positions and obstacle category information of each target, which can be directly used for target display, prompting, and path planning in guide devices for the visually impaired.

[0117] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. The invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A method for multi-target detection in small images based on super-resolution, characterized in that, Includes the following steps: Step S1: Acquire low-resolution images generated by the guide camera unit and obtain small image data in the low-resolution images. Specifically, the image is divided into several small image blocks using a block segmentation algorithm based on local variance. The gradient magnitude and contrast index of each small image block are calculated. If the local information content is lower than the set threshold, the background block is removed and only the small image data containing obstacle edges or significant texture features are retained. Step S2: Recover edge texture information from small image data, generate high-resolution feature maps, extract target response regions at different scales from the high-resolution feature maps, and construct a feature set for obstacles; wherein, step S2 includes: To recover the edge texture and detail information of small objects in small image data and generate high-resolution feature maps; this includes: Calculate the gray-level difference between adjacent pixels for each pixel of the small image data to generate an initial feature map; The initial feature map is enlarged layer by layer into a pixel grid, up to four times its original size; Calculate the grayscale change amplitude of each pixel in the magnified feature map, and retain pixels whose grayscale change amplitude is greater than a preset change threshold. The retained pixels are filled with the grayscale transition values ​​of adjacent pixels, which are used as the filled feature map; The padded feature map is then pixel-by-pixel weighted and merged with the initial feature map to generate a high-resolution feature map. Align feature layers of different scales on a high-resolution feature map to generate aligned multi-scale features; In the aligned multi-scale features, the target response region is extracted, and the feature intensity and edge continuity of each target response region are calculated. Determine the distribution range of the target response area and distinguish between distant and near obstacles based on the scale difference of the distribution range; Integrate the feature information of distant and near obstacles to construct a feature set containing both distant and near obstacles; Step S3: Perform multi-target detection on the feature set to generate obstacle detection results, and perform boundary deduplication on the obstacle detection results to obtain multi-target detection results for small images; Step S4: Backproject the multi-target detection results of the small image to the coordinate system of the original low-resolution image, and output the multi-target obstacle detection results used by the guide device.

2. The method for multi-target detection in small images based on super-resolution according to claim 1, characterized in that, Calculate the grayscale change amplitude of each pixel in the magnified feature map, and retain pixels whose grayscale change amplitude is greater than a preset change threshold, including: The magnified feature map is traversed pixel by pixel starting from the top left corner to obtain the absolute value of the grayscale difference between the current pixel and its right-side neighboring pixel. Get the absolute value of the grayscale difference between the current pixel and its adjacent pixel below; Add the absolute value of the gray level difference on the right to the absolute value of the gray level difference below to obtain the gray level change range of the current pixel; The average gray level difference is obtained by summing the gray level changes of all pixels in the magnified feature map and dividing by the total number of pixels. Determine the preset change threshold using the average grayscale difference; If the grayscale change of the current pixel is greater than the preset change threshold, the grayscale value and position coordinates of the pixel are retained.

3. The method for multi-target detection in small images based on super-resolution according to claim 1, characterized in that, The grayscale transition values ​​used to fill in the remaining pixels with adjacent pixels include: Centered on each retained pixel, select its four adjacent pixels in the four directions above, below, left, and right; When the position of an adjacent pixel exceeds the image boundary, the gray value of the boundary pixel is used as the replacement value. The grayscale values ​​of adjacent pixels in four directions are summed, and the average value is used as the transition fill value. The updated grayscale value is obtained by summing 70% of the original grayscale value of the retained pixel with 30% of the transition fill value. The gray values ​​of the original retained pixels are replaced with the updated gray values ​​to form the filled feature map.

4. The method for multi-target detection in small images based on super-resolution according to claim 1, characterized in that, The padded feature map is then pixel-by-pixel weighted and merged with the initial feature map to generate a high-resolution feature map, including: The filled feature map is enlarged to have the same number of rows and columns as the initial feature map; The weight allocation coefficients are calculated based on the pixel gradient difference between the magnified feature map and the initial feature map. For each pixel location, determine the weight allocation coefficients of the initial feature map and the weight allocation coefficients of the padded feature map, and make the sum of the two equal to 1; Starting from the top left corner, traverse the corresponding positions of the two feature maps pixel by pixel, multiply the gray value of the corresponding pixel in the initial feature map by its weight allocation coefficient, multiply the gray value of the corresponding pixel in the filled feature map by its weight allocation coefficient, and add the results to obtain the gray value of the fused pixel. The grayscale values ​​of the merged pixels are rearranged according to their original pixel positions to generate a high-resolution feature map.

5. The method for multi-target detection in small images based on super-resolution according to claim 1, characterized in that, In the aligned multi-scale features, the target response region is extracted, and the feature intensity and edge continuity of each response region are calculated, including: The aligned multi-scale features are divided into fixed-size grid cells layer by layer; Calculate the sum of all pixel grayscale values ​​within each grid cell, and use this as the grid intensity value; Traverse all grid cells. If the grid intensity value is greater than the preset intensity threshold, mark the grid cell as a candidate response region. For each candidate response region, the proportion of adjacent pixels belonging to the same category in the boundary pixels is counted to obtain the edge continuity value; Regions in the candidate response region whose edge continuity value is greater than a preset continuity threshold are identified as target response regions.

6. The method for multi-target detection in small images based on super-resolution according to claim 1, characterized in that, Step S3 includes: For each target response region in the feature set, a preset obstacle category benchmark is matched one by one; Calculate the similarity score between the target response region and the matching preset obstacle category benchmark, and record the obstacle category corresponding to the highest score; The center pixel coordinates of the target response region are used as the target location. The obstacle category is combined with the target location to generate the initial obstacle detection result. The initial obstacle detection results are iterated through, and the results with a confidence score greater than a preset confidence threshold are retained as the detection retention results. Iterate through all detection and retention results. If the center distance between any two results is less than a preset distance threshold, merge them into one result as the multi-target detection result for the small image.

7. The method for multi-target detection in small images based on super-resolution according to claim 6, characterized in that, Using the center pixel coordinates of the target response region as the target location, and combining the obstacle category with the target location, the initial obstacle detection results are generated, including: During the detection phase, the average row and column coordinates of the top-left and bottom-right corner coordinates of the target response region are calculated, and the average value is rounded to determine the row and column coordinates of the center pixel. Record the row and column coordinates of the determined center pixel as the target position; Based on the obstacle category information in the detection results, the recorded target location is matched with the obstacle category; The paired obstacle categories and their corresponding target location information are stored together in the results list to generate preliminary obstacle detection results.

8. The method for multi-target detection in small images based on super-resolution according to claim 6, characterized in that, Iterate through all detection and retention results. If the center distance between any two results is less than a preset distance threshold, merge them into one result, which is then used as the multi-object detection result for the small image. Iterate through all the retained detection results and compare the center coordinates of any two detection results in turn. Calculate the squared difference between the row coordinates and the squared difference between the column coordinates of the two center coordinates, and add the squared difference between the row coordinates and the squared difference between the column coordinates to generate the sum of squares of the coordinate interpolation; When the sum of squares is less than the square of a preset distance threshold, the two detection results are determined to belong to the same target area, and the detection result with higher confidence is retained. Delete detection results with low confidence and merge the detection results; The merged detection results are sorted in ascending order according to the row coordinate values ​​of the center coordinate to generate multi-target detection results.

Citation Information

Patent Citations

  • System and method for identifying foreground and background portions of digitized images

    US6263091B1

  • Method of and apparatus for detecting a human face and observer tracking display

    US6633655B1