A Low-Resolution Object Detection Method and System Based on Heterogeneous Computing
By adopting a heterogeneous calculation method in low-resolution object object detection, adjusting the execution efficiency of the recognition strategy based on image resolution and system energy consumption level, the problem of difficult balance between efficiency and energy consumption in the prior art is solved, and efficient detection in long-term running scenarios is achieved.
Patent Information
- Application Number
- CN202510156745.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-13
AI Technical Summary
The prior art is difficult to achieve a balance between efficiency and energy consumption in low-resolution object object detection, especially in long-term operation scenarios.
The low-resolution object object detection method based on heterogeneous calculation is adopted, and the image resolution level and system energy consumption level are determined by receiving image processing requests, the identification strategy of the corresponding level is called, and the execution efficiency of the identification strategy is adjusted according to the image resolution level to achieve a balance of efficiency and energy consumption.
It achieves a balance of efficiency and energy consumption in low-resolution object object detection, which is suitable for long-term operation scenarios, extends the service life of the equipment and improves the continuity and reliability of tasks.
Smart Images

Figure CN119625505B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition technology, and in particular, to a low-resolution object target detection method and system based on heterogeneous computing. Background Art
[0002] In the prior art, for low-resolution object target detection, the following several technical solutions are mainly adopted:
[0003] CPU solution: Currently, many low-resolution object target detection systems use a general-purpose CPU for image processing and target detection. The advantage of the CPU lies in its versatility and programming flexibility, and it can handle multiple tasks. However, the CPU has low efficiency in processing image data. Especially when processing a large number of low-resolution images, the computing speed is slow and the energy consumption is high. This makes the CPU solution perform poorly in scenarios that require long-term stable operation (such as industrial monitoring, deep space exploration, etc.). For example, in deep space exploration missions, the energy supply is limited, and a low-energy consumption solution is required to extend the service life of the equipment.
[0004] GPU solution: The GPU performs excellently in parallel computing and can significantly improve the image processing speed. However, the GPU has high energy consumption and is not suitable for scenarios with long-term operation, such as industrial monitoring, deep space exploration, etc. For example, in the automatic recognition task of an unmanned ship, continuous operation for several hours or even several days requires a low-energy consumption solution, otherwise, it will frequently need to replenish power, affecting the continuity and reliability of the task.
[0005] FPGA solution: The FPGA can implement specific computing tasks at the hardware level, with relatively low energy consumption, but it has high development complexity and poor flexibility, and it is difficult to quickly deploy new models and algorithms. This is an obvious disadvantage for scenarios that require rapid iteration and update (such as laboratory research). In addition, the FPGA has a large resource occupancy, which is not conducive to miniaturization and portable applications. Summary of the Invention
[0006] To at least overcome to some extent the problem that it is difficult to achieve a balance between efficiency and energy consumption in the prior art for target detection of low-resolution objects, this application provides a low-resolution object target detection method and system based on heterogeneous computing.
[0007] The solution of this application is as follows:
[0008] According to the first aspect of the embodiments of this application, a low-resolution object target detection method based on heterogeneous computing is provided, including:
[0009] Receiving an image processing request and an image to be processed sent by a user;
[0010] Determine the resolution level of the image to be processed, and obtain the energy consumption level of the current system;
[0011] Call the recognition strategy of the corresponding level according to the energy consumption level of the current system;
[0012] Based on the called recognition strategy, adjust the execution efficiency of the recognition strategy according to the resolution level of the image to be processed, and use it as the execution strategy;
[0013] Perform low-resolution object target detection on the image to be processed through the execution strategy, and output the target detection image;
[0014] Among them, the levels of the recognition strategy at least include: the first-level recognition strategy, the second-level recognition strategy, and the third-level recognition strategy;
[0015] The recognition algorithms adopted by the recognition strategies of each level are different, and the energy consumption and processing ability of the first-level recognition strategy are greater than those of the second-level recognition strategy, which are greater than those of the third-level recognition strategy.
[0016] Preferably, the method further includes:
[0017] Judge whether the recognition strategy level required by the user is included in the image processing request;
[0018] If the recognition strategy level required by the user is included in the image processing request, judge whether the recognition strategy level required by the user exceeds the currently called recognition strategy level;
[0019] If the recognition strategy level required by the user does not exceed the currently called recognition strategy level, maintain the current execution strategy;
[0020] If the recognition strategy level required by the user exceeds the currently called recognition strategy level, determine whether the execution efficiency corresponding to the current execution strategy exceeds the set efficiency threshold;
[0021] If the execution efficiency corresponding to the current execution strategy does not exceed the set efficiency threshold level, adjust the execution efficiency of the recognition strategy to the maximum value and update the execution strategy;
[0022] If the execution efficiency corresponding to the current execution strategy exceeds the set efficiency threshold, update the execution strategy according to the recognition strategy level required by the user.
[0023] Preferably, the method further includes:
[0024] If the recognition strategy level required by the user exceeds the preset recognition strategy level, judge whether the energy consumption level of the current system is lower than the preset energy consumption threshold;
[0025] If the energy consumption level of the current system is not lower than the preset energy consumption threshold, determine whether the execution efficiency corresponding to the current execution strategy is the maximum value of the first-level recognition strategy;
[0026] If the execution efficiency corresponding to the current execution strategy is the maximum value of the first-level recognition strategy, then on the basis of the full-efficiency execution of the first-level recognition strategy, synchronously execute the second-level recognition strategy with full efficiency to perform low-resolution object target detection on the to-be-processed image;
[0027] If the current execution strategy is the first-level recognition strategy, but the corresponding execution efficiency is not the maximum value of the first-level recognition strategy, then on the basis of the full-efficiency execution of the first-level recognition strategy, synchronously execute the third-level recognition strategy with full efficiency to perform low-resolution object target detection on the to-be-processed image.
[0028] Preferably, the method further includes:
[0029] If the energy consumption level of the current system is lower than the preset energy consumption threshold, then synchronously execute all recognition strategies with full efficiency to perform low-resolution object target detection on the to-be-processed image.
[0030] Preferably, the method further includes:
[0031] Assign weights to each recognition strategy;
[0032] Fuse the target detection results of all recognition strategies according to the weights assigned to each recognition strategy, and output a consistent detection result.
[0033] Preferably, the method further includes:
[0034] Obtain the parameters of the current display system, and determine the display image format according to the parameters of the current display system;
[0035] Convert the target detection image into the display image format;
[0036] Perform image enhancement processing on the target detection image in the display image format through histogram equalization, and perform noise filtering through Gaussian filtering;
[0037] Determine the resolution according to the parameters of the current display system;
[0038] Adjust the resolution of the processed target detection image according to the resolution of the current display system through bilinear interpolation;
[0039] Output the target detection image with adjusted resolution through the current display system.
[0040] Preferably, the first-level recognition strategy includes:
[0041] Based on the parallel computing ability and high-bandwidth memory access ability of the GPU, perform low-resolution object target detection on the to-be-processed image through the first recognition algorithm;
[0042] The first recognition algorithm includes:
[0043] Capture the global dependencies of the to-be-processed image through the self-attention mechanism of the Transformer encoder;
[0044] According to the global dependencies of the to-be-processed image, gradually add noise to the to-be-processed image. After T steps, obtain a completely randomized feature map;
[0045] Gradually reduce the noise in the completely randomized feature map. After T steps, restore the to-be-processed image;
[0046] Extract features from the to-be-processed image added with noise and denoised reversely at different scales, and fuse the extracted features; wherein, different weights are assigned to each scale;
[0047] Perform low-resolution object target detection on the image after feature fusion, and output the target detection image.
[0048] Preferably, the secondary recognition strategy includes:
[0049] Based on the multi-core parallel structure of ARM, dynamically adjust the frequency according to different loads;
[0050] Perform low-resolution object target detection on the to-be-processed image through the second recognition algorithm;
[0051] The second recognition algorithm includes:
[0052] Perform hierarchical multi-scale feature extraction on the to-be-processed image, and obtain feature maps of different scales on the to-be-processed image through convolution kernels of different scales;
[0053] Send the obtained feature maps of different scales into the fusion layer to obtain the fused comprehensive feature map;
[0054] Perform low-resolution object target detection on the fused comprehensive feature map, and output the target detection image;
[0055] Among them, the fusion process includes: splicing the feature maps of different scales, and integrating multi-scale features through a convolution operation;
[0056] Introduce a hierarchical attention mechanism on the basis of the comprehensive feature map to adaptively adjust the feature importance of different regions in the comprehensive feature map;
[0057] Fuse the spatial domain and frequency domain features, and upsample the comprehensive feature map through an interpolation method to expand the comprehensive feature map to the target resolution.
[0058] Preferably, the three-level recognition strategy includes:
[0059] Perform instruction scheduling based on the multi-level pipeline of the MIPS instruction set architecture;
[0060] Perform memory access based on the set number of memory accesses and memory access patterns;
[0061] Dynamically adjust the working frequencies and voltages of the CPU and the neural network accelerator according to the current task load;
[0062] Dynamically screen the feature regions in the to-be-processed image that meet the requirement of importance through the third recognition algorithm for recognition;
[0063] The third recognition algorithm includes:
[0064] Perform sparsification processing on the to-be-processed image;
[0065] Apply a convolution kernel on the sparsified image to perform adaptive sparse convolution operations, and dynamically skip zero-valued features during the convolution process;
[0066] Extract features and classify the convolved image through a pooling layer and a fully connected layer;
[0067] Output the target detection image, including the position, category, and confidence of the target.
[0068] According to the second aspect of the embodiments of the present application, a low-resolution object target detection system based on heterogeneous computing is provided, including:
[0069] A receiving module, configured to receive an image processing request and a to-be-processed image sent by a user;
[0070] A determining module, configured to determine the resolution level of the to-be-processed image and obtain the energy consumption level of the current system;
[0071] A calling module, configured to call a recognition module of a corresponding level according to the energy consumption level of the current system;
[0072] An adjustment module, configured to, based on the called recognition module, adjust the execution efficiency of the recognition module according to the resolution level of the to-be-processed image as an execution module;
[0073] An execution module, configured to perform low-resolution object target detection on the to-be-processed image and output a target detection image;
[0074] Among them, the levels of the recognition module at least include: a first-level recognition module, a second-level recognition module, and a third-level recognition module;
[0075] The recognition algorithms adopted by the recognition modules at each level are different, and the energy consumption and processing capacity of the first-level recognition module are greater than those of the second-level recognition module, which are greater than those of the third-level recognition module.
[0076] The technical solution provided by this application may include the following beneficial effects:
[0077] The low-resolution object target detection method based on heterogeneous computing in this application includes: receiving an image processing request and an image to be processed sent by a user; determining the resolution level of the image to be processed, and obtaining the energy consumption level of the current system; calling a recognition strategy corresponding to the energy consumption level of the current system; based on the called recognition strategy, adjusting the execution efficiency of the recognition strategy according to the resolution level of the image to be processed as an execution strategy; performing low-resolution object target detection on the image to be processed through the execution strategy, and outputting a target detection image; among them, the levels of the recognition strategy at least include: a first-level recognition strategy, a second-level recognition strategy, and a third-level recognition strategy; the recognition algorithms adopted by the recognition strategies at each level are different, and the energy consumption and processing capacity of the first-level recognition strategy are greater than those of the second-level recognition strategy, which are greater than those of the third-level recognition strategy.
[0078] In this technical solution, first, the energy consumption level of the current system is obtained, and a recognition strategy corresponding to the energy consumption level of the current system is called, so as to ensure that the called recognition strategy does not consume a large amount of energy of the current system. Then, according to the resolution level of the image to be processed, the execution efficiency of the recognition strategy is adjusted to further reduce the energy consumption of the system when performing target recognition. And adjusting the execution efficiency of the recognition strategy according to the resolution level of the image to be processed can also ensure the processing efficiency of the image to be processed, so as to achieve the balance between efficiency and energy consumption.
[0079] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0081] Figure 1 is a schematic flowchart of a low-resolution object target detection method based on heterogeneous computing provided by an embodiment of this application;
[0082] Figure 2 is a schematic structural diagram of a low-resolution object target detection system based on heterogeneous computing provided by an embodiment of this application.
[0083] Reference signs: receiving module - 21; determining module - 22; calling module - 23; adjusting module - 24; executing module - 25; primary recognition module - 26; secondary recognition module - 27; tertiary recognition module - 28. Detailed implementation manners
[0084] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0085] Embodiment 1
[0086] Figure 1 is a flowchart of a low-resolution object target detection method based on heterogeneous computing provided by an embodiment of the present application. Referring to Figure 1 , a low-resolution object target detection method based on heterogeneous computing includes:
[0087] S11: Receive an image processing request and an image to be processed sent by a user;
[0088] S12: Determine the resolution level of the image to be processed, and obtain the energy consumption level of the current system;
[0089] S13: Call a recognition strategy of a corresponding level according to the energy consumption level of the current system;
[0090] It should be noted that the levels of the recognition strategies at least include: a primary recognition strategy, a secondary recognition strategy, and a tertiary recognition strategy;
[0091] The recognition algorithms adopted by the recognition strategies of each level are different, and the energy consumption and processing capabilities of the primary recognition strategy are greater than those of the secondary recognition strategy, which are greater than those of the tertiary recognition strategy.
[0092] The primary recognition strategy is a high-energy consumption recognition strategy, the secondary recognition strategy is a recognition strategy with a balance between energy consumption and efficiency, and the tertiary recognition strategy is a low-energy consumption recognition strategy.
[0093] In this embodiment, the energy consumption occupied during the operation of the current system is used as the energy consumption level of the system. By doing so, when the energy consumption level of the current system is low, a high-energy consumption recognition strategy can be called; when the energy consumption level of the current system is normal, a recognition strategy with a balance between energy consumption and efficiency is called; when the energy consumption level of the current system is high, a low-energy consumption recognition strategy is called to reduce the total energy consumption of the system.
[0094] S14: Based on the invoked recognition strategy, adjust the execution efficiency of the recognition strategy according to the resolution level of the image to be processed, and use it as the execution strategy;
[0095] It should be noted that the lower the resolution level of the image, the higher the processing power required, and the higher the system energy consumption occupied during processing. Conversely, when the resolution level of the image is relatively high, the required processing power is lower, and the system energy consumption occupied during processing is also lower.
[0096] Illustrative example:
[0097] The resolution levels of the image to be processed are divided into 10 levels in total, with level 10 being the lowest and level 1 being the highest.
[0098] If the resolution level of the current image to be processed is level 1, adjust the execution efficiency of the recognition strategy to 10%. If the resolution level of the current image to be processed is level 10, adjust the execution efficiency of the recognition strategy to 100%, and so on.
[0099] S15: Perform low-resolution object target detection on the image to be processed through the execution strategy, and output the target detection image.
[0100] In this technical solution, first obtain the energy consumption level of the current system, and invoke the corresponding level of recognition strategy according to the energy consumption level of the current system, so as to ensure that the invoked recognition strategy does not occupy a large amount of energy of the current system. Then, according to the resolution level of the image to be processed, adjust the execution efficiency of the recognition strategy to further reduce the system energy consumption during target recognition. And adjusting the execution efficiency of the recognition strategy according to the resolution level of the image to be processed can also ensure the processing efficiency of the image to be processed, thus achieving a balance between efficiency and energy consumption.
[0101] Embodiment 2
[0102] It should be noted that the method further includes:
[0103] Judge whether the recognition strategy level required by the user is included in the image processing request;
[0104] If the recognition strategy level required by the user is included in the image processing request, judge whether the recognition strategy level required by the user exceeds the currently invoked recognition strategy level;
[0105] If the recognition strategy level required by the user does not exceed the currently invoked recognition strategy level, maintain the current execution strategy;
[0106] If the recognition strategy level required by the user exceeds the currently invoked recognition strategy level, determine whether the execution efficiency corresponding to the current execution strategy exceeds the set efficiency threshold;
[0107] If the execution efficiency corresponding to the current execution policy does not exceed the set efficiency threshold level, adjust the execution efficiency of the recognition policy to the maximum value and update the execution policy.
[0108] If the execution efficiency corresponding to the current execution policy exceeds the set efficiency threshold, update the execution policy according to the recognition policy level required by the user.
[0109] In this embodiment, considering that the user may propose the recognition policy level required by themselves, the recognition policy level required by the user is included in the image processing request. It is necessary to determine whether the recognition policy level required by the user exceeds the currently invoked recognition policy level. If the recognition policy level required by the user does not exceed the currently invoked recognition policy level, it means that the currently invoked recognition policy level can meet the user's requirements, and the current execution policy is maintained.
[0110] If the recognition policy level required by the user exceeds the currently invoked recognition policy level, it means that the currently invoked recognition policy level fails to meet the user's requirements. At this time, it is necessary to determine the specific execution plan according to the current situation:
[0111] Determine whether the execution efficiency corresponding to the current execution policy exceeds the set efficiency threshold. The set efficiency threshold can be 50%. If the execution efficiency corresponding to the current execution policy does not exceed the set efficiency threshold level, it means that the resolution of the currently to-be-processed image is relatively high and above the average line. Then adjust the execution efficiency of the recognition policy to the maximum value, that is, perform low-resolution object target detection on the to-be-processed image with the 100% execution efficiency of the current recognition policy. In this way, without changing the current recognition policy, the user's needs can be met to a certain extent.
[0112] If the execution efficiency corresponding to the current execution policy exceeds the set efficiency threshold, it means that the resolution of the currently to-be-processed image is relatively low and below the average line. At this time, even if the execution efficiency of the recognition policy is adjusted to the maximum value, it cannot effectively improve. It is necessary to update the execution policy according to the recognition policy level required by the user. For example, if the recognition policy required by the user is the first-level recognition policy and the current execution policy is 70% of the second-level recognition policy, then the execution policy needs to be updated to 70% of the third-level recognition policy.
[0113] Embodiment III
[0114] It should be noted that the method further includes:
[0115] If the recognition policy level required by the user exceeds the preset recognition policy level, determine whether the energy consumption level of the current system is lower than the preset energy consumption threshold;
[0116] If the energy consumption level of the current system is not lower than the preset energy consumption threshold, determine whether the execution efficiency corresponding to the current execution policy is the maximum value of the first-level recognition policy;
[0117] If the execution efficiency corresponding to the current execution policy is the maximum value of the primary recognition policy, then on the basis of the full-efficiency execution of the primary recognition policy, the secondary recognition policy is synchronously executed at full efficiency to perform low-resolution object target detection on the image to be processed.
[0118] If the current execution policy is the primary recognition policy, but the corresponding execution efficiency is not the maximum value of the primary recognition policy, then on the basis of the full-efficiency execution of the primary recognition policy, the tertiary recognition policy is synchronously executed at full efficiency to perform low-resolution object target detection on the image to be processed.
[0119] In this embodiment, considering that there may be a situation where the level of the recognition policy required by the user is too high, in this case, multiple recognition policies need to be run simultaneously to recognize the image to be processed. When this situation occurs, it is necessary to first consider whether the system supports the energy consumption of running multiple recognition policies simultaneously.
[0120] First, determine whether the energy consumption level of the current system is lower than the preset energy consumption threshold. If the energy consumption level of the current system is not lower than the preset energy consumption threshold, it indicates that the energy consumption level of the current system is relatively high and it is difficult to support the simultaneous operation of multiple recognition policies.
[0121] Since the level of the recognition policy required by the user is very high, the current image to be processed can only be processed by the primary recognition policy supplemented by other recognition policies.
[0122] At this time, determine whether the execution efficiency corresponding to the current execution policy is the maximum value of the primary recognition policy. If the execution efficiency corresponding to the current execution policy is the maximum value of the primary recognition policy, it means that the resolution of the current image to be processed is the lowest level, and two high-processing-capability policies, namely the primary recognition policy and the secondary recognition policy, are required to process the current image to be processed. At this time, on the basis of the full-efficiency execution of the primary recognition policy, the secondary recognition policy is synchronously executed at full efficiency to perform low-resolution object target detection on the image to be processed.
[0123] If the execution efficiency corresponding to the current execution policy is not the maximum value of the primary recognition policy, it means that the resolution of the current image to be processed is not the lowest level. At this time, the primary recognition policy has been called, so there is no need to call the secondary recognition policy to increase the system energy consumption, and only the primary recognition policy can be supplemented.
[0124] It should be noted that the method further includes:
[0125] If the energy consumption level of the current system is lower than the preset energy consumption threshold, then all recognition policies are synchronously executed at full efficiency to perform low-resolution object target detection on the image to be processed.
[0126] If the energy consumption level of the current system is lower than the preset energy consumption threshold, it indicates that the energy consumption level of the system is low and can support the simultaneous operation of multiple recognition strategies. At this time, directly synchronize and execute all recognition strategies with full efficiency to perform low-resolution object target detection on the image to be processed, which can better meet the user's needs.
[0127] In the technical solution of this embodiment, when the user has a need for a customized recognition strategy, a compromise execution strategy can be obtained among meeting the user's needs, reducing the system energy consumption, and improving the processing efficiency.
[0128] Furthermore, the method further includes:
[0129] Assign weights to each recognition strategy;
[0130] Fuse the target detection results of all recognition strategies according to the weights assigned to each recognition strategy, and output a consistent detection result.
[0131] Since in this embodiment, target detection is performed on the image to be processed through multiple recognition strategies, multiple target detection results will be generated.
[0132] In this embodiment, it is necessary to assign weights to each recognition strategy. For example, the first-level recognition strategy is assigned a weight of 0.5, the second-level recognition strategy is assigned a weight of 0.3, and the third-level recognition strategy is assigned a weight of 0.2. Fuse the target detection results of all recognition strategies according to the weights assigned to each recognition strategy, and output a consistent detection result.
[0133] Embodiment Four
[0134] It should be noted that the method further includes:
[0135] Obtain the parameters of the current display system, and determine the display image format according to the parameters of the current display system;
[0136] Convert the target detection image into the display image format;
[0137] Perform image enhancement processing on the target detection image in the display image format through histogram equalization, and filter out noise through Gaussian filtering;
[0138] Determine the resolution according to the parameters of the current display system;
[0139] Adjust the resolution of the processed target detection image according to the resolution of the current display system through bilinear interpolation;
[0140] Output the target detection image with adjusted resolution through the current display system.
[0141] In specific practices, the display system is based on augmented reality technology and adopts advanced array optical waveguide and LCOS projection technology, aiming to present the picture in front of the eyes with high precision while maintaining the user's clear vision of the real scene. As an effective auxiliary terminal, the display system is suitable for diverse scenarios that require high-precision recognition and display.
[0142] In this technical solution, first, image format conversion is performed to convert the internal format into the standard display format (such as RGB) of the current display system to ensure compatibility with the display requirements of the LCOS projection system. Then, enhancement processing such as histogram equalization is used to improve contrast and brightness, and Gaussian filtering is employed for noise filtering, thereby achieving a more delicate and smooth display effect. To adapt to the resolution of the display system, bilinear interpolation is used to scale the image to ensure the best display effect of the picture projected onto the optical waveguide array.
[0143] The display system is responsible for transmitting the processed image data to the augmented reality display device to ensure precise matching with the array optical waveguide and the LCOS projection optical engine. First, the image data is converted into a signal format suitable for LCOS projection (such as HDMI or MIPI) through a signal conversion chip. According to the characteristics of the display device, display parameters such as resolution, brightness, and contrast are set to ensure the clarity and color saturation of the picture. The display driver projects the optimized image content onto the optical waveguide lens to form a virtual picture in the user's field of view, while allowing the user to see through the real world, enabling the user to superimpose virtual information on the real scene, thereby enhancing the perception and understanding of the environment.
[0144] Example Five
[0145] In this example, a specific recognition strategy is described.
[0146] The first-level recognition strategy is applicable to high-performance usage scenarios, such as laboratories, computer rooms, data centers, etc.
[0147] The first-level recognition strategy includes:
[0148] Based on the parallel computing ability and high-bandwidth memory access ability of the GPU, low-resolution object target detection is performed on the image to be processed through the first recognition algorithm;
[0149] The first recognition algorithm includes:
[0150] The self-attention mechanism of the Transformer encoder is used to capture the global dependencies of the image to be processed;
[0151] According to the global dependencies of the image to be processed, noise is gradually added to the image to be processed. After T steps, a completely randomized feature map is obtained;
[0152] Gradually reduce the noise in the completely randomized feature map, and after T steps, restore the image to be processed;
[0153] Extract features from the image to be processed after adding noise and reverse denoising at different scales, and fuse the extracted features; among them, different weights are assigned to each scale;
[0154] Perform low-resolution object detection on the image after feature fusion, and output the object detection image.
[0155] A high-performance GPU is one of the core components of the first-level recognition strategy, with powerful parallel computing capabilities and high-bandwidth memory access capabilities. The high-performance GPU accelerates the computing tasks of the neural network through hardware, especially the forward propagation process of the convolutional neural network (CNN). To further improve the recognition accuracy and processing speed, the high-performance GPU combines a deeply customized neural network accelerator to optimize key computing tasks through the hardware accelerator. The design of the deeply customized neural network accelerator includes multiple computing units and storage units, and realizes the efficient transmission and processing of data through on-chip memory and inter-chip communication mechanisms. To further optimize power consumption and performance, the accelerator adopts a diffusion model based on transformer (TBDiff), which is optimized specifically for low-resolution images.
[0156] The first recognition algorithm improves the expression ability and processing speed of the model by combining the Transformer model and the diffusion model, and is optimized specifically for the image to be processed. Assume the input image to be processed , where H and W represent the height and width of the image to be processed respectively, and C represents the number of channels. The convolution kernel , where k represents the size of the convolution kernel and F represents the number of output channels. The traditional convolution operation can be expressed as:
[0157]
[0158] The TBDiff algorithm improves the expression ability and processing speed of the model by combining the Transformer model and the diffusion model, and is optimized specifically for low-resolution images. Specifically, the TBDiff algorithm is divided into two main steps: Transformer encoding and diffusion model generation.
[0159] 1. Transformer encoding: The Transformer encoder captures the global dependencies of the input feature map through the self-attention mechanism, which is especially suitable for processing the detailed information in low-resolution images. Assume the input feature map X is unfolded into N vectors , where . Each vector . The output of the Transformer encoder can be expressed as:
[0160]
[0161] Specifically, the self-attention mechanism of the Transformer encoder can be expressed as:
[0162]
[0163]
[0164] Among them, , , are learnable weight matrices, is the feature dimension. After the multi-head self-attention mechanism, the output Z can be expressed as:
[0165]
[0166]
[0167] Among them, h is the number of heads, is the final output weight matrix.
[0168] 2. The diffusion model generates high-quality feature maps by gradually adding noise and then denoising reversely, which is particularly suitable for processing detailed information in low-resolution images. Assuming the initial feature map , the diffusion process can be expressed as:
[0169]
[0170] Among them, is the noise addition coefficient, is the random noise of the standard normal distribution. The reverse denoising process can be expressed as:
[0171]
[0172] Among them, is the noise predicted by the neural network. The whole process of the diffusion model can be described as:
[0173] 1. Forward diffusion process:
[0174] Starting from the initial feature map , gradually add noise.
[0175] The diffusion process of each step can be expressed as:
[0176]
[0177] After T steps, a completely randomized feature map is obtained 。
[0178] 2. Backward denoising process:
[0179] Starting from the completely randomized feature map gradually reduce the noise.
[0180] The denoising process at each step can be expressed as:
[0181]
[0182] After T steps, the initial feature map is restored 。
[0183] To further optimize the processing of low-resolution images, the TBDiff algorithm introduces a multi-scale feature fusion mechanism. Specifically, the multi-scale feature fusion mechanism extracts features at different scales and fuses these features to improve the model's recognition ability for low-resolution images. Suppose the multi-scale feature maps are respectively , , , multi-scale feature fusion can be expressed as:
[0184]
[0185] where is the weight of the feature maps at different scales, is to upsample the feature map of the th scale to the same resolution as the initial feature map . In this way, the multi-scale feature fusion mechanism can make full use of information at different scales and improve the model's recognition ability for low-resolution images.
[0186] To further optimize the processing of low-resolution images, the TBDiff algorithm introduces a multi-scale feature fusion mechanism. Specifically, the multi-scale feature fusion mechanism extracts features at different scales and fuses these features to improve the model's recognition ability for low-resolution images. Suppose the multi-scale feature maps are respectively , the multi-scale feature fusion function can be expressed as:
[0187]
[0188] where is the weight at the th scale, which can be obtained through training.
[0189] The final feature map Further processing is performed by a high-performance GPU and a deeply customized neural network accelerator to generate high-precision object detection results.
[0190] The secondary recognition strategy includes:
[0191] An ARM-based multi-core parallel structure that dynamically adjusts the frequency according to different loads;
[0192] Performing low-resolution object target detection on the image to be processed through a second recognition algorithm;
[0193] The second recognition algorithm includes:
[0194] Performing hierarchical multi-scale feature extraction on the image to be processed, and obtaining feature maps of different scales on the image to be processed through convolutional kernels of different scales;
[0195] Feeding the obtained feature maps of different scales into a fusion layer to obtain a fused comprehensive feature map;
[0196] Performing low-resolution object target detection on the fused comprehensive feature map and outputting an object detection image;
[0197] Among them, the fusion process includes: splicing feature maps of different scales and integrating multi-scale features through a convolutional operation;
[0198] Introducing a hierarchical attention mechanism on the basis of the comprehensive feature map to adaptively adjust the feature importance of different regions in the comprehensive feature map;
[0199] Fusing spatial domain and frequency domain features, performing upsampling on the comprehensive feature map through an interpolation method, and expanding the comprehensive feature map to the target resolution.
[0200] The secondary recognition strategy aims to meet the requirements of intelligent devices for efficiently executing object detection tasks under low-power conditions. With the increasing requirements for long-time and high-performance operation in application scenarios such as home security monitoring systems and small drones, there is an urgent need for a module that not only has high-performance computing capabilities but also can strictly control power consumption to achieve long-time stable operation of the device and high-quality object detection. The secondary recognition strategy combines an efficient processor based on the ARM architecture with an optimized neural network accelerator, and at the same time incorporates an adaptive multi-scale super-resolution algorithm (HAMSR) at the algorithm level for efficient reconstruction and object detection of low-resolution images. The HAMSR algorithm ensures high precision and high real-time performance while significantly reducing power consumption when processing low-resolution images through technical means such as hierarchical attention, multi-scale feature extraction, and spatial-frequency domain fusion.
[0201] The hardware part of the secondary recognition strategy is based on an efficient processor with an ARM architecture and a specially designed neural network accelerator. This processor adopts a multi-core parallel structure and can dynamically adjust the frequency according to different loads to achieve dynamic balance between performance and power consumption. By combining the hardware optimization design of the neural network accelerator, the core calculations in the image feature extraction and target detection processes are accelerated. This hardware combination structure greatly improves the overall computing power and provides basic support for the efficient operation of the HAMSR algorithm.
[0202] At the algorithm level, the design focus of the HAMSR algorithm (the second recognition algorithm) lies in multi-scale adaptive feature extraction and spatial frequency domain feature fusion. The HAMSR algorithm first processes the input image to be processed for hierarchical multi-scale feature extraction. Let the image size be , and the number of channels be C . Features are extracted through three different scales. Specifically, three groups of convolution operations with a convolution kernel size of are used to extract feature maps of different scales in order to capture the detailed features in the image. The convolution operation for each scale is defined as:
[0203]
[0204] Among them, and represent the convolution kernel weights and biases respectively, is the convolution kernel size, is the stride, is the padding size, . Through convolution kernels of different scales, the module obtains feature maps with different resolutions on the same input image, ensuring the effective combination of detailed features and global features.
[0205] The multi-scale feature maps are sent to the fusion layer in the HAMSR algorithm to obtain the integrated comprehensive feature map . In order to retain multi-scale information during the fusion process, the HAMSR algorithm first concatenates the feature maps of each scale and then uses a convolution operation to integrate the multi-scale features, so that it maintains a high information density in the next processing step. The fusion operation expression is:
[0206]
[0207] Among them, represents the concatenation operation of the feature maps, and are the convolution kernel weights and biases of the fusion layer, k is the convolution kernel size, s is the stride,p For filling. This operation ensures that the fused feature map contains important information at various scales.
[0208] To further improve the reconstruction accuracy of low-resolution images, the HAMSR algorithm introduces a hierarchical attention mechanism based on the comprehensive feature map \(F_{\text{fuse}}\) to adaptively adjust the feature importance of different regions. The hierarchical attention mechanism generates an attention weight matrix through a combination of a convolution and an activation function A , and applies it to the element-wise multiplication of the comprehensive feature map, thereby highlighting the key regions in the image. The expression of the hierarchical attention mechanism is:
[0209]
[0210]
[0211] where, represents the activation function, and are the convolution kernel weights and biases, is the element-wise multiplication. The attention mechanism endows the module with the ability to adaptively identify the key regions of the image, effectively improving the local fineness of image reconstruction.
[0212] During the feature map fusion process, the features in the spatial domain and the frequency domain are further fused to enhance the expressiveness of image edges and details when processing low-resolution images. The fused feature map is upsampled to the high-resolution image size through an upsampling layer. The upsampling operation is implemented by interpolation methods, such as bilinear interpolation or sub-pixel convolution, to make the image size consistent with the high-resolution target image. The upsampled feature map is expressed as:
[0213]
[0214] where, r is the upsampling ratio, usually taking r = 2, indicating that the image size is doubled. After upsampling, the feature map is passed to the final reconstruction layer to generate the final high-resolution image :
[0215]
[0216] where, and are the convolution kernel weights and biases. This high-resolution image has sufficient details and high-quality resolution, and can provide accurate data support for the target detection module.
[0217] High-resolution image After generation, it will be input into the object detection module to complete object detection and recognition tasks. Through the high-quality images provided by the HAMSR algorithm, high-precision object localization and classification can still be achieved under low-resolution input, thus ensuring the detection accuracy under low-power conditions.
[0218] The three-level recognition strategy includes:
[0219] Perform instruction scheduling based on the multi-stage pipeline of the MIPS instruction set architecture;
[0220] Perform memory access based on the set number of memory accesses and memory access patterns;
[0221] Dynamically adjust the operating frequencies and voltages of the CPU and the neural network accelerator according to the current task load;
[0222] Dynamically screen the feature regions with the required importance in the images to be processed through the third recognition algorithm for recognition;
[0223] The third recognition algorithm includes:
[0224] Perform sparsification processing on the images to be processed;
[0225] Apply a convolution kernel to the sparsified image to perform adaptive sparse convolution operations, and dynamically skip zero-valued features during the convolution process;
[0226] Extract features and classify the convolved image through the pooling layer and the fully connected layer;
[0227] Output the object detection image, including the location, category, and confidence of the object.
[0228] The three-level recognition strategy is designed specifically for extremely low-power usage scenarios, such as long-term industrial monitoring, robots, automatic recognition of unmanned boats, and deep space exploration, etc. The three-level recognition strategy realizes efficient object detection and recognition under low-power conditions by combining a five-stage pipeline low-power CPU based on the MIPS instruction set architecture and a deeply customized general neural network accelerator.
[0229] The hardware architecture of the three-level recognition strategy is highly integrated, mainly including a five-stage pipeline low-power CPU based on the MIPS instruction set architecture and a deeply customized general neural network accelerator. The five-stage pipeline low-power CPU reduces unnecessary energy consumption by optimizing the instruction execution process while maintaining a high computing efficiency. The general neural network accelerator further reduces the power consumption by accelerating the key computing tasks in the neural network inference process through hardware-level optimization.
[0230] To achieve efficient object detection under low-power conditions, a three-level recognition strategy introduces an innovative lightweight neural network algorithm (ASCN) as the third recognition algorithm. The ASCN algorithm reduces redundant calculations in convolution operations through adaptive sparse convolution technology, thus significantly reducing computational complexity and power consumption while maintaining high detection accuracy.
[0231] The core idea of the ASCN algorithm is to dynamically select important feature map regions for calculation during convolution operations, avoiding comprehensive convolution operations on the entire feature map. Specifically, the ASCN algorithm first performs sparsification processing on the image to be processed. The input sparsified image , where H and W represent the height and width of the feature map respectively, and C represents the number of channels. Through a sparsification layer S processes the input feature map to generate a sparse feature map . The specific implementation of the sparsification layer S can be a threshold-based filter, that is:
[0232]
[0233] where, is a preset threshold used to determine which feature values are considered important.
[0234] Applying the convolution kernel on the sparse feature map , where k is the size of the convolution kernel, and F is the number of channels of the output feature map. The adaptive sparse convolution operation can be represented by the following formula:
[0235]
[0236] To further reduce the amount of computation, zero-valued features can be dynamically skipped during convolution, that is:
[0237]
[0238] where, is the set of indices of all non-zero feature values.
[0239] The convolved feature map Y is passed to subsequent pooling layers and fully connected layers for further feature extraction and classification. Finally, the object detection results are output, including the location, category, and confidence of the object.
[0240] Example Six
[0241] A low-resolution object target detection system based on heterogeneous computing, referring to Figure 2 , includes:
[0242] A receiving module 21, configured to receive an image processing request and an image to be processed sent by a user;
[0243] A determination module 22, configured to determine the resolution level of the image to be processed and obtain the energy consumption level of the current system;
[0244] An invocation module 23, configured to invoke an identification module of a corresponding level according to the energy consumption level of the current system;
[0245] An adjustment module 24, configured to, based on the invoked identification module, adjust the execution efficiency of the identification module according to the resolution level of the image to be processed, as an execution module;
[0246] An execution module 25, configured to perform low-resolution object target detection on the image to be processed and output a target detection image;
[0247] Wherein, the levels of the identification module at least include: a first-level identification module 26, a second-level identification module 27, and a third-level identification module 28;
[0248] The identification algorithms adopted by the identification modules of each level are different, and the energy consumption and processing capabilities of the first-level identification module are greater than those of the second-level identification module, which are greater than those of the third-level identification module.
[0249] In this system, first, the energy consumption level of the current system is obtained, and an identification module of a corresponding level is invoked according to the energy consumption level of the current system, so as to ensure that the invoked identification module does not consume a large amount of energy of the current system. Then, according to the resolution level of the image to be processed, the execution efficiency of the identification module is adjusted to further reduce the energy consumption of the system when performing target recognition. And adjusting the execution efficiency of the recognition strategy according to the resolution level of the image to be processed can also ensure the processing efficiency of the image to be processed, so as to achieve the balance of efficiency and energy consumption.
[0250] It can be understood that the same or similar parts in the above embodiments can be referred to each other, and the content not detailed in some embodiments can be referred to the same or similar content in other embodiments.
[0251] It should be noted that in the description of this application, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "a plurality" refers to at least two.
[0252] Any process or method description depicted in the flowchart or described otherwise herein can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations where functions may be executed not in the order shown or discussed, including in a substantially simultaneous manner according to the involved functions or in a reverse order, which should be understood by those skilled in the technical field of the embodiments of the present application.
[0253] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0254] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0255] In addition, in each embodiment of the present application, the functional units can be integrated into one processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0256] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.
[0257] In the description of this specification, the descriptions with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0258] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A low-resolution object detection method based on heterogeneous computing, characterized in that: include: Receive image processing requests and images to be processed sent by users; Determine the resolution level of the image to be processed, and obtain the energy consumption level of the current system; Call the corresponding level identification strategy according to the energy consumption level of the current system; Based on the called recognition strategy, adjusting the execution efficiency of the recognition strategy according to the resolution level of the image to be processed as the execution strategy; Performing low-resolution object detection on the image to be processed by the execution strategy, and outputting an object detection image; The levels of the identification strategy include at least: a primary identification strategy, a secondary identification strategy, and a tertiary identification strategy; The recognition algorithms adopted by the recognition strategies of different levels are different. The energy consumption and processing capability of the first-level recognition strategy are greater than those of the second-level recognition strategy and greater than those of the third-level recognition strategy.
2. The method according to claim 1, characterized in that The method further comprises: Determining whether the image processing request includes a recognition strategy level required by a user; If the image processing request includes a recognition strategy level required by the user, determining whether the recognition strategy level required by the user exceeds the currently called recognition strategy level; If the identification strategy level requested by the user does not exceed the currently called identification strategy level, the current execution strategy is maintained; If the recognition strategy level required by the user exceeds the recognition strategy level currently called, determine whether the execution efficiency corresponding to the current execution strategy exceeds the set efficiency threshold; If the execution efficiency corresponding to the current execution strategy does not exceed the set efficiency threshold level, the execution efficiency of the identification strategy is adjusted to the maximum value, and the execution strategy is updated; If the execution efficiency corresponding to the current execution strategy exceeds the set efficiency threshold, the execution strategy is updated according to the identification strategy level required by the user.
3. The method according to claim 2, characterized in that The method further comprises: If the identification strategy level required by the user exceeds the preset identification strategy level, determining whether the current system energy consumption level is lower than the preset energy consumption threshold; If the energy consumption level of the current system is not lower than the preset energy consumption threshold, determine whether the execution efficiency corresponding to the current execution strategy is the maximum value of the first-level identification strategy; If the execution efficiency corresponding to the current execution strategy is the maximum value of the first-level recognition strategy, then based on the full-efficiency execution of the first-level recognition strategy, the second-level recognition strategy is synchronously executed with full efficiency to perform low-resolution object target detection on the image to be processed; If the current execution strategy is a first-level recognition strategy, but the corresponding execution efficiency is not the maximum value of the first-level recognition strategy, then on the basis of the full-efficiency execution of the first-level recognition strategy, the third-level recognition strategy is synchronously executed with full efficiency to perform low-resolution object target detection on the image to be processed.
4. The method according to claim 3, characterized in that The method further comprises: If the energy consumption level of the current system is lower than the preset energy consumption threshold, all recognition strategies are synchronously executed with full efficiency to perform low-resolution object detection on the image to be processed.
5. The method according to any one of claims 3-4, characterized in that: The method further comprises: Assign weights to each identification strategy; According to the weights assigned to each recognition strategy, the target detection results of all recognition strategies are fused and the consistency detection results are output.
6. The method according to claim 1, characterized in that The method further comprises: Obtaining the parameters of the current display system, and determining the display image format according to the parameters of the current display system; Converting the target detection image into the display image format; The target detection image in the display image format is enhanced by histogram equalization and noise is filtered out by Gaussian filtering; Determine the resolution based on the parameters of the current display system; The resolution of the processed target detection image is adjusted according to the resolution of the current display system by bilinear interpolation; The target detection image with adjusted resolution is output through the current display system.
7. The method according to claim 1, characterized in that The primary identification strategy includes: Based on the parallel computing capability and high-bandwidth memory access capability of the GPU, low-resolution object target detection is performed on the image to be processed by a first recognition algorithm; The first recognition algorithm comprises: Capturing the global dependencies of the image to be processed through the self-attention mechanism of the Transformer encoder; According to the global dependency of the image to be processed, noise is gradually added to the image to be processed. After T steps, a completely randomized feature map is obtained. Step by step, reduce the noise in the completely randomized feature map, and after T steps, restore the image to be processed; Extract features at different scales from the image to be processed after adding noise and reverse denoising, and fuse the extracted features; different weights are assigned to each scale; Perform low-resolution object detection on the image after feature fusion and output the target detection image.
8. The method according to claim 1, characterized in that The secondary identification strategy includes: Based on ARM's multi-core parallel structure, the frequency can be dynamically adjusted according to different loads; Performing low-resolution object detection on the image to be processed by a second recognition algorithm; The second recognition algorithm comprises: Performing hierarchical multi-scale feature extraction on the image to be processed, and obtaining feature maps of different scales on the image to be processed through convolution kernels of different scales; The obtained feature maps of different scales are sent to the fusion layer to obtain the fused comprehensive feature map; Perform low-resolution object detection on the fused comprehensive feature map and output the target detection image; The fusion process includes: splicing feature maps of different scales and integrating multi-scale features through a convolution operation; A hierarchical attention mechanism is introduced based on the comprehensive feature map to adaptively adjust the feature importance of different regions in the comprehensive feature map; The spatial domain and frequency domain features are fused, and the comprehensive feature map is upsampled by interpolation method to expand the comprehensive feature map to the target resolution.
9. The method according to claim 1, characterized in that: The three-level identification strategy includes: Instruction scheduling based on the multi-stage pipeline of the MIPS instruction set architecture; Perform memory access based on a set number of memory accesses and a memory access pattern; Dynamically adjust the operating frequency and voltage of the CPU and neural network accelerator according to the current task load; Dynamically screening feature areas in the image to be processed whose importance meets the requirements for identification by a third identification algorithm; The third recognition algorithm comprises: Performing sparse processing on the image to be processed; Apply convolution kernels to the sparsely processed image to perform adaptive sparse convolution operations, dynamically skipping zero-value features during the convolution process; The convolutional image is passed through the pooling layer and the fully connected layer for feature extraction and classification; Output target detection image, including the location, category and confidence of the target.
10. A low-resolution object detection system based on heterogeneous computing, characterized in that: include: A receiving module, used for receiving an image processing request and an image to be processed sent by a user; A determination module, used to determine the resolution level of the image to be processed and obtain the energy consumption level of the current system; A calling module is used to call the corresponding level of identification module according to the energy consumption level of the current system; An adjustment module, used to adjust the execution efficiency of the recognition module based on the called recognition module according to the resolution level of the image to be processed, as an execution module; An execution module, configured to perform low-resolution object detection on the image to be processed and output an object detection image; The levels of the identification modules include at least: a primary identification module, a secondary identification module and a tertiary identification module; Different levels of recognition modules use different recognition algorithms. The energy consumption and processing capacity of the first-level recognition module are greater than those of the second-level recognition module, which are greater than those of the third-level recognition module.
Citation Information
Patent Citations
Image target detection method and system
CN115457363A
Inference task scheduling method and device and computer equipment
CN116841706A