Image Processing Method and Device for Low-Power MCU Vision Module
By performing image denoising, feature extraction, dimensionality reduction and clustering optimization and image compression on low-power MCU vision modules, the problem of low-power MCU vision modules with low classification accuracy in complex visual tasks is solved, and efficient image data processing is achieved.
Patent Information
- Application Number
- CN202510363149.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-03-26
AI Technical Summary
When dealing with complex visual tasks, low-power MCU vision modules have problems such as low classification accuracy and inability to effectively extract complex features.
By acquiring the original image data collected by the sensor for denoising, using deep learning models for target recognition and feature extraction, combining dimensionality reduction and clustering algorithms to optimize the target area, and reconstructing and compressing through the image compression algorithm to improve the quality and classification accuracy of the image data.
In a low-power environment, the classification accuracy and feature recognition capabilities of image data are significantly improved, data redundancy is reduced, and efficient processing of complex visual tasks is achieved.
Smart Images

Figure CN119888380B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of image data processing, and particularly to an image processing method and device for a low-power MCU vision module. Background Art
[0002] With the rapid development of Internet of Things (IoT) technology and the wide popularity of smart devices, the application demand for low-power MCU (Microcontroller Unit) vision modules in edge computing has increased significantly. These modules play an important role in scenarios such as smart homes, industrial monitoring, and medical devices. In order to achieve efficient processing of image data while meeting the requirements of low power consumption and cost, the research and development of related technologies have become the focus of the industry.
[0003] In related technical means, when an MCU vision module performs image processing, it usually obtains raw image data through a sensor and processes the image using basic image preprocessing techniques (such as filtering, edge detection, and simple threshold segmentation). Subsequently, a lightweight object detection algorithm or a rule-based method is used to identify the target area. These methods can achieve a certain degree of object detection and classification tasks and meet the application requirements of low complexity, such as simple environmental monitoring or the recognition of specific objects.
[0004] For the above technical solutions, although the basic image processing functions of low-power devices can be achieved through these traditional image preprocessing methods and lightweight object detection algorithms, when dealing with complex vision tasks, such as accurately identifying multi-target scenarios or classifying fine-grained target categories, there are problems of low classification accuracy and inability to effectively extract complex features. Summary of the Invention
[0005] In order to improve the problems of low classification accuracy and inability to effectively extract complex features when dealing with complex vision tasks, this application provides an image processing method and device for a low-power MCU vision module.
[0006] The present invention provides an image processing method for a low-power MCU vision module, including: acquiring the original image data collected by a sensor, and performing denoising processing on the original image data to obtain clear image data; performing target recognition based on the clear image data to obtain a target area, inputting the target area into a preset deep learning model to obtain feature data, performing dimensionality reduction processing on the feature data to obtain a dimensionality-reduced feature vector; identifying several category data and category confidence data of the target area according to the dimensionality-reduced feature vector, performing clustering processing on all the category data to generate a clustering result, using the category confidence data to optimize the clustering result to obtain an optimized clustering result, and optimizing the target area according to the optimized clustering result to obtain an optimized target area; using the optimized target area to reconstruct the original image data to obtain processed image data, and inputting the processed image data into an image compression algorithm for calculation to obtain compressed image data.
[0007] As a preferred solution, the step of acquiring the original image data collected by the sensor and performing denoising processing on the original image data to obtain clear image data includes: collecting the original image data output by the sensor and the noise feature data corresponding to the original image data, using a multi-scale filtering method to perform denoising on the original image data and the noise feature data to obtain preliminarily denoised image data and residual noise data; constructing a noise mapping relationship through the residual noise data, and performing adaptive noise reduction processing on the preliminarily denoised image data based on the noise mapping relationship to obtain target-denoised image data; removing artifacts from the target-denoised image data to obtain optimized image data and artifact feature data, calculating the gradient change information of the optimized image data, and performing edge refinement processing on the gradient change information and the artifact feature data to obtain clear image data.
[0008] As a preferred solution, the steps of removing artifacts from the denoised image data of the target to obtain optimized image data and artifact feature data, calculating the gradient change information of the optimized image data, and performing edge refinement processing on the gradient change information and the artifact feature data to obtain clear image data include: performing spectral decomposition on the denoised image data of the target based on a frequency-domain analysis method to obtain component data, extracting artifact feature data based on the component data, using a directional filtering method to suppress artifacts in the component data based on the artifact feature data to obtain artifact-suppressed image data and artifact residual data; calculating the local gradient change information of the artifact-suppressed image data, and adaptively compensating the artifact residual data based on the local gradient change information to obtain optimized image data; applying an interpolation method to reconstruct the edges of the optimized image data to obtain edge-optimized image data and edge error information, calculating the overall texture information of the edge-optimized image data, and performing texture enhancement processing on the edge-optimized image data based on the overall texture information and the edge error information to obtain clear image data.
[0009] As a preferred solution, the steps of performing target recognition on the clear image data to obtain a target area, inputting the target area into a preset deep learning model to obtain feature data, and performing dimensionality reduction processing on the feature data to obtain a dimensionality-reduced feature vector include: performing feature stratification on the clear image data to obtain global feature data and local feature data, applying a domain adaptive image segmentation method to extract regions from the global feature data and the local feature data to obtain multiple candidate target regions and region confidence data; screening the multiple candidate target regions based on the region confidence data using a preset feature fusion strategy to obtain a target area and target feature data; inputting the target area and the target feature data into a preset deep learning model, and performing feature mapping calculation on the target area and the target feature data using a feature mapping method of the Transformer self-attention mechanism in the deep learning model to obtain an original feature vector and spatial feature relationship data; based on the spatial feature relationship data, performing feature association on the original feature vector using a graph analysis method to obtain an associated feature vector and feature importance data, and performing dimensionality reduction calculation on the original feature vector using the associated feature vector and the feature importance data to obtain a dimensionality-reduced feature vector.
[0010] As a preferred solution, the steps of identifying a number of category data and category confidence data of the target area according to the dimension-reduced feature vector, performing clustering processing on all the category data to generate a clustering result, optimizing the clustering result by using the category confidence data, and optimizing the target area according to the optimized clustering result to obtain an optimized target area include: inputting the dimension-reduced feature vector into a preset target area feature model, and in the target area feature model, applying a multi-classification method to perform category identification on the target area based on the dimension-reduced feature vector to obtain a number of category data and category confidence data; performing hierarchical clustering on all the category data to obtain a clustering result and category center data, and optimizing and adjusting the clustering result by using the category confidence data and the category center data to obtain an optimized clustering result; calculating a feature distance between the dimension-reduced feature vector and the optimized clustering result to obtain matching error data, and adaptively correcting the optimized clustering result according to the matching error data to obtain corrected category data; identifying category distribution information of the target area based on the corrected category data, and inputting the category distribution information into the target area feature model for category optimization processing to obtain an optimized target area.
[0011] As a preferred solution, the steps of calculating a feature distance between the dimension-reduced feature vector and the optimized clustering result to obtain matching error data, and adaptively correcting the optimized clustering result according to the matching error data to obtain corrected category data include: calculating the Euclidean distance, cosine similarity, and Mahalanobis distance between the dimension-reduced feature vector and the optimized clustering result, calculating matching error data based on the Euclidean distance and the Mahalanobis distance, and generating feature similarity data by using the cosine similarity; applying a dynamic weight adjustment method to normalize the matching error data based on the feature similarity data to obtain normalized matching error data and feature contribution degree data, and performing error compensation calculation on the normalized matching error data by using the feature contribution degree data to obtain optimized matching error data and error distribution information; adaptively adjusting the optimized matching error data by using the error distribution information to obtain an error correction parameter, and performing category center migration adjustment on the optimized clustering result according to the error correction parameter to obtain corrected category data.
[0012] As a preferred solution, the step of using the optimized target region to reconstruct the original image data to obtain processed image data and inputting the processed image data into an image compression algorithm for calculation to obtain compressed image data includes: comparing the optimized target region with other regions in the original image data to obtain a comparison result, and reconstructing the target region in the original image data based on the comparison result through an image reconstruction algorithm to obtain a reconstructed target region; fusing the reconstructed target region with other regions in the original image data to obtain fused image data, preprocessing the fused image data to obtain processed image data, and compressing the processed image data using an image compression algorithm to obtain compressed image data.
[0013] This application also provides an image processing device for a low-power MCU vision module, including: an acquisition module, configured to acquire the original image data collected by a sensor and perform denoising processing on the original image data to obtain clear image data; a dimensionality reduction module, configured to perform target recognition based on the clear image data to obtain a target region, input the target region into a preset deep learning model to obtain feature data, and perform dimensionality reduction processing on the feature data to obtain a dimensionality-reduced feature vector; a clustering module, configured to identify several category data and category confidence data of the target region according to the dimensionality-reduced feature vector, perform clustering processing on all the category data to generate a clustering result, optimize the clustering result using the category confidence data to obtain an optimized clustering result, and optimize the target region according to the optimized clustering result to obtain an optimized target region; a compression module, configured to use the optimized target region to reconstruct the original image data to obtain processed image data, and input the processed image data into an image compression algorithm for calculation to obtain compressed image data.
[0014] Compared with the prior art, this application has the following beneficial effects: high classification accuracy and strong recognition features. The efficient removal of noise from the original image data is achieved by using median filtering and adaptive filtering, ensuring the quality of clear image data; the feature data of the target region is extracted by a lightweight deep learning model, and the subsequent calculation complexity is simplified by combining dimensionality reduction processing; the classification accuracy of the target region and the matching degree of the target boundary are significantly improved through clustering algorithms and confidence optimization; the efficient processing of complex image data in a low-power environment is achieved by optimizing the reconstruction and compression of the target region, effectively solving the problems of low classification accuracy and data redundancy in the low-power MCU vision module, and improving the problems of low classification accuracy and inability to effectively extract complex features when processing complex vision tasks. Description of the Drawings
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0016] The structures, ratios, sizes, etc. shown in the accompanying drawings of this specification are only used to cooperate with the content disclosed in the specification for those who are familiar with this technology to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have technical substantive significance. Any modification of the structure, change of the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope that can be covered by the technical content disclosed by the present invention.
[0017] Figure 1 It is a schematic flowchart of the image processing method of the low-power MCU vision module provided by the embodiment of the present invention;
[0018] Figure 2 It is a schematic block diagram of the structure of the image processing device of the low-power MCU vision module provided by the embodiment of the present invention.
[0019] Explanation of reference numerals:
[0020] 10. Image processing device of the low-power MCU vision module; 11. Acquisition module; 12. Dimensionality reduction module; 13. Clustering module; 14. Compression module. Specific embodiments
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present invention.
[0022] The flowchart shown in the accompanying drawings is only an example, and does not necessarily include all contents and operations / steps, nor does it necessarily execute in the described order. For example, some operations / steps can be decomposed, combined, or partially merged, so the actual execution order may be changed according to the actual situation.
[0023] It should also be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0024] It should be further understood that the term "and / or" used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0025] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and through specific embodiments.
[0026] Embodiment 1:
[0027] As Figure 1 shown, this application provides an image processing method for a low-power MCU vision module, including steps S100 to S400.
[0028] Step S100: Obtain the original image data collected by the sensor, and perform denoising processing on the original image data to obtain clear image data.
[0029] In this step, after obtaining the original image data through the sensor, use image denoising algorithms such as median filtering, mean filtering, and adaptive filtering to suppress the noise signals in the original image data, eliminate high-frequency noise and irregular interference signals, and obtain clear image data with a clean background and clear edges; specifically, combined with the limitations of hardware resources, select a low-complexity filtering method according to the noise characteristics of the original image data, and at the same time use multi-window filtering technology to further improve the detail expressiveness of the image.
[0030] For example, in a monitoring device in an industrial environment, through denoising processing, high-frequency noise caused by vibration or light changes can be effectively filtered out, ensuring the quality of the original image data, thereby providing a better input basis for subsequent processing.
[0031] Step S200: Perform target recognition based on the clear image data to obtain a target area, input the target area into a preset deep learning model to obtain feature data, and perform dimensionality reduction processing on the feature data to obtain a dimensionality-reduced feature vector.
[0032] In this step, the target region in the clear image data is accurately located by using an object recognition algorithm based on edge detection and region growing, and the deep learning model (such as convolutional neural network) is used to extract features from the target region. Specifically, a lightweight network structure (such as MobileNet) is used in the deep learning model to adapt to the computing power limitation of the low-power MCU, and the principal component analysis (PCA) method is used to reduce the dimension of the extracted feature data, remove redundant features, and obtain a more discriminative reduced-dimensional feature vector.
[0033] For example, in the application of smart cameras in smart homes, through the deep learning model, key feature points in the face region can be accurately extracted, and the data processing efficiency and device response speed can be improved through dimensionality reduction processing.
[0034] Step S300: Identify several category data and category confidence data of the target region according to the reduced-dimensional feature vector, perform clustering processing on all category data to generate a clustering result, optimize the clustering result by using the category confidence data to obtain an optimized clustering result, and optimize the target region according to the optimized clustering result to obtain an optimized target region.
[0035] In this step, the reduced-dimensional feature vector is input into a clustering algorithm (such as the K-Means algorithm) for category clustering processing, and the clustering result is weighted and optimized with confidence according to the category confidence data to make the clustering result more accurate. Specifically, the optimized clustering result is used to further adjust the boundary of the target region to ensure a higher matching degree between the optimized target region and the actual target.
[0036] For example, in the microscopic image analysis of medical devices, the cell target region can be located through the clustering result of category data, and its boundary can be optimized to more accurately identify specific cell categories.
[0037] Step S400: Use the optimized target region to reconstruct the original image data to obtain processed image data, and input the processed image data into an image compression algorithm for calculation to obtain compressed image data.
[0038] In this step, the local features in the original image data are enhanced and reconstructed by using the optimized target region, and at the same time, image compression algorithms such as discrete cosine transform (DCT) or discrete wavelet transform (DWT) are used to efficiently compress the processed image data to obtain compressed image data. Specifically, by dynamically adjusting the compression coefficient, the compression rate is maximized and the key information of the image is retained in a low-power environment.
[0039] For example, when transmitting surveillance images on edge devices, the compressed image data can significantly reduce bandwidth occupancy while maintaining the clarity of the target area to meet the requirements of real-time transmission.
[0040] In this embodiment, by acquiring the original image data collected by the sensor and performing denoising processing on the original image data, clear image data is obtained. Based on the clear image data, target recognition is performed, the target area is extracted, and the target area is input into a preset deep learning model to generate feature data. Subsequently, dimensionality reduction processing is performed on the feature data to obtain the dimensionality-reduced feature vector. According to the dimensionality-reduced feature vector, several category data and category confidence data of the target area are identified, and clustering processing is performed on all category data to generate a clustering result. Then, by using the category confidence data to optimize the clustering result, an optimized clustering result is obtained, and the target area is optimized by using the optimized clustering result to obtain an optimized target area. Finally, the original image data is reconstructed by using the optimized target area to generate processed image data, and the processed image data is input into an image compression algorithm for calculation to obtain compressed image data.
[0041] The denoising processing improves the quality of the original image data, laying a foundation for subsequent processing; the deep learning model combined with dimensionality reduction processing reduces the complexity of the feature data while retaining key information, improving the efficiency and accuracy of target recognition; the clustering processing and optimization steps further enhance the classification accuracy and reduce the interference of redundant data; finally, through the reconstruction and compression steps, not only can the data volume during storage and transmission be significantly reduced, but the main features of the processed image data are also retained, thus realizing the efficient processing of complex visual tasks with limited resources and power consumption, and improving the problems of low classification accuracy and inability to effectively extract complex features when dealing with complex visual tasks.
[0042] Embodiment 2:
[0043] In step S100, the original image data output by the sensor and the noise feature data corresponding to the original image data are collected, and a multi-scale filtering method is used to denoise the original image data and the noise feature data to obtain the preliminarily denoised image data and residual noise data.
[0044] Through the multi-scale filtering method, combined with the spectral characteristics of the original image data and the noise feature data output by the sensor, adaptive filters in multiple frequency bands are designed to perform targeted processing on low-frequency, medium-frequency, and high-frequency signals respectively; specifically, a low-pass filter is used to filter out low-frequency noise, a band-pass filter is used to suppress interference in the medium-frequency signal, and a high-pass filter is used to weaken the sharp noise in the high-frequency signal, realizing the layer-by-layer decomposition and reconstruction of the original image data to maximize the retention of image detail information and suppress noise interference.
[0045] For example, in the industrial equipment monitoring scenario, the periodic noise caused by mechanical vibration in the original image data can be effectively removed by the multi-scale filtering method, while retaining the clear texture details on the surface of gears or bearings, providing high-quality data for subsequent image analysis.
[0046] Construct a noise mapping relationship through the residual noise data, and perform adaptive noise reduction processing on the preliminarily denoised image data based on the noise mapping relationship to obtain the target denoised image data.
[0047] By analyzing the spatial distribution characteristics of the residual noise data, construct a noise mapping matrix, and dynamically adjust the denoising parameters based on this mapping matrix to make the noise reduction processing more in line with the actual needs of different regions of the image; specifically, apply a stronger filtering strategy to the high-noise regions and a weaker noise reduction operation to the low-noise regions to ensure that image details are retained while reducing noise.
[0048] For example, in the analysis of surgical microscope images, more details can be retained for the low-noise data in the blood vessel region, while stronger noise reduction is performed on the high-noise in the background region, ultimately improving the overall clarity of the image and the reliability of diagnosis.
[0049] Remove artifacts from the target denoised image data to obtain the optimized image data and artifact feature data, calculate the gradient change information of the optimized image data, and perform edge refinement processing on the gradient change information and the artifact feature data to obtain clear image data.
[0050] By performing frequency domain analysis on the target denoised image data, separate the high-frequency components that generate artifacts and extract the artifact feature data, and use a directional filtering algorithm to specifically process the artifact components to eliminate artifacts at the image edges and complex textures; specifically, use the component decomposition method to locally separate the artifacts, then perform multi-scale suppression on the residual artifact data, and combine the gradient information to enhance the edge refinement processing to ensure the continuity and realism of the image edges.
[0051] For example, in the processing of remote sensing satellite images, by eliminating artifacts at the cloud or ground object boundaries, the optimized images can be used for precise geomorphic analysis and resource assessment.
[0052] Among them, artifact removal is performed on the target denoised image data to obtain optimized image data and artifact feature data. The steps of calculating the gradient change information of the optimized image data and performing edge refinement processing on the gradient change information and the artifact feature data to obtain clear image data include: performing spectral decomposition on the target denoised image data based on a frequency domain analysis method to obtain component data, and extracting artifact feature data based on the component data. Using a directional filtering method, artifact suppression is performed on the component data based on the artifact feature data to obtain artifact-suppressed image data and artifact residual data.
[0053] By performing FFT (Fast Fourier Transform) on the target denoised image data, the artifact components in the spectrum are separated, and the artifact distribution is located using an artifact feature extraction algorithm; specifically, the filtering direction is adjusted in combination with the artifact feature data to dynamically eliminate high-frequency artifact residuals, ensuring effective suppression of artifacts without damaging image details.
[0054] For example, in the processing of medical images (such as X-ray films), removing artifacts can significantly improve the clarity of tissue boundaries, thereby improving the diagnostic accuracy of doctors for lesions.
[0055] Calculate the local gradient change information of the artifact-suppressed image data, and perform adaptive compensation on the artifact residual data based on the local gradient change information to obtain optimized image data.
[0056] By calculating the local gradient change information of the artifact-suppressed image data, the gradient amplitude and direction features in the image edge region are extracted to construct a gradient change mapping model. Using this model, analyze the distribution of the artifact residual data in terms of local gradient characteristics, and adjust the compensation parameters through adaptive weight allocation to perform directional suppression and intensity optimization processing on the artifact residual data, and finally generate optimized image data; specifically, in high-gradient change regions, an enhanced filtering algorithm is used to strengthen the retention of edge features, and at the same time, in low-gradient change regions, the interference degree of artifact residuals is reduced to ensure the improvement of the overall image quality.
[0057] For example, in medical images such as MRI scans, by accurately compensating for the artifacts at the tumor edge, the clarity of the edge contour can be improved, thereby providing a more reliable basis for clinical diagnosis.
[0058] Apply an interpolation method to perform edge reconstruction on the optimized image data to obtain edge-optimized image data and edge error information, calculate the overall texture information of the edge-optimized image data, and perform texture enhancement processing on the edge-optimized image data based on the overall texture information and the edge error information to obtain clear image data.
[0059] By applying interpolation methods (such as bicubic interpolation or adaptive interpolation), the pixel values in the edge region of the optimized image data are reconstructed to fill the discontinuous or blurred edge regions, generating edge-optimized image data, and recording the error information during the interpolation process as edge error information; specifically, texture features are extracted by combining the overall texture analysis algorithm (such as local binary pattern LBP), and the performance of the edge-optimized image data in local details and overall smoothness is improved through texture enhancement technology to ensure clear image data with high visual quality.
[0060] For example, in intelligent security, by performing edge reconstruction and texture enhancement on the images collected in low-light environments, the clarity and contrast of the face region can be effectively improved, thereby enhancing the accuracy of target recognition.
[0061] In step S200, the clear image data is subjected to feature stratification to obtain global feature data and local feature data. The domain adaptive image segmentation method is applied to extract regions from the global feature data and local feature data, obtaining multiple candidate target regions and region confidence data.
[0062] Through hierarchical analysis of the clear image data, the overall geometric features of the image are extracted as global feature data, while the detailed texture and shape characteristics are extracted as local feature data. Using the domain adaptive image segmentation method, adaptive segmentation strategies are applied to the global feature data and local feature data respectively. Combining edge detection and region growing algorithms, multiple candidate target regions in the image are automatically extracted, and region confidence data is calculated for each region to measure its target relevance; specifically, by dynamically adjusting the segmentation parameters, the segmentation results are adapted to the image characteristics in different scenarios.
[0063] For example, in the autonomous driving scenario, through feature stratification and domain adaptive segmentation, candidate target regions of road boundaries, vehicles, and pedestrians can be effectively extracted, providing highly reliable data support for perception tasks in complex environments.
[0064] Based on the region confidence data, a preset feature fusion strategy is used to screen the multiple candidate target regions, obtaining the target region and target feature data.
[0065] By analyzing the region confidence data of multiple candidate target regions, a preset feature fusion strategy is adopted to comprehensively consider the geometric shape, texture features, and position distribution of the target region to screen the target region; specifically, using a multi-layer perception mechanism, the region with the highest confidence is selected as the target region, and at the same time, key features are further extracted for the target region to generate target feature data.
[0066] For example, in drone monitoring, by using a fusion strategy to screen target regions of interest (such as a specific building or vehicle), the efficiency of target recognition can be improved and key inputs for subsequent in-depth analysis can be provided.
[0067] Input the target region and target feature data into a preset deep learning model. In the deep learning model, apply the feature mapping method of the Transformer self-attention mechanism to perform feature mapping calculations on the target region and target feature data, obtaining the original feature vector and spatial feature relationship data.
[0068] By inputting the target region and target feature data into a preset deep learning model (such as a Transformer-based structure), using the dynamic weight allocation method of the self-attention mechanism, perform feature mapping and relationship modeling on each pixel point of the target region to generate the original feature vector. At the same time, further capture the interaction relationship between internal features of the target region through spatial feature relationship modeling to generate spatial feature relationship data; specifically, this process reduces the model calculation complexity while ensuring feature accuracy.
[0069] For example, in high-resolution remote sensing images, through feature mapping and relationship modeling, the geometric morphology and internal characteristics of farmland or building areas can be accurately captured, providing support for precise map drawing.
[0070] Based on the spatial feature relationship data, adopt a graph spectral analysis method to perform feature association on the original feature vector, obtaining the associated feature vector and feature importance data. Use the associated feature vector and feature importance data to perform dimensionality reduction calculation on the original feature vector, obtaining the dimensionality-reduced feature vector.
[0071] By constructing a feature graph based on the spatial feature relationship data, taking the original feature vector as the graph node, calculate feature association through the relationship weight value between nodes, obtaining the associated feature vector and feature importance data; specifically, combine a dimensionality reduction algorithm (such as principal component analysis or graph embedding method) to perform dimensionality reduction calculation on the feature graph, ensuring that the dimensionality-reduced feature vector reduces data redundancy while retaining the main features.
[0072] For example, in low-power intelligent camera devices, through feature dimensionality reduction processing, the computing burden of the device can be reduced, and at the same time, the detection efficiency of target features in dynamic scenes can be improved.
[0073] In step S300, input the dimensionality-reduced feature vector into a preset target region feature model. In the target region feature model, apply a multi-classification method to perform class recognition on the target region based on the dimensionality-reduced feature vector, obtaining several items of class data and class confidence data.
[0074] By inputting the feature vectors after dimensionality reduction into the target region feature model, this model performs multi-classification processing on the target region based on a neural network structure (such as a Softmax classifier). Specifically, by constructing a multi-layer classification network, the feature vectors after dimensionality reduction are subjected to layer-by-layer feature mapping and classification scoring, and finally several items of category data and corresponding category confidence data are output. The category confidence data represents the credibility of the classification result.
[0075] For example, in an intelligent transportation system, by classifying the feature vectors of a target vehicle, the type of the vehicle (such as a small car, a large car, a motorcycle, etc.) and the confidence of this classification can be identified, providing high-precision support for traffic flow analysis.
[0076] Hierarchical clustering is performed on all the category data to obtain the clustering result and the category center data. The clustering result is optimized and adjusted using the category confidence data and the category center data to obtain the optimized clustering result.
[0077] By performing hierarchical clustering on several items of category data, a bottom-up clustering method is used to gradually merge categories to generate a clustering tree structure, and the clustering result is optimized according to the Euclidean distance of the category center data and the category confidence data, adjusting the accuracy of the clustering level and category division. Specifically, combining the geometric distribution characteristics of the category center and the weight of the category confidence data, the demarcation point of the clustering result is finely adjusted to improve the rationality and consistency of the clustering.
[0078] For example, in biological microscopic image processing, through clustering analysis and optimization, multiple similar cell categories can be integrated to generate a clustering result that is more in line with the actual biological significance for subsequent pathological analysis.
[0079] Calculate the feature distance between the feature vectors after dimensionality reduction and the optimized clustering result to obtain the matching error data, and adaptively correct the optimized clustering result according to the matching error data to obtain the corrected category data.
[0080] Calculate the Euclidean distance, cosine similarity, and Mahalanobis distance of the feature distance between the feature vectors after dimensionality reduction and the optimized clustering result, comprehensively analyze their matching degree, generate the matching error data, and adjust the category attribution relationship in the optimized clustering result based on the matching error data. Specifically, use the feature distance to dynamically adjust the position of the category center and correct the category boundary to achieve the optimal performance of the corrected category data in terms of accuracy and consistency.
[0081] For example, in an autonomous driving scenario, by performing feature matching and correction on dynamic pedestrian targets, the feature categories of different pedestrians can be accurately distinguished to ensure the recognition ability of the system in a high-density environment.
[0082] Based on the corrected category data, the category distribution information of the target area is identified, and the category distribution information is input into the target area feature model for category optimization processing to obtain the optimized target area.
[0083] By analyzing the corrected category data, a category distribution probability map of the target area is generated and input into the target area feature model for category optimization processing. The model further improves the accuracy of category division through convolution operations and iterative learning, and refines the category attribution of the boundary area to generate the optimized target area. Specifically, the parameters of the category recognition model are dynamically optimized by combining the category distribution information and the feature space information, so that the optimized target area is closer to the visual target distribution of the actual scene.
[0084] For example, in agricultural precision monitoring, through the category optimization of the field crop area, the types and distribution ranges of different crops can be accurately distinguished, providing support for crop management and planning.
[0085] Among them, the steps of calculating the feature distance between the dimension-reduced feature vector and the optimized clustering result to obtain the matching error data and adaptively correcting the optimized clustering result according to the matching error data to obtain the corrected category data include: calculating the Euclidean distance, cosine similarity, and Mahalanobis distance between the dimension-reduced feature vector and the optimized clustering result, calculating the matching error data based on the Euclidean distance and Mahalanobis distance, and generating feature similarity data using the cosine similarity.
[0086] By comprehensively using the Euclidean distance, cosine similarity, and Mahalanobis distance to describe the relationship between the dimension-reduced feature vector and the optimized clustering result, the matching error data is generated by combining the geometric distance (Euclidean distance), direction similarity (cosine similarity), and distribution consistency (Mahalanobis distance) of the features. At the same time, the feature similarity data is generated based on the cosine similarity to comprehensively measure the reliability and strength of the matching.
[0087] For example, in medical imaging, by calculating the matching error and similarity data between the tumor area and the features of the standard model, more detailed error analysis can be provided for determining potential problem areas in image diagnosis.
[0088] Apply the dynamic weight adjustment method to normalize the matching error data based on the feature similarity data to obtain the normalized matching error data and feature contribution data, and use the feature contribution data to perform error compensation calculation on the normalized matching error data to obtain the optimized matching error data and error distribution information.
[0089] Through the dynamic weight adjustment method, the matching error data is processed by multi-factor normalization according to the feature similarity data to generate standardized normalized matching error data, and the error of different feature dimensions is compensated by using the feature contribution degree data to ensure that the error proportion of high-contribution features is smaller. At the same time, the error distribution information is output for guiding subsequent category adjustment.
[0090] For example, in space exploration tasks, by adjusting the feature attribution of the celestial body target through the error distribution information, the category and position features of unknown celestial bodies can be accurately identified.
[0091] The error distribution information is used to adaptively adjust the optimized matching error data to obtain an error correction parameter, and the category center of the optimized clustering result is adjusted by the error correction parameter to obtain the corrected category data.
[0092] By dynamically calculating the error correction parameter using the error distribution information, combining the error correction parameter to adjust the category center in the optimized clustering result, redrawing the category boundary, and correcting the feature distribution within the category, the obtained corrected category data is made more representative and consistent. Specifically, increase the correction intensity in the high-error area and retain the original distribution in the low-error area to ensure the overall correction effect.
[0093] For example, in low-power edge computing devices, by adjusting the category boundary of the target area through the error correction parameter, efficient target recognition and classification in complex scenarios can be achieved.
[0094] In step S400, the optimized target area is compared with other areas in the original image data to obtain a comparison result, and the target area in the original image data is reconstructed based on the comparison result through an image reconstruction algorithm to obtain the reconstructed target area.
[0095] By performing pixel-level comparison between the optimized target area and other areas in the original image data, calculating the similarity difference between the two in terms of light intensity, texture, and edge features, a comparison result is generated. According to the comparison result, an image reconstruction algorithm based on sparse representation is used to accurately reconstruct the target area in the original image data; specifically, combining feature point alignment technology and pixel interpolation method to enhance the detail characteristics of the target area while ensuring the natural transition with the surrounding background area to obtain the reconstructed target area.
[0096] For example, in industrial defect detection, by comparing and analyzing the optimized defect area on the metal plate surface with the original image data, the smooth and defect-free metal surface area can be effectively reconstructed for subsequent quality assessment.
[0097] Fuse the reconstructed target region with other regions in the original image data to obtain the fused image data. Preprocess the fused image data to obtain the processed image data. Use an image compression algorithm to compress the processed image data to obtain the compressed image data.
[0098] By fusing the reconstructed target region with other regions in the original image data, a multi-resolution fusion method is adopted to dynamically match each part of the image to eliminate edge stitching traces. At the same time, the visual consistency of the fusion region is optimized through a color balance algorithm to generate the fused image data. Subsequently, preprocess the fused image data, including quantization, redundancy removal, and encoding processing, to ensure the compactness and high fidelity of the processed image data during storage and transmission. Finally, further compress the processed image data through an image compression algorithm (such as JPEG compression or a deep learning-based compression algorithm) to obtain the compressed image data.
[0099] For example, in an intelligent monitoring scenario, by fusing and compressing the moving target regions detected in a dynamic scene, high-definition but highly compressed video frame data can be generated to adapt to the real-time monitoring requirements with limited bandwidth.
[0100] In this embodiment, by collecting the original image data and noise feature data output by the sensor, denoising and adaptive noise reduction processing are performed using a multi-scale filtering method combined with the noise mapping relationship, effectively improving the quality of the image data. And the artifacts and edge regions are optimized and reconstructed through artifact removal, edge refinement, and interpolation methods to finally obtain clear image data. At the same time, during the image feature extraction and region analysis process, a feature hierarchical and domain adaptive image segmentation method is adopted to combine the global feature data and local feature data, extract multiple candidate target regions and region confidence data, screen out the target region and target feature data through a feature fusion strategy, and use the Transformer self-attention mechanism and graph analysis method in the deep learning model to complete feature mapping, feature association, and dimensionality reduction calculation to generate the dimensionality-reduced feature vector. For the classification and optimization of the target region, a multi-classification method and a hierarchical clustering strategy are applied. Through the optimization adjustment of the category data, adaptive correction, and the optimization processing of the category distribution information, the optimized target region is finally obtained. In the data fusion and reconstruction link, by reconstructing the target region and fusing it with other regions, high-quality image compression is achieved in combination with an image compression algorithm, significantly improving the efficiency of data storage and transmission. The entire solution, through a multi-level and multi-step processing method, not only significantly improves the quality and classification accuracy of the image data but also shows outstanding advantages in optimizing efficiency and reducing power consumption, providing an efficient solution for the image processing of low-power MCU vision modules.
[0101] Embodiment 3:
[0102] As shown Figure 2 in the figure, the present application also provides an image processing apparatus 10 for a low-power MCU vision module, including an acquisition module 11, a dimensionality reduction module 12, a clustering module 13, and a compression module 14.
[0103] The acquisition module 11 is mainly used to acquire the original image data collected by the sensor, and perform denoising processing on the original image data to obtain clear image data.
[0104] The dimensionality reduction module 12 is mainly used to perform target recognition based on the clear image data to obtain a target area, input the target area into a preset deep learning model to obtain feature data, and perform dimensionality reduction processing on the feature data to obtain a dimensionality-reduced feature vector.
[0105] The clustering module 13 is mainly used to identify several category data and category confidence data of the target area according to the dimensionality-reduced feature vector, perform clustering processing on all the category data to generate a clustering result, optimize the clustering result by using the category confidence data to obtain an optimized clustering result, and optimize the target area according to the optimized clustering result to obtain an optimized target area.
[0106] The compression module 14 is mainly used to reconstruct the original image data by using the optimized target area to obtain processed image data, and input the processed image data into an image compression algorithm for calculation to obtain compressed image data.
[0107] In this embodiment, the acquisition module 11 accurately processes the original image data collected by the sensor, uses multi-scale denoising technology to eliminate the noise in the data, and retains the detailed characteristics of the image, thereby generating clear image data, laying a foundation for subsequent processing; the dimensionality reduction module 12 completes the recognition of the target area and the feature extraction of the deep learning model based on the clear image data, combines the dimensionality reduction algorithm to remove redundant feature data, and finally generates a dimensionality-reduced feature vector, which not only improves the efficiency of data processing, but also retains the key characteristics of the target; the clustering module 13 performs multi-dimensional analysis on the dimensionality-reduced feature vector, generates category data and performs clustering optimization, and further optimizes the characteristic performance of the target area by dynamically adjusting the matching relationship between the clustering result and the category confidence data to obtain an accurate optimized target area; finally, the compression module 14 reconstructs the original image data by combining the optimized characteristics of the target area, and compresses the processed image data by using an efficient image compression algorithm to reduce the burden of data storage and transmission, while retaining the key image information.
[0108] Through modular design, an efficient and precise image processing process is achieved, significantly enhancing the image processing ability and power consumption efficiency of the low-power MCU vision module in complex scenarios, providing reliable technical support for vision tasks in edge computing.
[0109] It should be noted that those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described device and each module can refer to the corresponding processes in the foregoing Embodiment 1 and will not be elaborated herein.
[0110] The structures, proportions, sizes, etc. shown in the drawings of this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limiting conditions under which the present invention can be implemented. Therefore, they do not have a substantial technical meaning. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope that can be covered by the technical content disclosed in the present invention.
[0111] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. An image processing method for a low-power MCU vision module, characterized in that, Including: Collecting the original image data output by the sensor and the noise feature data corresponding to the original image data, and denoising the original image data and the noise feature data by using a multi-scale filtering method to obtain the preliminarily denoised image data and the residual noise data; Constructing a noise mapping relationship through the residual noise data, and performing adaptive noise reduction processing on the preliminarily denoised image data based on the noise mapping relationship to obtain the target denoised image data; Performing spectral decomposition on the target denoised image data by a method based on frequency domain analysis to obtain component data, extracting artifact feature data based on the component data, and suppressing artifacts of the component data by using a directional filtering method based on the artifact feature data to obtain the artifact-suppressed image data and the artifact residual data; calculating the local gradient change information of the artifact-suppressed image data, and adaptively compensating the artifact residual data based on the local gradient change information to obtain the optimized image data; applying an interpolation method to perform edge reconstruction on the optimized image data to obtain the edge-optimized image data and the edge error information, calculating the overall texture information of the edge-optimized image data, and performing texture enhancement processing on the edge-optimized image data based on the overall texture information and the edge error information to obtain the clear image data; Performing target recognition based on the clear image data to obtain a target area, inputting the target area into a preset deep learning model to obtain feature data, and performing dimensionality reduction processing on the feature data to obtain a dimensionality-reduced feature vector; Identifying a number of category data and category confidence data of the target area according to the dimensionality-reduced feature vector, performing clustering processing on all the category data to generate a clustering result, optimizing the clustering result by using the category confidence data to obtain an optimized clustering result, and optimizing the target area according to the optimized clustering result to obtain an optimized target area; Reconstructing the original image data by using the optimized target area to obtain the processed image data, and inputting the processed image data into an image compression algorithm for calculation to obtain the compressed image data.
2. The image processing method of the low-power MCU vision module according to claim 1, wherein The step of performing target recognition based on the clear image data to obtain a target area, inputting the target area into a preset deep learning model to obtain feature data, and performing dimensionality reduction processing on the feature data to obtain a dimensionality-reduced feature vector includes: Performing feature stratification on the clear image data to obtain global feature data and local feature data, and applying a domain adaptive image segmentation method to perform region extraction on the global feature data and the local feature data to obtain a plurality of candidate target regions and region confidence data; Screening the plurality of candidate target regions based on the region confidence data by using a preset feature fusion strategy to obtain a target region and target feature data; Input the target region and the target feature data into a preset deep learning model, and apply the feature mapping method of the Transformer self-attention mechanism in the deep learning model to perform feature mapping calculations on the target region and the target feature data to obtain an original feature vector and spatial feature relationship data; Based on the spatial feature relationship data, use a graph analysis method to perform feature association on the original feature vector to obtain an associated feature vector and feature importance data, and use the associated feature vector and feature importance data to perform dimensionality reduction calculations on the original feature vector to obtain a dimensionality-reduced feature vector.
3. The image processing method of the low-power MCU vision module according to claim 1, wherein, The steps of identifying several category data and category confidence data of the target region according to the dimensionality-reduced feature vector, performing clustering processing on all the category data to generate a clustering result, using the category confidence data to optimize the clustering result to obtain an optimized clustering result, and optimizing the target region according to the optimized clustering result to obtain an optimized target region include: Input the dimensionality-reduced feature vector into a preset target region feature model. In the target region feature model, apply a multi-classification method to perform category recognition on the target region based on the dimensionality-reduced feature vector to obtain several category data and category confidence data; Perform hierarchical clustering on all the category data to obtain a clustering result and category center data, and use the category confidence data and the category center data to optimize and adjust the clustering result to obtain an optimized clustering result; Calculate the feature distance between the dimensionality-reduced feature vector and the optimized clustering result to obtain matching error data, and perform adaptive correction on the optimized clustering result according to the matching error data to obtain corrected category data; Based on the corrected category data, identify the category distribution information of the target region, and input the category distribution information into the target region feature model for category optimization processing to obtain an optimized target region.
4. The image processing method of the low-power MCU vision module according to claim 3, wherein The steps of calculating the feature distance between the dimensionality-reduced feature vector and the optimized clustering result to obtain matching error data, and performing adaptive correction on the optimized clustering result according to the matching error data to obtain corrected category data include: Calculate the Euclidean distance, cosine similarity, and Mahalanobis distance between the dimensionality-reduced feature vector and the optimized clustering result, calculate the matching error data based on the Euclidean distance and the Mahalanobis distance, and generate feature similarity data using the cosine similarity; Apply a dynamic weight adjustment method to normalize the matching error data based on the feature similarity data to obtain normalized matching error data and feature contribution data, and use the feature contribution data to perform error compensation calculations on the normalized matching error data to obtain optimized matching error data and error distribution information; Use the error distribution information to adaptively adjust the optimized matching error data to obtain an error correction parameter, and perform category center migration adjustment on the optimized clustering result according to the error correction parameter to obtain corrected category data.
5. The image processing method of the low-power MCU vision module according to claim 1, wherein The step of reconstructing the original image data by using the optimized target region to obtain the processed image data, and inputting the processed image data into an image compression algorithm for calculation to obtain the compressed image data includes: Comparing the optimized target region with other regions in the original image data to obtain a comparison result, and reconstructing the target region in the original image data based on the comparison result through an image reconstruction algorithm to obtain the reconstructed target region; Fusing the reconstructed target region with other regions in the original image data to obtain the fused image data, preprocessing the fused image data to obtain the processed image data, and compressing the processed image data by using an image compression algorithm to obtain the compressed image data.
6. An image processing device for a low-power MCU vision module, characterized in that, including: An acquisition module, configured to collect the original image data output by a sensor and the noise feature data corresponding to the original image data, and perform denoising on the original image data and the noise feature data by using a multi-scale filtering method to obtain the preliminarily denoised image data and the residual noise data; Constructing a noise mapping relationship through the residual noise data, and performing adaptive noise reduction processing on the preliminarily denoised image data based on the noise mapping relationship to obtain the target denoised image data; Performing spectral decomposition on the target denoised image data by using a frequency domain analysis method to obtain component data, extracting artifact feature data based on the component data, and performing artifact suppression on the component data by using a directional filtering method based on the artifact feature data to obtain the artifact-suppressed image data and the artifact residual data; calculating the local gradient change information of the artifact-suppressed image data, and performing adaptive compensation on the artifact residual data based on the local gradient change information to obtain the optimized image data; applying an interpolation method to perform edge reconstruction on the optimized image data to obtain the edge-optimized image data and the edge error information, calculating the overall texture information of the edge-optimized image data, and performing texture enhancement processing on the edge-optimized image data based on the overall texture information and the edge error information to obtain the clear image data; A dimensionality reduction module, configured to perform target recognition based on the clear image data to obtain a target region, input the target region into a preset deep learning model to obtain feature data, and perform dimensionality reduction processing on the feature data to obtain the dimensionality-reduced feature vector; A clustering module, configured to identify several category data and category confidence data of the target region according to the dimensionality-reduced feature vector, perform clustering processing on all the category data to generate a clustering result, optimize the clustering result by using the category confidence data to obtain an optimized clustering result, and optimize the target region according to the optimized clustering result to obtain the optimized target region; A compression module, which is used to reconstruct the original image data by using the optimized target region to obtain processed image data, and input the processed image data into an image compression algorithm for calculation to obtain compressed image data.
Citation Information
Patent Citations
Depth compressed sensing network for expanding iterative optimization algorithm
CN112884851A
Double denoising method and device for medical image based on deep learning
CN117095074A
Training data set labeling method and device for AI visual identification
CN119399579A