A visual optimization processing method and system based on real-time image processing
A visual optimization method that dynamically adjusts model parameters through multimodal feature extraction, deep feature fusion, and cluster analysis solves the problems of insufficient real-time performance and flexibility in existing image processing methods, and achieves high-quality image optimization results.
Patent Information
- Application Number
- CN202511124476.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-08-12
AI Technical Summary
Existing image processing methods are inadequate in terms of real-time performance, flexibility, and effectiveness, and cannot fully capture the rich information in images, resulting in poor image processing performance in complex scenes.
A visual optimization method based on real-time image processing is adopted. Through multimodal feature extraction, deep feature fusion and cluster analysis, the model parameters are dynamically adjusted, visual optimization rules are formulated, and high-quality optimized images are generated.
It enables flexible and comprehensive image optimization in complex and ever-changing scenarios, improving the visual effect of images and the accuracy of subsequent analysis and processing, and meeting the high-quality real-time image optimization requirements of practical applications.
Smart Images

Figure CN120634893B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to a visual optimization processing method based on real-time image processing, a control device, a visual optimization processing system and a computer readable storage medium. BACKGROUND
[0002] With the widespread popularity and continuous improvement of camera devices, it becomes easier to obtain real-time image data. However, due to the complexity and diversity of actual application scenarios, the collected images often have various quality problems, such as noise interference, uneven lighting, blurring, etc., which seriously affect the visual effect of the images and subsequent analysis and processing. Therefore, it is of great practical significance to effectively optimize the visual effect of real-time collected images.
[0003] Traditional image processing methods mainly focus on the extraction and processing of single features, such as only focusing on color features or texture features of images. The limitation of this method is that it cannot comprehensively capture the rich information of images, resulting in poor performance when processing images in complex scenarios. For example, in the security monitoring scenario, when encountering dramatic changes in light or occlusions, relying solely on color features for image enhancement may result in the loss of other important information of the image, affecting the identification and analysis of the target.
[0004] And the existing image processing methods have many shortcomings in real-time performance, flexibility and effect, etc., and cannot meet the requirements of high-quality real-time image optimization processing in actual applications. Therefore, a more efficient, flexible and comprehensive visual optimization processing method is needed, which can dynamically adjust the processing strategy according to the real-time features of the image, improve the quality of image optimization, and adapt to various complex application scenarios.
[0005] The above content is only used to assist in understanding the technical solutions of the present application, and does not represent the acknowledgement of the above content as prior art. SUMMARY
[0006] The main purpose of the present application is to provide a visual optimization processing method based on real-time image processing, a control device, a visual optimization processing system and a computer readable storage medium, which aims to optimize the visual effect of real-time images to improve the flexibility of processing and image quality.
[0007] To achieve the above purpose, the present application provides a visual optimization processing method based on real-time image processing, comprising the following steps:
[0008] Based on the real-time image data collected by the camera, image preprocessing is performed;
[0009] The pre-processed image is subjected to multi-modal feature extraction, and a pre-trained machine learning model is used to extract deep features of the image; wherein the multi-modal features include color features, texture features and shape features; the machine learning model has a dynamic feature extraction architecture based on meta-learning, so as to dynamically adjust the network parameters of the model according to different image resolutions and / or complexities;
[0010] The extracted multi-modal features and deep learning features are fused, and the fused features are screened through a feature selection algorithm;
[0011] The screened features are subjected to cluster analysis by using a cluster algorithm;
[0012] According to the results of the feature analysis, visual optimization rules are formulated for different feature types and feature value ranges;
[0013] According to the visual optimization rules, the image is subjected to visual optimization processing to generate an optimized real-time image.
[0014] To achieve the above-mentioned purposes, the application further provides a control device, comprising:
[0015] The acquisition module is used for acquiring image data in real time based on a camera and performing image preprocessing;
[0016] The extraction module is used for multi-modal feature extraction of the pre-processed image, and a pre-trained machine learning model is used to extract deep features of the image; wherein the multi-modal features include color features, texture features and shape features; the machine learning model has a dynamic feature extraction architecture based on meta-learning, so as to dynamically adjust the network parameters of the model according to different image resolutions and / or complexities;
[0017] The fusion module is used for fusing the extracted multi-modal features and deep learning features, and screening the fused features through a feature selection algorithm;
[0018] The cluster module is used for cluster analysis of the screened features by using a cluster algorithm;
[0019] The rule module is used for formulating visual optimization rules for different feature types and feature value ranges according to the results of the feature analysis;
[0020] The optimization module is used for visual optimization processing of the image according to the visual optimization rules, to generate an optimized real-time image.
[0021] To achieve the above object, the application further provides a visual optimization processing system, comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the visual optimization processing method based on real-time image processing.
[0022] To achieve the above object, the application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program, when executed by a processor, implements the steps of the visual optimization processing method based on real-time image processing.
[0023] The visual optimization processing method based on real-time image processing, the control device, the visual optimization processing system and the computer readable storage medium provided by the application can not only extract multi-modal features in the feature extraction stage, but also extract deep features by using a machine learning model designed based on meta-learning, and can dynamically adjust network parameters according to image resolution and complexity, thereby comprehensively and flexibly capturing image information; and can formulate visual optimization rules based on feature analysis results, thereby accurately optimizing different feature types and ranges; and finally processing images according to the rules, thereby effectively improving image visual effects and the accuracy of subsequent analysis and processing, meeting the requirements of actual applications for high-quality real-time image optimization processing, and having stronger adaptability in complex and variable scenes. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 A step schematic diagram of the visual optimization processing method based on real-time image processing in an embodiment of the application;
[0025] Figure 2 A schematic diagram of the control device in an embodiment of the application;
[0026] Figure 3 An internal architecture schematic diagram of the visual optimization processing system in an embodiment of the application.
[0027] The implementation, functional features and advantages of the application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0028] The embodiments of the application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference signs represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the application, and cannot be understood as a limitation on the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.
[0029] In addition, if the description in the present application involves "first", "second", etc., it is only for the purpose of description (such as for distinguishing the same or similar features), and cannot be understood as indicating or implying the relative importance of the indicated technical features or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first", "second" can be explicitly or implicitly included at least one of the features. In addition, the technical solutions of various embodiments can be combined with each other, but it must be based on the fact that a person skilled in the art can realize it, and when the combination of technical solutions appears contradictory or cannot be realized, it should be considered that the combination of technical solutions does not exist, nor within the protection scope required by the present application.
[0030] Reference Figure 1 In an embodiment, the visual optimization processing method based on real-time image processing includes:
[0031] Step S10, based on the camera real-time acquisition of image data, and image preprocessing is carried out;
[0032] Step S20, the preprocessed image is subjected to multi-modal feature extraction, and a pre-trained machine learning model is used to extract deep features of the image; wherein the multi-modal features include color features, texture features and shape features; the machine learning model has a dynamic feature extraction architecture based on meta-learning, so as to dynamically adjust the network parameters of the model according to the different image resolution and / or complexity;
[0033] Step S30, the extracted multi-modal features and deep learning features are fused, and the fused features are screened through a feature selection algorithm;
[0034] Step S40, the screened features are subjected to cluster analysis by using a clustering algorithm;
[0035] Step S50, according to the results of feature analysis, visual optimization rules are formulated for different feature types and feature value ranges;
[0036] Step S60, according to the visual optimization rules, the image is subjected to visual optimization processing, and the optimized real-time image is generated.
[0037] In the embodiment, the execution terminal of the embodiment can be a visual optimization processing system, or other devices or apparatuses (such as control apparatuses) for controlling the visual optimization processing system.
[0038] As described in step S10, the camera is used to continuously capture image data in the real scene at a certain frame rate (such as 30 frames per second, 60 frames per second, etc.). The type of camera can be selected according to the specific application scenario, such as a general RGB camera, a depth camera, etc.
[0039] To improve the effectiveness and efficiency of subsequent processing, the collected raw image data needs to be pre-processed. Optional pre-processing operations include:
[0040] (1) Grayscale: Convert color images to grayscale images to reduce data volume, while also helping some feature extraction based on grayscale information;
[0041] (2) Filtering: Use filters (such as Gaussian filters, median filters, etc.) to remove noise in the image, making the image smoother;
[0042] (3) Normalization: Normalize the pixel values of the image, usually scaling the pixel values to the range of [0, 1] or [0, 255] to ensure consistency between different images.
[0043] As described in step S20, feature extraction is performed from different angles to obtain more comprehensive image information. Multi-modal features include color features, texture features, and shape features.
[0044] Among them, color features can be extracted by calculating the color histogram of the image, color moments, etc. The color histogram reflects the distribution of different colors in the image, and the color moments can describe the mean, variance, and skewness of the color statistics.
[0045] Optional texture feature extraction methods include gray level co-occurrence matrix and local binary pattern (LBP). The gray level co-occurrence matrix describes the gray level spatial relationship between pixels in the image, and the LBP extracts texture information by comparing the gray level of the center pixel with the gray level of the neighborhood pixels.
[0046] Shape features can be extracted by calculating the perimeter, area, circularity, and other geometric features of objects in the image. For complex shapes, contour analysis and Fourier descriptors can also be used for description.
[0047] Optionally, a pre-trained machine learning model is used for deep feature extraction, which is designed based on meta-learning with a dynamic feature extraction architecture.
[0048] Optionally, a convolutional neural network (CNN) pre-trained on a large-scale image dataset (such as ImageNet) is used, such as VGG, ResNet, etc. These models have learned rich image feature representation capabilities on large-scale data.
[0049] Meta-learning, also known as "learning to learn", aims to enable a model to quickly learn and adapt in different task environments with a small amount of data. In this real-time image processing-based visual optimization scenario, meta-learning is used to design a feature extraction architecture that can dynamically adjust network parameters based on different image attributes (resolution and / or complexity), enabling the model to maintain efficient and accurate feature extraction capabilities under various image conditions.
[0050] When the input image has low resolution or low complexity, the model can reduce the number of convolution kernels, reduce the size of the convolution kernel, or reduce the number of network layers to reduce the amount of computation and the number of parameters; when the input image has high resolution or high complexity, the model increases the corresponding computing resources to extract more rich feature information. For example, a controller network is trained through meta-learning, which can dynamically adjust the hyperparameters of the CNN model, such as the stride of the convolution layer and the number of channels, based on the features of the input image (such as resolution, variance, etc.).
[0051] Among them, the image resolution directly affects the degree of detail that can be obtained in the image. High-resolution images contain more details and require more complex feature extraction processes; low-resolution images are relatively simple. Therefore, the image resolution is input as meta-data into the meta-learning module, which can adjust the complexity of the feature extraction architecture based on the resolution.
[0052] Among them, the image complexity can be measured in various ways, such as the richness of the image texture, the diversity of the color distribution, the complexity of the object number and shape, etc. Some pre-defined indicators or another small neural network can be used to evaluate the image complexity, and the evaluation results are passed to the meta-learning module as meta-data.
[0053] Optionally, a controller network is designed as the core component of meta-learning. The controller network receives the meta-data (resolution, complexity, etc.) of the image as input and outputs a series of control signals, which are used to dynamically adjust the network parameters of the feature extraction model. For example, the controller network can decide whether to enable certain convolution layers, adjust the size and number of convolution kernels, change the parameters of the pooling layer, etc. based on the image resolution and complexity.
[0054] Optionally, during the meta-training phase, a large number of image data with different resolutions and complexities are used to train the controller network. During the training process, the controller network tries to find the best parameter adjustment strategy to enable the feature extraction model to obtain the optimal feature extraction effect under various image conditions. Reinforcement learning or gradient descent methods are usually used to optimize the parameters of the controller network.
[0055] Optionally, the specific manner of dynamically adjusting the network parameters includes adjusting the convolution layer parameters and / or the pooling layer parameters.
[0056] The convolution layer parameters include the number of convolution kernels, the size of the convolution kernel, and the depth of the convolution layer.
[0057] Optionally, for low-resolution or low-complexity images, reducing the number of convolution kernels can reduce the computational load and the complexity of the model. For example, a convolution layer originally using 64 convolution kernels on a high-resolution image can be reduced to 32 when processing a low-resolution image.
[0058] Optionally, the size of the convolution kernel is adjusted according to the level of detail of the image. For high-resolution images that require more local details to be captured, larger convolution kernels (such as 5x5 or 7x7) can be used; for low-resolution images, smaller convolution kernels (such as 3x3) may be sufficient.
[0059] Optionally, in some cases, it can be determined according to the image properties whether to enable certain convolution layers. For simple images, some deep convolution layers can be skipped to reduce the computational load; for complex images, the complete convolution layer structure is used to extract more rich features.
[0060] The pooling layer parameters include the size of the pooling window and the type of pooling.
[0061] The pooling operation can reduce the size of the feature map and reduce the computational load. For high-resolution images, larger pooling windows (such as 2x2 or 3x3) can be used; for low-resolution images, smaller pooling windows (such as 1x1 or no pooling) can be used to preserve more detail information.
[0062] Optionally, different pooling types can be selected according to the image situation, such as maximum pooling, average pooling, etc. Maximum pooling focuses more on extracting prominent features in the image, while average pooling emphasizes the smoothness of the features. In processing images with rich textures, maximum pooling may be more appropriate; while for images with uniform colors, average pooling may work better.
[0063] On this basis, the preprocessed image is input into the pre-trained CNN model, and the deep features of the image are gradually extracted through convolutional layers, pooling layers and other operations. The output of the middle layer or the last layer of the CNN model is selected as the deep feature vector of the image. These deep feature vectors have high abstraction and representativeness, and can well describe the semantic information of the image. And by dynamically adjusting the network parameters according to the image resolution and complexity, the use of the same complex model on all images is avoided, thereby reducing unnecessary computational load and improving the efficiency of real-time processing; and the feature extraction process can be optimized for different image conditions, so that the model can extract more accurate and representative features on various images, thereby improving the effect of subsequent visual optimization processing.
[0064] As described in step S30, after completing step S20, the multi-modal features of the image (including color features, texture features and shape features) and the deep features extracted by the pre-trained machine learning model have been obtained. These features describe the information of the image from different angles, and in order to comprehensively utilize these information, they need to be fused.
[0065] Optionally, considering that the importance of different types of features to the final result may be different, different weights can be given to the multi-modal features and the deep learning features, and then the weighted features are added to obtain the fused features. The weights can be optimized through model training based on the initial setting by relevant engineers through experience (i.e., optimization based on prior knowledge learned by the model).
[0066] The fused feature vector may contain a large number of features, some of which may have no actual effect on subsequent clustering analysis and visual optimization rule formulation, and even increase the computational complexity and noise. Therefore, a feature selection algorithm needs to be used to screen the fused features, and the most representative and discriminative features are retained.
[0067] Optionally, the importance of a feature is evaluated based on its variance. The greater the variance, the greater the variation of the feature in different samples, and the more likely it is to help classification or clustering. For numerical features, the variance of each feature can be calculated, and then the features with a variance greater than a certain threshold are selected.
[0068] For non-numerical features, an L1 regularization term can be added to the regression model, so that the coefficients of some unimportant features become 0, thereby realizing feature selection. It should be understood that by adjusting the regularization parameter, the number of selected features can be controlled.
[0069] As described in step S40, a suitable clustering algorithm (such as K-means clustering, hierarchical clustering, etc.) is selected to perform clustering analysis on the filtered features. The purpose of clustering is to divide image data points with similar features into the same category, so that subsequent processing of images in different categories can be targeted.
[0070] Taking K-means clustering as an example, first randomly select K cluster centers, then assign each data point to the category where the nearest cluster center is located, then update the position of the cluster center, repeat this process until the cluster center no longer changes or reaches the maximum number of iterations.
[0071] As described in step S50, after completing the feature clustering analysis (step S40), according to the results of the analysis, formulate corresponding visual optimization rules for different feature types (such as color features, texture features, shape features, etc.) and feature value ranges.
[0072] After step S40 completes the clustering analysis, different feature clusters can be obtained, and the features within each cluster have similarity. First, these features need to be carefully sorted out to determine different feature types and their corresponding feature value ranges.
[0073] Optionally, classify features by color, texture, shape, etc. For example, color features may include RGB values, HSV values, etc.; texture features may include roughness, contrast, etc.; shape features may include perimeter, area, aspect ratio, etc.
[0074] Optionally, for each feature type, statistics of its feature value range in each cluster are calculated. The minimum, maximum, mean, standard deviation, etc. can be calculated to better understand the distribution of the features. For example, for the red channel value in the color feature, its range in a certain cluster may be [100, 200].
[0075] Among them, the formulation of color feature optimization rules:
[0076] Optionally, if the color feature shows that the brightness of some areas of the image is low, rules can be formulated to increase the brightness of these areas; for example, when the brightness value is lower than a certain threshold (such as 50, the value range is 0-255), increase the RGB value of the area by a certain percentage (such as 1.2 times). For areas with low color contrast, the contrast can be enhanced by adjusting the saturation and brightness difference of the color; for example, calculate the color difference of adjacent pixels, if the difference is less than a certain threshold, appropriately increase the difference in saturation and brightness.
[0077] Optionally, if the color feature analysis shows that the image has a color cast (e.g., overall reddish), the values of other color channels can be adjusted according to the degree and direction of the color cast; for example, if the red channel value is generally too high, the red channel value can be appropriately reduced, and the values of the green and blue channels can be increased. For regions that are not bright enough in color, the saturation value can be increased according to the characteristics of the HSV color space; for example, when the saturation is less than 30%, the saturation value can be increased to 50%.
[0078] In the establishment of the optimization rules for the texture feature:
[0079] Optionally, if the texture feature shows that the texture of some regions is blurred, a sharpening filter algorithm (such as the Laplacian operator or the Sobel operator) can be used to process these regions to enhance the edge information of the texture and improve the clarity. For regions with overly rough texture, a smoothing filter algorithm (such as Gaussian filtering or mean filtering) can be used to process them to make the texture more delicate.
[0080] Optionally, when the texture features of adjacent regions differ too much, interpolation or fusion methods can be used to make the texture transition of these regions more natural. For example, a bilinear interpolation algorithm can be used to process the texture boundary.
[0081] In the establishment of the optimization rules for the shape feature:
[0082] Optionally, if the shape feature shows that the shape of an object is missing, the missing part can be supplemented according to the overall shape of the object and the context information; for example, for a circular object with a missing part of the edge, the missing part can be fitted and repaired according to the center and radius information. For an object with a deformed shape, the original shape characteristics (such as the aspect ratio and symmetry) can be used for correction; for example, an inclined rectangle can be corrected to a right rectangle.
[0083] Optionally, for shapes that are important (such as logos or key objects), the edges can be enhanced, and the color can be adjusted to make them more prominent; for example, the outline color of an important shape can be set to a prominent color (such as yellow). For overly complex shapes, some small edges and details can be removed to simplify the representation of the shape; for example, an irregular polygon can be approximated to a polygon with a reduced number of vertices.
[0084] Optionally, after the basic optimization strategies are established, the rules can be further refined in combination with the feature value range to achieve more accurate optimization: the feature value range can be divided into different intervals, and different optimization parameters can be established for each interval; for example, for the brightness feature, when the brightness value is in the [0, 50] interval, the brightness can be increased by 30%, and when the brightness value is in the [51, 100] interval, the brightness can be increased by 20%.
[0085] The optimization rules are dynamically adjusted according to the distribution density and variation trend of the eigenvalues. For example, finer optimization parameters are used in regions with dramatic eigenvalue variation, and relatively uniform optimization strategies are used in regions with uniform eigenvalue distribution.
[0086] Through the above steps, a comprehensive, detailed and effective visual optimization rule set for different feature types and eigenvalue ranges can be developed based on the results of feature analysis.
[0087] As described in step S60, before image optimization, the visual optimization rules developed in step S50 need to be analyzed and converted into specific operation parameters. According to the feature types (such as color, texture, shape, etc.) targeted by the rules, the corresponding rules are read respectively. For example, for color optimization rules, information such as brightness adjustment ratio and color balance correction parameters is extracted.
[0088] The abstract descriptions in the rules are converted into specific numerical parameters. For example, the rule states "when the brightness value is less than 50, increase the brightness by 30%", in actual operation, parameters such as 50 and 30% need to be extracted for subsequent calculation.
[0089] In order to apply the optimization rules more accurately, the image needs to be partitioned, and the features of each region are matched with the rules.
[0090] Optionally, the image can be divided into different regions according to the clustering results of step S40, and the features in each region have similarity; or a fixed grid division method can be used to divide the image into small blocks. For each partition, its color, texture, shape and other features are extracted and matched with the feature types and eigenvalue ranges in the rules to determine the specific optimization rules applicable to the partition. For example, if the brightness value of a certain partition is within the low brightness range defined in the rules, the corresponding brightness enhancement rule is applied.
[0091] Among them, for color feature optimization:
[0092] Optionally, the brightness of the image partition is adjusted according to the brightness adjustment ratio set in the rules. This can be achieved by modifying the RGB values of the image. For example, for low brightness regions, multiply the RGB values of each pixel in the region by the brightness adjustment coefficient.
[0093] Optionally, histogram equalization, adaptive histogram equalization and other methods are used to enhance the contrast of the image. These methods can adjust the brightness distribution of the image, making the bright parts brighter and the dark parts darker, thereby improving the contrast.
[0094] Optionally, the color channel values of the image are adjusted according to the color correction parameters determined in the rules. For example, if the image is found to be reddish, the value of the red channel is reduced, and the values of the green and blue channels are appropriately increased.
[0095] Optionally, the saturation, hue, and other parameters of the image are adjusted to make the colors of the image more vivid. The HSV color space can be used to operate, and the saturation value is increased to enhance the vividness of the colors.
[0096] Among them, the texture feature optimization is:
[0097] Optionally, a sharpening filter (such as a Laplace filter, a Sobel filter, etc.) is used to perform convolution operation on the image to enhance the edge information of the texture and make the texture clearer. If the texture feature display image has noise, Gaussian filtering, median filtering, etc. can be used to remove the noise and make the texture smoother.
[0098] Optionally, for the case where the texture difference between adjacent regions is large, interpolation or fusion method is used to make the texture transition more natural. For example, the bilinear interpolation algorithm is used to process the texture boundary.
[0099] Among them, the shape feature optimization is:
[0100] Optionally, according to the shape repair strategy formulated in the rules, the missing part of the object shape is supplemented. For example, for the missing edge of a circular object, the missing edge can be fitted and repaired according to the center and radius information.
[0101] Optionally, the deformed object shape is corrected by affine transformation, perspective transformation, etc. For example, an inclined rectangle is corrected to a right rectangle.
[0102] Optionally, the important shape is made more prominent by enhancing its edge and adjusting its color. For example, the contour color of the important shape is set to a prominent color (such as yellow).
[0103] Optionally, the small edges and details in the shape are removed by using polygon approximation and contour simplification algorithms, so that the shape is more concise.
[0104] After the optimization processing of each partition is completed, the optimized partitions need to be merged to generate a complete optimized image. That is, the optimized partitions are spliced according to the original positions to form a complete image. In the process of partition splicing, the problem of discontinuous boundary may occur, and the boundary can be processed by using smooth transition method to make the image look more natural.
[0105] Through the above steps, the image can be comprehensively and carefully optimized according to the visual optimization rules to generate high-quality optimized real-time image.
[0106] In an embodiment, in the feature extraction stage, not only multi-modal features are extracted, but also a machine learning model designed based on meta-learning is used for deep feature extraction, and network parameters can be dynamically adjusted according to image resolution and complexity, so as to comprehensively and flexibly capture image information; and visual optimization rules are formulated based on feature analysis results, so as to accurately optimize different feature types and ranges; finally, the image is processed according to the rules, which can effectively improve the image visual effect and the accuracy of subsequent analysis and processing, meet the requirements of real-time image optimization processing of high-quality images in actual applications, and have stronger adaptability in complex and variable scenes.
[0107] In an embodiment, on the basis of the above embodiment, before the steps of performing multi-modal feature extraction on the preprocessed image and using a pre-trained machine learning model to perform deep feature extraction on the image, the method further comprises:
[0108] Performing multi-scale decomposition on the preprocessed image to decompose the image into sub-band images of different scales and directions for subsequent feature extraction.
[0109] In this embodiment, before performing multi-modal feature extraction and deep feature extraction on the preprocessed image, multi-scale decomposition is performed first to more comprehensively and meticulously analyze the feature information of the image. Images of different scales can reflect the features of the image at different resolutions, while sub-band images of different directions can capture the texture and structure information of the image in each direction. This helps to more accurately extract multi-modal features and deep features subsequently, thereby improving the performance of the entire visual optimization processing method.
[0110] The optional multi-scale decomposition method includes wavelet transform, Laplacian pyramid decomposition, etc. Taking wavelet transform as an example, it is a time-frequency analysis method that has good localization characteristics in both time and frequency domains, and is suitable for multi-scale decomposition of images.
[0111] The preprocessed image is decomposed in scale using the selected decomposition method. Generally, the image is gradually decomposed into multiple sub-images of different scales from the original scale. For example, in wavelet transform, after one level of decomposition, the image is decomposed into four sub-bands: one low-frequency sub-band (LL) and three high-frequency sub-bands (LH, HL, HH). The low-frequency sub-band represents the smooth part of the image and contains the main contour information of the image; while the high-frequency sub-bands represent the detail information in the horizontal, vertical and diagonal directions, respectively.
[0112] Optionally, multi-level decomposition can be performed, and each level of decomposition is a further decomposition of the low-frequency sub-band obtained in the previous level of decomposition. In this way, sub-band images at different scales can be obtained, and each sub-band image at a scale contains specific information of the image at that scale.
[0113] Optionally, in addition to the decomposition in scale, the image can also be decomposed in direction. In wavelet transform, through different filter combinations, high-frequency sub-band images in horizontal, vertical and diagonal directions can be obtained. These sub-band images in different directions can capture the texture and edge information of the image in each direction. For example, the high-frequency sub-band image in the horizontal direction can highlight the edges and textures in the horizontal direction of the image, and the high-frequency sub-band image in the vertical direction can highlight the edges and textures in the vertical direction.
[0114] The role of sub-band images after multi-scale decomposition in subsequent feature extraction:
[0115] (1) For color features, sub-band images of different scales can reflect the distribution of colors at different resolutions. For example, in sub-band images of larger scales, the overall distribution trend of colors can be analyzed; while in sub-band images of smaller scales, the detailed changes of colors can be focused on. For texture features, sub-band images in different directions can more clearly show the texture information of the image in each direction, which helps to more accurately extract texture features. For shape features, sub-band images of different scales can provide shape contour information of different accuracies, which can be combined to more comprehensively describe the shape of objects in the image.
[0116] (2) The pre-trained machine learning model can take the sub-band images after multi-scale decomposition as input for deep feature extraction. Sub-band images of different scales and directions contain rich information, and the model can learn the association and feature representation between these information. For example, the model can learn the hierarchical structure features of the image from sub-band images of different scales, and learn the directional features of the image from sub-band images of different directions. Through deep feature extraction on these sub-band images, more representative and discriminative deep features can be obtained, thereby improving the effect of subsequent visual optimization processing.
[0117] In an embodiment, multi-scale decomposition is an important step before feature extraction on the pre-processed image. It provides more rich and comprehensive information for subsequent multi-modal feature extraction and deep feature extraction by decomposing the image into sub-band images of different scales and directions. This can make the entire visual optimization processing method more accurately capture the features of the image, thereby achieving better visual optimization effect.
[0118] In an embodiment, on the basis of the above embodiment, the visual optimization processing method based on real-time image processing further comprises:
[0119] In the fusion of multi-modal features and deep learning features, a multi-level fusion method is adopted, that is, feature fusion is performed at different feature extraction stages, and all features obtained after the first fusion are secondarily fused.
[0120] In this embodiment, during the multi-modal feature extraction process, color features and texture features often have close connections in different local regions and scales of the image. For example, in some images, a specific color region can be accompanied by a specific texture pattern. At an early stage of extracting these features, color features and texture features can be fused. For example, while performing color histogram statistics based on the local region of the image, the texture energy features of the region are combined, and the two are linearly combined according to certain weights to form a new fused feature vector. Such fusion can better describe the comprehensive features of the local region of the image.
[0121] Mid-level fusion of shape features with other features: After the shape features are preliminarily extracted, they can be fused with the color-texture features that have been fused at the mid-level. Shape features generally reflect the contour information of objects in the image, while color-texture features describe the properties of the object surface. By fusing shape features with color-texture fusion features, a more representative feature combination can be obtained. For example, the moment invariants of shape features can be spliced with the color-texture fusion feature vector to form a new feature vector containing object contour and surface property information.
[0122] Fusion of shallow deep learning features with multi-modal features: At the shallow stage of deep feature extraction by a pre-trained machine learning model, the features output by the model usually contain some basic edge and local structure information of the image. At this time, these shallow deep learning features can be fused with part of the multi-modal features (such as color features). For example, the feature maps output by the shallow convolutional layer can be element-wise added or multiplied with the color feature maps to obtain a new fused feature map. This fusion can combine the sensitivity of the deep learning model to image structure and the ability of color features to describe the visual appearance of the image.
[0123] Further fusion of mid-level deep learning features with multi-modal features: As the deep learning model goes deeper, mid-level features begin to capture some more abstract semantic information of the image. At this stage, mid-level deep learning features can be fused with color-texture-shape features that have been fused in the multi-modal features. For example, the feature vector output by the mid-level convolutional layer can be weighted summed with the multi-modal fusion feature vector to obtain a more rich fusion feature vector that integrates the detailed information of multi-modal features and the semantic information of mid-level deep learning features.
[0124] After completing the fusion of different feature extraction stages, all the features obtained after the first fusion are fused again. The purpose of the second fusion is to integrate the information contained in each first fusion feature, further eliminate redundancy, and enhance the representativeness and discriminativeness of the features.
[0125] Optionally, all the feature vectors after the first fusion can be spliced to form a longer feature vector. Then, according to the importance and correlation of different features, different weights are assigned to each feature for weighted summation. For example, features that are more discriminative in image classification tasks can be assigned higher weights; while features with more redundancy or noise can be assigned lower weights.
[0126] Optionally, in the secondary fusion process, a nonlinear transformation can also be introduced, such as using a fully connected layer of a neural network to perform nonlinear mapping on the spliced feature vector. In this way, the model can learn the complex nonlinear relationships between different features, further improving the quality of the fused features.
[0127] Multi-level fusion can fully exploit the information of multi-modal features and deep learning features at different extraction stages, so that the fused features can more comprehensively and accurately describe the characteristics of the image. This helps to improve the accuracy of subsequent feature selection, clustering analysis and visual optimization rule formulation.
[0128] Different images may exhibit different feature advantages at different feature extraction stages. The multi-level fusion method can flexibly fuse features at different stages according to the specific situation of the image, thereby enhancing the adaptability of the entire visual optimization processing method to different images.
[0129] High-quality features obtained through multi-level fusion can provide more reliable basis for subsequent visual optimization rule formulation. Optimization rules based on these features can more accurately optimize different features of the image, thereby improving the visual quality and clarity of the image.
[0130] In an embodiment, the multi-level fusion method provides a more effective feature integration method for real-time image processing visual optimization by separately fusing features at different feature extraction stages and then performing secondary fusion. This method can fully exploit the advantages of multi-modal features and deep learning features, improve the expressiveness of features and the adaptability of the model, and ultimately optimize the visual effect of the image.
[0131] In an embodiment, on the basis of the above embodiment, the clustering algorithm includes a dynamic clustering algorithm, which is used to adjust the clustering result according to the dynamic change of the object between the previous image and the current image.
[0132] In this embodiment, before processing the current image, the features of the objects in the previous image need to be tracked. This can be achieved by some feature tracking algorithms, such as optical flow, Kalman filtering, etc. The optical flow method can calculate the motion vector of each pixel point in the image between adjacent frames, thereby determining the motion direction and speed of the object; Kalman filtering can predict and update the motion state of the object, improving the accuracy of feature tracking.
[0133] Alternatively, by comparing the features of the object in the previous image and the current image, the dynamic changes of the object are detected. These changes can include position movement, shape deformation, color change, etc. For example, the change of the center of gravity position of the object in two frames of images is calculated to determine whether the object has moved; the change of the contour shape of the object is calculated to determine whether the object has deformed.
[0134] Before performing the dynamic clustering algorithm on the current image, a traditional clustering algorithm (such as K-means clustering, hierarchical clustering, etc.) is first used to preliminarily cluster the features of the current image, and an initial clustering result is obtained (this initial result is based on the static features of the current image), and then the dynamic clustering algorithm is performed.
[0135] Optionally, according to the dynamic changes of the objects, the positions of the clustering centers can be adjusted. If a certain object has moved, the clustering center containing the object can need to move to the new position of the object. For example, if an object moves from one clustering area to the vicinity of another clustering area, the positions of the two clustering centers can be adjusted appropriately to better adapt to the new distribution of the object.
[0136] Optionally, it is checked whether the dynamic changes of each object make it more suitable to belong to other clusters. If the features of an object have changed greatly in the current image, resulting in a decrease in similarity to the cluster it currently belongs to and an increase in similarity to other clusters, the object can be reassigned to a more suitable cluster. For example, the color of an object has changed significantly between two frames of images, making the difference in color features between the object and the original cluster increase, and the color features of another cluster closer to the object, at which time the object can be removed from the original cluster and added to the new cluster.
[0137] Optionally, according to the dynamic changes of the objects, merging or splitting operations of the clusters can also be needed. If the objects in two clusters become more and more similar due to dynamic changes, the two clusters can be merged into one; if the objects in a cluster have obvious differences due to dynamic changes, the cluster can be split into multiple sub-clusters.
[0138] Optionally, by considering the dynamic changes of objects, the dynamic clustering algorithm can more accurately cluster the objects in the image, avoiding clustering errors caused by object movement or changes. This helps to more accurately identify different object categories in the image, providing a more reliable basis for subsequent visual optimization rule making.
[0139] In real-time image processing, image content is constantly changing. The dynamic clustering algorithm can adjust the clustering results in real time to adapt to such changes, ensuring the real-time and effectiveness of visual optimization processing. For example, in a video monitoring scenario, when an object enters or leaves the monitoring area, the dynamic clustering algorithm can timely adjust the clustering results to correctly cluster and analyze the newly appearing object.
[0140] In this way, the dynamic changes of objects in the image can be adapted to, and accurate clustering analysis can be performed in different dynamic scenarios, improving the robustness and adaptability of the clustering algorithm; more accurate clustering results help to develop more reasonable visual optimization rules, thereby improving the effect of visual optimization processing and making the optimized image more meet the actual needs; in real-time image processing, the clustering results can be adjusted in real time, ensuring the real-time of the entire visual optimization processing flow, and being applicable to application scenarios with high real-time requirements.
[0141] In an embodiment, on the basis of the above-mentioned embodiments, the visual optimization rule is further dynamically adjusted according to the real-time application scenario of the image:
[0142] When the image is used for a video conference scenario, the skin color and clarity of the human face are optimized, background interference is reduced, and the overall brightness and contrast of the image are improved;
[0143] When the image is used for a security monitoring scenario, the edge clarity and detail information of the image are enhanced, and the recognition of the target object is improved;
[0144] When the image is used for an advertising promotion scenario, the brightness and saturation of the color are strengthened.
[0145] In this embodiment, in the visual optimization process of real-time image processing, different application scenarios have different requirements for images. Simply optimizing according to the features of the image is not comprehensive enough, because the same image has great differences in aspects that need to be highlighted and optimized in different scenarios. Therefore, dynamically adjusting the visual optimization rule according to the real-time application scenario of the image can make the optimized image better meet the needs of specific scenarios, and improve the visual effect and application value.
[0146] In video conferencing, the facial information of the person is the key to communication. In order to enable the participants to clearly see each other's expressions and gestures, it is necessary to focus on optimizing the skin color and clarity of the person's face. The skin color of the person's face can be adjusted by color correction algorithm, so that it is more natural and real. For example, use skin color model to detect face area, and then fine-tune the color of the area according to the preset skin color standard. At the same time, image sharpening technology is adopted to enhance the edges and details of the face, improve the clarity of the face, and make the facial features more obvious.
[0147] In video conferencing, the background may distract the participants. Therefore, it is necessary to reduce the background interference. The focus of the background can be blurred by the background blurring algorithm, so that the audience's attention is more focused on the person's face. Background replacement technology can also be used to replace the cluttered background with a simple and unified background, such as a solid color background or a virtual background.
[0148] Optionally, appropriate brightness and contrast can make the image clearer and brighter, and improve the viewing experience. By adjusting the brightness and contrast parameters of the image, the details of the person's face and the surrounding environment can be clearly displayed. For example, use histogram equalization or adaptive histogram equalization method to enhance the contrast of the image, and adjust the brightness according to the ambient light conditions to avoid over-bright or over-dark situations.
[0149] In security monitoring, accurate identification of target objects is crucial. In order to improve the recognition of target objects, it is necessary to enhance the edge clarity and detail information of the image. Edge detection algorithms such as Sobel operator, Canny operator, etc. can be used to highlight the edge information in the image, making the outline of the object more clear. At the same time, image enhancement techniques such as Laplacian operator enhancement, wavelet transform enhancement, etc. can be used to enhance the details of the image, making the features of the target object more obvious.
[0150] In addition to enhancing edges and details, the color, contrast, and other parameters of the image can also be adjusted to make the target object more prominent in the image. For example, for the target object that may be dark in night monitoring, its brightness and contrast can be increased to make it easier to be identified. At the same time, real-time analysis and processing of the monitoring image is carried out, and according to the characteristics of the target object (such as shape, size, color, etc.), classification and labeling are carried out, which is convenient for monitoring personnel to quickly identify and locate the target object.
[0151] The purpose of advertising is to attract the attention of the audience, and the image with bright and high saturation color can achieve this purpose more effectively. By using color enhancement algorithm, the color of the image is adjusted to improve the brightness and saturation of the color. For example, using color lookup table (LUT) technology, the color in the image is mapped to a more bright and saturated color space. At the same time, the overall color tone of the image is adjusted to make it more consistent with the theme and style of the advertisement, creating a more attractive visual effect.
[0152] In order to realize the dynamic adjustment of visual optimization rules according to real-time application scenarios, a real-time scene judgment mechanism can be established. The application scenario of the current image can be judged through the metadata of the image (such as shooting device, shooting time, shooting place, etc.), user input scene information or intelligent analysis algorithm. Once the scene changes, the system can quickly switch to the corresponding visual optimization rule to optimize the image in real time, ensuring that the optimized image always meets the needs of the current scene.
[0153] In an embodiment, specific visual optimization rules are formulated according to different application scenarios, which can highlight the key information of images in different scenarios, meet the special needs of each scenario, and improve the pertinence and effectiveness of visual optimization.
[0154] In the video conference scenario, the optimized image makes the participants communicate more smoothly; in the security monitoring scenario, it helps to more accurately identify target objects; in the advertising scenario, it can attract more attention of the audience, thereby improving the user experience in different scenarios.
[0155] In this way, the optimization rules can be adjusted in real time according to the changes of the scene, ensuring that the images in different scenarios can be optimized in time and accurately, adapting to the dynamic requirements of real-time image processing.
[0156] In an embodiment, based on the above embodiment, the visual optimization processing method based on real-time image processing further comprises:
[0157] When formulating the visual optimization rule, the user's individual preference factor is introduced;
[0158] Among them, the system records the user's historical optimization operation and feedback on different types of images, learns the user's individual optimization preference; in the subsequent image optimization process, according to the type of image currently processed by the user, combined with the user's individual preference, the visual optimization rule is fine-tuned.
[0159] In this embodiment, the system records in detail the user's historical optimization operations on different types of images (e.g. landscape, portrait, art, etc.). These operations include but are not limited to adjustments to image brightness, contrast, saturation, sharpness, etc., as well as local optimization operations on specific regions. For example, when the user processes a landscape photo and increases the brightness by 20% and reduces the contrast by 10%, the system accurately records the specific values and operation objects (i.e. image type) of these operations.
[0160] Optionally, in addition to recording the user's operations, the system also collects the user's feedback on the optimized image. The feedback can be the user's direct evaluation (e.g. satisfied, dissatisfied, needs further adjustment, etc.), or the user's behavior of modifying the image again. For example, if the user adjusts the skin color again on an optimized portrait, indicating that the user is not satisfied with the previous skin color optimization result, the system will record this feedback information.
[0161] Optionally, the system will mine and analyze the recorded user historical optimization operations and feedback data. By using machine learning algorithms (such as clustering analysis, association rule mining, etc.), the user's preference patterns in different types of image optimization are found out. For example, through clustering analysis, it can be found that the user always tends to adjust the skin color to a certain specific color range when processing a portrait; through association rule mining, it can be found that the user usually increases the saturation of the image when increasing the brightness of the image.
[0162] Optionally, according to the results of data analysis, the system will establish a personalized optimization preference model for each user. This model can describe the user's preference degree and value range of various optimization parameters for different types of images. For example, for a portrait, the user's preference model may show that the user likes to keep the brightness between 50% - 60% and the saturation between 70% - 80%.
[0163] In the subsequent image optimization process, the system first identifies the type of the image being processed. For example, through image recognition algorithms, it determines whether the image is a landscape, a portrait or other types of images.
[0164] Then, according to the identified image type, the system calls the user's personalized preference model for that type of image to fine-tune the general visual optimization rules. For example, if the general rule suggests increasing the brightness of a portrait by 15%, but according to the user's personalized preference model, the user usually likes to keep the brightness of the portrait between 50% - 60%, the system will adjust the brightness based on the actual brightness of the current image and the general rule, possibly only increasing the brightness by 10% to meet the user's personalized preferences.
[0165] Optionally, the system continuously updates the user's historical operations and feedback data, continuously learning the user's personalized preference changes. As the user uses the system for a longer time, the system's understanding of the user's preferences becomes more and more accurate, enabling more accurate fine-tuning of the visual optimization rules and providing the user with more personalized image optimization services.
[0166] In an embodiment, the user's personalized preference factor is introduced, enabling image optimization to be customized according to each user's unique preferences, meeting the user's personalized needs and improving the user's satisfaction with the optimized images. By providing personalized image optimization services, the interaction and dependence between the user and the system are enhanced, and the user's usage stickiness is improved.
[0167] In addition, the system's continuous learning mechanism can continuously adapt to changes in user preferences, ensuring the accuracy and effectiveness of personalized optimization services in the long-term use process, so that the system can always provide image optimization results that meet the user's latest preferences.
[0168] In an embodiment, on the basis of the above embodiment, the visual optimization rules also consider the time series information of the image, analyze the feature change trend between adjacent images for a continuously captured image sequence, and dynamically adjust the visual optimization rules according to this trend.
[0169] When there is a dynamic scene in the image, different blurring and sharpening processing is applied to images at different times according to the speed and direction of motion, to simulate real visual effects and enhance the dynamic feeling of the picture.
[0170] In this embodiment, for a continuously captured image sequence, the system first extracts features from each image. These features can include color features (such as color histogram, color distribution), texture features (such as gray level co-occurrence matrix, local binary pattern), shape features (such as edge information, contour shape), etc. By extracting these features, the visual characteristics of the image can be quantified, providing a basis for subsequent analysis.
[0171] Optionally, after extracting the features of adjacent images, the system compares and analyzes these features to determine the trend of the features. For example, by comparing the color histograms of adjacent frames of images, the direction and degree of color change can be determined; by analyzing the changes in edge information, the motion and shape changes of objects can be understood. Specific analysis methods can use difference calculation, correlation analysis, etc. For example, the gray level difference of corresponding pixel points between adjacent frames of images is calculated, and if the difference is large, it indicates that there has been a significant change in that area between adjacent frames.
[0172] Optionally, according to the obtained feature change trend, the system can dynamically adjust the visual optimization rules. For example, if it is found that the color gradually becomes brighter between adjacent images, then the brightness adjustment amplitude can be appropriately reduced in subsequent optimization to avoid over-brightening; if it is found that the object gradually becomes clear in the image, then the sharpening parameter can be adjusted accordingly to make the optimized image more consistent with this change trend.
[0173] Optionally, for blur and sharpening processing of dynamic scenes: the system needs to first determine whether there is a dynamic scene in the image. This can be achieved by analyzing the feature change amplitude between adjacent images. If the feature change between adjacent images exceeds a certain threshold, it is considered that there is a dynamic change in the region. For example, when the motion of the object causes the pixel gray value difference of the corresponding region in the adjacent frame image to exceed the preset threshold, it is determined that there is a dynamic change in the region.
[0174] Once a dynamic scene is detected, the system further calculates the speed and direction of the object motion. Motion estimation techniques such as optical flow can be used to achieve this. The optical flow method calculates the displacement of corresponding pixel points in adjacent frame images to obtain the motion vector of the object, thereby determining the speed and direction of the motion.
[0175] Optionally, for objects with high motion speed, in order to simulate the real visual effect, blur processing is applied in the direction of motion. The faster the motion speed, the greater the degree of blur. This is because when the object moves quickly, the human eye will produce a visual persistence phenomenon, causing the imaging of the object on the retina to become blurred. For example, when processing an image of a fast-moving car, a certain degree of blur is added in the direction of the car's motion to make the picture look more realistic.
[0176] Optionally, for relatively static or slow-moving regions, or the edge parts of the object, sharpening processing is applied to enhance the clarity and details. This can highlight the contours and features of the object, making the picture more lively. For example, when processing an image containing moving people and static background, the edges of the background and the people are sharpened, while the moving part of the people is appropriately blurred, thereby enhancing the dynamic sense of the picture.
[0177] In an embodiment, by considering the time sequence information of the image and simulating the real visual effect, the optimized picture is more consistent with the visual perception of the human eye, enhancing the realism and dynamic sense of the picture; and the optimization rules can be dynamically adjusted according to the actual change of the image, suitable for various complex dynamic scenes, improving the flexibility and adaptability of visual optimization, providing users with a more lively and realistic visual experience, and having important application value in video playback, monitoring and other fields.
[0178] In addition, with reference to Figure 2The embodiment of the application also provides a control device Z10, which comprises
[0179] The acquisition module Z11 is configured to acquire image data in real time based on the camera and perform image preprocessing.
[0180] The extraction module Z12 is configured to perform multi-modal feature extraction on the preprocessed image and perform deep feature extraction on the image by using a pre-trained machine learning model; wherein the multi-modal features include color features, texture features and shape features; the machine learning model has a dynamic feature extraction architecture based on meta-learning, so as to dynamically adjust the network parameters of the model according to different image resolutions and / or complexities.
[0181] The fusion module Z13 is configured to fuse the extracted multi-modal features and deep learning features, and screen the fused features by using a feature selection algorithm.
[0182] The clustering module Z14 is configured to perform clustering analysis on the screened features by using a clustering algorithm.
[0183] The rule module Z15 is configured to formulate visual optimization rules for different feature types and feature value ranges according to the results of the feature analysis.
[0184] The optimization module Z16 is configured to perform visual optimization processing on the image according to the visual optimization rules, to generate an optimized real-time image.
[0185] Optionally, the control device Z10 can be a virtual control device (such as a virtual machine), or can be a physical device (such as a physical device that can execute the corresponding method except the visual optimization processing system).
[0186] In addition, the embodiment of the application also provides a visual optimization processing system, and the internal architecture of the visual optimization processing system can be as shown in Figure 3 The visual optimization processing system comprises a processor, a memory, a communication interface and an input interface connected by a system bus. The processor is configured to provide computing and control capabilities. The memory comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The database is configured to store data called by the computer program. The communication interface is configured to perform data communication with an external terminal. The input interface is configured to receive signals input by an external device. The computer program is executed by the processor to implement a visual optimization processing method based on real-time image processing as described in the above embodiment.
[0187] Those skilled in the art can understand that Figure 3The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the visual optimization processing system to which the scheme of the present application is applied. For example, in some optional embodiments, the visual optimization processing system can also include an output interface (not shown in the figure), and the output interface is also connected to the system bus and is used to output corresponding signals to peripherals.
[0188] In addition, the present application also provides a computer readable storage medium, which comprises a computer program, and the computer program is executed by a processor to implement the steps of the visual optimization processing method based on real-time image processing as described in the above embodiments. It can be understood that the computer readable storage medium in the present embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0189] To sum up, for the visual optimization processing method based on real-time image processing, the control device, the visual optimization processing system and the computer readable storage medium provided in the embodiments of the present application, in the feature extraction stage, not only multi-modal features are extracted, but also a machine learning model designed based on meta-learning is used for deep feature extraction, and the network parameters can be dynamically adjusted according to the image resolution and complexity, so that the image information can be comprehensively and flexibly captured; and the visual optimization rules are formulated based on the feature analysis results, so that the image can be accurately optimized for different feature types and ranges; finally, the image is processed according to the rules, which can effectively improve the image visual effect and the accuracy of subsequent analysis and processing, meet the requirements of real-time image optimization processing with high quality in actual application, and has stronger adaptability in complex and variable scenes.
[0190] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, databases, or other media in this application and in embodiments used herein, unless specifically stated otherwise, can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct RAM bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0191] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusions, so that processes, devices, articles or methods including a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such processes, devices, articles or methods. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of other identical elements in the process, device, article or method including the element.
[0192] The above description is only the preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation using the content of the specification and drawings, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.
Claims
1. A visual optimization processing method based on real-time image processing, characterized in that, The method comprises the following steps: Real-time image data is collected by a camera and image preprocessing is performed; Multi-modal feature extraction is performed on the preprocessed image, and a pre-trained machine learning model is used to extract deep features from the image; wherein the multi-modal features include color features, texture features and shape features; the machine learning model has a dynamic feature extraction architecture based on meta-learning, which dynamically adjusts the network parameters of the model according to the different resolutions and / or complexities of the image; The extracted multi-modal features and deep features are fused, and the fused features are filtered by a feature selection algorithm; wherein the color features and texture features are fused in the early stage of multi-modal feature extraction, and the shape features are fused with the color-texture features in the middle stage; in the shallow stage of deep feature extraction by the machine learning model, the shallow deep features are fused with the color features in the multi-modal features; in the middle stage of deep feature extraction by the machine learning model, the middle deep features are fused with the color-texture-shape features in the multi-modal features; after one fusion in different feature extraction stages, all features obtained after one fusion are fused again; Cluster analysis is performed on the filtered features using a clustering algorithm; wherein after the features of the current image are preliminarily clustered to obtain an initial clustering result, a dynamic clustering algorithm is executed based on the initial clustering result; the dynamic clustering algorithm is used to adjust the clustering result according to the dynamic changes of the object between the previous image and the current image, including adjusting the position of the clustering center according to the dynamic changes of the object; According to the results of feature clustering analysis, visual optimization rules for different feature types and feature value ranges are formulated; According to the visual optimization rules, the image is processed for visual optimization to generate an optimized real-time image.
2. The real-time image processing based visual optimization process method according to claim 1, wherein, Before the step of performing multi-modal feature extraction on the preprocessed image and using a pre-trained machine learning model to extract deep features from the image, the method further comprises: Multi-scale decomposition is performed on the preprocessed image to decompose the image into sub-band images of different scales and directions for subsequent feature extraction.
3. The real-time image processing based visual optimization process method according to claim 1, wherein, The visual optimization rules are also dynamically adjusted according to the real-time application scenarios of the image: When the image is used in a video conference scenario, the skin color and clarity of the human face are optimized, background interference is reduced, and the overall brightness and contrast of the image are improved; When the image is used in a security monitoring scenario, the edge clarity and detail information of the image are enhanced to improve the recognition of the target object; When the image is used in an advertising scenario, the brightness and saturation of the color are enhanced.
4. The real-time image processing based visual optimization method of claim 1, wherein, The real-time image processing-based visual optimization processing method further comprises: When formulating the visual optimization rules, user individual preference factors are introduced; Wherein, the system records the user's historical optimization operations and feedback on different types of images, learns the user's individual optimization preferences; in the subsequent image optimization process, according to the type of image currently processed by the user, the visual optimization rules are fine-tuned in combination with the user's individual preferences.
5. The real-time image processing based visual optimization method of claim 1, wherein, The visual optimization rule also considers the time sequence information of the image, analyzes the feature change trend between adjacent images for a continuously captured image sequence, and dynamically adjusts the visual optimization rule according to the trend; When there is a dynamic scene in the image, different blurring and sharpening processing is applied to images at different times according to the speed and direction of movement to simulate a real visual effect and enhance the dynamic feeling of the picture.
6. A control device characterized by comprising: It comprises: The acquisition module is configured to acquire image data in real time based on the camera and perform image preprocessing; The extraction module is configured to extract multi-modal features from the preprocessed image, and extract deep features from the image using a pre-trained machine learning model; wherein the multi-modal features include color features, texture features, and shape features; the machine learning model has a dynamic feature extraction architecture based on meta-learning, so as to dynamically adjust the network parameters of the model according to the image resolution and / or complexity; The fusion module is configured to fuse the extracted multi-modal features and deep features, and select the fused features through a feature selection algorithm; wherein the color features and texture features are fused in the early stage of multi-modal feature extraction, and the shape features are fused with the color-texture features in the middle stage; in the shallow stage of deep feature extraction by the machine learning model, the shallow deep features are fused with the color features in the multi-modal features; in the middle stage of deep feature extraction by the machine learning model, the middle deep features are fused with the color-texture-shape features in the multi-modal features; after one fusion in different feature extraction stages, all features obtained after one fusion are subjected to secondary fusion; The clustering module is configured to perform clustering analysis on the selected features using a clustering algorithm; wherein after the features of the current image are preliminarily clustered to obtain an initial clustering result, a dynamic clustering algorithm is executed based on the initial clustering result; the dynamic clustering algorithm is configured to adjust the clustering result according to the dynamic change of the object between the previous image and the current image, including adjusting the position of the clustering center according to the dynamic change of the object; The rule module is configured to formulate visual optimization rules for different feature types and feature value ranges according to the results of feature clustering analysis; The optimization module is configured to perform visual optimization processing on the image according to the visual optimization rules to generate an optimized real-time image.
7. A visual optimization processing system characterized by comprising: The visual optimization processing system comprises a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the real-time image processing-based visual optimization processing method according to any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer program is stored on the computer readable storage medium, and when executed by the processor, implements the steps of the real-time image processing-based visual optimization processing method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Remote sensing image classification method and system
CN117422936A
Image processing method based on AI algorithm
CN117934354A
Defect image enhancement method based on traditional data enhancement
CN118967474A
Image transmission method and system based on meta learning
CN120223907A