Visual optimization processing method and system based on real-time image processing

Through multimodal feature extraction, deep feature fusion and cluster analysis, the visual optimization processing method that dynamically adjusts model parameters solves the problems of insufficient real-time and flexibility in image processing in existing technologies, and achieves high-quality image optimization and target recognition.

CN120634893AActive Publication Date: 2025-09-12SHENZHEN QINUO TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511124476.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-09-12
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

Existing image processing methods have deficiencies in real-time performance, flexibility, and effectiveness, and are unable to fully capture image information in complex scenes, resulting in poor image quality and affecting target recognition and analysis.

Method used

A visual optimization processing method based on real-time image processing is adopted. Through multimodal feature extraction, deep feature fusion and cluster analysis, model parameters are dynamically adjusted, visual optimization rules are formulated, and high-quality optimized images are generated.

Benefits of technology

It improves the flexibility and quality of image processing, adapts to complex and changing application scenarios, and enhances the visual effects of images and the accuracy of subsequent analysis and processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634893A_ABST
    Figure CN120634893A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing technology, and discloses a visual optimization processing method and system based on real-time image processing, and the method comprises the steps: collecting image data in real time based on a camera, and carrying out the image preprocessing; carrying out multi-modal feature extraction on the preprocessed image, and carrying out deep feature extraction on the image by using a pre-trained machine learning model; fusing the extracted multi-modal features and deep learning features, and screening the fused features through a feature selection algorithm; performing clustering analysis on the screened features by using a clustering algorithm; according to a feature analysis result, making a visual optimization rule for different feature types and feature value ranges; and performing visual optimization processing on the image according to the visual optimization rule to generate an optimized real-time image. The invention further discloses a control device and a computer readable storage medium. The invention aims to perform visual optimization processing on the real-time image so as to improve the processing flexibility and the image quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a visual optimization processing method, a control device, a visual optimization processing system and a computer-readable storage medium based on real-time image processing. Background Art

[0002] With the widespread adoption of camera devices and the continuous improvement of their performance, acquiring real-time image data has become increasingly easier. However, due to the complexity and diversity of practical application scenarios, the captured images often suffer from various quality issues, such as noise, uneven lighting, and blur. These issues seriously affect the visual quality and subsequent analysis and processing. Therefore, effective visual optimization processing of real-time captured images is of great practical significance.

[0003] Traditional image processing methods primarily focus on extracting and processing a single feature, such as color or texture. This approach is limited in that it fails to fully capture the rich information in an image, resulting in poor performance when processing images in complex scenarios. For example, in security surveillance scenarios, when encountering dramatic lighting changes or occlusions, relying solely on color features for image enhancement can result in the loss of other important image information, hindering target recognition and analysis.

[0004] Existing image processing methods also have many deficiencies in terms of real-time performance, flexibility, and effectiveness, and are unable to meet the requirements for high-quality, real-time image optimization processing in practical applications. Therefore, a more efficient, flexible, and comprehensive visual optimization processing method is needed that can dynamically adjust the processing strategy based on the real-time characteristics of the image, improve the quality of image optimization, and adapt to various complex and changing application scenarios.

[0005] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention

[0006] The main purpose of this application is to provide a visual optimization processing method, a control device, a visual optimization processing system and a computer-readable storage medium based on real-time image processing, aiming to perform visual optimization processing on real-time images to improve processing flexibility and image quality.

[0007] To achieve the above objectives, the present application provides a visual optimization processing method based on real-time image processing, comprising the following steps: Collect image data in real time based on the camera and perform image preprocessing; Perform multimodal feature extraction on the preprocessed image and perform deep feature extraction on the image using a pretrained machine learning model. The multimodal features include color, texture, and shape features. The machine learning model is designed with a dynamic feature extraction architecture based on meta-learning to dynamically adjust the model's network parameters based on image resolution and / or complexity. The extracted multimodal features and deep learning features are fused, and the fused features are screened through feature selection algorithms; Use clustering algorithm to perform cluster analysis on the filtered features; Based on the results of feature analysis, formulate visual optimization rules for different feature types and feature value ranges; According to the visual optimization rules, the image is visually optimized to generate an optimized real-time image.

[0008] To achieve the above objectives, the present application further provides a control device, comprising: The acquisition module is used to collect image data in real time based on the camera and perform image preprocessing; An extraction module is used to extract multimodal features from preprocessed images and to extract deep features from images using a pretrained machine learning model. The multimodal features include color, texture, and shape features. The machine learning model is designed with a dynamic feature extraction architecture based on meta-learning to dynamically adjust the model's network parameters based on image resolution and / or complexity. The fusion module is used to fuse the extracted multimodal features and deep learning features, and screen the fused features through the feature selection algorithm; Clustering module, used to perform cluster analysis on the filtered features using clustering algorithms; The rule module is used to formulate visual optimization rules for different feature types and feature value ranges based on the results of feature analysis; The optimization module is used to perform visual optimization processing on the image according to visual optimization rules to generate an optimized real-time image.

[0009] To achieve the above-mentioned purpose, the present application also provides a visual optimization processing system, which includes: a memory, a processor, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, the steps of the visual optimization processing method based on real-time image processing as described above are implemented.

[0010] To achieve the above objectives, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the visual optimization processing method based on real-time image processing are implemented.

[0011] The visual optimization processing method, control device, visual optimization processing system and computer-readable storage medium based on real-time image processing provided by this application not only extract multimodal features in the feature extraction stage, but also use a machine learning model based on meta-learning design to perform deep feature extraction, and can dynamically adjust network parameters according to image resolution and complexity, comprehensively and flexibly capturing image information; and, based on the feature analysis results, formulate visual optimization rules, which can accurately optimize for different feature types and ranges; finally, process the image according to the rules, which can effectively improve the image visual effect and the accuracy of subsequent analysis and processing, meet the requirements of practical applications for high-quality real-time image optimization processing, and have stronger adaptability in complex and changing scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 This is a schematic diagram of the steps of a visual optimization processing method based on real-time image processing in one embodiment of the present application; Figure 2 This is a schematic diagram of a control device in an embodiment of the present application; Figure 3 Schematic diagram of the internal architecture of a visual optimization processing system according to an embodiment of the present application.

[0013] The realization of the objectives, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0014] The embodiments of the present application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be understood as limiting the present application. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application without making any creative efforts shall fall within the scope of protection of the present application.

[0015] In addition, any descriptions of "first," "second," etc., in this application are for descriptive purposes only (e.g., to distinguish identical or similar features) and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined as "first" or "second" may explicitly or implicitly include at least one such feature. Furthermore, the technical solutions of various embodiments may be combined with each other, but this must be based on the ability of a person of ordinary skill in the art to implement them. If the combination of technical solutions contradicts or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.

[0016] Reference Figure 1In one embodiment, a visual optimization processing method based on real-time image processing includes: Step S10: collecting image data in real time based on the camera and performing image preprocessing; Step S20: performing multimodal feature extraction on the preprocessed image, and performing deep feature extraction on the image using a pretrained machine learning model; wherein the multimodal features include color features, texture features, and shape features; the machine learning model is designed with a dynamic feature extraction architecture based on meta-learning to dynamically adjust the model's network parameters according to the image resolution and / or complexity; Step S30: Fusing the extracted multimodal features and deep learning features, and screening the fused features through a feature selection algorithm; Step S40: performing cluster analysis on the filtered features using a clustering algorithm; Step S50: formulating visual optimization rules for different feature types and feature value ranges based on the results of feature analysis; Step S60: Perform visual optimization processing on the image according to the visual optimization rules to generate an optimized real-time image.

[0017] In this embodiment, the execution terminal of the embodiment may be a visual optimization processing system, or may be other equipment or devices (such as a control device) that controls the visual optimization processing system.

[0018] As described in step S10, a camera is used to continuously capture image data of a real scene at a certain frame rate (e.g., 30 frames per second, 60 frames per second, etc.). The type of camera can be selected according to the specific application scenario, such as a common RGB camera, a depth camera, etc.

[0019] In order to improve the effect and efficiency of subsequent processing, the collected raw image data needs to be preprocessed. Optional preprocessing operations include: (1) Grayscale conversion: Converting color images into grayscale images reduces the amount of data and also helps extract some features based on grayscale information; (2) Filtering: Use filters (such as Gaussian filters, median filters, etc.) to remove noise from the image and make the image smoother; (3) Normalization: Normalize the pixel values ​​of the image, usually scaling the pixel values ​​to the range of [0, 1] or [0, 255] to ensure consistency between different images.

[0020] As described in step S20, feature extraction is performed on the image from different angles to obtain more comprehensive image information. Multimodal features include color features, texture features, and shape features.

[0021] Among them, color features can be extracted by calculating the color histogram and color moment of the image. The color histogram reflects the distribution of different colors in the image, and the color moment can describe the statistical characteristics of the color such as mean, variance and skewness.

[0022] Optional texture feature extraction methods include gray-level co-occurrence matrix and local binary pattern (LBP). The gray-level co-occurrence matrix describes the gray-level spatial relationship between pixels in an image, while LBP extracts texture information by comparing the gray-level values ​​of the central pixel with those of its neighboring pixels.

[0023] Shape features can be extracted by calculating geometric features such as the perimeter, area, and circularity of objects in the image. For complex shapes, methods such as contour analysis and Fourier descriptors can also be used to describe them.

[0024] Optionally, use a pre-trained machine learning model for deep feature extraction, which has a dynamic feature extraction architecture designed based on meta-learning.

[0025] Alternatively, use convolutional neural networks (CNNs) pre-trained on large-scale image datasets (such as ImageNet), such as VGG and ResNet. These models have learned rich image feature representation capabilities on large-scale data.

[0026] Meta-learning, also known as "learning to learn," aims to enable models to quickly learn and adapt to diverse tasks using small amounts of data. In this vision optimization scenario based on real-time image processing, meta-learning is used to design a feature extraction architecture that dynamically adjusts network parameters based on image properties (resolution and / or complexity). This allows the model to maintain efficient and accurate feature extraction across a wide range of image conditions.

[0027] When the input image has low resolution or low complexity, the model can reduce the number of convolutional kernels in the convolutional layer, reduce the size of the convolution kernels, or reduce the number of network layers to reduce the amount of computation and the number of parameters. When the input image has high resolution or high complexity, the model increases the corresponding computing resources to extract richer feature information. For example, a controller network trained through meta-learning can dynamically adjust the hyperparameters of the CNN model, such as the stride and number of channels of the convolutional layer, based on the characteristics of the input image (such as resolution and variance).

[0028] Image resolution directly affects the level of detail that can be captured in an image. High-resolution images contain more detail, requiring a more complex feature extraction process; low-resolution images are relatively simpler. Therefore, image resolution is input as metadata into the meta-learning module, which can adjust the complexity of the feature extraction architecture based on the resolution.

[0029] Image complexity can be measured in many ways, such as the richness of image texture, the diversity of color distribution, the number of objects, and the complexity of shapes. Image complexity can be evaluated using some predefined metrics or through another small neural network, and the evaluation results can be passed to the meta-learning module as metadata.

[0030] Optionally, a controller network is designed as the core component of meta-learning. This controller network receives image metadata (resolution, complexity, etc.) as input and outputs a series of control signals that dynamically adjust the network parameters of the feature extraction model. For example, the controller network can decide whether to enable certain convolutional layers, adjust the size and number of convolution kernels, or change the parameters of the pooling layer based on image resolution and complexity.

[0031] Optionally, during the meta-training phase, the controller network is trained using a large amount of image data of varying resolutions and complexities. During training, the controller network attempts to find the best parameter tuning strategy to ensure that the feature extraction model achieves optimal feature extraction under a variety of image conditions. Reinforcement learning or gradient descent methods are typically used to optimize the controller network's parameters.

[0032] Optionally, the specific method of dynamically adjusting network parameters includes adjusting convolution layer parameters and / or pooling layer parameters.

[0033] Among them, the convolution layer parameters include the number of convolution kernels, the size of the convolution kernel and the depth of the convolution layer.

[0034] Alternatively, for low-resolution or less complex images, reducing the number of convolution kernels can reduce computational complexity and model complexity. For example, a convolution layer that originally used 64 convolution kernels on a high-resolution image can be reduced to 32 when processing a low-resolution image.

[0035] Optionally, adjust the convolution kernel size based on the level of detail in the image. For high-resolution images that need to capture more local details, a larger convolution kernel (such as 5x5 or 7x7) can be used; for low-resolution images, a smaller convolution kernel (such as 3x3) may be sufficient.

[0036] Optionally, in some cases, you can decide whether to enable certain convolutional layers based on image properties. For simple images, you can skip some deep convolutional layers to reduce the amount of computation; for complex images, you use the full convolutional layer structure to extract richer features.

[0037] Among them, the pooling layer parameters include pooling window size and pooling type.

[0038] Pooling can reduce the size of feature maps and reduce computational complexity. For high-resolution images, a larger pooling window (such as 2x2 or 3x3) can be used; for low-resolution images, a smaller pooling window (such as 1x1 or no pooling) can be used to retain more detailed information.

[0039] Optionally, you can choose different pooling types based on the image, such as max pooling and average pooling. Max pooling focuses on extracting prominent features in the image, while average pooling emphasizes smoothness. Max pooling may be more appropriate for images with rich textures, while average pooling may be more effective for images with uniform colors.

[0040] On this basis, the preprocessed image is input into a pretrained CNN model. Through operations such as convolutional layers and pooling layers, the image's deep features are gradually extracted. The output of the intermediate or final layer of the CNN model is selected as the image's deep feature vector. These deep feature vectors are highly abstract and representative, and can well describe the image's semantic information. Furthermore, by dynamically adjusting network parameters based on image resolution and complexity, the same complex model is avoided for all images, reducing unnecessary computation and improving real-time processing efficiency. Furthermore, the feature extraction process can be optimized for different image conditions, allowing the model to extract more accurate and representative features for a variety of images, thereby improving the effectiveness of subsequent visual optimization processing.

[0041] As described in step S30, after completing step S20, the image's multimodal features (including color, texture, and shape features) and deep features extracted by the pre-trained machine learning model have been obtained. These features describe the image information from different perspectives. To comprehensively utilize this information, they need to be fused.

[0042] Optionally, considering that different types of features may have different importance to the final result, different weights can be assigned to multimodal features and deep learning features, and then the weighted features can be added together to form a fused feature. The weights can be initially set based on the experience of relevant engineers and then optimized through model training (i.e., optimized based on prior knowledge learned by the model).

[0043] The fused feature vector may contain a large number of features, some of which may not be useful for subsequent clustering analysis and visual optimization rule formulation, and may even increase computational complexity and noise. Therefore, a feature selection algorithm is needed to filter the fused features and retain the most representative and discriminative features.

[0044] Optionally, you can evaluate feature importance based on its variance. A larger variance indicates a greater variability in the feature's values ​​across samples, making it more likely to be helpful for classification or clustering. For numerical features, you can calculate the variance of each feature and select features with variances greater than a threshold.

[0045] For non-numerical features, an L1 regularization term can be added to the regression model to reduce the coefficients of unimportant features to 0, thereby achieving feature selection. It should be understood that by adjusting the regularization parameter, the number of features selected can be controlled.

[0046] As described in step S40, a suitable clustering algorithm (such as K-means clustering or hierarchical clustering) is selected to perform cluster analysis on the filtered features. The purpose of clustering is to group image data points with similar features into the same category, so that images of different categories can be processed in a targeted manner.

[0047] Taking K-means clustering as an example, first randomly select K cluster centers, then assign each data point to the category of the cluster center closest to it, then update the position of the cluster center, and repeat this process until the cluster center no longer changes or the maximum number of iterations is reached.

[0048] As described in step S50 , after completing the feature cluster analysis (step S40 ), corresponding visual optimization rules are formulated for different feature types (such as color features, texture features, shape features, etc.) and feature value ranges based on the analysis results.

[0049] After the cluster analysis is completed in step S40, different feature clusters can be obtained, and the features in each cluster have similarities. First, these features need to be carefully sorted to clarify different feature types and their corresponding feature value ranges.

[0050] Optionally, features may be classified according to types such as color, texture, and shape. For example, color features may include RGB values, HSV values, etc.; texture features may include roughness, contrast, etc.; shape features may include perimeter, area, aspect ratio, etc.

[0051] Optionally, for each feature type, calculate the range of its feature values ​​across clusters. You can calculate statistics such as minimum, maximum, mean, and standard deviation to gain a more comprehensive understanding of the feature distribution. For example, the red channel value in a color feature might be in the range [100, 200] within a cluster.

[0052] Among them, the formulation of color feature optimization rules: Optionally, if color features indicate low brightness in certain areas of the image, rules can be developed to increase the brightness of these areas. For example, when the brightness value falls below a certain threshold (such as 50, with a range of 0-255), the RGB value of that area is increased by a certain ratio (such as 1.2 times). For areas with low color contrast, the contrast can be enhanced by adjusting the color saturation and brightness difference. For example, the color difference between adjacent pixels can be calculated. If the difference is less than a set threshold, the saturation and brightness difference can be appropriately increased.

[0053] Optionally, if color analysis reveals a color cast in the image (e.g., an overall reddish tint), the values ​​of other color channels can be adjusted based on the degree and direction of the cast. For example, if the red channel values ​​are generally too high, the red channel values ​​can be appropriately lowered while the green and blue channel values ​​can be increased. For areas with less vivid colors, the saturation value can be increased based on the characteristics of the HSV color space. For example, if the saturation is below 30%, the saturation value can be increased to 50%.

[0054] Among them, the formulation of texture feature optimization rules: Optionally, if the texture features indicate blurry textures in certain areas, you can use a sharpening filter algorithm (such as the Laplacian operator or the Sobel operator) to enhance the texture's edge information and improve clarity. For areas with excessively coarse textures, you can use a smoothing filter algorithm (such as the Gaussian filter or the mean filter) to refine the texture.

[0055] Optionally, when the texture features of adjacent regions differ greatly, interpolation or fusion methods can be used to make the texture transition between these regions more natural. For example, a bilinear interpolation algorithm can be used to process the texture boundaries.

[0056] Among them, the formulation of shape feature optimization rules: Optionally, if shape features indicate an object's shape is missing, it can be supplemented based on the object's overall shape and contextual information. For example, if a circular object's edges are missing, it can be repaired by fitting based on its center and radius information. For deformed objects, corrections can be made based on the original shape's characteristics (such as aspect ratio and symmetry). For example, a tilted rectangle can be corrected to a regular rectangle.

[0057] Optionally, for important shapes (such as logos and key objects), their edges can be enhanced or their colors adjusted to make them stand out. For example, the outline color of important shapes can be set to a striking color (such as yellow). For overly complex shapes, the representation can be simplified by removing some small edges and details. For example, irregular polygons can be approximated using polygons to reduce the number of vertices.

[0058] Optionally, after developing a basic optimization strategy, you can further refine the rules based on the eigenvalue range to achieve more precise optimization: divide the eigenvalue range into different intervals and develop different optimization parameters for each interval. For example, for the brightness feature, when the brightness value is in the range [0, 50], increase the brightness by 30%; when the brightness value is in the range [51, 100], increase the brightness by 20%.

[0059] Dynamically adjust the optimization rules by taking into account the distribution density and changing trends of eigenvalues. For example, in areas where eigenvalues ​​vary dramatically, more refined optimization parameters are used; in areas where eigenvalues ​​are more evenly distributed, a relatively unified optimization strategy is adopted.

[0060] Through the above steps, a comprehensive, detailed and effective set of visual optimization rules for different feature types and feature value ranges can be formulated based on the results of feature analysis.

[0061] As described in step S60, before image optimization, the visual optimization rules developed in step S50 need to be parsed and converted into specific operational parameters. Depending on the feature type targeted by the rule (e.g., color, texture, shape, etc.), the corresponding rule is retrieved. For example, for color optimization rules, information such as brightness adjustment ratio and color balance correction parameters is extracted.

[0062] Convert the abstract descriptions in the rules into specific numerical parameters. For example, if the rule states "When the brightness value is lower than 50, increase the brightness by 30%", in actual operation, the parameters such as 50 and 30% need to be extracted for subsequent calculations.

[0063] In order to apply the optimization rules more accurately, the image needs to be partitioned and the features of each region need to be matched with the rules.

[0064] Optionally, the image can be divided into different regions based on the clustering results of step S40, with similar features within each region. Alternatively, a fixed grid partitioning method can be used to divide the image into several small blocks. For each partition, its color, texture, shape, and other features are extracted and matched with the feature type and feature value range specified in the rule to determine the specific optimization rule applicable to that partition. For example, if the brightness value of a partition falls within the low brightness area defined in the rule, the corresponding brightness boost rule is applied.

[0065] Among them, the color feature is optimized: Optionally, the brightness of the image partition is adjusted according to the brightness adjustment ratio set in the rule. This can be achieved by modifying the RGB values ​​of the image. For example, for a low-brightness area, the RGB value of each pixel in the area is multiplied by the brightness adjustment coefficient.

[0066] Optionally, methods such as histogram equalization and adaptive histogram equalization can be used to enhance the contrast of the image. These methods can adjust the brightness distribution of the image to make the bright parts brighter and the dark parts darker, thereby improving the contrast.

[0067] Optionally, the color channel values ​​of the image are adjusted according to the color correction parameters determined in the rule. For example, if the image is found to be reddish, the value of the red channel is reduced, while the values ​​of the green and blue channels are appropriately increased.

[0068] Optionally, adjust the image's saturation, hue, and other parameters to make the image's colors more vivid. You can use the HSV color space to increase the saturation value to enhance the vividness of the colors.

[0069] Among them, for texture feature optimization: Optionally, convolve the image using a sharpening filter (such as a Laplacian filter or a Sobel filter) to enhance edge information and make the texture clearer. If the texture features indicate that the image is noisy, use methods such as Gaussian filtering and median filtering to remove the noise and make the texture smoother.

[0070] Optionally, when there are large differences in textures between adjacent regions, interpolation or fusion methods can be used to make the texture transition more natural. For example, a bilinear interpolation algorithm can be used to process texture boundaries.

[0071] Among them, for shape feature optimization: Optionally, the missing parts of the object shape can be supplemented according to the shape repair strategy specified in the rule. For example, the missing edge of a circular object can be repaired by fitting it based on the center and radius information.

[0072] Optionally, the deformed object shape is corrected by affine transformation, perspective transformation, etc. For example, a tilted rectangle can be corrected to a regular rectangle.

[0073] Optionally, make important shapes stand out more by enhancing their edges, adjusting their colors, etc. For example, set the outline color of important shapes to a striking color (such as yellow).

[0074] Optionally, algorithms such as polygon approximation and contour simplification are used to remove small edges and details in the shape to make the shape more concise.

[0075] After optimizing each partition, you need to merge them to create a complete optimized image. This involves stitching the optimized partitions back together to create a complete image. During the partition stitching process, discontinuities may occur at the edges. This can be achieved by smoothing the edges to create a more natural-looking image.

[0076] Through the above steps, the image can be comprehensively and carefully optimized according to the visual optimization rules to generate a high-quality optimized real-time image.

[0077] In one embodiment, during the feature extraction stage, not only multimodal features are extracted, but also deep feature extraction is performed using a machine learning model designed based on meta-learning. The network parameters can be dynamically adjusted according to the image resolution and complexity, so as to comprehensively and flexibly capture image information. Furthermore, visual optimization rules are formulated based on the feature analysis results, which can be accurately optimized for different feature types and ranges. Finally, the image is processed according to the rules, which can effectively improve the image visual effect and the accuracy of subsequent analysis and processing, meet the requirements of practical applications for high-quality real-time image optimization processing, and have stronger adaptability in complex and changing scenarios.

[0078] In one embodiment, based on the above embodiment, before the steps of performing multimodal feature extraction on the preprocessed image and performing deep feature extraction on the image using a pretrained machine learning model, the method further includes: The preprocessed image is subjected to multi-scale decomposition to decompose the image into sub-band images of different scales and directions for subsequent feature extraction.

[0079] In this embodiment, multi-scale decomposition is performed before multimodal feature extraction and deep feature extraction on the preprocessed image to more comprehensively and meticulously analyze the image's feature information. Images of different scales can reflect the image's features at different resolutions, while sub-band images in different orientations can capture the image's texture and structural information in various directions. This facilitates more accurate subsequent extraction of multimodal and deep features, thereby improving the performance of the entire visual optimization method.

[0080] Optional multi-scale decomposition methods include wavelet transform, Laplace pyramid decomposition, etc. Taking wavelet transform as an example, it is a time-frequency analysis method that has good localization characteristics in both the time domain and the frequency domain, and is suitable for multi-scale decomposition of images.

[0081] The preprocessed image is decomposed at different scales using the selected decomposition method. Generally, the image is decomposed from its original scale into multiple sub-images of varying scales. For example, in a wavelet transform, after one level of decomposition, the image is decomposed into four sub-bands: a low-frequency sub-band (LL) and three high-frequency sub-bands (LH, HL, and HH). The low-frequency sub-band represents the smooth portion of the image and contains the primary contour information, while the high-frequency sub-bands represent horizontal, vertical, and diagonal details, respectively.

[0082] Optionally, multiple levels of decomposition can be performed, where each level further decomposes the low-frequency subbands obtained at the previous level. This allows for the generation of subband images at different scales, each containing specific information about the image at that scale.

[0083] Optionally, in addition to decomposing the image in terms of scale, the image can also be decomposed in terms of direction. In the wavelet transform, different filter combinations can be used to generate high-frequency sub-band images in the horizontal, vertical, and diagonal directions. These sub-band images in different directions can capture the texture and edge information of the image in various directions. For example, high-frequency sub-band images in the horizontal direction can highlight the horizontal edges and textures in the image, while high-frequency sub-band images in the vertical direction can highlight the vertical edges and textures.

[0084] The role of the sub-band image after multi-scale decomposition in subsequent feature extraction: (1) For color features, sub-band images of different scales can reflect the distribution of colors at different resolutions. For example, in a larger-scale sub-band image, the overall distribution trend of colors can be analyzed; while in a smaller-scale sub-band image, the detailed changes in colors can be focused on. For texture features, sub-band images of different directions can more clearly show the texture information of the image in various directions, which helps to extract texture features more accurately. For shape features, sub-band images of different scales can provide shape contour information of different precisions, and combined together, they can more comprehensively describe the shape of the object in the image.

[0085] (2) The pre-trained machine learning model can use the sub-band images after multi-scale decomposition as input to perform deep feature extraction. Sub-band images of different scales and orientations contain rich information, and the model can learn the associations and feature representations between these information. For example, the model can learn the hierarchical structural features of the image from sub-band images of different scales, and the directional features of the image from sub-band images of different orientations. By performing deep feature extraction on these sub-band images, more representative and discriminative deep features can be obtained, thereby improving the effect of subsequent visual optimization processing.

[0086] In one embodiment, multi-scale decomposition is an important step before feature extraction on pre-processed images. By decomposing the image into sub-band images of different scales and orientations, it provides richer and more comprehensive information for subsequent multimodal feature extraction and deep feature extraction. This enables the entire visual optimization method to more accurately capture image features, thereby achieving better visual optimization results.

[0087] In one embodiment, based on the above embodiment, the visual optimization processing method based on real-time image processing further includes: When fusing multimodal features and deep learning features, a multi-level fusion approach is adopted. Feature fusion is performed separately at different feature extraction stages, and then all features obtained after the first fusion are fused for the second time.

[0088] In this embodiment, during the multimodal feature extraction process, color features and texture features are often closely related in different local areas and scales of the image. For example, in some images, specific color areas may be accompanied by specific texture patterns. In the early stages of extracting these features, color features and texture features can be fused. For example, while performing color histogram statistics based on a local area of ​​the image, the texture energy features of the area are combined and the two are linearly combined according to certain weights to form a new fused feature vector. Such fusion can better describe the comprehensive features of the local area of ​​the image.

[0089] Mid-term fusion of shape features with other features: After the shape features are initially extracted, they can be fused with the already fused color-texture features in the mid-term. Shape features typically reflect the outline of objects in the image, while color-texture features describe the properties of the object's surface. By fusing shape features with color-texture fusion features, a more representative feature combination can be obtained. For example, the moment invariant of the shape feature can be concatenated with the color-texture fusion feature vector to form a new feature vector that contains information about the object's outline and surface properties.

[0090] Fusion of shallow deep learning features with multimodal features: During the shallow stages of deep feature extraction in a pre-trained machine learning model, the features output by the model typically contain some basic edge and local structural information from the image. These shallow deep learning features can then be fused with some multimodal features (such as color). For example, the feature map output by a shallow convolutional layer can be element-wise added or multiplied with the color feature map to generate a new fused feature map. This fusion combines the deep learning model's sensitivity to image structure with the color feature's ability to describe the image's visual appearance.

[0091] Further integration of mid-level deep learning features with multimodal features: As the deep learning model progresses, mid-level features begin to capture more abstract semantic information about the image. At this stage, mid-level deep learning features can be integrated with the already integrated color, texture, and shape features from the multimodal features. For example, the feature vector output by the mid-level convolutional layer can be weighted and summed with the multimodal fused feature vector to produce a richer fused feature vector that combines the detailed information of the multimodal features with the semantic information of the mid-level deep learning features.

[0092] After completing the first fusion of different feature extraction stages, all the features obtained from the first fusion are fused again. The purpose of the second fusion is to integrate the information contained in each first fusion feature, further eliminate redundancy, and enhance the representativeness and discrimination of the features.

[0093] Optionally, all fused feature vectors can be concatenated to form a longer feature vector. Then, based on the importance and relevance of each feature, different weights are assigned to each feature and a weighted sum is performed. For example, features that are more discriminative in image classification tasks can be given higher weights, while redundant or noisy features can be given lower weights.

[0094] Optionally, nonlinear transformations can be introduced during the secondary fusion process, such as using a fully connected layer of a neural network to perform nonlinear mapping on the concatenated feature vectors. This allows the model to learn the complex nonlinear relationships between different features, further improving the quality of the fused features.

[0095] Multi-level fusion can fully exploit the information of multimodal features and deep learning features at different extraction stages, making the fused features more comprehensive and accurate in describing the characteristics of the image. This helps improve the accuracy of subsequent feature screening, clustering analysis, and visual optimization rule formulation.

[0096] Different images may exhibit different feature advantages at different feature extraction stages. Multi-level fusion can flexibly fuse features at different stages based on the specific conditions of the image, thereby enhancing the adaptability of the entire visual optimization processing method to different images.

[0097] The high-quality features obtained through multi-level fusion can provide a more reliable basis for the subsequent formulation of visual optimization rules. The optimization rules formulated based on these features can more accurately optimize the different features of the image, thereby improving the visual quality and clarity of the image.

[0098] In one embodiment, a multi-level fusion approach provides a more efficient feature integration method for visual optimization in real-time image processing by fusing features separately at different feature extraction stages and then performing a secondary fusion. This method fully leverages the advantages of multimodal features and deep learning features, improving feature expressiveness and model adaptability, ultimately optimizing image visual effects.

[0099] In one embodiment, based on the above embodiment, the clustering algorithm includes a dynamic clustering algorithm, and the dynamic clustering algorithm is used to adjust the clustering result according to the dynamic changes of the object between the previous image and the current image.

[0100] In this embodiment, before processing the current image, it is necessary to track the object features in the previous image. This can be achieved using feature tracking algorithms such as optical flow and Kalman filtering. Optical flow can calculate the motion vector of each pixel in the image between adjacent frames, thereby determining the object's direction and speed of motion. Kalman filtering can predict and update the object's motion state, improving feature tracking accuracy.

[0101] Alternatively, dynamic changes in an object can be detected by comparing its features in the previous image with the current image. These changes may include movement, deformation, or color change. For example, the change in the center of gravity of an object between two frames can be calculated to determine whether the object has shifted, or the change in the object's outline shape can be calculated to determine whether the object has deformed.

[0102] Before executing the dynamic clustering algorithm on the current image, the features of the current image are first preliminarily clustered using traditional clustering algorithms (such as K-means clustering, hierarchical clustering, etc.). After obtaining the initial clustering result (this initial result is obtained based on the static features of the current image), the dynamic clustering algorithm is then executed.

[0103] Optionally, adjust the positions of cluster centers based on the dynamic changes of objects. If an object moves, the cluster center containing that object may need to be moved to the object's new location. For example, if an object moves from one cluster area to the vicinity of another cluster area, the positions of the two cluster centers can be adjusted appropriately to better adapt to the new distribution of objects.

[0104] Optionally, check whether the dynamic changes of each object make it more suitable to belong to a different cluster. If the characteristics of an object have changed significantly in the current image, resulting in a decrease in similarity with the current cluster and an increase in similarity with other clusters, then the object can be reassigned to a more appropriate cluster. For example, if the color of an object has changed significantly between two frames, making the color characteristics of the original cluster more different and closer to the color characteristics of another cluster, then the object can be removed from the original cluster and added to the new cluster.

[0105] Optionally, clusters may need to be merged or split depending on the dynamic changes of objects. If the objects in two clusters become increasingly similar due to dynamic changes, the two clusters can be merged into one; if the objects in a cluster become significantly different due to dynamic changes, the cluster can be split into multiple subclusters.

[0106] Alternatively, by taking into account the dynamic changes of objects, dynamic clustering algorithms can more accurately cluster objects in an image, avoiding clustering errors caused by object movement or changes. This helps to more accurately identify different object categories in an image and provides a more reliable basis for the subsequent formulation of visual optimization rules.

[0107] In real-time image processing, image content is constantly changing. Dynamic clustering algorithms can adjust clustering results in real time to adapt to these changes, ensuring the real-time and effectiveness of visual optimization processing. For example, in a video surveillance scenario, when objects enter or leave the monitoring area, dynamic clustering algorithms can promptly adjust clustering results to accurately cluster and analyze the newly appeared objects.

[0108] In this way, it can adapt to the dynamic changes of objects in the image and accurately perform clustering analysis in different dynamic scenes, thereby improving the robustness and adaptability of the clustering algorithm; more accurate clustering results help to formulate more reasonable visual optimization rules, thereby improving the effect of visual optimization processing and making the optimized image more in line with actual needs; in real-time image processing, the clustering results can be adjusted in real time, ensuring the real-time performance of the entire visual optimization processing process, which is suitable for application scenarios with high real-time requirements.

[0109] In one embodiment, based on the above embodiment, the visual optimization rules are further dynamically adjusted according to the real-time application scenario of the image: When images are used in video conferencing scenarios, focus on optimizing facial skin tone and clarity, reducing background interference, and improving the overall brightness and contrast of the image. When images are used in security monitoring scenarios, they enhance image edge clarity and detail information, improving the recognition of target objects. When the image is used in advertising, enhance the vividness and saturation of the colors.

[0110] In this embodiment, during the visual optimization process of real-time image processing, different application scenarios have different image requirements. Optimizing solely based on image features is insufficiently comprehensive, as the same image may require significantly different aspects to be highlighted and optimized in different scenarios. Therefore, dynamically adjusting visual optimization rules based on the image's real-time application scenario allows the optimized image to better meet the needs of specific scenarios, enhancing visual quality and application value.

[0111] In video conferencing, facial information is crucial for communication. To ensure participants can clearly see each other's expressions and demeanor, it's crucial to optimize facial skin tone and clarity. Color correction algorithms can be used to adjust facial skin tone for a more natural and realistic look. For example, a skin tone model can be used to detect facial areas and then fine-tune the color of that area based on preset skin tone standards. Image sharpening technology can also be used to enhance facial edges and details, improving clarity and making facial features more distinct.

[0112] During video conferencing, background distractions can be distracting. Therefore, reducing background interference is crucial. Background blur algorithms can blur the background, focusing the viewer's attention on the subject's face. Alternatively, background replacement technology can replace a cluttered background with a simple, unified background, such as a solid color or a virtual background.

[0113] Optionally, appropriate brightness and contrast can make images clearer and more vivid, enhancing the viewing experience. By adjusting the brightness and contrast parameters of an image, details of both the person's face and the surrounding environment can be clearly displayed. For example, histogram equalization or adaptive histogram equalization can be used to enhance image contrast, while adjusting brightness based on ambient light conditions to avoid overly bright or dark images.

[0114] Accurately identifying target objects is crucial in security surveillance. To improve object recognition, it's necessary to enhance image edge clarity and detail. Edge detection algorithms, such as the Sobel operator and the Canny operator, can be used to highlight image edges and make object outlines clearer. Image enhancement techniques, such as Laplace operator enhancement and wavelet transform enhancement, can also be used to enhance image detail and make the features of the target object more distinct.

[0115] In addition to enhancing edges and details, the system can also adjust image parameters such as color and contrast to make the target object stand out more clearly. For example, for targets that may appear dark during nighttime surveillance, their brightness and contrast can be increased to make them easier to identify. Furthermore, real-time analysis and processing of surveillance images are performed, classifying and labeling target objects based on their characteristics (such as shape, size, and color), making it easier for monitoring personnel to quickly identify and locate them.

[0116] The goal of advertising is to capture the viewer's attention, and vibrant, highly saturated images are particularly effective in achieving this goal. Color enhancement algorithms adjust image colors to enhance their vividness and saturation. For example, using a color lookup table (LUT) technique, the colors in an image are mapped to a more vivid and saturated color space. Simultaneously, the overall image hue is adjusted to better align with the advertisement's theme and style, creating a more engaging visual experience.

[0117] To dynamically adjust visual optimization rules based on real-time application scenarios, a real-time scene determination mechanism can be established. This can be determined based on image metadata (such as the capture device, capture time, and capture location), user-entered scene information, or intelligent analysis algorithms. Once the scene changes, the system can quickly switch to the corresponding visual optimization rules and optimize the image in real time, ensuring that the optimized image always meets the requirements of the current scenario.

[0118] In one embodiment, specific visual optimization rules are formulated according to different application scenarios, which can highlight key information of images in different scenarios, meet the special needs of each scenario, and improve the pertinence and effectiveness of visual optimization.

[0119] In video conferencing scenarios, the optimized images allow participants to communicate more smoothly; in security monitoring scenarios, they help to more accurately identify target objects; in advertising scenarios, they can attract more audience attention, thereby improving the user experience in different scenarios.

[0120] In this way, the optimization rules can be adjusted in real time according to scene changes, ensuring that images can be optimized and processed promptly and accurately in different scenes, adapting to the dynamic requirements of real-time image processing.

[0121] In one embodiment, based on the above embodiment, the visual optimization processing method based on real-time image processing further includes: When formulating visual optimization rules, user personalized preference factors should be introduced; Among them, the system records the user's historical optimization operations and feedback on different types of images, and learns the user's personalized optimization preferences; in the subsequent image optimization process, the visual optimization rules are fine-tuned according to the user's current image type and the user's personalized preferences.

[0122] In this embodiment, the system records in detail the user's historical optimization operations on different image types (e.g., landscape photos, portraits, and artistic photos). These operations include, but are not limited to, adjustments to image parameters such as brightness, contrast, saturation, and sharpness, as well as local optimization operations on specific areas. For example, if a user increases the brightness by 20% and decreases the contrast by 10% on a landscape photo, the system will accurately record the specific values ​​and the objects (i.e., image types) of these operations.

[0123] Optionally, in addition to recording user actions, the system also collects user feedback on the optimized image. This feedback can be a direct user evaluation (e.g., satisfied, dissatisfied, needs further adjustment, etc.) or the user's actions to modify the image again. For example, if a user adjusts the skin tone of an optimized portrait again, indicating that the user is dissatisfied with the previous skin tone optimization result, the system will record this feedback.

[0124] Optionally, the system will mine and analyze recorded user optimization operations and feedback data. By applying machine learning algorithms (such as cluster analysis and association rule mining), it can identify user preferences for different types of image optimization. For example, cluster analysis can reveal that users tend to adjust skin tones to a specific tonal range when processing portraits; while association rule mining can reveal that users typically increase image saturation while increasing image brightness.

[0125] Optionally, based on the results of data analysis, the system will create a personalized optimization preference model for each user. This model can describe the user's preferences and value ranges for various optimization parameters for different image types. For example, for portraits, the user's preference model might indicate that the user prefers brightness between 50% and 60% and saturation between 70% and 80%.

[0126] During the subsequent image optimization process, the system first identifies the type of image being processed. For example, it uses image recognition algorithms to determine whether the image is a landscape photo, a portrait photo, or another type of image.

[0127] Then, based on the identified image type, the system will call upon the user's personalized preference model for that type of image to fine-tune the general visual optimization rules. For example, if the general rule recommends increasing the brightness of a portrait by 15%, but the user's personalized preference model shows that the user generally prefers to keep the brightness of portraits between 50% and 60%, the system will adjust the general rule based on the actual brightness of the current image, perhaps only increasing the brightness by 10% to meet the user's personalized preference.

[0128] Optionally, the system will continuously update the user's historical operation and feedback data to continuously learn from the user's personalized preferences. As the user uses the system for a longer time, the system's understanding of the user's preferences will become more and more accurate, allowing it to more accurately fine-tune the visual optimization rules and provide users with more personalized image optimization services.

[0129] In one embodiment, user personalized preference factors are introduced so that image optimization can be customized according to each user's unique preferences, meeting the user's personalized needs and improving the user's satisfaction with the optimized image; and, by providing personalized image optimization services, the interaction and dependence between the user and the system are enhanced, thereby improving the user's usage stickiness.

[0130] In addition, the system's continuous learning mechanism can continuously adapt to changes in user preferences, ensuring the accuracy and effectiveness of personalized optimization services during long-term use, so that the system can always provide users with image optimization results that meet their latest preferences.

[0131] In one embodiment, based on the above embodiment, the visual optimization rule further considers the time series information of the image, analyzes the feature change trend between adjacent images for a continuously acquired image sequence, and dynamically adjusts the visual optimization rule according to this trend; Among them, when there are dynamic scenes in the image, different blurring and sharpening processes are applied to the images at different moments according to the speed and direction of movement to simulate real visual effects and enhance the dynamic sense of the picture.

[0132] In this embodiment, for a continuously acquired image sequence, the system first extracts features from each frame. These features can include color features (such as color histograms and color distribution), texture features (such as gray-level co-occurrence matrices and local binary patterns), and shape features (such as edge information and contour shapes). By extracting these features, the visual characteristics of the image can be quantified, providing a basis for subsequent analysis.

[0133] Optionally, after extracting features from adjacent images, the system compares and analyzes these features to determine trends in feature change. For example, by comparing the color histograms of adjacent frames, the direction and extent of color change can be determined; by analyzing changes in edge information, the motion and shape of objects can be understood. Specific analysis methods include difference calculation and correlation analysis. For example, the grayscale difference between corresponding pixels in adjacent frames can be calculated. A large difference indicates that the region has undergone significant changes between the adjacent frames.

[0134] Optionally, the system can dynamically adjust visual optimization rules based on the analyzed feature change trends. For example, if the colors between adjacent images are gradually brightening, the brightness adjustment can be appropriately reduced in subsequent optimizations to avoid over-brightening. If the objects in the image gradually become clearer from blurry, the sharpening parameters can be adjusted accordingly to make the optimized image more consistent with this trend.

[0135] Optional blurring and sharpening for dynamic scenes: The system first determines whether there is dynamic scene in the image. This can be achieved by analyzing the magnitude of feature changes between adjacent images. If the feature changes between adjacent images exceed a certain threshold, the image is considered to contain dynamic scene. For example, if the motion of an object causes the grayscale difference of the corresponding pixel in adjacent frames to exceed a preset threshold, the area is considered to have dynamic changes.

[0136] Once a dynamic scene is detected, the system further calculates the speed and direction of the object's motion. This can be achieved using motion estimation techniques such as optical flow. Optical flow calculates the displacement of corresponding pixels in adjacent image frames to obtain the object's motion vector, thereby determining the speed and direction of the movement.

[0137] Optionally, for fast-moving objects, blurring can be applied in the direction of their motion to simulate realistic visual effects. The faster the motion, the greater the blur. This is because when objects move quickly, the human eye experiences persistence of vision, causing the image on the retina to become blurred. For example, when processing an image of a fast-moving car, adding a certain degree of blur in the direction of the car's motion can make the image appear more realistic.

[0138] Optionally, sharpening can be applied to relatively still or slow-moving areas, or to the edges of objects, to enhance clarity and detail. This can highlight the outlines and features of objects, making the image more vivid. For example, when processing an image containing a moving person and a static background, sharpening the edges of the background and person, while appropriately blurring the moving parts of the person, can enhance the sense of motion in the image.

[0139] In one embodiment, by considering the time series information of the image and simulating real visual effects, the optimized picture is more in line with the visual perception of the human eye, enhancing the realism and dynamics of the picture; and it can dynamically adjust the optimization rules according to the actual changes in the image, and is applicable to various complex dynamic scenes, thereby improving the flexibility and adaptability of visual optimization, providing users with a more vivid and realistic visual experience, and having important application value in video playback, monitoring and other fields.

[0140] In addition, refer to Figure 2 In an embodiment of the present application, a control device Z10 is further provided, comprising: The acquisition module Z11 is used to acquire image data in real time based on the camera and perform image preprocessing; Extraction module Z12 is used to extract multimodal features from the preprocessed image and to extract deep features from the image using a pretrained machine learning model. The multimodal features include color, texture, and shape features. The machine learning model has a dynamic feature extraction architecture based on meta-learning to dynamically adjust the model's network parameters based on image resolution and / or complexity. Fusion module Z13 is used to fuse the extracted multimodal features and deep learning features, and screen the fused features through feature selection algorithm; Clustering module Z14, used to perform cluster analysis on the filtered features using clustering algorithms; Rule module Z15 is used to formulate visual optimization rules for different feature types and feature value ranges based on the results of feature analysis; The optimization module Z16 is used to perform visual optimization processing on the image according to the visual optimization rules to generate an optimized real-time image.

[0141] Optionally, the control device Z10 may be a virtual control device (such as a virtual machine) or a physical device (such as a physical device other than the visual optimization processing system that can execute the corresponding method).

[0142] In addition, a visual optimization processing system is also provided in the embodiment of the present application. The internal architecture of the visual optimization processing system can be as follows: Figure 3As shown, it includes a processor, a memory, a communication interface and an input interface connected via a system bus. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database is used to store data called by the computer program. The communication interface is used to communicate data with an external terminal. The input interface is used to receive signals input by an external device. When the computer program is executed by the processor, a visual optimization processing method based on real-time image processing as described in the above embodiment is implemented.

[0143] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the present invention and does not limit the visual optimization processing system to which the present invention is applied. For example, in some optional embodiments, the visual optimization processing system may further include an output interface (not shown in the figure), which is also connected to the system bus and is used to output corresponding signals to peripheral devices.

[0144] In addition, this application also provides a computer-readable storage medium, which includes a computer program. When executed by a processor, the computer program implements the steps of the visual optimization method based on real-time image processing as described in the above embodiment. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0145] In summary, the visual optimization processing method, control device, visual optimization processing system and computer-readable storage medium based on real-time image processing provided in the embodiments of the present application, in the feature extraction stage, not only extract multimodal features, but also use a machine learning model based on meta-learning design to perform deep feature extraction, and can dynamically adjust network parameters according to image resolution and complexity, and comprehensively and flexibly capture image information; and, based on the feature analysis results, formulate visual optimization rules, which can accurately optimize for different feature types and ranges; finally, process the image according to the rules, which can effectively improve the image visual effect and the accuracy of subsequent analysis and processing, meet the requirements of practical applications for high-quality real-time image optimization processing, and have stronger adaptability in complex and changing scenes.

[0146] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media provided in this application and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct RAM bus dynamic RAM (DRDRAM), and RAM bus dynamic RAM (RDRAM).

[0147] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, apparatus, article, or method. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, apparatus, article, or method comprising the element.

[0148] The above description is only a preferred embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A visual optimization processing method based on real-time image processing, characterized in that: include: Collect image data in real time based on the camera and perform image preprocessing; Perform multimodal feature extraction on the preprocessed image and perform deep feature extraction on the image using a pretrained machine learning model. The multimodal features include color, texture, and shape features. The machine learning model is designed with a dynamic feature extraction architecture based on meta-learning to dynamically adjust the model's network parameters based on image resolution and / or complexity. The extracted multimodal features and deep learning features are fused, and the fused features are screened through feature selection algorithms; Use clustering algorithm to perform cluster analysis on the filtered features; Based on the results of feature analysis, formulate visual optimization rules for different feature types and feature value ranges; According to the visual optimization rules, the image is visually optimized to generate an optimized real-time image.

2. The visual optimization processing method based on real-time image processing according to claim 1, characterized in that: Before the steps of performing multimodal feature extraction on the preprocessed image and performing deep feature extraction on the image using the pretrained machine learning model, the following steps are further included: The preprocessed image is subjected to multi-scale decomposition to decompose the image into sub-band images of different scales and directions for subsequent feature extraction.

3. The visual optimization processing method based on real-time image processing according to claim 1, characterized in that: The visual optimization processing method based on real-time image processing also includes: When fusing multimodal features and deep learning features, a multi-level fusion approach is adopted. Feature fusion is performed separately at different feature extraction stages, and then all features obtained after the first fusion are fused for the second time.

4. The visual optimization processing method based on real-time image processing according to claim 1, characterized in that: The clustering algorithm includes a dynamic clustering algorithm, which is used to adjust the clustering result according to the dynamic changes of the object between the previous image and the current image.

5. The visual optimization processing method based on real-time image processing according to claim 1, characterized in that: The visual optimization rules are also dynamically adjusted according to the real-time application scenario of the image: When images are used in video conferencing scenarios, focus on optimizing facial skin tone and clarity, reducing background interference, and improving the overall brightness and contrast of the image. When images are used in security monitoring scenarios, they enhance image edge clarity and detail information, improving the recognition of target objects. When the image is used in advertising, enhance the vividness and saturation of the colors.

6. The visual optimization processing method based on real-time image processing according to claim 1, characterized in that: The visual optimization processing method based on real-time image processing also includes: When formulating visual optimization rules, user personalized preference factors should be introduced; Among them, the system records the user's historical optimization operations and feedback on different types of images, and learns the user's personalized optimization preferences; in the subsequent image optimization process, the visual optimization rules are fine-tuned according to the user's current image type and the user's personalized preferences.

7. The visual optimization processing method based on real-time image processing according to claim 1, characterized in that: The visual optimization rule also takes into account the time series information of the image. For a sequence of continuously acquired images, the feature change trend between adjacent images is analyzed and the visual optimization rule is dynamically adjusted according to this trend. Among them, when there are dynamic scenes in the image, different blurring and sharpening processes are applied to the images at different moments according to the speed and direction of movement to simulate real visual effects and enhance the dynamic sense of the picture.

8. A control device, characterized in that: include: The acquisition module is used to collect image data in real time based on the camera and perform image preprocessing; An extraction module is used to extract multimodal features from preprocessed images and to extract deep features from images using a pretrained machine learning model. The multimodal features include color, texture, and shape features. The machine learning model is designed with a dynamic feature extraction architecture based on meta-learning to dynamically adjust the model's network parameters based on image resolution and / or complexity. The fusion module is used to fuse the extracted multimodal features and deep learning features, and screen the fused features through the feature selection algorithm; Clustering module, used to perform cluster analysis on the filtered features using clustering algorithms; The rule module is used to formulate visual optimization rules for different feature types and feature value ranges based on the results of feature analysis; The optimization module is used to perform visual optimization processing on the image according to visual optimization rules to generate an optimized real-time image.

9. A visual optimization processing system, characterized in that: The visual optimization processing system includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the visual optimization processing method based on real-time image processing as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the visual optimization processing method based on real-time image processing according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Rail surface state identification method based on multi-feature fusion

    CN112308111A

  • Remote sensing image classification method and system

    CN117422936A

  • Flotation process working condition identification method based on deep and shallow layer feature fusion

    CN117496436A

  • Image processing method based on AI algorithm

    CN117934354A

  • Defect image enhancement method based on traditional data enhancement

    CN118967474A