Machine vision-oriented image preprocessing method and apparatus, device, and storage medium
By blurring the original image and enhancing the semantic feature, the target image suitable for machine vision analysis is generated, which solves the problem that the prior art cannot be effectively applied to machine vision image analysis, and achieves the effect of improving analysis performance and reducing computing costs.
Patent Information
- Application Number
- PCT/CN2024/139136
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-19
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-26
AI Technical Summary
The existing image preprocessing technology cannot be effectively applied to machine vision-oriented image analysis tasks, resulting in low analysis performance.
An image preprocessing method for machine vision is proposed, which generates an image to be enhanced by blurring the original image and reduces its clarity; then the semantic features of the image to be enhanced are enhanced to generate a target image, and the target image is input into the image processing neural network to perform image analysis tasks.
On the basis of not reducing the intensity of the semantic features of the original image, the original image code rate is reduced, the calculation amount and cost are reduced, and the analysis performance of the image processing neural network is improved.
Smart Images

Figure CN2024139136_26062025_PF_FP_ABST
Abstract
Description
Image preprocessing method, device, equipment and storage medium for machine vision Technical Field
[0001] The present application belongs to the field of image processing technology, and specifically relates to an image preprocessing method, device, equipment and storage medium for machine vision. Background Art
[0002] Image analysis tasks for machine vision (also known as computer vision) can include tasks such as image segmentation and image matching based on neural networks. The raw image data volume is generally large and contains some image data that is not particularly relevant to the image analysis task. Therefore, in some embodiments, to reduce data processing volume and conserve resources, the raw image should be preprocessed and then fed into the neural network to perform the corresponding image analysis task. It can be seen that the quality of the preprocessed image has a decisive influence on the neural network's analysis performance.
[0003] However, existing image preprocessing technologies mostly process image pixels based on color and brightness, making them unsuitable for machine vision-based image analysis tasks. Other image processing technologies, when applied to machine vision-based image analysis tasks, also suffer from poor analysis performance. Therefore, a preprocessing method suitable for machine vision-based image analysis tasks, which can improve analysis performance after preprocessing, is urgently needed in this field. Summary of the Invention
[0004] This application proposes an image preprocessing method, device, equipment and storage medium for machine vision, which are suitable for image preprocessing for image analysis tasks for machine vision, and make the preprocessed images conducive to improving analysis performance.
[0005] The first embodiment of the present application proposes an image preprocessing method for machine vision, comprising:
[0006] Performing blur processing on the original image to generate an image to be enhanced, wherein the clarity of the image to be enhanced is lower than that of the original image;
[0007] Performing enhancement processing on the semantic features of the image to be enhanced to generate a target image;
[0008] The target image is input into an image processing neural network to trigger the image processing neural network to perform an image analysis task based on the semantic features of the target image.
[0009] An embodiment of the second aspect of the present application provides an image preprocessing device, including:
[0010] A blur processing module, configured to perform blur processing on the original image to generate an image to be enhanced, wherein the clarity of the image to be enhanced is lower than that of the original image;
[0011] An enhancement module, configured to enhance the semantic features of the image to be enhanced to generate a target image;
[0012] An input module is used to input the target image into an image processing neural network to trigger the image processing neural network to perform an image analysis task based on the semantic features of the target image.
[0013] An embodiment of the third aspect of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect above.
[0014] An embodiment of the fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the program is executed by a processor to implement the method described in the first aspect above.
[0015] The technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:
[0016] In an embodiment of the present application, the original image can first be blurred to generate an image to be enhanced, where the clarity of the enhanced image is lower than that of the original image. This reduces the bitrate of the original image by reducing the clarity of the original image. Furthermore, the semantic features of the image to be enhanced are enhanced to generate a target image, and the target image is input into an image processing neural network to trigger the image processing neural network to perform an image analysis task based on the semantic features of the target image. The semantic features can be feature data used by the image processing neural network to perform the image analysis task. As can be seen, compared to traditional image preprocessing based on the color and brightness of each pixel in the image, the technical solution of the embodiment of the present application achieves the effect of reducing the bitrate of the original image without reducing the strength of the semantic features of the original image by comprehensively reducing the clarity of the original image and then specifically enhancing the semantic features in the image, thereby reducing the computational complexity and saving costs. Furthermore, compared to traditional image preprocessing based on the color and brightness of each pixel in the image, the technical solution of the embodiment of the present application performs preprocessing based on the semantic features of the image, allowing the target image to be used as the input image of the image processing neural network and helping to maintain the analysis performance of the image processing neural network at a relatively high level.
[0017] Additional aspects and advantages of the present application will be given in part in the description below and in part will become apparent from the description below or learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present application. Throughout the accompanying drawings, the same reference numerals are used to denote the same components.
[0019] In the attached figure:
[0020] FIG1 shows a schematic block diagram of a CV-oriented analysis system provided in one embodiment of the present application;
[0021] FIG2 is a schematic diagram showing a scenario of an image preprocessing method for machine vision provided by an embodiment of the present application;
[0022] FIG3 shows a flowchart of an image preprocessing method for machine vision provided by an embodiment of the present application;
[0023] FIG4A is a schematic diagram showing data flow of an image preprocessing method for machine vision provided by an embodiment of the present application;
[0024] FIG4B is a schematic diagram showing data flow of another image preprocessing method for machine vision provided by an embodiment of the present application;
[0025] FIG5 shows a schematic diagram of the architecture of an image preprocessing network training system provided in one embodiment of the present application;
[0026] FIG6a and FIG6b are schematic diagrams showing a comparison of the effects of image preprocessing provided by an embodiment of the present application;
[0027] FIG7 shows a schematic structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0028] The technical solutions of the embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.
[0029] The terms used in the following embodiments of the present application are for the purpose of describing specific embodiments and are not intended to be limiting of the technical solutions of the present application. As used in the specification of the present application and the appended claims, the singular expressions "a", "an", "said", "above", "the" and "this" are intended to also include plural expressions, unless the context clearly indicates otherwise. It should also be understood that although the terms first, second, etc. may be used to describe a certain class of objects in the following embodiments, the objects should not be limited to these terms. These terms are used to distinguish the specific implementation objects of this class of objects. For example, the terms first, second, etc. are used in the following embodiments to describe semantic features, but semantic features are not limited to these terms. These terms are only used to distinguish the semantic features of different images. The same applies to other classes of objects that may be described using the terms first, second, etc. in the following embodiments, and they will not be repeated here.
[0030] The following is an introduction to the relevant technologies involved in the embodiments of this application.
[0031] The embodiments of the present application relate to the field of image processing technology, and disclose a method for preprocessing an image based on artificial intelligence (AI) so that the preprocessed image supports image analysis tasks oriented to computer vision (CV).
[0032] AI is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. AI software technologies primarily encompass computer vision (CV), speech processing (Speech Technology), natural language processing (NLP), and machine learning (ML) / deep learning.
[0033] The technical solution of this application mainly involves CV technology, which is a science that studies how to make machines "see". More specifically, it refers to using cameras and computers to replace human eyes to identify, track and measure targets, and other machine vision, and further perform graphic processing to make computer processing into images that are more suitable for human eye observation or transmission to instruments for detection. CV technology generally includes image processing (including image encryption, etc.), image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and map construction, and also includes common biometric recognition technologies such as face recognition. The CV-based image analysis tasks involved in the embodiments of this application include, for example, at least one of the following: image matching, image recognition, image segmentation, image content extraction, and face recognition, etc.
[0034] Taking the CV-based image analysis task of image content extraction as an example, the CV network can extract the target objects (e.g., target faces, target lesions, target vehicles, etc.) contained in the input image through a series of processes such as image encoding and feature extraction. Other applications can then be performed based on the extracted target objects. Taking target vehicle extraction as an example, the CV network's analysis performance is characterized by the proportion of all target vehicles extracted from the input image and the accuracy with which the extracted vehicles are identified as target vehicles. For example, the higher the proportion of all target vehicles extracted by the CV network and the higher the accuracy with which the extracted vehicles are identified as target vehicles, the better the CV network's analysis performance; conversely, a lower proportion indicates relatively poor analysis performance.
[0035] It should be understood that when the CV-based image analysis task is other image analysis tasks, the characterization of the CV network analysis performance can be reflected by the expression degree of the CV network analysis results on the target object to be analyzed, which will not be expanded here.
[0036] In existing CV processing, encoding the original image facilitates transmission and analysis of the image data. However, original images are typically lossless images, often with high bitrates. Most CV processing equipment may have limited computing power, resulting in reduced clarity or even distortion of the encoded image, which in turn reduces the performance of the CV analysis network. To reduce the data processing workload of the CV network, the original image can be preprocessed and then fed into the CV network. However, traditional image preprocessing is based on the three-component information of the image pixels (red, green, and blue, RGB), grayscaling the image, and performing geometric transformations to obtain the preprocessed image. Specifically, traditional image preprocessing focuses on the grayscale and brightness of the image pixels and does not involve processing or understanding the semantic features contained in the image. CV, as a deep learning technique, processes images based on the semantic features represented by the distribution of image pixels. Therefore, traditional image preprocessing is not suitable for CV-based image analysis tasks.
[0037] In view of this, an embodiment of the present application provides an image preprocessing technology for CV, which first blurs the original image, thereby reducing the bit rate of the original image by reducing the clarity of the original image. Furthermore, the semantic features of the image to be enhanced obtained after the blurring are enhanced so that the intensity of the semantic features of the enhanced image is greater than or equal to the intensity of the semantic features of the original image. In this way, the effect of reducing the bit rate is achieved by reducing the bit rate of the content in the original image that is less relevant to the image analysis task without reducing the intensity of the semantic features of the original image. Moreover, the above processing process is not limited to the smaller aspect of the grayscale of the pixels in the image, but is performed in terms of the semantic features of the image, so that the enhanced image can be used as the input image of the image processing neural network, so that the image processing neural network performs the image analysis task based on the semantic features, which is conducive to maintaining the analysis performance of the image processing neural network at a better level.
[0038] The following is an introduction to the technical scenarios and system architecture involved in the embodiments of this application.
[0039] Referring to Figure 1, Figure 1 shows a schematic block diagram of a CV-oriented analysis system provided by an embodiment of the present application. As shown in Figure 1, the CV-oriented analysis system may include an image source 11, an image preprocessing network 12, an encoding network 13, and an analysis network 14. In a specific implementation form, the image source 11, the image preprocessing network 12, the encoding network 13, and the analysis network 14 can be implemented as hardware components, software components, or a combination of hardware and software components. Among them, the image preprocessing network 12, the encoding network 13, and the analysis network 14 can be an algorithm network based on deep learning. The encoding network 13 and the analysis network 14 can be used to form the above-mentioned CV network. They are described as follows:
[0040] Image source 11 can be used to provide raw images for a CV-oriented analysis system and can include or be any type of image capture device, for example, for capturing images of the real world, images of real objects, and / or any other type of image. Image source 11 can be a camera for capturing images or a memory for storing images. When image source 11 is a camera, it can be, for example, a local camera or an integrated camera integrated into the source device; when image source 11 is a memory, it can be, for example, a local memory or an integrated memory integrated into the source device.
[0041] The image preprocessing network 12 can be used to preprocess the original image from the image source 11 to minimize the image bitrate while effectively maintaining the performance of analyzing the content to be analyzed. In the embodiment of the present application, the image preprocessing network 12 can reduce the amount of data in the original image that is less relevant to the content to be analyzed, while maintaining or enhancing the portion of the original image that represents the content to be analyzed. For example, the preprocessing performed by the image preprocessing network 12 can include blurring and directional enhancement of the blurred image.
[0042] In one embodiment, the image preprocessing network 12 may include algorithm modules, deep learning-based image processing networks, or models. These algorithm modules, deep learning-based image processing networks, or models can be combined in various ways to implement the image preprocessing method for machine vision in the embodiments of the present application. For example, a degradation algorithm, a U-net network, etc.
[0043] In this way, the image preprocessing method and principle of the image preprocessing network 12 are compatible with CV technology, making the image preprocessing network 12 applicable to CV analysis systems. Moreover, the preprocessed images output by the image preprocessing network 12 can be directly used for image analysis tasks in the CV network.
[0044] The encoding network 13 can be used to receive the image preprocessed by the image preprocessing network 12 and process the preprocessed image using a calculation module preset by the encoding network 13, thereby providing image data containing the semantic features of the original image. In some embodiments, the encoding network 13 can be implemented as a full neural network.
[0045] Analysis network 14 can be used to perform image analysis tasks on the encoded data image and output analysis results. Analysis network 14 can be a neural network that performs at least one image analysis task, including image matching, image recognition, image segmentation, image content extraction, and face recognition. The analysis results output by analysis network 14 can be implemented as predictions such as probability distributions and confidence parameters. These predictions can represent the degree to which analysis network 14 understands the semantic features of the original image.
[0046] It should be understood that although the image preprocessing network 12 is integrated into the CV-oriented analysis system in Figure 1, in the device embodiment, the image preprocessing network 12 can be deployed in an exemplary CV-oriented analysis system and coupled with other functional devices in the CV-oriented analysis system; the functions of the image preprocessing network 12 can also be integrated into an independent computer device, so that the computer device has the image preprocessing function of the embodiment of the present application, and the computer device can be used in different CV analysis application scenarios.
[0047] In this way, the image preprocessing method for machine vision in the embodiment of the present application can be flexibly applied to CV image analysis tasks in different scenarios, and there is no need to change the basic CV network structure in any application scenario, so it has good scalability.
[0048] The following describes the machine vision-oriented image preprocessing method, device, equipment, and storage medium of the embodiments of the present application in combination with the aforementioned embodiments.
[0049] First, an embodiment of the present application provides an image preprocessing method for machine vision. The image preprocessing method for machine vision can be used in the CV-oriented analysis system shown in FIG1 . The image preprocessing network 12 in FIG1 can be executed by the image preprocessing network 12. The image preprocessing network 12 in FIG1 may include:
[0050] Performing blur processing on the original image to generate an image to be enhanced, wherein the clarity of the image to be enhanced is lower than that of the original image;
[0051] Performing enhancement processing on the semantic features of the image to be enhanced to generate a target image;
[0052] The target image is input into an image processing neural network to trigger the image processing neural network to perform an image analysis task based on the semantic features of the target image.
[0053] The strength of the semantic features of the target image may be greater than or equal to the strength of the semantic features of the original image. The image processing neural network may be implemented as the analysis network 14 in FIG1 .
[0054] In some embodiments, as shown in FIG2 , in order to make the intensity of the semantic features of the target image no less than the intensity of the semantic features of the original image, after obtaining the image to be enhanced, the image preprocessing network 12 may generate enhancement parameters based on the original image and the image to be enhanced, and then enhance the image to be enhanced according to the enhancement parameters to generate the target image.
[0055] The following uses examples to illustrate the specific processing procedures related to the image preprocessing method for machine vision.
[0056] As shown in FIG3 , FIG3 illustrates an exemplary machine vision-oriented image preprocessing method according to an embodiment of the present application. The machine vision-oriented image preprocessing method specifically includes the following steps:
[0057] Step S101 : down-sampling the original image according to a preset sampling ratio.
[0058] In some embodiments, the original image may be an image from the image source 11 in FIG1 . The original image may include a target object to be analyzed by the analysis network 14 in FIG1 . The target object may be, for example, an image of the target object to be analyzed. The target object to be analyzed may include, for example, a target object, a target face, a target building, etc., without limitation herein.
[0059] Exemplarily, the preset sampling magnification may be related to the resolution of the original image. For example, if the resolution of at least one of the width and height of the original image is greater than or equal to a preset value, the preset sampling magnification is determined to be a first preset sampling magnification; if the resolution of both the width and height of the original image is less than the preset value, the preset sampling magnification is determined to be a second preset sampling magnification. Here, the first preset sampling magnification is greater than the second preset sampling magnification.
[0060] For example, if the resolution of at least one direction of the width and height of the original image is greater than or equal to 1080 pixels (pixel, P), the preset sampling magnification can be determined to be 8; if the resolution of the width and height of the original image are both less than 1080P, the preset sampling magnification can be determined to be 4.
[0061] Step S102 : upsampling the downsampled image according to the preset sampling ratio to generate an image to be enhanced.
[0062] The resolution of the image to be enhanced may be the same as the resolution of the original image.
[0063] It should be noted that step S101 and step S102 are exemplary implementation processes of performing degradation processing on the original image. In step S102, the up-sampled image can also be called the image obtained by degradation processing, that is, the blurred image.
[0064] It should be understood that degradation processing is only an exemplary method for blurring the original image and does not limit the blurring processing in the embodiments of the present application. In the embodiments of the present application, the original image can also be blurred by any of the following methods: adding noise to the original image, blurring the original image using an image blurring algorithm.
[0065] In some embodiments, if an image blur algorithm is used to blur the original image, any of the following blur algorithms can be used: Gaussian Blur, Box Blur, Kawase Blur, Dual Blur, Bokeh Blur, Tilt Shift Blur, Iris Blur, Grainy Blur, Radial Blur, Directional Blur, etc.
[0066] As can be seen, this implementation method blurs the original image, not just the grayscale angles of the pixels within the image, but rather applies indiscriminate blurring to the entire image distribution. This helps maintain the integrity of the original image's semantic features while reducing the bitrate, thereby improving the analysis performance of the CV network.
[0067] Step S103: generating a residual image between the original image and the image to be enhanced.
[0068] Step S104 : extracting semantic features of the original image to obtain first semantic features, and extracting features of the residual image.
[0069] Step S105 , obtaining the enhancement parameter by fusing the first semantic feature with the feature of the residual image.
[0070] The enhancement parameter is used to characterize the degree of enhancement of the image to be enhanced, and in particular, is used to characterize the degree of enhancement of the semantic feature portion of the image to be enhanced.
[0071] Step S106 : performing enhancement processing on the semantic features of the image to be enhanced according to the enhancement parameters to generate the target image.
[0072] It should be noted that the semantic features of a target image can refer to the features of the target object being analyzed by the image processing neural network. The strength of the semantic features can include the dimensionality of the target object's features, the number of eigenvalues contained in each dimension, and the size of the receptive field corresponding to each dimensional feature.
[0073] Among them, the dimensions of semantic features can include visual dimension, object dimension and concept dimension. The features of the visual dimension can include the color, texture and shape of the target object; the object dimension can include the attribute features of the target object (for example, animals, plants, scenery), etc.; the concept dimension can represent the meaning expressed by the target object. For example, the target objects include beaches, blue sky and sea water, etc. The features of the visual dimension can include the outline, color, texture and shape of each part of the beach, blue sky and sea water, the features of the object dimension include the attribute features of each part of the sand, blue sky and sea water; the features of the concept dimension can represent the beach.
[0074] In some embodiments, the more feature dimensions a semantic feature contains, the more feature values each dimension of the feature contains, and the smaller the receptive field corresponding to the feature of each dimension, the stronger the strength of the semantic feature can be considered; conversely, the weaker the strength of the semantic feature can be considered.
[0075] Combined with the aforementioned blurring process of the original image, it can be seen that the image to be enhanced is obtained by indiscriminately blurring the original image as a whole. That is, the clarity of the target object and other content outside the target object contained in the image to be enhanced is reduced. In order to improve the analytical performance of the image processing neural network while reducing the data processing volume of the image processing neural network, in one embodiment of the present application, the image preprocessing stage can perform targeted enhancement on the target object in the image to be enhanced to increase the strength of the semantic features.
[0076] Step S107: input the target image into an image processing neural network to trigger the image processing neural network to perform an image analysis task based on the semantic features of the target image.
[0077] Furthermore, image processing neural networks can perform image analysis tasks on target images based on the semantic features of the target image. For example, if the image analysis task is image segmentation, the image processing neural network can process the semantic features to determine the target object to be segmented in the target image, and then separate the target object from the target image.
[0078] In summary, the technical solution of the embodiment of the present application reduces the bit rate of the original image by comprehensively reducing the clarity of the original image. Subsequently, the semantic features of the image are specifically enhanced, thereby achieving the effect of reducing the bit rate of the original image without reducing the strength of the semantic features of the original image, which is beneficial for reducing the amount of computation and saving costs. Furthermore, preprocessing is performed based on the semantic features of the image, which helps maintain the analytical performance of the image processing neural network at a relatively good level.
[0079] It should be noted that step S101 and step S102 are exemplary implementation processes of performing degradation processing on the original image. In step S102, the up-sampled image can also be called the image obtained by degradation processing, that is, the blurred image.
[0080] It should be understood that steps S103 to S105 are merely an exemplary method for calculating enhancement parameters and do not limit the enhancement processing of the embodiments of the present application. In other embodiments of the present application, the image preprocessing network may also calculate enhancement parameters according to other methods. For example, the enhancement parameters may be determined based on the semantic features of the original image and the semantic features of the image to be enhanced.
[0081] Exemplarily, the image preprocessing network can extract the semantic features of the original image to obtain a first semantic feature, and extract the semantic features of the image to be enhanced to obtain a second semantic feature. In combination with the above description of the semantic features, it can be seen that the first semantic feature and the second semantic feature both contain features of at least two dimensions. Any semantic feature can, for example, include features of the visual dimension, features of the object dimension, and features of the conceptual dimension. Furthermore, for the features of any dimension, the image preprocessing network can calculate the difference between the features of the dimension in the first semantic feature and the features of the dimension in the second semantic feature to obtain at least two differences, and then calculate the enhancement parameters based on the at least two differences (as shown in the embodiment illustrated in Figure 4B).
[0082] It should be noted that the feature of any dimension may contain multiple eigenvalues, and the number of eigenvalues contained in the feature of this dimension may be determined according to the setting of the enhancement algorithm.
[0083] Furthermore, in some embodiments, the difference corresponding to the feature of the dimension may be the difference between the average value corresponding to the first semantic feature and the average value corresponding to the second semantic feature. The average value corresponding to the first semantic feature is the average value of the feature values of the dimension in the first semantic feature, and the average value corresponding to the second semantic feature is the average value of the feature values of the dimension in the second semantic feature.
[0084] In some other embodiments, the difference value corresponding to the feature of the dimension may be the variance between the feature value of the dimension in the first semantic feature and the feature value of the dimension in the second semantic feature.
[0085] After obtaining at least two differences, for any difference, the image preprocessing network can multiply the difference by the weight of the difference, and the multiplication result is the enhancement factor corresponding to the difference to obtain at least two enhancement factors. Thereafter, the sum of the at least two enhancement factors is calculated, and the sum result is the enhancement parameter.
[0086] In some embodiments, the weights may be pre-set. The weights are used to characterize the degree of enhancement of features corresponding to the dimensions of the corresponding differences. Taking the example of semantic features including features of the visual dimension, features of the object dimension, and features of the conceptual dimension, the weight corresponding to the features of the visual dimension is, for example, 0.15, the weight corresponding to the features of the object dimension is, for example, 0.35, and the weight corresponding to the features of the conceptual dimension is, for example, 0.5. It can be characterized that the degree of enhancement of the features of the visual dimension is relatively the weakest, the degree of enhancement of the features of the object dimension is relatively medium, and the degree of enhancement of the features of the conceptual dimension is relatively the strongest.
[0087] It should be noted that the image preprocessing network can use a U-net model to extract the first semantic features of the original image and enhance the second semantic features of the image to be enhanced according to the enhancement parameters.
[0088] In some embodiments, the image preprocessing network can extract global semantic features from the original image. The global features include image features of the target object and image features of non-target objects in the original image. Furthermore, the image preprocessing network can perform at least one dimensionality reduction process on the global semantic features to obtain initial image features of the target object. The initial image features can then be subjected to the at least one dimensionality increase process to obtain the first semantic features.
[0089] Any dimensionality reduction process may be downsampling of features. Exemplarily, downsampling may be performed through convolution operations and pooling.
[0090] In some embodiments, global semantic features include image features of the target object and image features of non-target objects in the original image, and the image features of the target object in the global semantic features are relatively unimportant. By gradually reducing the dimensionality, more compact semantic information in the original image can be obtained, thereby facilitating the emphasis of the image features of the target object.
[0091] For example, the number of dimensionality increase processing and the number of dimensionality reduction processing may be the same. Any dimensionality increase processing may be upsampling of features.
[0092] Specifically, the image preprocessing network can upsample the initial image features and splice the upsampled features with reduced-dimensionality features of the same size, where the reduced-dimensionality features of the same size refer to features of the same size as the upsampled features obtained during at least one dimensionality reduction process. If the size of the spliced features is the same as the size of the global semantic features, the spliced features are used as the first semantic features. If the size of the spliced features is smaller than the size of the global semantic features, the spliced features are used as new initial image features, and the new initial image features are upsampled.
[0093] For example, the global semantic features of the original image have a size of 224*224. The U-net model can perform three convolution and pooling operations on the global semantic features. After the first convolution and pooling operation, the resulting reduced-dimensionality features have a size of 112*112, the second convolution and pooling operation, the resulting reduced-dimensionality features have a size of 56*56, and the third convolution and pooling operation, the resulting reduced-dimensionality features have a size of 28*28. The 28*28 features can be considered the initial image features of the target object. Furthermore, the 28*28 features are upsampled to obtain increased-dimensional features of 56*56. These increased-dimensional features are concatenated with the 56*56 features obtained after dimensionality reduction. The concatenated features are convolved and upsampled again to obtain increased-dimensional features of 112*112. These increased-dimensional features are then concatenated with the 112*112 features obtained after dimensionality reduction. After convolution operation on the spliced features, upsampling is performed for the third time to obtain a 224*224 dimension-enhanced feature. This dimension-enhanced feature is concatenated with the 224*224 dimension-enhanced feature obtained after dimensionality reduction, and convolution operation is performed on the spliced features. The resulting 224*224 dimension-enhanced feature is used as the first semantic feature of the original image.
[0094] This gradually increases the dimension of features that highlight the importance of the target object's image features, helping to gradually improve the detailed features of the target object, so that the resulting first semantic features can accurately and completely represent the semantics of the target image. Furthermore, the enhancement parameters derived from the first semantic features have strong reliability.
[0095] The following takes the image preprocessing network including the degradation module and the U-net model as an example, and introduces the image preprocessing method for machine vision in an embodiment of the present application in combination with examples.
[0096] For example, as shown in Figures 4A and 4B, the original image includes an image of a circular area and an image of a house. The image of the circular area may be, for example, a lawn (not shown in the figures). The display effect of the image of the circular area is similar to that of the ring in the figures, and is not illustrated in the figures. The image of the circular area serves as the background of the image of the house. The image of the house obscures part of the image of the circular area and is displayed at the forefront of the original image. The image of the house is, for example, the target object to be analyzed by the image processing neural network.
[0097] See Figure 4A, which shows an exemplary data flow diagram of an image preprocessing method for machine vision. In this example, after receiving the original image, the image preprocessing network can call a pre-deployed degradation module to degrade the original image and perform blurring to obtain a degraded image (i.e., the aforementioned image to be enhanced).
[0098] For example, if the width of the original image is greater than 1080p, the degradation module may select a sampling rate of 8, first downsample the original image by 8, and then upsample the downsampled image by 8. The resolution of the upsampled image is, for example, the same as the resolution of the original image.
[0099] Referring again to Figure 4A , the image of the circular area and the image of the house in the original image both have high definition, as evidenced by the finer lines of their outlines. However, the image of the circular area and the image of the house in the degraded image have relatively low definition, as evidenced by the grainy pixels of their outlines.
[0100] Furthermore, the degradation module can generate a residual image between the original image and the degraded image, and input the original image, the residual image, and the degraded image into a U-net model. The image preprocessing network can call the U-net model to enhance the degraded image to obtain a preprocessed image (i.e., the aforementioned target image).
[0101] For example, the U-net model can extract a first semantic feature from the original image. In this example, the first semantic feature can refer to a feature representing the image of the house in the original image.
[0102] For example, the U-net model can first extract the global semantic features of the original image, which may include the features of the circular area and the features of the house. Afterwards, the U-net model can, for example, perform three dimensionality reduction processes on the global semantic features in succession, and each dimensionality reduction process achieves the effect of downsampling through convolution operations and pooling. For details, please refer to the aforementioned exemplary description of dimensionality reduction, which will not be repeated here. Among the features obtained after the three dimensionality reduction processes, the intensity of the house feature is greater than the intensity of the circular area feature, and the weight of the house feature can also be greater than the weight of the circular area feature. The weight here can be used to characterize the importance of the feature. Then, the features obtained after the three dimensionality reduction processes can be used as the initial image features of the house. Further, the initial image features of the house are successively subjected to three dimensionality increase processes, and the features obtained after the three dimensionality increase processes are the first semantic features of the original image.
[0103] As mentioned above, among the initial image features of the house, the house features have greater strength and importance. Therefore, in the process of dimensionality upgrading, most of the enriched and detailed features are also the features of the house. It can be considered that the first semantic feature is the final image feature of the house.
[0104] The residual image can represent the difference between the degraded image and the original image. Therefore, the U-net model can extract features from the residual image and fuse the first semantic features with the features of the residual image. The fusion result can be used as an enhancement parameter, which represents the degree of enhancement for the degraded image.
[0105] The U-net model can extract a second semantic feature from the degraded image, which may be a characteristic of a house in the degraded image. The U-net model can then enhance the second semantic feature based on an enhancement parameter to obtain the target image. For example, the U-net model can perform a series of convolution operations on the enhancement parameter and the second semantic feature to output the enhanced semantic features. These semantic features can be represented as feature values between 0 and 1, with larger feature values indicating greater semantic strength.
[0106] Referring again to Figure 4A, the house in the target image has high image clarity, while the outline of the circular area still exhibits a grainy pixel texture. This indicates that, compared to the original image, the bitrate of the preprocessed target image is reduced in areas less relevant to the image analysis task, while the strength of semantic features in areas more relevant to the task remains unchanged. This reduces computational complexity while maintaining the performance of the CV network when inputting the target image into the network.
[0107] It should be understood that Figure 4A is an exemplary implementation of image preprocessing in the present application and does not limit the image preprocessing method for machine vision in the embodiments of the present application. In other implementations, the image preprocessing network may also use other methods to enhance the semantic features of the house. For example, see Figure 4B for another exemplary image preprocessing method for machine vision.
[0108] As shown in FIG4B , in this example, after receiving the original image, the image preprocessing network also calls the degradation module to blur the original image by degrading it to obtain a degraded image. This implementation process can be found in the description of the embodiment illustrated in FIG4A and will not be repeated here.
[0109] Different from FIG4A , in this example, after obtaining the degraded image, the degradation module no longer generates a residual image, but instead inputs both the original image and the degraded image into the U-net model.
[0110] As shown in Figure 4B, in this example, the U-net model can extract semantic features from the original image and the degraded image, respectively, obtaining a first semantic feature for the original image and a second semantic feature for the degraded image. Combining the aforementioned processing of the U-net model, we can see that the first semantic feature represents the final image characteristics of the house in the original image, and the second semantic feature represents the final image characteristics of the house in the degraded image. Therefore, the U-net model can calculate the difference between the first and second semantic features to obtain the enhancement parameters.
[0111] For example, in this example, the U-net model can separate the first semantic feature to obtain the features of the visual dimension, the features of the object dimension, and the features of the concept dimension; and separate the second semantic feature to obtain the features of the visual dimension, the features of the object dimension, and the features of the concept dimension. For the features of each dimension, the feature value of the dimension in the first semantic feature and the variance of the feature value in the corresponding first semantic feature are calculated, for example, the visual dimension feature variance, the object dimension feature variance, and the concept dimension feature variance are obtained respectively. Further, the feature variance corresponding to each dimension is multiplied by the weight value corresponding to the dimension to obtain the enhancement factor of the visual dimension, the enhancement factor of the object dimension, and the enhancement factor of the concept dimension. In this example, the sum of the enhancement factor of the visual dimension, the enhancement factor of the object dimension, and the enhancement factor of the concept dimension is the enhancement parameter.
[0112] Afterwards, the U-net model can enhance the second semantic feature according to the enhancement parameter to obtain the target image, which is not described here.
[0113] It should be understood that the embodiments illustrated in FIG4A and FIG4B are introduced using the degradation algorithm and the U-net model as examples, and do not constitute a limitation on the image preprocessing network of the present application. In other embodiments of the present application, the image preprocessing network may include other algorithm models with the same or similar functions, or a combination network, etc., and the image preprocessing network may also include more algorithm models than those shown in the figure. It will be apparent to those skilled in the art that many modifications and variations are possible based on the above teachings.
[0114] In summary, the technical solution of the embodiment of the present application achieves the effect of reducing the bit rate of the original image without reducing the strength of the semantic features of the original image by comprehensively reducing the clarity of the original image and then specifically enhancing the semantic features of the image. As shown in the effect comparison diagram of Figures 6a and 6b, Figure 6a is the original image. The grass and the racket in Figure 6a are both clearer, indicating that the bit rates of the grass and the racket are both high, while Figure 6b is the image of Figure 6a after being processed by this technical solution. The racket in Figure 6b is still clearer, while the grass is relatively blurred, that is, the bit rate of the grass part is reduced. It can be seen that this is conducive to reducing the amount of calculation and saving costs. In addition, preprocessing is performed in terms of the semantic features of the image, so that the target image can be used as the input image of the image processing neural network, and it is conducive to maintaining the analysis performance of the image processing neural network at a better level.
[0115] In conjunction with the foregoing description, the image preprocessing network used to perform the image preprocessing method for machine vision of the embodiment of the present application can be obtained by training the network to be trained using a proxy neural network, and then a connection can be established with the image processing neural network to obtain the CV-oriented analysis system shown in Figure 1. The proxy neural network is a neural network used to perform the image analysis task of the aforementioned CV network.
[0116] In some embodiments, the image preprocessing network used to perform blurring is typically a pre-set algorithm, while the image enhancement function is typically performed using a deep learning network. In this example, the network to be trained can be a network model used to perform the enhancement function. For example, the network to be trained is the U-net model to be trained in Figures 4A and 4B.
[0117] Furthermore, it should be noted that since the image preprocessing network can be independent of the image processing neural network, in some embodiments, the network to be trained can be a pre-built initial network; in other embodiments, the network to be trained can be a preprocessing network to be trained, which is suitable for another image analysis task that is different from the image analysis task. If the network to be trained is the preprocessing network to be trained, the network to be trained can be disconnected from the original image processing neural network before using the proxy neural network to train the network to be trained to obtain the image preprocessing network. The original image processing neural network here is the image processing neural network that performs the other image analysis task.
[0118] Among them, using a proxy neural network to train the network to be trained to obtain an image preprocessing network can include: calling the network to be trained to preprocess a sample image, the preprocessed image being the image to be analyzed, and then calling the proxy neural network to perform the image analysis task on the image to be analyzed to obtain a predictive analysis result. Calculating a loss value and determining whether the loss value has converged, the loss value includes preprocessing loss, analysis loss, and proxy loss. If the loss value converges, the network to be trained is used as the image preprocessing network; if the image preprocessing network does not converge, adjusting the parameters of the network to be trained, and using the model with adjusted parameters as the new network to be trained, and again calling the network to be trained to preprocess the sample image.
[0119] Exemplarily, adjusting the parameters of the network to be trained may be adjusting the parameters of the enhancement model to be trained, such as adjusting the parameters of the U-net model to be trained.
[0120] In some embodiments, the preprocessing loss is used to characterize the loss between the image to be analyzed and the sample image. For example, it can be implemented as the mean square error (MSE) between the image to be analyzed and the sample image to constrain the enhancement bias. The analysis loss is used to characterize the loss between the predicted analysis result and the sample image annotation result. The proxy loss is used to characterize the loss of the proxy neural network during the sample image processing process.
[0121] It should be noted that the arithmetic function included in the proxy loss can be related to the processing of the proxy neural network. For example, if the proxy neural network includes a proxy encoding network and an execution network for the image analysis task, the proxy loss can include the encoding loss and the discrete cosine transform (DCT) loss. The encoding loss is used to represent the bit rate loss between the encoded data and the original data, and the DCT loss is used to represent the coding complexity loss.
[0122] Taking the example of a proxy neural network including a proxy encoding network and an execution network for image analysis tasks, an image preprocessing network training system is shown in Figure 5. Referring to Figure 5, the image preprocessing network training system includes a network to be trained, a proxy encoding network, and an execution network for image analysis tasks.
[0123] During training, a sample image is input into the network to be trained. After processing by the trained network, the image to be analyzed is obtained. Furthermore, the mean square error (MSE) between the image to be analyzed and the sample image can be calculated and used as a preprocessing loss for the network to be trained. The image to be analyzed is then input into the proxy encoding network. The proxy encoding network is then invoked to estimate the DCT loss and coding loss of the image to be analyzed. The proxy encoding network is then invoked to encode the image to be analyzed, obtaining encoded image data. Furthermore, the image data is input into the image analysis task execution network, which then invokes the image analysis task execution network to perform the image analysis task on the image data. After obtaining the analysis results, the analysis loss is calculated.
[0124] Furthermore, the sum of MSE, DCT loss, coding loss and analysis loss is taken as the loss value of this round of training. If the loss value converges, the network to be trained can be used as the image preprocessing network corresponding to the image analysis task; otherwise, the parameters of the network to be trained are adjusted and the above training process is continued until the loss value converges.
[0125] It can be seen that since the image preprocessing network in the embodiment of the present application is a neural network based on deep learning, it can not only be used as an image analysis task for the CV network, but also can be trained using an agent neural network based on deep learning, so that the image preprocessing network of the technical solution of the present application can be flexibly and widely applicable to image analysis tasks in various scenarios.
[0126] Corresponding to the above-mentioned image preprocessing method for machine vision, an embodiment of the present application further provides an image preprocessing device for machine vision, which can be deployed in the image preprocessing network of the CV-oriented analysis system illustrated in FIG1 . The CV-oriented image preprocessing device can modularize the functions of the algorithms and enhancement models in the above-mentioned image preprocessing network through software, hardware, or a combination of software and hardware, and can be used to execute the image preprocessing method for machine vision provided in any of the above-mentioned embodiments.
[0127] The device may include: a blur processing module and an enhancement module. The blur processing module is configured to blur an original image, and the processed image is the image to be enhanced, wherein the clarity of the image to be enhanced is lower than that of the original image; and the enhancement module is configured to enhance the image to be enhanced, wherein the image after the enhancement is the target image, wherein the intensity of the semantic features of the target image is greater than or equal to the intensity of the semantic features of the original image, and the target image is used as the input image for a neural network to perform an image analysis task based on the semantic features.
[0128] The image preprocessing device provided in the embodiment of the present application and the image preprocessing method for machine vision provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, operated or implemented by them.
[0129] The embodiment of the present application also provides a computer device for use in a CV-oriented analysis system to perform the above-mentioned image preprocessing method for machine vision. Please refer to Figure 7, which shows a schematic diagram of a computer device provided by some embodiments of the present application. As shown in Figure 7, the computer device 7 includes: a processor 700, a memory 701, a bus 702 and a communication interface 703, and the processor 700, the communication interface 703 and the memory 701 are connected via the bus 702; the memory 701 stores a computer program that can be run on the processor 700, and when the processor 700 runs the computer program, it executes the image preprocessing method for machine vision provided by any of the aforementioned embodiments of the present application.
[0130] The memory 701 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage. The communication connection between the device network element and at least one other network element is achieved through at least one communication interface 703 (which may be wired or wireless), and may use the Internet, a wide area network, a local area network, a metropolitan area network, etc.
[0131] The bus 702 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory 701 is used to store programs, and the processor 700 executes the programs after receiving execution instructions. The image preprocessing method for machine vision disclosed in any of the aforementioned embodiments of the present application may be applied to the processor 700 or implemented by the processor 700.
[0132] The processor 700 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor 700 or by software instructions. The processor 700 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 701 , and the processor 700 reads the information in the memory 701 and completes the steps of the above method in combination with its hardware.
[0133] The computer device provided in the embodiment of the present application and the image preprocessing method for machine vision provided in the embodiment of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, operated or implemented by them.
[0134] An embodiment of the present application also provides a computer-readable storage medium corresponding to the image preprocessing method for machine vision provided by the aforementioned embodiment, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it will execute the image preprocessing method for machine vision provided by any of the aforementioned embodiments.
[0135] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.
[0136] The computer-readable storage medium provided in the above-mentioned embodiments of the present application and the image preprocessing method for machine vision provided in the embodiments of the present application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the application programs stored therein.
[0137] It should be noted that:
[0138] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some examples, well-known structures and technologies are not shown in detail so as not to obscure the understanding of this description.
[0139] Similarly, it should be understood that in order to streamline the present application and aid in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present application, various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, this disclosed method should not be interpreted as reflecting the following schematic diagram: the claimed application requires more features than the features expressly recited in each claim. Rather, as reflected in the claims below, inventive aspects lie in less than all the features of the individual embodiments disclosed above. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim itself serving as a separate embodiment of the present application.
[0140] Furthermore, those skilled in the art will appreciate that although some embodiments described herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of this application and to form different embodiments. For example, in the claims below, any of the claimed embodiments may be used in any combination.
[0141] The above description is merely a preferred embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. An image preprocessing method for machine vision, characterized in that: include: Performing blur processing on the original image to generate an image to be enhanced, wherein the clarity of the image to be enhanced is lower than the clarity of the original image; Performing enhancement processing on the semantic features of the image to be enhanced to generate a target image; The target image is input into an image processing neural network to trigger the image processing neural network to perform an image analysis task based on the semantic features of the target image.
2. The method according to claim 1, characterized in that The step of enhancing the semantic features of the image to be enhanced to generate a target image includes: generating enhancement parameters according to the original image and the image to be enhanced, wherein the enhancement parameters are used to characterize the degree of enhancement of the image to be enhanced; The target image is generated by performing enhancement processing on the semantic features of the image to be enhanced according to the enhancement parameters.
3. The method according to claim 2, characterized in that The step of generating enhancement parameters according to the original image and the image to be enhanced includes: generating a residual image between the original image and the image to be enhanced; Extracting semantic features of the original image to obtain a first semantic feature, and extracting features of the residual image; The enhancement parameter is obtained by performing feature fusion on the first semantic feature and the feature of the residual image.
4. The method according to claim 2, characterized in that: The step of generating enhancement parameters according to the original image and the image to be enhanced includes: Extracting semantic features of the original image to obtain a first semantic feature, and extracting semantic features of the image to be enhanced as a second semantic feature, wherein both the first semantic feature and the second semantic feature include features of at least two dimensions; For a feature of any dimension, calculating a difference between a feature of the dimension in the first semantic feature and a feature of the dimension in the second semantic feature to obtain at least two differences; The enhancement parameter is calculated based on the at least two difference values.
5. The method according to claim 4, characterized in that The calculating the enhancement parameter according to the at least two differences comprises: For any difference, multiply the difference by the weight of the difference, and the multiplication result is the enhancement factor corresponding to the difference, so as to obtain at least two enhancement factors; wherein the weight is used to characterize the degree of enhancement of the feature of the dimension corresponding to the difference; A sum of the at least two enhancement factors is calculated, and the sum is the enhancement parameter.
6. The method according to claim 3 or 4, characterized in that: The step of extracting the semantic feature of the original image to obtain a first semantic feature includes: Extracting global semantic features of the original image; Performing at least one dimensionality reduction process on the global semantic features to obtain initial image features of a target object, wherein the target object refers to an object related to an image analysis task of the image processing neural network; The initial image feature is subjected to at least one dimensionality increase process to obtain the first semantic feature.
7. The method according to claim 6, characterized in that The performing at least one dimensionality upgrading process on the initial image features comprises: Upsampling the initial image features; The upsampled features are concatenated with the reduced-dimensionality features of the same size, wherein the reduced-dimensionality features of the same size refer to features of the same size as the upsampled features obtained during the at least one dimensionality reduction process; If the size of the spliced feature is the same as the size of the global semantic feature, use the spliced feature as the first semantic feature; If the size of the spliced feature is smaller than the size of the global semantic feature, the spliced feature is used as a new initial image feature, and the new initial image feature is upsampled.
8. The method according to claim 1, characterized in that The blurring of the original image comprises: The original image is blurred by any of the following methods: Noise points are added to the original image, degradation processing is performed on the original image, and blurring processing is performed on the original image using an image blurring algorithm.
9. The method according to claim 8, characterized in that The performing degradation processing on the original image comprises: Downsampling the original image according to a preset sampling ratio; The down-sampled image is up-sampled according to the preset sampling ratio, and the up-sampled image is an image obtained by degradation processing.
10. The method according to claim 9, characterized in that If the resolution of at least one direction of the width and height of the original image is greater than or equal to a preset value, determining that the preset sampling magnification is a first preset sampling magnification; If the resolutions of the original image in width and height directions are both smaller than the preset values, determining the preset sampling magnification to be a second preset sampling magnification; Wherein, the first preset sampling ratio is greater than the second preset sampling ratio.
11. The method according to claim 1, characterized in that: Also includes: Using a proxy neural network to train the network to be trained to obtain an image preprocessing network, wherein the proxy neural network is used to perform the image analysis task; The image preprocessing network is used to execute the image preprocessing method for machine vision, and the image preprocessing network is a deep learning network; The image preprocessing network is connected to the image processing neural network.
12. The method according to claim 11, characterized in that The method of using a proxy neural network to train the network to be trained to obtain an image preprocessing network includes: Calling the network to be trained to preprocess the sample image, the preprocessed image being the image to be analyzed; Calling the proxy neural network to perform the image analysis task on the image to be analyzed to obtain a prediction analysis result; Calculating a loss value, wherein the loss value includes a preprocessing loss, an analysis loss, and a proxy loss; the preprocessing loss is used to characterize the loss between the image to be analyzed and the sample image, the analysis loss is used to characterize the loss between the prediction analysis result and the sample image annotation result, and the proxy loss is used to characterize the loss of the sample image processing process by the proxy neural network; Determining whether the loss value converges; If the loss value converges, using the network to be trained as the image preprocessing network; If the image preprocessing network has not converged, the parameters of the network to be trained are adjusted, and the model after the parameter adjustment is used as a new network to be trained, and the operation of calling the network to be trained to preprocess the sample image is executed again.
13. The method according to claim 12, characterized in that The proxy neural network includes a proxy encoding network and an execution network of the image analysis task, and calling the proxy neural network to perform the image analysis task on the image to be analyzed to obtain a prediction analysis result includes: Calling the proxy coding network to encode the image to be analyzed, and obtaining a coding loss and a discrete cosine transform (DCT) loss, wherein the proxy loss includes the coding loss and the DCT loss, the coding loss is used to characterize the bit rate loss between the encoded data and the pre-encoded data, and the DCT loss is used to characterize the coding complexity loss; The execution network is called to perform the image analysis task on the encoded image data, and the analysis loss is obtained.
14. The method according to claim 12 or 13, characterized in that The network to be trained includes a fuzzy algorithm module and an enhanced model to be trained, and the adjusting of the parameters of the network to be trained includes: Adjust the parameters of the enhanced model to be trained.
15. The method according to claim 1, characterized in that After inputting the target image into the image processing neural network, the method further includes: Calling the image processing neural network to encode the target image to obtain encoded data; The analysis module of the image processing neural network is called to perform the image analysis task on the encoded data to output an analysis result by analyzing the semantic features of the target image.
16. The method according to claim 11, characterized in that The network to be trained is a pre-constructed initial network or a pre-processed network to be trained, and the pre-processed network to be trained is suitable for another image analysis task, and the other image analysis task is different from the image analysis task; If the network to be trained is the preprocessed network to be trained, before adopting the proxy neural network to train the network to be trained to obtain the image preprocessed network, the method further includes: The network to be trained is disconnected from the original image processing neural network, where the original image processing neural network is the image processing neural network that performs the other image analysis task.
17. The method according to claim 16, characterized in that The image analysis task includes at least one of the following: image matching, image recognition, image segmentation, image content extraction and face recognition.
18. An image preprocessing device for machine vision, characterized in that: include: A blur processing module, used for blurring the original image to generate an image to be enhanced, wherein the clarity of the image to be enhanced is lower than the clarity of the original image; An enhancement module, used for enhancing the semantic features of the image to be enhanced to generate a target image; An input module is used to input the target image into an image processing neural network to trigger the image processing neural network to perform an image analysis task based on the semantic features of the target image.
19. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The processor runs the computer program to implement the method according to any one of claims 1-17.
20. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the method according to any one of claims 1 to 17.
Citation Information
Patent Citations
Image enhancement method and device
CN112446834A
Image enhancement method and device and automatic driving control method and device
CN113744141A
Image target retrieval method and device, electronic device and storage medium
CN114462490A
Image preprocessing method and device for machine vision, equipment and storage medium
CN117422855A
Center-biased machine learning techniques to determine saliency in digital images
US20210012201A1
Cited By
Infrared super lens image enhancement design method and device
CN121921192A
Image enhancement method, device and equipment for OCR (Optical Character Recognition) system and medium
CN122023190A