Image Enhancement Method, Device, Electronic Device and Storage Medium
Through the image enhancement method combining edge detection and label features, the problem of lack of authenticity and diversity of generated images in the prior art is solved, high-quality data enhancement is achieved, and the performance of the object detection model and its ability to adapt to complex scenarios is improved.
Patent Information
- Application Number
- CN202510377749.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-28
AI Technical Summary
Existing object detection data enhancement methods are difficult to accurately control the diversity and authenticity of generated images, resulting in insufficient generalization capabilities of the model in complex scenarios.
The contour features of the target image are obtained through edge detection, and combined with the label features of the target detection label for image enhancement, generate target enhanced images with the same target detection label as the original image, and use diffusion model and deep learning technology to simulate the image style and environment.
High-quality and diverse image enhancement is achieved, and the categories and locations of objects in the generated image can be accurately controlled, providing high-quality data sets, and improving the performance and generalization capabilities of the object detection model.
Smart Images

Figure CN119904374B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to an image enhancement method, device, electronic device, and storage medium. Background Art
[0002] In recent years, object detection technology has developed towards multi-scale, small object, real-time, etc., such as RetinaNet, Feature Pyramid Network, etc. Although the proposed object detection algorithms have improved the efficiency and accuracy of object detection, they still face problems such as data scarcity, class imbalance, and single scene, which seriously restrict the generalization ability and practical application effect of the model.
[0003] Current object detection data augmentation methods, such as Simple Copy-Paste, although alleviating the problem of data scarcity to a certain extent and improving model performance, their limitations are also obvious. Most of these methods rely on simple image processing techniques, such as cutting and pasting target objects into a new background environment. Although simple and fast, they often ignore the complex interaction relationship between the target object and the background environment, resulting in the newly generated images lacking authenticity and naturalness.
[0004] In addition, these methods also have deficiencies in the control of data diversity. It is difficult for them to precisely control the diversity of the generated images and effectively simulate different lighting conditions, weather conditions, and background environments, etc., thus restricting the scene adaptability of the augmented data. This leads to the performance of the model often being unsatisfactory when facing complex and changing actual scenes. Summary of the Invention
[0005] The present invention provides an image enhancement method, device, electronic device, and storage medium to solve the problems in the prior art that the data augmentation process lacks authenticity and diversity, resulting in the generated images not being real and natural enough, with limited application scenarios, and further affecting the generalization ability and application effect of the model, and realizing intelligent, refined, and diverse image enhancement.
[0006] The present invention provides an image enhancement method, including:
[0007] Determine a target image, where the target image carries an object detection label;
[0008] Perform edge detection on the target image to obtain contour features;
[0009] Perform image enhancement on the target image based on the contour features and label features to obtain a target enhanced image;
[0010] Among them, the target enhanced image and the target image carry the same target detection label, the target detection label includes the coordinates and categories of the detection frames of each target in the corresponding image, and the label feature is constructed based on the target detection label.
[0011] According to an image enhancement method provided by the present invention, the label feature is determined based on the following steps:
[0012] Encode the categories of the detection frames of each target in the target detection label to obtain the category prompt word features of each target;
[0013] Embed the coordinates of the detection frames of each target in the target detection label to obtain the detection frame position features of each target;
[0014] Fuse the category prompt word features and detection frame position features of each target to obtain the label feature of the target image.
[0015] According to an image enhancement method provided by the present invention, the step of embedding the coordinates of the detection frames of each target in the target detection label to obtain the detection frame position features of each target includes:
[0016] Extract the diagonal coordinates of the detection frames of each target in the target detection label;
[0017] Perform Fourier embedding on the diagonal coordinates to obtain the detection frame position features of each target;
[0018] Among them, the diagonal coordinates are the coordinates of two vertices on the diagonal of the corresponding detection frame.
[0019] According to an image enhancement method provided by the present invention, the step of performing image enhancement on the target image based on the contour feature and the label feature to obtain a target enhanced image includes:
[0020] Perform image enhancement on the target image based on the contour feature, label feature, and multiple noise images to obtain multiple initial enhanced images; each initial enhanced image carries the target detection label; each noise image is obtained by random sampling;
[0021] Based on the detection frames of each target in any one of the initial enhanced images, determine the enhanced regions corresponding to each target;
[0022] Based on the enhanced regions corresponding to each target in any one of the initial enhanced images and the label feature of the target image, determine the generation score of any one of the initial enhanced images;
[0023] Based on the generation scores of each initial enhanced image, screen out the target enhanced image from each initial enhanced image.
[0024] An image enhancement method provided by the present invention, which performs image enhancement on the target image based on the contour feature and the label feature to obtain a target enhanced image, includes:
[0025] Input the contour feature and the label feature into an image enhancement model to obtain the target enhanced image output by the image enhancement model;
[0026] Among them, the image enhancement model is trained based on a diffusion model, using a sample image and the sample object detection label of the sample image.
[0027] An image enhancement method provided by the present invention, the image enhancement model is trained based on the following steps:
[0028] Based on the diffusion model, construct a first enhancement model and a second enhancement model, and based on the first enhancement model and the second enhancement model, construct an initial enhancement model;
[0029] Train the initial enhancement model based on the sample image, the sample object detection label of the sample image, and the sample prompt text of each object in the sample image to obtain the image enhancement model;
[0030] Among them, the first enhancement model and the second enhancement model in the initial enhancement model are connected by a convolutional layer; the first enhancement model is constructed by adding layers to the attention mechanism in the diffusion model, and except for the parameters of the gated self-attention layer added by the layer addition, other parameters are not updated during training.
[0031] According to an image enhancement method provided by the present invention, the layer addition is to add a gated self-attention layer between each self-attention layer and each cross-attention layer in the attention mechanism of the diffusion model, and the input of the gated self-attention layer includes the sample label feature corresponding to the sample object detection label and the sample contour feature corresponding to the sample image;
[0032] The second enhancement model is constructed by removing the decoding layer based on the diffusion model.
[0033] The present invention also provides an image enhancement device, including:
[0034] A determination unit for determining a target image, the target image carrying an object detection label;
[0035] A detection unit for performing edge detection on the target image to obtain a contour feature;
[0036] An enhancement unit for enhancing the target image based on the contour feature and the label feature to obtain a target enhanced image;
[0037] Wherein, the target enhanced image and the target image carry the same target detection label, the target detection label includes the coordinates and categories of the detection frames of each target in the corresponding image, and the label feature is constructed based on the target detection label.
[0038] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, the image enhancement method as described in any one of the above is implemented.
[0039] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the image enhancement method as described in any one of the above is implemented.
[0040] The image enhancement method, device, electronic device, and storage medium provided by the present invention perform edge detection on a target image to obtain a contour feature, and perform image enhancement based on the contour feature and the label feature of the target detection label carried by the target image to obtain a target enhanced image with the same target detection label as the target image; the target detection label includes the coordinates and categories of the detection frames of each target in the corresponding image, overcoming the defects in the traditional solution that it is difficult to precisely control the generated image, and the generated image lacks authenticity and naturalness. It not only realizes high-quality and diverse image enhancement, but also can precisely control the categories and positions of objects in the generated image, thereby providing a large number of high-quality data sets for the target detection task, further helping to improve the performance of the target detection model and enhance its generalization ability, making it more adaptable to complex tasks and performing better in specific tasks. Description of the Drawings
[0041] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 It is a flowchart of the image enhancement method provided by the present invention;
[0043] Figure 2 It is an overall framework diagram of the image enhancement method provided by the present invention;
[0044] Figure 3 It is a structural diagram of the initial enhancement model provided by the present invention;
[0045] Figure 4 It is a schematic diagram of the addition in the attention mechanism layer provided by the present invention;
[0046] Figure 5 It is a schematic structural diagram of the image enhancement device provided by the present invention;
[0047] Figure 6 It is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0048] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0049] In the field of object detection, with the continuous evolution of algorithms, the proposal of advanced models such as RetinaNet and Feature Pyramid Network has significantly improved the accuracy and efficiency of object detection, especially making important progress in dealing with multi-scale objects and small objects. In addition, it has also contributed to the improvement of the real-time requirement, enabling object detection technology to show broad application potential in multiple fields such as autonomous driving, video surveillance, and medical image analysis.
[0050] However, although these algorithms have made significant progress in theory, they still face many challenges in practical applications. Among them, problems such as data scarcity, class imbalance, and single scene are particularly prominent. These problems seriously restrict the generalization ability of the object detection model and affect the performance of the model in different environments and conditions. To solve these problems, researchers have continuously explored effective data enhancement methods in order to improve the performance of the model by increasing the diversity and richness of training data.
[0051] Current object detection data augmentation methods, such as Simple Copy-Paste, etc., although alleviating the problem of data scarcity to a certain extent and improving the performance of the model, still have many limitations. First of all, these methods mostly rely on simple image processing techniques, such as cutting and pasting, by pasting the target object into a new background to generate a new image. Although this method is simple and fast, it ignores the complex interaction relationship between the target object and the background environment, resulting in the generated images often lacking authenticity and naturalness and being difficult to effectively simulate complex situations in real scenes. Secondly, these methods also have deficiencies in controlling data diversity, that is, it is difficult to precisely control the diversity of the generated images and effectively simulate different lighting conditions, weather conditions, background environments, etc., thus limiting the scene adaptability of the augmented data and causing the model's performance to often be unsatisfactory when facing complex and changeable actual scenes.
[0052] In response to this, the present invention provides an image enhancement method, aiming to overcome the limitations of the current data augmentation methods. By respectively obtaining the contour features and label features of the target image and its corresponding object detection label, and based on this, performing image enhancement to obtain the target enhanced image. In this way, the category and position of the target object in the generated image can be precisely controlled, high-quality and diverse images can be realized, thereby providing a large number of high-quality data sets for the object detection task, further effectively improving the performance of the object detection model, and enhancing its generalization ability, providing more solid data support for the application of object detection technology.
[0053] Figure 1 is a schematic flowchart of the image enhancement method provided by the present invention, as Figure 1 shown, this method includes:
[0054] Step 110, determining a target image, the target image being accompanied by an object detection label;
[0055] Step 120, performing edge detection on the target image to obtain contour features;
[0056] Step 130, performing image enhancement on the target image based on the contour features and label features to obtain a target enhanced image;
[0057] Among them, the target enhanced image and the target image carry the same object detection label, and the object detection label includes the coordinates and categories of the detection frames of each target in the corresponding image, and the label features are constructed based on the object detection label.
[0058] Specifically, considering the two aspects of problems existing in the current object detection data augmentation methods, namely, it is difficult to precisely control the generated images, the diversity of the generated images cannot be guaranteed, different light background environments cannot be simulated, the scene adaptability is limited, and the generated images lack authenticity and naturalness. In the embodiments of the present invention, it is proposed to perform image augmentation through the target image and its associated object detection labels, so that the categories and positions of the objects in the generated images can be precisely controlled, and various real and complex scenes (such as lighting conditions, weather conditions, background environments, etc.) can be simulated, thereby realizing intelligent and diversified data augmentation for the original images, and real, diverse, and natural generated images can be obtained, solving the problems of difficult precise control and lack of authenticity and naturalness in the current data augmentation methods, and being able to provide a large number of high-quality data sets for the object detection task, and further significantly improving the performance and generalization ability of the object detection model.
[0059] It should be noted that the method provided by the present invention can not only be applied to the object detection task to provide training data for the object detection model, but also be applied to other image perception tasks, such as semantic segmentation, image classification, object tracking, etc., and can provide data support for many types of image perception tasks, and contribute to the continuous progress of the image perception tasks.
[0060] In detail, in the embodiments of the present invention, before performing image augmentation, it is first necessary to determine the image to be augmented, which is called the target image here. The target image can be acquired through various image acquisition devices, such as ordinary cameras, wide-angle cameras, depth cameras, infrared cameras, multi-spectral cameras, etc.; it can also be obtained through network search and download, and can also be generated through image generation algorithms, models, etc. The embodiments of the present invention do not make specific limitations.
[0061] Here, the target image can be images in various fields, such as security monitoring, medical imaging, industrial production, agricultural production, retail warehouses, etc. In the embodiments of the present invention, specifically when obtaining the target image, it can be to first obtain the original image and its corresponding annotation information. The annotation information here can include the categories of each target in the original image and the coordinates of the detection boxes, and then the original image and its annotation information can be subjected to image transformation, such as random cropping, size transformation, etc., so as to obtain the transformed image and its corresponding annotation information. At this time, the image and its annotation information are the target image and its corresponding object detection labels.
[0062] After obtaining the target image, in the embodiments of the present invention, this target image can be processed to extract the edge information of the objects therein, so as to obtain the contour features. Here, specifically, through the visual prior generator, the target image can be processed to extract the edge contour information, so as to obtain the contour features corresponding to the target image.
[0063] Among them, the visual prior generator can be an edge detector, such as HED (Holistically-Nested Edge Detection). The process of contour feature extraction is to calculate the edge information of the objects in the target image through the edge detection operator of the edge detector, and the overall contour features can be obtained. Here, contour feature extraction through the edge detector can better balance visual diversity and the quality of the bounding box, making the extracted contour features more obvious.
[0064] Furthermore, after obtaining the contour features, in the embodiments of the present invention, based on these contour features, combined with the label features, image enhancement can be performed to obtain the generated enhanced image, that is, the enhanced image corresponding to the target image, which is called the target enhanced image.
[0065] Specifically, here it can be to first draw a contour according to the contour features to draw an image contour consistent with the current contour features; then, according to the category and position of the target represented by the label features, color assignment, texture filling, etc. can be performed for each target in the contour, and thus the preliminarily generated image can be obtained; after that, the preliminarily generated image can also be enhanced, such as color enhancement, contrast enhancement, detail enhancement, etc., to make it clearer, more vivid, and more detailed, and finally the target enhanced image can be obtained.
[0066] Among them, the label features are constructed based on the object detection labels corresponding to the target image. Since the object detection labels include the categories of each object in the image and the coordinates of the detection boxes, when constructing the label features that can represent the information contained in the object detection labels, it can be to extract features from these two types of information respectively, and on this basis, the overall label features are constructed.
[0067] Specifically, it can be to perform image enhancement on the target image according to the label features and the contour features to obtain the generated target enhanced image. That is, based on the contour features, guided by the label features, image generation is carried out to generate an image with the same contour and consistent with the object category and position represented by the label features, but with different styles, environments, etc.
[0068] It should be noted that although the target enhanced image obtained through the above image enhancement process has the same contour features as the original target image and carries the same object detection labels, the visual features of the two are significantly different. For example, it can have different lighting conditions, weather conditions, background environments, etc., or be an image directly simulating other scenes, or have undergone style transformation, migration, etc. The embodiments of the present invention do not make specific limitations.
[0069] In the embodiments of the present invention, taking images and object detection labels as inputs, image enhancement processing is performed through deep learning techniques, and a series of target-enhanced images are output, which keep the positions and categories of the original targets unchanged, but are diverse in terms of image style, environmental conditions, etc., thus realizing intelligent and diverse image enhancement. It not only retains the key target information in the original target image, but also can simulate various complex real scenarios, greatly improving the diversity, authenticity, and naturalness of the generated images, and is a high-quality data reserve for object detection tasks.
[0070] In the embodiments of the present invention, it can flexibly adapt to different types of image inputs. By optimizing the input images, enhanced images are obtained, which can be applied to multiple fields that require high-precision object detection, such as autonomous driving, security monitoring, medical image analysis, industrial quality inspection, agricultural intelligence, retail, etc., providing high-quality data support for them.
[0071] In addition, the method provided in the embodiments of the present invention can also be combined with other advanced object detection techniques, such as multi-scale detection, feature pyramid network, attention mechanism, etc., which can further improve the detection performance and accuracy. For example, if the method provided in the present invention is integrated into the training process of object detection tasks, dynamic generation and update of training samples can be achieved, enabling the model to continuously learn and adapt to new scenarios and object categories, thus laying a solid foundation for the construction of a more intelligent and adaptable computer vision system.
[0072] The image enhancement method provided by the present invention performs edge detection on the target image to obtain contour features, and performs image enhancement based on the contour features and the label features of the object detection labels carried by the target image, obtaining a target-enhanced image with the same object detection label as the target image; the object detection label includes the coordinates and categories of the detection frames of each object in the corresponding image, overcoming the defects in the traditional scheme that it is difficult to precisely control the generated image, and the generated image lacks authenticity and naturalness. It not only realizes high-quality and diverse image enhancement, but also can precisely control the categories and positions of objects in the generated image, thus providing a large number of high-quality data sets for object detection tasks, further helping to improve the performance of object detection models and enhance their generalization ability, making them more adaptable to complex tasks and performing better in specific tasks.
[0073] Based on the above embodiments, the label features are determined based on the following steps:
[0074] Encode the categories of the detection frames of each object in the object detection label to obtain the category prompt word features of each object;
[0075] Embed the coordinates of the detection frames of each object in the object detection label to obtain the detection frame position features of each object;
[0076] Fuse the category prompt word features and detection box position features of each target to obtain the label features of the target image.
[0077] Specifically, the determination of the label features of the target image can be achieved through the following steps:
[0078] Since the target detection label is composed of detection box-category pairs, and the category and detection box in it represent two different types of information respectively. The detection box reflects the position of the target, and the category represents the object type to which the target belongs. Therefore, in order to better reflect the information of the target, in the embodiments of the present invention, these two types of information can be processed separately to obtain the features corresponding to each of the two types of information, and then the overall features can be constructed on this basis to obtain the final features, that is, the label features of the target image.
[0079] In detail, here can be to first construct prompt words based on the categories of the detection boxes of each target in the target detection label to obtain the category prompt word features. That is, for each target in the target image, first obtain its category, and then encode the category of the target, such as encoding it using the multi-modal model CLIP (Contrastive Language–Image Pre-training) into features, so as to obtain the category prompt word features, which can be expressed as , represents the th category prompt word feature of the
[0080] At the same time, the position features can be constructed based on the coordinates of the detection boxes of each target in the target detection label to obtain the detection box position features of each target. That is, for each target in the target image, first obtain the coordinates of the detection box, and then perform an embedding representation on the coordinates to represent it as an embedding vector of the coordinates, so as to obtain the detection box position features of each target, which can be expressed as , represents the th detection box position feature of the
[0081] After that, the label features of the target image can be determined according to the category prompt word features and detection box position features of each target in the target image; that is, the features obtained by separately extracting the features of the two types of information in the target detection label can be fused to obtain the features at the overall level, that is, the category prompt word features and detection box position features of each target in the target image are fused to obtain the label features of the target image.
[0082] Based on the above embodiments, the label features of the target image can be represented by the following formula:
[0083]
[0084]
[0085] In the formula, represents the label features of the th target image, represents the number of targets in the th target image, represents the label features of the th target, and respectively represent the category prompt word features of the th target and the detection box position features of the th target in the th target image, is Multilayer Perceptron, representing a multi-layer perceptron.
[0086] In the embodiments of the present invention, based on the target detection labels, through prompt word construction, position feature extraction, and feature fusion, the overall label features are obtained, and on this basis, combined with the contour features of the target image obtained through the visual prior generator, the target image is enhanced together to obtain multiple target enhanced images that are the same as the contour features and target detection labels. In this way, high-quality data enhancement is achieved, diverse images can be generated while the key target information remains unchanged, and different lighting conditions, weather conditions, background environments, etc. can be simulated, greatly improving the scene adaptability of the enhanced data.
[0087] Based on the above embodiments, the coordinates of the detection boxes of each target in the target detection labels are embedded to obtain the detection box position features of each target, including:
[0088] Extract the diagonal coordinates of the detection boxes of each target in the target detection labels;
[0089] Perform Fourier embedding on the diagonal coordinates to obtain the detection box position features of each target;
[0090] Among them, the diagonal coordinates are the coordinates of two vertices on the diagonal of the corresponding detection box.
[0091] Specifically, the process of embedding the coordinates of the detection boxes of each target in the target detection labels to obtain the detection box position features of each target may specifically include:
[0092] First, according to the object detection labels, the diagonal coordinates of the detection boxes of each object in the target image can be determined. Here, the diagonal coordinates refer to the coordinates of the two endpoints of the diagonal of the detection box. Since the object detection labels already contain the coordinates of the detection box, here, the coordinates of the upper left corner and the lower right corner can be directly obtained, or the coordinates of the upper right corner and the lower left corner can be obtained, and directly used as the diagonal coordinates of the detection boxes of each object.
[0093] Then, the Fourier embedding representation can be performed on the diagonal coordinates of each detection box to represent it as an embedding vector of the coordinates, so as to obtain the detection box position features of each object; that is, Fourier embedding can be performed on the basis of the diagonal coordinates of each detection box, and an embedding vector of the position of each detection box can be obtained after the Fourier embedding representation.
[0094] Based on the above embodiments, step 130 includes:
[0095] Based on the contour features, label features, and multiple noise images, perform image enhancement on the target image to obtain multiple initial enhanced images; each initial enhanced image is accompanied by object detection labels; each noise image is obtained by random sampling;
[0096] Based on the detection boxes of each object in any one of the initial enhanced images, determine the enhanced regions corresponding to each object;
[0097] Based on the enhanced regions corresponding to each object in the initial enhanced image and the label features of the target image, determine the generation score of the initial enhanced image;
[0098] Based on the generation scores of each initial enhanced image, screen out the target enhanced image from each initial enhanced image.
[0099] Specifically, in the above step 130, the process of performing image enhancement on the target image according to the contour features and label features to obtain the target enhanced image may specifically include the following steps:
[0100] Figure 2 is the overall framework diagram of the image enhancement method provided by the present invention. As Figure 2 shown, after obtaining the contour features and label features of the target image according to the visual prior generator and the prompt constructor respectively, specifically when performing image enhancement, an input image can also be provided. The input image can be a blank image or a noise image. Preferably, in the embodiments of the present invention, a noise image is selected as the input image, and the image enhancement process is performed on the basis of this noise image. The specific goal is to restore the noise image to a clear image that is consistent with the contour and object detection labels of the target image, so as to obtain the generated diverse target enhanced images.
[0101] Specifically, here, it may be possible to first obtain a noise image, which can be obtained by random sampling. Specifically, it can be randomly sampled from a Gaussian distribution, a standard normal distribution, etc., so as to obtain a noise image, and this can be used as an input for image enhancement. Specifically, the noise image can be preprocessed first. For example, the noise can be reduced by Gaussian filtering or other means to improve the visual effect. Then, the contour features can be matched with the preprocessed noise image. For example, contour matching can be performed on the basis of the preprocessed noise image through image registration, feature point matching, etc. After that, according to the label features, color filling, texture replacement, etc. can be performed on the objects in the matched contours. Finally, enhancement processing can be performed, such as color enhancement, contrast enhancement, detail enhancement, etc., to make it clearer, more vivid, and more detailed, and a diversified target-enhanced image can be obtained, that is, an image that has the same contour features as the target image and is consistent with the object category and position represented by the label features but has different visual features.
[0102] However, it is worth noting that since the final image is generated on the basis of the noise image, and the noise image is obtained by random sampling, when the sampled noises are different, even if the given conditions (contour features and label features) are the same, the finally generated images will necessarily be different. In order to achieve diversified image enhancement, usually in the sampling process, multiple noises are randomly sampled. Therefore, multiple images are often generated, and the quality of these multiple images cannot be ensured one by one, and further screening and confirmation are required. Therefore, the initially generated images can be called initial enhanced images, and these initial enhanced images carry the same object detection labels as the target image.
[0103] Furthermore, after obtaining multiple initial enhanced images, in the embodiments of the present invention, these multiple initial enhanced images also need to be screened to filter out inferior images and obtain high-quality images as reserve data for subsequent object detection tasks, so as to improve the performance and generalization ability of the object detection model, and further optimize its performance in object detection tasks.
[0104] Specifically, for any one of the initially generated enhanced images, first, the initial enhanced image can be cropped according to the detection frames of each object to crop out the images corresponding to each detection frame from it, that is, the enhanced regions corresponding to each object are cropped from the initial enhanced image according to the detection frames.
[0105] Immediately afterwards, according to the enhanced regions corresponding to each object cropped from the initial enhanced image and the categories of each object in the object detection label, a score for the generation quality of this part of the target image in the image enhancement task can be performed to judge whether the enhanced regions of each generated object match and correspond to the originally recorded categories, so as to obtain the similarity matching scores corresponding to each object.
[0106] Preferably, similarity matching scoring can be performed here by means of the multi-modal model CLIP, that is, similarity calculation can be performed according to the enhanced region and category prompt word features corresponding to each target, so as to obtain the similarity matching score. , , denotes the th enhanced region corresponding to the th target in the initial enhanced image corresponding to the th target image, and denotes the category prompt word feature of the
[0107] After that, the similarity matching scores of each target in the initial enhanced image can be statistically analyzed to obtain the score at the image level. Here, specifically, the similarity matching scores of each target in the initial enhanced image can be fused. The specific fusion method can be averaging, weighted summation, or other methods. The embodiments of the present invention do not make specific limitations. Finally, the score of the initial enhanced image can be obtained, that is, the score is generated.
[0108] Furthermore, the image can be screened according to this generated score to screen out high-quality images and remove low-quality images from all the generated initial enhanced images.
[0109] Specifically, here, all the initial enhanced images can be sorted in descending order (or ascending order) according to the generated score to obtain an image sequence; then, according to a preset filtering ratio, low-quality images can be filtered out from this image sequence, and high-quality images can be retained as the final target enhanced images. Here, specifically, according to the filtering ratio, such as 10:3 (0.3) or 5:1 (0.2), the initial enhanced images ranked in the last 70% or 80% (or the first 70% or 80%) in the image sequence can be filtered out, and the first 30% or 20% (the last 30% or 20%) can be retained, and the retained images are used as the final target enhanced images.
[0110] Based on the above embodiments, step 130 includes:
[0111] Input the contour feature and label feature into the image enhancement model to obtain the target enhanced image output by the image enhancement model;
[0112] Among them, the image enhancement model is trained on the basis of the diffusion model by applying the sample image and the sample object detection label of the sample image.
[0113] Specifically, in step 130, the process of image enhancement for the target image based on the contour feature and the label feature can be implemented through an image enhancement model. Specifically, the contour feature and the label feature of the target image, along with the noise image, can be input into the image enhancement model, so that under the guidance of the input information, the image enhancement model restores the noise to a clear image and outputs it, thereby obtaining the target enhanced image.
[0114] It should be noted that in the embodiment of the present invention, the adopted image enhancement model is constructed based on diffusion models (such as Stable Diffusion, Denoising Diffusion Probabilistic Models, etc.). After constructing the initial model during the training process through the diffusion model, in order to better improve the image enhancement ability of the model and make it perform better in practical applications, in the embodiment of the present invention, the initial model also needs to be trained, that is, applying the pre-collected sample images and the sample object detection labels corresponding to the sample images to perform parameter iteration on the initial model, so as to obtain the trained image enhancement model. In this way, the authenticity and diversity of the finally obtained target enhanced image can be better guaranteed.
[0115] Among them, the sample images can include images in multiple scenarios and multiple fields.
[0116] In the embodiment of the present invention, taking images and object detection labels as inputs, through deep learning technology and an image enhancement model constructed based on a diffusion model, a series of target enhanced images are output that maintain the positions and categories of the objects in the original image but are diverse in terms of image style, environmental conditions, etc., realizing intelligent and diverse image enhancement. It not only retains the key target information in the original image but also can simulate various complex real scenarios, greatly enhancing the diversity and authenticity of the generated images.
[0117] Based on the above embodiments, the image enhancement model is trained based on the following steps:
[0118] Based on the diffusion model, construct a first enhancement model and a second enhancement model, and based on the first enhancement model and the second enhancement model, construct an initial enhancement model;
[0119] Based on the sample images, the sample object detection labels of the sample images, and the sample prompt texts of each target in the sample images, train the initial enhancement model to obtain the image enhancement model;
[0120] Among them, in the initial enhanced model, the first enhanced model and the second enhanced model are connected by a convolutional layer; the first enhanced model is constructed by adding layers to the attention mechanism in the diffusion model, and except for the parameters of the gated self-attention layer added by the layer addition in the first enhanced model, other parameters are not updated during the training process.
[0121] Specifically, the training process of the image enhancement model can specifically include:
[0122] First, two models can be constructed based on the diffusion model, namely the first enhanced model and the second enhanced model. Here, specifically, the Stable Diffusion model is selected as the basis, and the first enhanced model and the second enhanced model are constructed based on this.
[0123] Then, based on the first enhanced model and the second enhanced model, the initial model during the training process can be constructed, that is, the initial enhanced model; here, specifically, a convolutional layer can be used to connect the first enhanced model and the second enhanced model, and the first enhanced model and the second enhanced model are connected through multiple convolutional layers, thereby constructing the initial enhanced model.
[0124] Among them, it should be noted that most of the parameters in the first enhanced model here do not participate in the training, that is, they are not updated during the training process; the parameters in the second enhanced model participate in the training.
[0125] Figure 3 is a schematic structural diagram of the initial enhanced model provided by the present invention, as Figure 3 shown, the initial enhanced model retains two copies of the original diffusion model, one copy (the first enhanced model) freezes most of the parameters; the other copy (the second enhanced model) only contains the encoding layer and the intermediate layer in the original diffusion model, and the parameters are trainable; the two are connected by a convolution with the parameters initialized to zero, specifically by convolving to connect the decoding layer in the first enhanced model and the encoding layer in the second enhanced model, and connecting the intermediate layer in the first enhanced model and the intermediate layer in the second enhanced model through another convolution, so as to realize the overall image construction. It should be noted that Figure 3 in the figure, the red path is the path of gradient update during the training process, while the black path is the flow path during actual use. Figure 3 in represents an image that is all noise, then represents the image after one-step iteration of .
[0126] However, it should be noted that compared with the original diffusion model that can only receive text input and generate images based on text guidance, in the embodiments of the present invention, object detection labels and contour features are added to the input. And precisely because of the added input, it is necessary to improve the structure of the original diffusion model, and this improvement specifically occurs in the first enhanced model. Specifically, on the basis of the original diffusion model, its attention mechanism is improved by adding a gated self-attention layer to integrate the information contained in the object detection labels, resulting in the first enhanced model. Moreover, except for the parameters of the newly added gated self-attention layer in the first enhanced model, the parameters of all other layers are frozen and do not participate in the update during the training process. That is, Figure 3 The partial trainable layers in Figure 3 mean that the parameters of the gated self-attention layer are trainable.
[0127] After the initial enhanced model is constructed, in the embodiments of the present invention, the initial enhanced model can be trained to obtain a trained image enhancement model.
[0128] Specifically, here, the initial enhanced model can be trained using sample images, sample object detection labels of the sample images, and sample prompt texts of each object in the sample images; that is, on the basis of the sample images, guided by the sample label features of the sample object detection labels and the sample prompt texts, the model is trained to enable it to have good image enhancement capabilities and be able to output high-quality images, and finally a trained image enhancement model can be obtained.
[0129] It should be noted that the sample prompt text is just a simple descriptive text. For example, when the sample image contains an object table, the sample prompt text can be a simple descriptive statement such as "generate a table". The process of image enhancement mainly relies on the visual features provided by the sample object detection labels and the sample images. The sample prompt text only ensures that the model can execute in accordance and does not make obvious mistakes.
[0130] Based on the above embodiments, the layer addition is to add a gated self-attention layer between each self-attention layer and each cross-attention layer in the attention mechanism of the diffusion model. The input of the gated self-attention layer includes the sample label features corresponding to the sample object detection labels and the sample contour features corresponding to the sample images;
[0131] The second enhanced model is constructed by removing the decoding layer on the basis of the diffusion model.
[0132] Figure 4 is a schematic diagram of layer addition in the attention mechanism provided by the present invention, as Figure 4As shown, when improving the attention mechanism in the diffusion model, a gated self-attention layer is added at all places where text information needs to be fused in the original diffusion model. That is, in the attention mechanism, a gated self-attention layer is added between the self-attention layer and the cross-attention layer. Here, the gated self-attention layer is actually a copy of the cross-attention layer, except that the input text becomes the label. That is, the input of the gated self-attention layer includes the sample label features corresponding to the sample target detection labels and the sample contour features corresponding to the sample images.
[0133] The second enhanced model is obtained by removing the decoding layer on the basis of the diffusion model, that is, only the encoding layer and the intermediate layer in the original diffusion model are retained, and the second enhanced model is thus obtained.
[0134] Based on the above embodiments, see Figure 4 It can be seen that the gated self-attention layer can be expressed as:
[0135]
[0136] In the formula, represents the sample contour features, is the model parameter, represents the sample label features, represents the self-attention layer.
[0137] Next, the image enhancement device provided by the present invention will be described. The image enhancement device described below can be correspondingly referred to the image enhancement method described above.
[0138] Figure 5 is a schematic structural diagram of the image enhancement device provided by the present invention. As Figure 5 shown, the device includes:
[0139] A determination unit 510, configured to determine a target image, where the target image carries a target detection label;
[0140] A detection unit 520, configured to perform edge detection on the target image to obtain contour features;
[0141] An enhancement unit 530, configured to perform image enhancement on the target image based on the contour features and the label features to obtain a target enhanced image;
[0142] Wherein, the target enhanced image and the target image carry the same target detection label, the target detection label includes the coordinates and categories of the detection frames of each target in the corresponding image, and the label features are constructed based on the target detection label.
[0143] The image enhancement device provided by the present invention obtains contour features by performing edge detection on a target image, and performs image enhancement based on the contour features and the label features of the target detection labels carried by the target image to obtain a target enhanced image with the same target detection labels as the target image; the target detection labels include the coordinates and categories of the detection frames of each target in the corresponding image, overcoming the defects in the traditional solution that it is difficult to precisely control the generated image, and the generated image lacks authenticity and naturalness. It not only realizes high-quality and diverse image enhancement, but also can precisely control the categories and positions of objects in the generated image, so as to provide a large number of high-quality data sets for the target detection task, and further helps to improve the performance of the target detection model and enhance its generalization ability, making it more adaptable to complex tasks and performing better in specific tasks.
[0144] Based on the above embodiment, the device further includes a label feature determination unit for:
[0145] Encoding the categories of the detection frames of each target in the target detection label to obtain the category prompt word features of each target;
[0146] Embedding the coordinates of the detection frames of each target in the target detection label to obtain the detection frame position features of each target;
[0147] Fusing the category prompt word features and detection frame position features of each target to obtain the label features of the target image.
[0148] Based on the above embodiment, the label feature determination unit is used for:
[0149] Extracting the diagonal coordinates of the detection frames of each target in the target detection label;
[0150] Performing Fourier embedding on the diagonal coordinates to obtain the detection frame position features of each target;
[0151] Wherein, the diagonal coordinates are the coordinates of two vertices on the diagonal of the corresponding detection frame.
[0152] Based on the above embodiment, the enhancement unit 530 is used for:
[0153] Performing image enhancement on the target image based on the contour features, label features, and multiple noise images to obtain multiple initial enhanced images; each initial enhanced image carries the target detection label; each noise image is obtained by random sampling;
[0154] Based on the detection frames of each target in any one of the initial enhanced images, determining the enhanced regions corresponding to each target;
[0155] Determine the generation score of any one of the initial enhanced images based on the enhanced regions corresponding to the targets in any one of the initial enhanced images and the label features of the target image;
[0156] Based on the generation scores of the initial enhanced images, screen out the target enhanced image from the initial enhanced images.
[0157] Based on the above embodiments, the enhancement unit 530 is used for:
[0158] Input the contour feature and the label feature into an image enhancement model to obtain the target enhanced image output by the image enhancement model;
[0159] Wherein, the image enhancement model is trained based on a diffusion model, using a sample image and the sample object detection labels of the sample image.
[0160] Based on the above embodiments, the device further includes a model training unit, which is used for:
[0161] Based on the diffusion model, construct a first enhancement model and a second enhancement model, and based on the first enhancement model and the second enhancement model, construct an initial enhancement model;
[0162] Train the initial enhancement model based on the sample image, the sample object detection labels of the sample image, and the sample prompt texts of the targets in the sample image to obtain the image enhancement model;
[0163] Wherein, the first enhancement model and the second enhancement model in the initial enhancement model are connected by a convolutional layer; the first enhancement model is constructed by adding layers to the attention mechanism in the diffusion model, and except for the parameters of the gated self-attention layer added by the layer addition, other parameters are not updated during the training process.
[0164] Based on the above embodiments, the layer addition is to add a gated self-attention layer between each self-attention layer and each cross-attention layer in the attention mechanism of the diffusion model, and the input of the gated self-attention layer includes the sample label features corresponding to the sample object detection labels and the sample contour features corresponding to the sample image;
[0165] The second enhancement model is constructed by removing the decoding layer on the basis of the diffusion model.
[0166] Figure 6 Illustrate a schematic diagram of the physical structure of an electronic device, as Figure 6As shown in the figure, the electronic device may include: a processor 610, a communications interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communications interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 can call the logical instructions in the memory 630 to execute an image enhancement method, which includes: determining a target image with a target detection label; performing edge detection on the target image to obtain contour features; performing image enhancement on the target image based on the contour features and label features to obtain a target enhanced image; where the target enhanced image has the same target detection label as the target image, the target detection label includes the coordinates and categories of the detection boxes of each target in the corresponding image, and the label features are constructed based on the target detection label.
[0167] In addition, when the logical instructions in the above-mentioned memory 630 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0168] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the image enhancement method provided by the above-mentioned various methods. The method includes: determining a target image with a target detection label; performing edge detection on the target image to obtain contour features; performing image enhancement on the target image based on the contour features and label features to obtain a target enhanced image; where the target enhanced image has the same target detection label as the target image, the target detection label includes the coordinates and categories of the detection boxes of each target in the corresponding image, and the label features are constructed based on the target detection label.
[0169] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements an image enhancement method provided by the above-mentioned various methods. The method includes: determining a target image with a target detection label; performing edge detection on the target image to obtain contour features; performing image enhancement on the target image based on the contour features and label features to obtain a target enhanced image; wherein the target enhanced image has the same target detection label as the target image, and the target detection label includes the coordinates and categories of the detection frames of each target in the corresponding image, and the label features are constructed based on the target detection label.
[0170] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative work.
[0171] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0172] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image enhancement method, characterized in that, Including: Determine a target image, where the target image carries target detection labels; Perform edge detection on the target image to obtain contour features; Perform image enhancement on the target image based on the contour features and label features to obtain a target enhanced image; Wherein, the target enhanced image and the target image carry the same target detection labels, the target detection labels include the coordinates and categories of the detection frames of each target in the corresponding image, and the label features are constructed based on the target detection labels; The label features are determined based on the following steps: Encode the categories of the detection frames of each target in the target detection labels to obtain the category prompt word features of each target; Embed the coordinates of the detection frames of each target in the target detection labels to obtain the detection frame position features of each target; Fuse the category prompt word features and detection frame position features of each target to obtain the label features of the target image.
2. The image enhancement method according to claim 1, wherein The step of embedding the coordinates of the detection frames of each target in the target detection labels to obtain the detection frame position features of each target includes: Extract the diagonal coordinates of the detection frames of each target in the target detection labels; Perform Fourier embedding on the diagonal coordinates to obtain the detection frame position features of each target; Wherein, the diagonal coordinates are the coordinates of two vertices on the diagonal of the corresponding detection frame.
3. The image enhancement method according to claim 1 or 2, characterized in that, The step of performing image enhancement on the target image based on the contour features and label features to obtain a target enhanced image includes: Perform image enhancement on the target image based on the contour features, label features, and multiple noise images to obtain multiple initial enhanced images; each initial enhanced image carries the target detection labels; each noise image is obtained by random sampling; Determine the enhanced regions corresponding to each target based on the detection frames of each target in any one of the initial enhanced images; Determine the generation score of any one of the initial enhanced images based on the enhanced regions corresponding to each target in any one of the initial enhanced images and the label features of the target image; Select the target enhanced image from the initial enhanced images based on the generation scores of the initial enhanced images.
4. The image enhancement method according to claim 1 or 2, characterized in that The step of performing image enhancement on the target image based on the contour features and label features to obtain a target enhanced image includes: Input the contour features and the label features into an image enhancement model to obtain the target enhanced image output by the image enhancement model; Wherein, the image enhancement model is trained based on a diffusion model, a sample image, and the sample target detection labels of the sample image.
5. The image enhancement method according to claim 4, wherein The image enhancement model is trained based on the following steps: Based on the diffusion model, construct a first enhancement model and a second enhancement model, and construct an initial enhancement model based on the first enhancement model and the second enhancement model; Train the initial enhancement model based on the sample image, the sample target detection labels of the sample image, and the sample prompt texts of each target in the sample image to obtain the image enhancement model; Among them, in the initial enhanced model, the first enhanced model and the second enhanced model are connected by a convolutional layer; the first enhanced model is constructed by adding layers to the attention mechanism in the diffusion model, and except for the parameters of the gated self-attention layer added by the layer addition in the first enhanced model, other parameters are not updated during the training process.
6. The image enhancement method according to claim 5, wherein the layer addition is to add a gated self-attention layer between each self-attention layer and each cross-attention layer in the attention mechanism of the diffusion model, and the input of the gated self-attention layer includes the sample label features corresponding to the sample target detection labels and the sample contour features corresponding to the sample images; the second enhanced model is constructed by removing the decoding layer on the basis of the diffusion model.
7. An image enhancement device, characterized in that, It includes: a determination unit for determining a target image with a target detection label; a detection unit for performing edge detection on the target image to obtain contour features; an enhancement unit for performing image enhancement on the target image based on the contour features and label features to obtain a target enhanced image; wherein, the target enhanced image has the same target detection label as the target image, the target detection label includes the coordinates and categories of the detection frames of each target in the corresponding image, and the label features are constructed based on the target detection label; the label features are determined based on the following steps: encoding the categories of the detection frames of each target in the target detection label to obtain the category prompt word features of each target; embedding the coordinates of the detection frames of each target in the target detection label to obtain the detection frame position features of each target; fusing the category prompt word features and detection frame position features of each target to obtain the label features of the target image.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the image enhancement method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the image enhancement method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Image enhancement method, text detection model training method and equipment
CN118657686A