An image detection method, device, electronic equipment, storage medium and program product
By using an image detection model to perform multi-scale processing and deep learning on images, the problem of low accuracy in image quality detection caused by manual evaluation is solved, and accurate detection and interpretation of image quality and abnormal regions are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2024-11-29
- Publication Date
- 2026-05-29
AI Technical Summary
Current image quality detection technologies rely on manual evaluation, resulting in low accuracy.
The image to be detected is processed by an image detection model to determine the image quality quantification information, image degradation areas and their textual description information, and image quality detection is performed using multi-scale image processing, feature processing and deep learning models.
It improves the accuracy of image detection and enables detailed detection and explanation of the causes of image anomalies.
Smart Images

Figure CN122115307A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more particularly to an image detection method, apparatus, electronic device, storage medium, and program product. Background Technology
[0002] Image quality inspection refers to the process of evaluating image quality. It assesses the overall quality of an image, and image artifacts are one of the factors affecting image quality. Image artifacts are unexpected visual phenomena that appear in an image or video frame, i.e., abnormal problems present in the image.
[0003] Currently, image quality is assessed by evaluators; however, this method is limited by the subjective evaluation of the evaluators, resulting in low accuracy in image quality detection. Summary of the Invention
[0004] This disclosure provides an image detection method, apparatus, electronic device, storage medium, and program product to achieve quality detection of the image to be detected through an image detection model.
[0005] In a first aspect, embodiments of this disclosure provide an image detection method, including:
[0006] Acquire an image to be detected, wherein the image to be detected includes an image to be subjected to image quality detection;
[0007] The image to be detected is processed by an image detection model to determine the image quality quantification information, the image degradation region in the image to be detected, and the text description information associated with the image degradation region. The image quality quantification information includes information characterizing whether there is a set anomaly in the image to be detected. The image degradation region includes the region in the image to be detected where there is a set anomaly. The text description information includes the description information of the set anomaly corresponding to the image degradation region.
[0008] Secondly, embodiments of this disclosure also provide an image detection apparatus, comprising:
[0009] The acquisition module is used to acquire the image to be detected, which includes the image to be subjected to image quality detection;
[0010] The processing module is used to process the image to be detected using an image detection model to determine the image quality quantification information of the image to be detected, the image degradation region in the image to be detected, and the text description information associated with the image degradation region. The image quality quantification information includes information characterizing whether there is a set anomaly in the image to be detected. The image degradation region includes the region in the image to be detected where the set anomaly exists. The text description information includes the description information of the set anomaly corresponding to the image degradation region.
[0011] Thirdly, embodiments of this disclosure also provide an electronic device, including:
[0012] One or more processing devices;
[0013] Storage device for storing one or more programs.
[0014] When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement any of the image detection methods described in this disclosure.
[0015] Fourthly, embodiments of this disclosure also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform any of the image detection methods described in this disclosure.
[0016] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the image detection method described in any one of the present disclosures.
[0017] The technical solution of this disclosure first acquires an image to be detected, which includes an image to be subjected to image quality detection. Then, it processes the image to be detected using an image detection model. Finally, it determines the image quality quantification information of the image to be detected, the image degradation region in the image to be detected, and the text description information associated with the image degradation region. The image quality quantification information includes information characterizing whether a specified anomaly exists in the image to be detected. The image degradation region includes the region in the image to be detected where the specified anomaly exists. The text description information includes the description information of the specified anomaly corresponding to the image degradation region. The image detection method provided by this disclosure processes the acquired image to be detected using an image detection model. The application of the image detection model improves the accuracy of image detection. By obtaining the image quality quantification information characterizing whether a specified anomaly exists in the image to be detected, the image degradation region characterizing the region in the image to be detected where the specified anomaly exists, and the text description information of the image degradation region, the detailed detection results of the image to be detected are determined. By identifying the anomaly and degradation region in the image to be detected, accurate image detection is achieved, and the reasons for the determination of the image quality quantification information of the image to be detected are also explained. Attached Figure Description
[0018] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0019] Figure 1 This is a schematic flowchart of an image detection method provided in an embodiment of this disclosure;
[0020] Figure 2 This is a schematic diagram of an image detection model provided in an embodiment of this disclosure;
[0021] Figure 3 This is a schematic diagram of an image detection model training method provided in an embodiment of this disclosure;
[0022] Figure 4 This is a schematic diagram of the structure of an image detection device provided in an embodiment of this disclosure;
[0023] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0024] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0025] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0026] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0027] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0028] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0029] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0030] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0031] Figure 1 This is a schematic flowchart of an image detection method provided in this disclosure. This disclosure is applicable to situations where quality detection of an image to be detected is performed using an image detection model. The method can be executed by an image detection device, which can be implemented in software and / or hardware, or optionally by an electronic device, such as a mobile terminal, a PC, or a server.
[0032] like Figure 1 As shown, the method includes:
[0033] S110. Obtain the image to be detected, wherein the image to be detected includes the image to be subjected to image quality detection.
[0034] In this embodiment, the image to be detected can be understood as an image that needs to be image quality detected. There may be abnormal problems in the image to be detected. When performing quality detection on the image to be detected, this disclosure can detect abnormal problems in the image to be detected.
[0035] Anomalies can be understood as unexpected or abnormal phenomena. Anomalies can be image artifacts, such as screen tearing or pixelation. They can also include distortion of objects in the image.
[0036] Image quality detection can be understood as a method used to evaluate the image quality of an image to be detected, and it can be used to detect abnormal problems in the image to be detected.
[0037] The image to be inspected for image quality testing can be an image with abnormalities, such as an image from a video scene or an image from a text-based image scene. The text-based image scene can be considered an application scenario where a corresponding image is generated from an input text description; this scenario can use a model to generate images based on text.
[0038] For example, this operation can acquire images to be detected transmitted by other electronic devices, or determine image frames included in the video of this electronic device as images to be detected; it can also use images generated by this electronic device in Chinese raw image scenes as images to be detected.
[0039] S120. The image to be detected is processed by an image detection model to determine the image quality quantification information of the image to be detected, the image degradation region in the image to be detected, and the text description information associated with the image degradation region. The image quality quantification information includes information characterizing whether there is a set anomaly in the image to be detected. The image degradation region includes the region in the image to be detected where there is a set anomaly. The text description information includes the description information of the set anomaly corresponding to the image degradation region.
[0040] In this embodiment, the image detection model can be understood as a model used to perform image detection on the image to be detected. The image detection model can be used to detect anomalies in the image to be detected and output information describing these anomalies.
[0041] Image quality quantification information can be understood as quantitative information used to describe defined anomalies in the image to be detected. This quantification information can be obtained by scoring the image to be detected. The score of the image to be detected can be a score within a preset score range. The preset score range is not limited here and can be set according to needs. For example, it can be 0-5, 0-10, or 0-100, etc. A higher score indicates that the probability of the image to be detected being an anomaly is lower, or it can indicate that the probability of the image to be detected being an anomaly is higher.
[0042] Image degradation regions can be understood as areas in the image to be detected that exhibit specific abnormalities. Image detection models can use convolution and deconvolution methods to output image degradation regions of the image to be detected.
[0043] Text description information can be understood as a text description used to describe the specified abnormal problems that occur in the image to be detected. Text description information can be obtained by decoding the degraded areas of the image.
[0044] Setting anomalies can be understood as abnormal phenomena present in a pre-defined image. Setting anomalies indicate that unexpected or abnormal phenomena have occurred in the image. These could be image quality degradation issues caused by noise or optical characteristics, such as image artifacts.
[0045] The description information can be understood as information used to describe the setting anomalies of the image to be detected in text form.
[0046] In this embodiment, the image to be detected is input into the image detection model to output image quality quantification information, image degradation regions in the image, and text description information associated with the image degradation regions. The image detection model first transforms the image to be detected input into the image detection model, then performs feature processing on the transformed image, and finally extracts image quality quantification information, image degradation regions, and text description information associated with the image degradation regions from the feature-processed image.
[0047] For example, when the image to be detected is input into the image detection model, the image processing module first transforms the image to obtain the transformed image. Then, the feature processing module performs feature processing on the transformed image to obtain a fused feature map. Finally, the image quality quantization information and image degradation regions are extracted from the feature-processed image, and the text description information associated with the image degradation regions is obtained by decoding the image degradation regions.
[0048] The image processing module can be considered as a module that preprocesses the image to be detected, such as a multi-scale image processing module. The multi-scale image processing module can be a module that performs multi-scale transformations on the image to be detected. After processing by the multi-scale image processing module, images at multiple scales can be obtained.
[0049] The feature processing module can be understood as a module used to extract and process features from an image. After processing, a feature map of the image after multi-level fusion can be obtained.
[0050] The technical solution of this disclosure first acquires an image to be detected, which includes an image to be subjected to image quality detection. Then, it processes the image to be detected using an image detection model. Finally, it determines the image quality quantification information of the image to be detected, the image degradation region in the image to be detected, and the text description information associated with the image degradation region. The image quality quantification information includes information characterizing whether a specified anomaly exists in the image to be detected. The image degradation region includes the region in the image to be detected where the specified anomaly exists. The text description information includes the description information of the specified anomaly corresponding to the image degradation region. The image detection method provided by this disclosure processes the acquired image to be detected using an image detection model. The application of the image detection model improves the accuracy of image detection. By obtaining the image quality quantification information characterizing whether a specified anomaly exists in the image to be detected, the image degradation region characterizing the region in the image to be detected where the specified anomaly exists, and the text description information of the image degradation region, the detailed detection results of the image to be detected are determined. By identifying the anomaly and degradation region in the image to be detected, accurate image detection is achieved, and the reasons for the determination of the image quality quantification information of the image to be detected are also explained.
[0051] Based on the above embodiments, modified embodiments of the above embodiments are proposed. It should be noted that, in order to keep the description brief, only the differences from the above embodiments are described in the modified embodiments.
[0052] In one embodiment, the step of processing the image to be detected using an image detection model to determine the image quality quantification information of the image to be detected, the image degradation regions in the image to be detected, and the text description information associated with the image degradation regions includes:
[0053] The image to be detected is input into the multi-scale image processing module in the image detection model;
[0054] The multi-scale image processing module performs multi-scale transformation on the image to be detected to obtain multiple transformed images.
[0055] The feature processing module processes the image to be detected, the transformed image, and the attribute information associated with the image to obtain a fused feature map.
[0056] The fused feature map is input to the post-processing module to obtain the image quality quantification information of the image to be detected, the image degradation region in the image to be detected, and the text description information associated with the image degradation region.
[0057] In this embodiment, the multi-scale image processing module can be understood as a module used to perform multi-scale transformation on the image to be detected. Multi-scale transformation can involve transforming the image to be detected multiple times according to its own scale, resulting in multiple images with transformed dimensions. This facilitates the image detection model in obtaining local details and global information of the image to be detected. After processing by the multi-scale image processing module, images at multiple scales can be obtained. These multiple-scale images can be considered as representations of the same image based on different scales.
[0058] The converted image can be understood as the image obtained after the image to be detected has been processed by a multi-scale image processing module. The converted image may include multiple images at different scales than the image to be detected.
[0059] The feature processing module can extract and fuse features from the image to be detected, the transformed image, and their associated attribute information.
[0060] The image-related attribute information can be understood as attribute information associated with both the image to be detected and the transformed image, respectively. This can include information such as the location of each image block after it has been segmented into image blocks and / or the scale of the image it belongs to. An image block can be understood as a local region formed by dividing an image into a series of blocks or regions. The shape of the image block is not limited and can be a regular shape, such as a square or rectangle.
[0061] The fused feature map can be understood as the feature map output by the feature processing module, which integrates the image to be detected and the transformed image and their associated attribute information.
[0062] The post-processing module can be understood as a module used to output image quality quantification information of the image to be detected, image degradation areas in the image to be detected, and text description information associated with the image degradation areas.
[0063] In this embodiment, the image detection model includes three modules: a multi-scale image processing module, a feature processing module, and a post-processing module. First, the image to be detected is input to the multi-scale image processing module, which performs multi-scale transformation on the image. This transformation can be based on the scale of the image to be detected, resulting in multiple transformed images. Next, the feature processing module processes the image to be detected, the multiple transformed images, and the attribute information associated with the images. Each image and its associated attribute information can be converted into a vector, thus obtaining multiple vectors corresponding to the image to be detected and the transformed images. Alternatively, each image and its associated attribute information can be converted into a vector, and these vectors can be fused to obtain a fused feature map. The fused feature map is then input to the post-processing module to obtain the image quality quantification information of the image to be detected, the image degradation regions in the image to be detected, and the text description information associated with the image degradation regions.
[0064] For example, Figure 2 This is a schematic diagram of an image detection model provided in an embodiment of this disclosure. Figure 2 As shown, the image detection model comprises three modules: a multi-scale image processing module 1, a feature processing module 2, and a post-processing module 3. The multi-scale image processing module 1 performs multi-scale transformations on the image to be detected, resulting in multiple transformed images. The transformation can be performed based on the length, width, and aspect ratio of the image to be detected; different lengths and widths result in different transformations. The transformed images are then input into the feature processing module 2, where the image to be detected, the multiple transformed images, and the associated attribute information are processed. Through feature fusion and information interaction on the vectors, a new vector is output after processing, thus obtaining the fused feature map. Finally, the fused feature map is input into the post-processing module 3 to obtain the image quality quantification information of the image to be detected, the image degradation regions in the image to be detected, and the textual description information associated with the image degradation regions.
[0065] Figure 2 The feature map can reflect semantic information at different levels of the image to be detected, including deep and shallow semantic information. The fused feature map is obtained by combining the shallow and deep semantic features.
[0066] The post-processing module may include a multilayer perceptron, convolutional and deconvolutional submodules, and a decoder. The multilayer perceptron processes the fused feature map to obtain image quality quantification information. The convolutional and deconvolutional submodules process the fused feature map to identify image degradation regions. The decoder processes the fused feature map to obtain textual description information.
[0067] In one embodiment, the step of performing multi-scale transformation on the image to be detected by the multi-scale image processing module to obtain multiple transformed images includes:
[0068] The multi-scale image processing module adjusts the length of the set side of the image to be detected a set number of times to obtain a set number of transformed images. The set side of each transformed image has a different length, and each transformed image has the same aspect ratio.
[0069] In this embodiment, the set number can be understood as data used to describe the number of converted images, or it can represent the number of times multi-scale image processing is performed on the image to be detected. The set side can be understood as a set side in the image, such as a long side or a short side, or a side representing the length of the image or a side representing the width of the image. If the image is rectangular, the set side can be any side of the rectangle, such as the long side. The length of the adjusted set side is not limited here and can be a preset value.
[0070] In this embodiment, the image to be detected is input into the multi-scale image processing module. During the generation of a converted image, the length of a predetermined side of the image to be detected is first determined. The predetermined side of the image to be processed is then adjusted, maintaining the aspect ratio unchanged during the conversion process. The length of the predetermined side after conversion is not limited here and can be preset. After adjustment, a converted image is obtained. The method for adjusting the predetermined number of times is the same and will not be described in detail here.
[0071] By keeping the aspect ratio constant and performing a set number of transformations on the image to be detected, the image after the set number of transformations can be obtained.
[0072] For example, assuming the image to be detected is 448*448 pixels, the aspect ratio is 1. If the number of adjustments is set twice, with the longer side being adjusted, the adjusted side lengths can be 384 and 224. The image to be detected is input into the multi-scale image processing module, which adjusts the image size according to an aspect ratio of 1 and a side length of 384, resulting in a transformed image. Similarly, the image to be detected is adjusted according to an aspect ratio of 1 and a side length of 224, resulting in another transformed image. Thus, two different transformed images are obtained.
[0073] In one embodiment, the step of processing the image to be detected, the transformed image, and the attribute information associated with the image through the feature processing module to obtain a fused feature map includes:
[0074] The feature processing module divides the image to be detected and the converted image into multiple image blocks and determines the encoding information of each image block.
[0075] The feature processing module determines the location and scale information corresponding to the image patch.
[0076] The feature processing module determines the fused feature map corresponding to the encoded information, the encoded position information, and the scale information.
[0077] In this embodiment, the encoded information can be understood as a vector representation of the image patch information after linear projection. The location information can be understood as the spatial location of the image patch within its respective image to be detected or its transformed image. The scale information can be understood as the scale feature representation of the image patch within its respective image to be detected or its transformed image.
[0078] In this embodiment, the image to be detected and multiple transformed images are first divided into multiple image blocks by a feature processing module. These image blocks can be of fixed size, and each block is linearly projected into a fixed-length vector, which is then used as the encoding information for that image block. Next, position embedding and scale embedding are added to the encoding information of each image block. Position embedding encodes the location information of the image block, while scale embedding distinguishes image blocks of different scales. Finally, the fused feature map corresponding to the encoding information, the encoded position information, and the scale information is determined.
[0079] For example, each image to be detected and multiple converted images can be divided into multiple 16*16 image blocks, and each 16*16 image block can be converted into a fixed-length vector, which is used as the encoding information of the image block.
[0080] In one embodiment, determining the location information corresponding to the image patch through the feature processing module includes:
[0081] The feature processing module determines the position of the image block in the corresponding image;
[0082] Determine the hash value corresponding to the location;
[0083] The location information of the image patch is obtained by mapping the hash value into a two-dimensional matrix.
[0084] In this embodiment, a hash value can be understood as a fixed-length string or number generated after the hash function processes the data, which can be used to encode the location information of image blocks.
[0085] In this embodiment, the feature processing module divides the image to be detected and multiple transformed images into multiple image blocks. For each image block, its position in the corresponding image is determined. The position of each image block is represented by a hash value. Each hash value representing the position information of the image block is mapped to a fixed-size two-dimensional matrix in the image corresponding to the image block, which is used to represent the position information of the image block in the corresponding image.
[0086] For example, hash-based two-dimensional spatial embedding is used to encode the positional information of each image patch, thereby obtaining the positional information of each image patch in its corresponding image. For a target image and a transformed image, their corresponding image patches are mapped into a fixed two-dimensional matrix. Each target image and each transformed image are images at different scales; by determining the positional information of each image patch, the spatial relationship between images at different scales can be represented.
[0087] When image detection models perform quality checks on images, image resizing (i.e., the resize operation) can lead to a loss of image quality. Currently, most models use a fixed input of 448x448. To reduce the image quality loss caused by resizing, this disclosure uses a multi-scale image processing module to process the images to be detected.
[0088] The detection of the image to be detected includes the following steps:
[0089] 1. The image to be detected is converted into a multi-scale representation, including the original resolution image and two converted images with their long sides adjusted to 224 and 384, to obtain images at different resolutions. This allows for the simultaneous capture of both local details and global information of the image to be detected.
[0090] 2. The image at each scale is divided into fixed-size image blocks, and each image block is encoded as a token. The divided image blocks can be local regions of the image, still presented as image pixels. Computers cannot directly use Transformers to extract semantic relationships and other feature information from these blocks. Therefore, this disclosure encodes each image block into a token. A token is a vector representation that summarizes the information contained in the image block and exists in a format suitable for Transformer processing. After image block division and token encoding, the image to be detected transforms from a two-dimensional image into a sequence of tokens, i.e., encoded information.
[0091] 3. A hash-based two-dimensional spatial embedding is introduced to encode the location information of each image patch, enabling the image detection model to effectively utilize the spatial relationships between different scales. When introducing location information, the location information of image patches at different resolutions can be mapped to a fixed two-dimensional matrix, allowing the Transformer network of the image detection model's feature processing module to determine the spatial relationships between different scales. During the mapping process, the width and height of the image patches can be mapped.
[0092] 4. At the same time, a scale embedding is introduced to help the image detection model distinguish image patches from different scales.
[0093] 5. Input these encoded image patch sequences into the Transformer network of the feature processing module of the image detection model, and use the self-attention mechanism to capture multi-scale features of image quality.
[0094] Figure 3 This is a schematic diagram of an image detection model training method provided in this embodiment of the disclosure. This embodiment of the disclosure is applicable to the training of image detection models.
[0095] like Figure 3 As shown, the method includes:
[0096] S210. Obtain a training sample set, wherein the training sample set includes sample images and sample quality quantification information corresponding to the sample images, sample degradation regions corresponding to the sample images and degradation description information of the sample degradation regions, the sample quality quantification information includes information characterizing whether there is a set anomaly in the sample images, the sample degradation regions include regions in the sample images where there is a set anomaly, and the degradation description information includes description information of the set anomaly corresponding to the sample degradation regions.
[0097] In this embodiment, the training sample set can be understood as a collection of data used to store the data required for training the image detection model. A sample image can be understood as an image that can be used as a training sample for the image detection model; some regions of the sample image may exhibit image anomalies. Sample quality quantification information can be understood as quantification information used to describe the specified anomalies appearing in the sample image. Sample degradation region can be understood as a region in the sample image where a specified anomaly occurs. Degradation description information can be understood as a textual description of the specified anomalies appearing in the sample image.
[0098] Sample quality quantization information corresponds to image quality quantization information. Sample quality quantization information is the output of the image detection model during the training phase. Image quality quantization information is the output of the image detection model during the application phase.
[0099] The sample degradation region corresponds to the image degradation region. The sample degradation region is the output of the image detection model during the training phase, while the image degradation region is the output of the image detection model during the application phase.
[0100] The degradation description information corresponds to the text description information. The degradation description information is the output of the image detection model during the training phase, while the text description information is the output of the image detection model during the application phase.
[0101] In this embodiment, both images with and without the specified abnormality are used as sample images, and multiple sample images are collected and placed into the training sample set.
[0102] The training sample set includes sample quality quantification information corresponding to the sample images, sample degradation regions corresponding to the sample images, and degradation description information of the sample degradation regions. Among them, the sample quality quantification information corresponding to the sample images, sample degradation regions corresponding to the sample images, and degradation description information of the sample degradation regions can be the result of manual annotation, which is used for training the image detection model.
[0103] S220. The deep learning model is trained using the training sample set to obtain an image detection model.
[0104] In this embodiment, the deep learning model learns the complex features of the data through a multi-layer neural network. It is a machine learning model that can be used to learn the features of sample images in the training sample set, thereby generating an image detection model.
[0105] This operation inputs image samples from the training sample set into a deep learning model. Based on the output of the deep learning model and the corresponding sample quality quantification information, sample degradation region, and degradation description information, the model parameters are adjusted to train the deep learning model. The sample quality quantification information, sample degradation region, and degradation description information corresponding to the image samples can be the expected labels of the deep learning model. The model is trained separately based on these three types of information from the sample images, and after multiple training sessions through multi-layer neural networks, an image detection model is obtained.
[0106] S230. Obtain the image to be detected, wherein the image to be detected includes the image to be subjected to image quality detection.
[0107] S240. The image to be detected is processed by an image detection model to determine the image quality quantification information of the image to be detected, the image degradation region in the image to be detected, and the text description information associated with the image degradation region. The image quality quantification information includes information characterizing whether there is a set anomaly in the image to be detected. The image degradation region includes the region in the image to be detected where there is a set anomaly. The text description information includes the description information of the set anomaly corresponding to the image degradation region.
[0108] The technical solution of this disclosure first obtains a training sample set, which includes sample images and corresponding sample quality quantification information, sample degradation regions corresponding to the sample images, and degradation description information of the sample degradation regions. The sample quality quantification information includes information characterizing whether a specified anomaly exists in the sample image. The sample degradation regions include regions in the sample image where the specified anomaly exists. The degradation description information includes description information of the specified anomaly corresponding to the sample degradation regions. Finally, a deep learning model is trained using the training sample set to obtain an image detection model. By obtaining a training sample set composed of sample images and their corresponding sample quality quantification information, sample degradation regions, and degradation description information of the sample degradation regions, and training a deep learning model to obtain an image detection model, the training process of the image detection model is refined, resulting in an image detection model that can be used by image detection methods.
[0109] In one embodiment, training the deep learning model using the training sample set to obtain an image detection model includes:
[0110] Using the training sample set, the deep learning model is jointly trained by the quality quantization task, image segmentation task and description task to obtain the image detection model.
[0111] The quality quantization task includes the task of quantifying image quality, the image segmentation task includes the task of segmenting degraded regions in an image, and the description task includes the task of describing the setting anomalies of degraded regions in an image.
[0112] In this embodiment, during joint training of the image detection model, each task corresponds to a specific objective that the model is expected to achieve. For example, the quality quantization task aims to enable the model to determine the result of image quality quantization. The image segmentation task aims to segment degraded regions in the image at the pixel level. The description task aims to enable the model to determine the type of quality problem in the degraded region.
[0113] In this embodiment, the training sample set includes sample quality quantification information, sample degradation regions, and degradation description information. When training the deep learning model based on the training sample set, three training tasks can be differentiated according to these three different types of information: a quality quantification task trained using sample quality quantification information, an image segmentation task trained using sample degradation regions, and a description task trained using degradation description information. Joint training of these three tasks yields an image detection model that simultaneously possesses the capabilities of image quality quantification, degradation region segmentation, and quality problem classification.
[0114] For example, image segmentation tasks include segmenting degraded regions in an image so that image detection models can identify these degraded regions. This can be achieved through... Figure 2 The convolutional and deconvolutional submodules process the fused feature map to obtain degraded regions. Image segmentation tasks can be optimized using smoothing loss functions, such as smoothed L1 loss. When the amount of data in the image segmentation task is N, the loss can be expressed as:
[0115] l(x,y)=L={l1,...,l N} T
[0116] Each element is represented by the smoothing loss function as follows:
[0117]
[0118] Where x n The predicted value output by the image segmentation task, y n If the true value of the image sample is given, then the loss is expressed as:
[0119]
[0120] This embodiment can be achieved through... Figure 2 The decoder processes the fused image to obtain degraded descriptive information. A cross-entropy loss function is used to optimize the descriptive task; the loss function corresponding to the descriptive task can be expressed as:
[0121]
[0122] The loss value for joint training of the three tasks can be expressed as:
[0123] final_loss=0.3×score_loss+0.3×grounding_loss+0.4×test_loss
[0124] Here, `score_loss` is the loss value for the quality quantization task, `grounding_loss` is the loss value for the image segmentation task, and `test_loss` is the loss value for the descriptive task. The weight coefficients are not limited here and can be adjusted according to requirements.
[0125] In one embodiment, the loss function corresponding to the quality quantification task is a loss function jointly determined by the comparison loss function and the regression loss function.
[0126] In this embodiment, the loss function can be understood as a function used to evaluate the difference between the predicted value and the actual value of the quality quantification task, and can output the magnitude of the prediction error.
[0127] A contrastive loss function is a loss function applied to contrastive tasks. Contrastive tasks can be those that explore the similarities and differences between data. The goal is for the model to learn the relative relationships between samples, so that the model's predicted ranking results closely approximate the true labeled ranking results.
[0128] Regression loss functions can be considered as loss functions applied to regression tasks. Regression tasks can involve predicting a continuous numerical value. The role of regression loss functions is to measure the degree of error between the model's predicted value and the true value. By minimizing this error, the model is optimized to make its predictions as close to the true value as possible.
[0129] The contrastive loss function represents the relative ranking loss between image pairs. If the predicted ranking is consistent with the true ranking, the output value of the contrastive loss function is relatively small. The regression loss function represents the sum of the squares of the differences between the predicted and true values. If the predicted and true values are close, the output value of the regression loss function is relatively small.
[0130] The loss function for a quality quantification task can be jointly determined by a comparative loss function and a regression loss function, so that the training results of the quality quantification task have a better ability to perceive relative quality and evaluate absolute branches.
[0131] This embodiment does not restrict the way the comparative loss function and the regression loss function are combined; for example, they can be combined using a weighted summation method. The weight coefficients are not limited here.
[0132] For example, the loss function corresponding to the quality quantification task is:
[0133] loss total =loss reg +loss rank
[0134] Where loss reg This represents the regression-type loss function, loss. rank This represents the contrastive loss function.
[0135]
[0136] Where y i It is the true score of the sample image, that is, the quantitative information of sample quality, while It is the score estimated from the quality quantification task.
[0137] This embodiment can explicitly utilize the relative ranking of image pairs using pairwise ranking loss. We will not specify the formula for the comparison-based loss function here, as long as the deep learning model can learn the ranking of sample quality quantification information, making the model's predicted ranking consistent with the actual ranking.
[0138] In one embodiment, the contrastive loss function includes a loss function determined based on a first size relationship and a second size relationship. The first size relationship includes the size relationship of the sample quality quantification information corresponding to every two sample images in the training sample set, and the second size relationship includes the size relationship of the sample quality quantification information output by the deep learning model from the two sample images.
[0139] In this embodiment, the first size relationship can be considered as a size relationship determined based on the labeled sample quality quantification information. The second size relationship can be considered as a size relationship determined based on the sample quality quantification information output by the deep learning model. The contrastive loss function uses the first and second size relationships to train the deep learning model, so that the deep learning model maintains a consistency between the first and second size relationships after training.
[0140] In the training sample set, a first size relationship and a second size relationship can be determined for every two images. Then, the first size relationships that satisfy a predetermined relationship are selected. Finally, the second size relationships corresponding to the first size relationships are summarized to obtain the contrastive loss function. The predetermined relationship is not limited here; it can be that the quality quantization information of the sample labeled for the i-th image is less than the quality quantization information of the sample labeled for the (i+1)-th image.
[0141] The method of aggregation is not limited here; it could be a determination of the mean, etc.
[0142] The formula for the contrast loss function is as follows:
[0143]
[0144] Where y i and y i+1 It is the true score of two sample images in the training sample set, that is, the quantitative information of the labeled sample quality, while and These are the scores estimated from two sample images through a quality quantization task, representing the sample quality quantization information output by the model. The first magnitude relationship can be y. i <y i+1 The second size relationship can be The contrast class loss function is determined by the first and second size relationships.
[0145] Figure 4This is a schematic diagram of the structure of an image detection device provided in an embodiment of this disclosure, as shown below. Figure 4 As shown, the device includes an acquisition module 310 and a processing module 320.
[0146] The acquisition module 310 is used to acquire the image to be detected, the image to be detected including the image to be subjected to image quality detection;
[0147] The processing module 320 is used to process the image to be detected through an image detection model to determine the image quality quantification information of the image to be detected, the image degradation region in the image to be detected, and the text description information associated with the image degradation region. The image quality quantification information includes information characterizing whether there is a set anomaly in the image to be detected. The image degradation region includes the region in the image to be detected where there is a set anomaly. The text description information includes the description information of the set anomaly corresponding to the image degradation region.
[0148] The technical solution of this disclosure first acquires an image to be detected through an acquisition module. This image includes the image to be subjected to image quality detection. Then, a processing module processes the image to be detected using an image detection model. Finally, it determines the image quality quantification information of the image to be detected, the image degradation regions in the image to be detected, and the text description information associated with the image degradation regions. The image quality quantification information includes information characterizing whether a specified anomaly exists in the image to be detected. The image degradation regions include areas in the image to be detected where specified anomalies exist. The text description information includes description information of the specified anomalies corresponding to the image degradation regions. Through the cooperation between the modules, the image detection model is applied to process the acquired image to be detected, improving the accuracy of image detection. By obtaining the image quality quantification information characterizing whether a specified anomaly exists in the image to be detected, the image degradation regions characterizing areas in the image to be detected where specified anomalies exist, and the text description information of the image degradation regions, detailed detection results for the image to be detected are determined. The anomalies and degradation regions in the image to be detected are identified, achieving accurate image detection and providing explanation for the reasons derived from the image quality quantification information of the image to be detected.
[0149] In one embodiment, the processing module 320 includes:
[0150] The first input unit is used to input the image to be detected into the multi-scale image processing module in the image detection model;
[0151] The conversion unit is used to perform multi-scale conversion on the image to be detected through the multi-scale image processing module to obtain multiple converted images;
[0152] The processing unit is used to process the image to be detected, the transformed image, and the attribute information associated with the image through the feature processing module to obtain a fused feature map;
[0153] The second input unit is used to input the fused feature map to the post-processing module to obtain the image quality quantification information of the image to be detected, the image degradation region in the image to be detected, and the text description information associated with the image degradation region.
[0154] In one embodiment, the conversion unit is specifically used for:
[0155] The multi-scale image processing module adjusts the length of the set side of the image to be detected a set number of times to obtain a set number of transformed images. The set side of each transformed image has a different length, and each transformed image has the same aspect ratio.
[0156] In one embodiment, the processing unit includes:
[0157] The segmentation unit is used to divide the image to be detected and the converted image into multiple image blocks through the feature processing module, and to determine the encoding information of each image block;
[0158] The first determining subunit is used to determine the position information and scale information corresponding to the image block through the feature processing module;
[0159] The second determining subunit is used to determine the fused feature map corresponding to the encoded information, the encoded position information, and the scale information through the feature processing module.
[0160] In one embodiment, the first determined subunit is specifically used for:
[0161] The feature processing module determines the position of the image block in the corresponding image;
[0162] Determine the hash value corresponding to the location;
[0163] The location information of the image patch is obtained by mapping the hash value into a two-dimensional matrix.
[0164] In one embodiment, the training operation of the image detection model includes:
[0165] An acquisition unit is used to acquire a training sample set, the training sample set including sample images and sample quality quantification information corresponding to the sample images, sample degradation regions corresponding to the sample images and degradation description information of the sample degradation regions, the sample quality quantification information including information characterizing whether there is a set anomaly in the sample images, the sample degradation regions including regions in the sample images where there is a set anomaly, and the degradation description information including description information of the set anomaly corresponding to the sample degradation regions;
[0166] The training unit is used to train the deep learning model using the training sample set to obtain the image detection model.
[0167] In one embodiment, the training unit is specifically used for:
[0168] Using the training sample set, the deep learning model is jointly trained by the quality quantization task, image segmentation task and description task to obtain the image detection model.
[0169] The quality quantization task includes the task of quantifying image quality, the image segmentation task includes the task of segmenting degraded regions in an image, and the description task includes the task of describing the setting anomalies of degraded regions in an image.
[0170] In one embodiment, the loss function corresponding to the quality quantification task is a loss function jointly determined by the comparison loss function and the regression loss function.
[0171] In one embodiment, the contrastive loss function includes a loss function determined based on a first size relationship and a second size relationship. The first size relationship includes the size relationship of the sample quality quantification information corresponding to every two sample images in the training sample set, and the second size relationship includes the size relationship of the sample quality quantification information output by the deep learning model from the two sample images.
[0172] The image detection apparatus provided in this disclosure can execute the image detection method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the method.
[0173] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0174] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Reference is made below. Figure 5It illustrates an electronic device suitable for implementing embodiments of the present disclosure (e.g., Figure 5 A structural diagram of the terminal device or server in the 500.
[0175] Electronic equipment 500, including:
[0176] One or more processing devices 501;
[0177] Storage device 508, for storing one or more programs,
[0178] When the one or more programs are executed by the one or more processing devices 501, the one or more processing devices 501 implement any of the methods provided in this disclosure.
[0179] The terminal devices in this disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle terminals (e.g., vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0180] like Figure 5 As shown, electronic device 500 may include a processing unit (e.g., central processing unit, graphics processor, etc.) 501, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 502 or a program loaded from storage device 508 into random access memory (RAM) 503. The RAM 503 also stores various programs and data required for the operation of electronic device 500. The processing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. An edit / output (I / O) interface 505 is also connected to bus 504.
[0181] Typically, the following devices can be connected to I / O interface 505: input devices 506 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 507 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 508 including, for example, magnetic tapes, hard disks, etc.; and communication devices 509. Communication device 509 allows electronic device 500 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 500 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0182] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by the processing device 501, it performs the functions defined in the methods of embodiments of this disclosure.
[0183] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0184] The electronic device provided in this embodiment and the image detection method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0185] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the image detection method provided in the above embodiments.
[0186] It should be noted that the computer-readable medium described above in this disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination thereof.
[0187] The computer storage medium may be a storage medium for computer-executable instructions, which, when executed by a computer processor, are used to perform the methods provided in this disclosure.
[0188] Computer-readable storage media can be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0189] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0190] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0191] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire at least two Internet Protocol (IP) addresses; send a node evaluation request including the at least two IP addresses to a node evaluation device, wherein the node evaluation device selects an IP address from the at least two IP addresses and returns it; and receive the IP address returned by the node evaluation device; wherein the acquired IP address indicates an edge node in a content delivery network.
[0192] Alternatively, the aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: receive a node evaluation request including at least two Internet Protocol (IP) addresses; select an IP address from the at least two IP addresses; and return the selected IP address; wherein the received IP address indicates an edge node in the content delivery network.
[0193] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0194] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0195] The modules or units described in the embodiments of this disclosure can be implemented in software or hardware. The names of modules or units do not necessarily limit the specific unit; for example, an acquisition module can also be described as an "image acquisition module."
[0196] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0197] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0198] A computer program product includes a computer program that, when executed by a processor, implements the image detection method provided in this disclosure.
[0199] According to one or more embodiments of this disclosure, [Example 1] provides an image detection method, including:
[0200] Acquire an image to be detected, wherein the image to be detected includes an image to be subjected to image quality detection;
[0201] The image to be detected is processed by an image detection model to determine the image quality quantification information, the image degradation region in the image to be detected, and the text description information associated with the image degradation region. The image quality quantification information includes information characterizing whether there is a set anomaly in the image to be detected. The image degradation region includes the region in the image to be detected where there is a set anomaly. The text description information includes the description information of the set anomaly corresponding to the image degradation region.
[0202] According to one or more embodiments of this disclosure, [Example 2] provides the method of Example 1, wherein processing the image to be detected using an image detection model to determine image quality quantification information of the image to be detected, image degradation regions in the image to be detected, and text description information associated with the image degradation regions includes:
[0203] The image to be detected is input into the multi-scale image processing module in the image detection model;
[0204] The multi-scale image processing module performs multi-scale transformation on the image to be detected to obtain multiple transformed images.
[0205] The feature processing module processes the image to be detected, the transformed image, and the attribute information associated with the image to obtain a fused feature map.
[0206] The fused feature map is input to the post-processing module to obtain the image quality quantification information of the image to be detected, the image degradation region in the image to be detected, and the text description information associated with the image degradation region.
[0207] According to one or more embodiments of this disclosure, [Example 3] provides the method of Example 2, wherein the multi-scale image processing module performs multi-scale transformation on the image to be detected to obtain multiple transformed images, including:
[0208] The multi-scale image processing module adjusts the length of the set side of the image to be detected a set number of times to obtain a set number of transformed images. The set side of each transformed image has a different length, and each transformed image has the same aspect ratio.
[0209] According to one or more embodiments of this disclosure, [Example 4] provides the method of Example 2, wherein the feature processing module processes the image to be detected, the transformed image, and attribute information associated with the image to obtain a fused feature map, including:
[0210] The feature processing module divides the image to be detected and the converted image into multiple image blocks and determines the encoding information of each image block.
[0211] The feature processing module determines the location and scale information corresponding to the image patch.
[0212] The feature processing module determines the fused feature map corresponding to the encoded information, the encoded position information, and the scale information.
[0213] According to one or more embodiments of this disclosure, [Example 5] provides the method described in Example 4, which determines the location information corresponding to the image patch through the feature processing module, including:
[0214] The feature processing module determines the position of the image block in the corresponding image;
[0215] Determine the hash value corresponding to the location;
[0216] The location information of the image patch is obtained by mapping the hash value into a two-dimensional matrix.
[0217] According to one or more embodiments of this disclosure, [Example 6] provides the method of Example 1, wherein the training operation of the image detection model includes:
[0218] A training sample set is obtained, which includes sample images and corresponding sample quality quantification information, sample degradation regions corresponding to the sample images and degradation description information of the sample degradation regions. The sample quality quantification information includes information characterizing whether there is a specified abnormal problem in the sample image. The sample degradation region includes the region in the sample image where the specified abnormal problem exists. The degradation description information includes the description information of the specified abnormal problem corresponding to the sample degradation region.
[0219] The deep learning model is trained using the training sample set to obtain an image detection model.
[0220] According to one or more embodiments of this disclosure, [Example 7] provides the method of Example 6, wherein training a deep learning model using the training sample set to obtain an image detection model includes:
[0221] Using the training sample set, the deep learning model is jointly trained by the quality quantization task, image segmentation task and description task to obtain the image detection model.
[0222] The quality quantization task includes the task of quantifying image quality, the image segmentation task includes the task of segmenting degraded regions in an image, and the description task includes the task of describing the setting anomalies of degraded regions in an image.
[0223] According to one or more embodiments of this disclosure, [Example 8] provides the method described in Example 7, wherein the loss function corresponding to the quality quantification task is a loss function jointly determined by a comparison loss function and a regression loss function.
[0224] According to one or more embodiments of this disclosure, [Example 9] provides the method described in Example 8, wherein the contrastive loss function includes a loss function determined based on a first size relationship and a second size relationship, wherein the first size relationship includes the size relationship of sample quality quantification information corresponding to every two sample images in the training sample set, and the second size relationship includes the size relationship of sample quality quantification information output by the deep learning model from the two sample images.
[0225] According to one or more embodiments of this disclosure, [Example 10] provides an image detection apparatus, including:
[0226] The acquisition module is used to acquire the image to be detected, which includes the image to be subjected to image quality detection;
[0227] The processing module is used to process the image to be detected using an image detection model to determine the image quality quantification information of the image to be detected, the image degradation region in the image to be detected, and the text description information associated with the image degradation region. The image quality quantification information includes information characterizing whether there is a set anomaly in the image to be detected. The image degradation region includes the region in the image to be detected where the set anomaly exists. The text description information includes the description information of the set anomaly corresponding to the image degradation region.
[0228] According to one or more embodiments of this disclosure, [Example 11] an electronic device is provided, the electronic device comprising:
[0229] One or more processing devices;
[0230] Storage device for storing one or more programs.
[0231] When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the image detection method as described in any of Examples 1-9.
[0232] According to one or more embodiments of this disclosure, [Example 12] provides a storage medium containing computer-executable instructions that, when executed by a computer processor, are used to perform an image detection method as described in any of Examples 1-9.
[0233] According to one or more embodiments of this disclosure, [Example 13] provides a computer program product including a computer program that, when executed by a processor, implements the image detection method according to any one of Examples 1-9.
[0234] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0235] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0236] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An image detection method, characterized in that, include: Acquire an image to be detected, wherein the image to be detected includes an image to be subjected to image quality detection; The image to be detected is processed by an image detection model to determine the image quality quantification information, the image degradation region in the image to be detected, and the text description information associated with the image degradation region. The image quality quantification information includes information characterizing whether there is a set anomaly in the image to be detected. The image degradation region includes the region in the image to be detected where there is a set anomaly. The text description information includes the description information of the set anomaly corresponding to the image degradation region.
2. The method according to claim 1, characterized in that, The step of processing the image to be detected using an image detection model to determine the image quality quantification information of the image to be detected, the image degradation regions in the image to be detected, and the text description information associated with the image degradation regions includes: The image to be detected is input into the multi-scale image processing module in the image detection model; The multi-scale image processing module performs multi-scale transformation on the image to be detected to obtain multiple transformed images. The feature processing module processes the image to be detected, the transformed image, and the attribute information associated with the image to obtain a fused feature map. The fused feature map is input to the post-processing module to obtain the image quality quantification information of the image to be detected, the image degradation region in the image to be detected, and the text description information associated with the image degradation region.
3. The method according to claim 2, characterized in that, The multi-scale image processing module performs multi-scale transformation on the image to be detected to obtain multiple transformed images, including: The multi-scale image processing module adjusts the length of the set side of the image to be detected a set number of times to obtain a set number of transformed images. The set side of each transformed image has a different length, and each transformed image has the same aspect ratio.
4. The method according to claim 2, characterized in that, The step involves processing the image to be detected, the transformed image, and the attribute information associated with the image through a feature processing module to obtain a fused feature map, including: The feature processing module divides the image to be detected and the converted image into multiple image blocks and determines the encoding information of each image block. The feature processing module determines the location and scale information corresponding to the image patch. The feature processing module determines the fused feature map corresponding to the encoded information, the encoded position information, and the scale information.
5. The method according to claim 4, characterized in that, The feature processing module determines the location information corresponding to the image patch, including: The feature processing module determines the position of the image block in the corresponding image; Determine the hash value corresponding to the location; The location information of the image patch is obtained by mapping the hash value into a two-dimensional matrix.
6. The method according to claim 1, characterized in that, The training operations for the image detection model include: A training sample set is obtained, which includes sample images and corresponding sample quality quantification information, sample degradation regions corresponding to the sample images and degradation description information of the sample degradation regions. The sample quality quantification information includes information characterizing whether there is a set anomaly in the sample image. The sample degradation region includes the region in the sample image where the set anomaly exists. The degradation description information includes the description information of the set anomaly corresponding to the sample degradation region. The deep learning model is trained using the training sample set to obtain the image detection model.
7. The method according to claim 6, characterized in that, The step of training the deep learning model using the training sample set to obtain the image detection model includes: Using the training sample set, the deep learning model is jointly trained by the quality quantization task, image segmentation task and description task to obtain the image detection model. The quality quantization task includes the task of quantifying image quality, the image segmentation task includes the task of segmenting degraded regions in an image, and the description task includes the task of describing the setting anomalies of degraded regions in an image.
8. The method according to claim 7, characterized in that, The loss function corresponding to the quality quantification task is a loss function jointly determined by the comparative loss function and the regression loss function.
9. The method according to claim 8, characterized in that, The contrast loss function includes a loss function determined based on a first size relationship and a second size relationship. The first size relationship includes the size relationship of the sample quality quantification information corresponding to every two sample images in the training sample set, and the second size relationship includes the size relationship of the sample quality quantification information output by the deep learning model from the two sample images.
10. An image detection device, characterized in that, include: The acquisition module is used to acquire the image to be detected, which includes the image to be subjected to image quality detection; The processing module is used to process the image to be detected using an image detection model to determine the image quality quantification information of the image to be detected, the image degradation region in the image to be detected, and the text description information associated with the image degradation region. The image quality quantification information includes information characterizing whether there is a set anomaly in the image to be detected. The image degradation region includes the region in the image to be detected where the set anomaly exists. The text description information includes the description information of the set anomaly corresponding to the image degradation region.
11. An electronic device, characterized in that, The electronic device includes: One or more processing devices; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processing devices, the one or more processing devices perform the method as described in any one of claims 1-9.
12. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the method as described in any one of claims 1-9.
13. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-9.