Method, device and equipment for acquiring compressed labeled image and storage medium
By compressing and adjusting images, obtaining target adjustment probability information, and determining the target value increment, the generalization and efficiency problems of deep learning methods in image compression are solved, achieving better image preprocessing and codec compression effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2022-04-08
- Publication Date
- 2026-05-19
AI Technical Summary
Existing deep learning methods have poor generalization ability in image compression, and decoding is time-consuming and the model size is large, making it difficult to promote their use in industry.
By compressing the original image, target adjustment probability information is obtained. Based on the adjustment probability information, the image is adjusted and compressed to determine the target value increment. If the set conditions are met, the adjusted image is used as a compressed labeled image to train the image processing model; otherwise, the adjustment is repeated until the conditions are met.
It improves the generalization and compression efficiency of neural networks in image compression, enhances image preprocessing effects, and adapts to codec compression in various scenarios.
Smart Images

Figure CN116934878B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and in particular to a method, apparatus, device, and storage medium for acquiring compressed labeled images. Background Technology
[0002] Image compression is the application of data compression technology to digital images. Its purpose is to reduce redundant information in image data, thereby storing and transmitting data in a more efficient format.
[0003] In recent years, with the rapid development of deep learning technology, how to improve compression efficiency using deep neural networks has become a hot research topic in the field of image compression. Many research results have been achieved, which can be summarized into two main aspects: First, using deep neural networks directly for end-to-end compression. This type of method is time-consuming to decode, has a large model size, and poor compatibility, making it difficult to directly promote and use in industry at present. Second, applying deep learning to the image preprocessing stage, such as image denoising, image decompression distortion, and image super-resolution. These methods all use specific prior methods for data annotation. For example, in image denoising, various noises are randomly added to the original image to obtain a noisy image; in image super-resolution, a high-resolution image is subjected to various degradation operations to obtain a low-resolution image. Models obtained by annotating training data using specific prior methods are only effective for the corresponding images and have poor generalization ability. Summary of the Invention
[0004] This disclosure provides a method, apparatus, device, and storage medium for acquiring compressed labeled images, which can not only improve the generalization of neural networks when compressing images, but also improve the compression efficiency of images.
[0005] In a first aspect, embodiments of this disclosure provide a method for obtaining compressed labeled images, including:
[0006] The original image is compressed to obtain the first compressed image;
[0007] Obtain target adjustment probability information for the current image; wherein, the current image is the original image or an image after at least one adjustment of the original image;
[0008] The current image is adjusted based on the adjustment probability information, and the adjusted image is compressed to obtain a second compressed image;
[0009] Determine the target value increment of the first compressed image and the second compressed image;
[0010] If the target value increment meets the set conditions, the adjusted image is determined as a compressed labeled image, and the set image processing model is trained based on the compressed labeled image;
[0011] If the target value increment does not meet the set condition, the adjusted image is used as the new current image, and the operation of obtaining the target adjustment probability information of the current image is returned until the target value increment meets the set condition.
[0012] Secondly, embodiments of this disclosure also provide an apparatus for acquiring compressed labeled images, comprising:
[0013] The first compressed image acquisition module is used to compress the original image to obtain the first compressed image;
[0014] The target adjustment probability information acquisition module is used to acquire target adjustment probability information of the current image; wherein, the current image is the original image or an image after the original image has been adjusted at least once;
[0015] An image adjustment module is used to adjust the current image based on the adjustment probability information, and to compress the adjusted image to obtain a second compressed image;
[0016] The target value increment determination module is used to determine the target value increment of the first compressed image and the second compressed image;
[0017] The compressed labeled image determination module is used to determine the adjusted image as a compressed labeled image if the target value increment meets the set conditions, so as to train the set image processing model based on the compressed labeled image;
[0018] The return execution module is configured to, if the target value increment does not meet the set conditions, use the adjusted image as the new current image and return to execute the operation of obtaining the target adjustment probability information of the current image until the target value increment meets the set conditions.
[0019] Thirdly, embodiments of this disclosure also provide an electronic device, the electronic device comprising:
[0020] One or more processing devices;
[0021] Storage device for storing one or more programs;
[0022] When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the method for acquiring compressed labeled images as described in the embodiments of this disclosure.
[0023] Fourthly, embodiments of this disclosure also provide a computer-readable medium having a computer program stored thereon, characterized in that, when executed by a processing device, the program implements the method for acquiring compressed labeled images as described in embodiments of this disclosure.
[0024] This disclosure provides a method, apparatus, device, and storage medium for acquiring compressed labeled images. The process involves compressing an original image to obtain a first compressed image; acquiring target adjustment probability information for a current image, wherein the current image is either the original image or an image after at least one adjustment of the original image; adjusting the current image based on the adjustment probability information and compressing the adjusted image to obtain a second compressed image; determining the target value increment of the first and second compressed images; if the target value increment meets a set condition, determining the adjusted image as the compressed labeled image for training a set image processing model; if the target value increment does not meet the set condition, using the adjusted image as the new current image and returning to the operation of acquiring the target adjustment probability information of the current image until the target value increment meets the set condition. Training a set neural network model based on the automatically labeled dataset obtained by the above method can not only improve the model's generalization and preprocessing effect but also improve the image compression effect of the set compression method. Attached Figure Description
[0025] Figure 1 This is a flowchart of a method for obtaining compressed labeled images according to an embodiment of this disclosure;
[0026] Figure 2 This is a schematic diagram of the structure of the multi-task neural network in an embodiment of this disclosure;
[0027] Figure 3 This is a schematic diagram illustrating the principle of acquiring compressed labeled images in an embodiment of this disclosure;
[0028] Figure 4 This is a schematic diagram of the structure of a device for acquiring compressed labeled images according to an embodiment of this disclosure;
[0029] Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of this disclosure. Detailed Implementation
[0030] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0031] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0032] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0033] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0034] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0035] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0036] Figure 1 This is a flowchart illustrating a method for acquiring compressed labeled images according to an embodiment of this disclosure. This embodiment is applicable to situations involving the acquisition of compressed labeled data. The method can be executed by a device for acquiring compressed labeled images. This device can consist of hardware and / or software and is generally integrated into a device with compressed labeled image acquisition capabilities. This device can be an electronic device such as a server, mobile terminal, or server cluster. Figure 1 As shown, the method specifically includes the following steps:
[0037] S110, compress the original image to obtain the first compressed image.
[0038] In this embodiment, any conventional compression method can be used to compress the original image; no limitation is made here. Examples include JPEG (Joint Photographic Experts Group), JPEG-2000, MPEG (Moving Picture Experts Group), VP8, HEVC (High Efficiency Video Coding), etc.
[0039] S120, Obtain the target adjustment probability information of the current image.
[0040] The current image can be the original image or an image that has been adjusted at least once from the original image. The target adjustment probability information contains the probability that each pixel in the current image will be adjusted according to a set method, represented by a matrix of the same size as the current image. The set method can be increasing or decreasing the pixel value by a set value. The set value can be set to 1, meaning the setting method is to either increase or decrease the pixel value by 1.
[0041] In this embodiment, the target adjustment probability information of the current image can be obtained by using a set optimization strategy to determine the target adjustment probability information, or by combining a multi-task neural network with a set optimization strategy to determine the target adjustment probability information. The set optimization strategy can be any of the following: temporal difference strategy, dynamic programming strategy, and Monte Carlo tree search.
[0042] Specifically, the method for obtaining the target adjustment probability information of the current image can be as follows: input the current image into a multi-task neural network, output the initial value increment and the initial adjustment probability information; optimize the initial adjustment probability information based on the set optimization strategy to obtain the target adjustment probability information.
[0043] The multi-task neural network includes a feature extraction sub-network, a value sub-network, and a policy sub-network, with the feature extraction sub-network connected to the value sub-network and the policy sub-network, respectively. Figure 2 This is a schematic diagram of the structure of the multi-task neural network in this embodiment, as shown below. Figure 2 As shown, the process of inputting the current image into a multi-task neural network and outputting the initial value increment and initial adjustment probability information can be as follows: input the current image into the feature extraction sub-network and output the feature map; input the feature map into the value sub-network and the policy sub-network respectively and output the initial value increment and initial adjustment probability information.
[0044] Among them, the value subnetwork is used to determine the value increment, and the strategy subnetwork is used to determine the adjustment probability.
[0045] In this embodiment, taking Monte Carlo Tree Search (MCTS) as an example, the process of optimizing the initial adjustment probability information based on the set optimization strategy is as follows:
[0046] In MCTS, each node represents an image state, and each edge represents an adjustment action performed from one image state to another. Each edge stores four pieces of information: average reward Q, number of visits N, total reward W, and adjustment probability p. The nodes at both ends of an edge are either parent or child nodes. The image state corresponding to the parent node adjusts the pixel value of one of its pixels to obtain the image state corresponding to the child node; that is, the adjustment action is adjusting the pixel value of a certain pixel. In this embodiment, the number of edges corresponding to the current image is related to the image size and the number of adjustment methods (e.g., incrementing or decrementing the pixel value by 1). If the image size is m*n and the number of adjustment methods is 2, then the number of edges is m*n*2.
[0047] Let the initial value increment be V. Taking one edge as an example, the reward for performing an adjustment action corresponding to that edge is: q = p * V + Re, where Re represents the value increment between the two ends of the image after compression.
[0048] During each search, the edges corresponding to the adjustment action are visited starting from the root node and continuing until a leaf node is reached. The root node is the node corresponding to the current image. Adjustment action a t The selection is based on the following formula:
[0049] a t =max(q+u); where, Where N is the number of times the current edge has been visited. This represents the sum of the number of visits to all edges searching downwards from the root node corresponding to the current image, where p is the probability of performing the adjustment action corresponding to the current edge, which can be obtained from the initial adjustment probability information. That is, calculate q+u for each edge, and then visit the edge with the largest q+u.
[0050] After each search is completed, update four pieces of information for each edge: average reward Q, number of visits N, total reward W, and adjustment probability p. The adjustment probability p is updated according to the following formula: Where π represents the updated adjustment probability, and τ is a hyperparameter. This formula represents the relationship between π and... The relationship is directly proportional. The search is iteratively performed in the manner described above until the termination condition is met, yielding the optimized target adjustment probability information Π, where Π is a matrix composed of all π edges. The termination condition can be that the number of visits to any edge reaches a set threshold.
[0051] Optionally, after determining the target value increments of the first compressed image and the second compressed image, the method further includes the following steps: determining a first difference information between the initial value increment and the target value increment; determining a second difference information between the initial adjustment probability information and the target adjustment probability information; determining a loss function based on the first difference information and the second difference information; and training a multi-task neural network based on the loss function.
[0052] The first difference information can be the square of the difference between the initial value increment and the target value increment, denoted as (ZV). 2 The second difference information can be the logarithm of the target adjustment probability information multiplied by the initial adjustment probability information, expressed as: Π·logP.
[0053] One way to determine the loss function based on the first and second difference information is to subtract the second difference information from the first difference information to obtain the loss function. This can be represented as: l = (ZV) 2 -Π·lo. Finally, the multi-task neural network is subjected to inverse parameter tuning based on the loss function.
[0054] Accordingly, the adjusted image is input into the pre-tuned multi-task neural network, which outputs initial value increment and initial adjustment probability information. Then, the initial adjustment probability information is optimized based on the set optimization strategy to obtain the target adjustment probability information. This process continues until the target value increment meets the set conditions.
[0055] S130, adjust the current image based on the target adjustment probability information, compress the adjusted image, and obtain the second compressed image.
[0056] Specifically, after obtaining the target adjustment probability information, the pixel that needs to be adjusted is determined based on the target adjustment probability information, and the pixel value of that pixel is adjusted according to the adjustment probability of that pixel.
[0057] In this embodiment, the method of adjusting the current image based on the adjustment probability information can be: obtaining the pixel point corresponding to the maximum probability in the adjustment probability information as the target pixel point; and adjusting the pixel value of the target pixel point according to the set method.
[0058] The setting method includes increasing or decreasing the pixel value by a set value. Specifically, the maximum probability is obtained from the probability matrix corresponding to the target adjustment probability information, and the pixel value of the pixel with the maximum probability is increased or decreased by 1 to obtain the adjusted image.
[0059] In this embodiment, any conventional compression method can be used to compress the adjusted image; no limitation is made here. Examples include JPEG (Joint Photographic Experts Group), JPEG-2000, MPEG (Moving Picture Experts Group), VP8, and HEVC (High Efficiency Video Coding). In this embodiment, the same algorithm is used for compressing both the original and adjusted images.
[0060] S140, determine the target value increment of the first compressed image and the second compressed image.
[0061] The target value increment can be determined by two parts: the peak signal-to-noise ratio information and the compression ratio information between the first compressed image and the second compressed image.
[0062] Specifically, the target value increment of the first compressed image and the second compressed image can be determined by: determining the peak signal-to-noise ratio (PSNR) difference information between the first compressed image and the second compressed image; determining the compression ratio difference information between the first compressed image and the second compressed image; and performing a weighted summation of the PSNR difference information and the compression ratio difference information to obtain the target value increment.
[0063] The process of determining the Peak Signal-to-Noise Ratio (PSNR) difference between the first compressed image and the second compressed image can be as follows: calculate the PSNR between the first compressed image and the original image, calculate the PSNR between the second compressed image and the adjusted image, and finally calculate the difference between the two PSNRs to obtain the PSNR difference information. The process of determining the compression ratio difference between the first compressed image and the second compressed image can be as follows: calculate the compression ratio between the first compressed image and the original image, calculate the compression ratio between the second compressed image and the adjusted image, and subtract the two compression ratios to obtain the compression ratio difference information.
[0064] Specifically, the process of obtaining the target value increment by weighted summation of PSNR and compression ratio information can be as follows: determine the weights corresponding to PSNR and compression ratio information respectively, and perform weighted summation based on these weights. The formula for calculating the target value increment can be expressed as: V = a(F2-F1) + b(H2-H1), where a and b are weights, F1 is the PSNR information between the first compressed image and the original image, F2 is the PSNR information between the second compressed image and the adjusted image, H1 is the compression ratio between the first compressed image and the original image, and H2 is the compression ratio between the second compressed image and the adjusted image.
[0065] S150, determine whether the target value increment meets the set conditions. If it does, proceed to step 160; otherwise, proceed to step 170.
[0066] The set condition can be that the target value increment exceeds a set threshold. That is, if the target value increment exceeds the set threshold, then S160 is executed; if the target value increment does not exceed the set threshold, then S170 is executed.
[0067] S160, the adjusted image is identified as a compressed labeled image, and the image processing model is trained based on the compressed labeled image.
[0068] In this embodiment, if the target value increment exceeds a set threshold, it indicates that the adjusted image can be used as a compressed annotation image. Specifically, S120-S150 are performed on a large number of original images to obtain a large number of compressed annotation images. These compressed annotation images can be used as samples to train a set image processing model. The trained model has generalizable image preprocessing capabilities. The trained model is used to preprocess the images, and then the preprocessed images are compressed to obtain compressed images. The images output by the model can adapt to compression by traditional codecs better than the original images, and are effective for images with various characteristics, thereby improving the compression effect of existing codecs in various scenarios.
[0069] S170: Use the adjusted image as the new current image and return to execute S120 until the target value increment meets the set conditions.
[0070] For example, Figure 3 This is a schematic diagram illustrating the principle of obtaining compressed labeled images in this embodiment. For example... Figure 3 As shown, the original image is compressed to obtain a first compressed image; the original image is then used as the current image and input into a multi-task neural network to obtain an initial value increment V and initial adjustment probability information P; then, the initial adjustment probability information is optimized based on a set optimization strategy to obtain target adjustment probability information Π; then, the current image is adjusted based on Π to obtain an adjusted image; the adjusted image is compressed to obtain a second compressed image; the target value increment Z between the first and second compressed images is calculated; a first difference information between V and Z is determined, a second difference information between P and Π is determined, a loss function is determined based on the first and second difference information, and a multi-task neural network is trained based on the loss function; the adjusted image is used as the new current image and input into the trained multi-task neural network, and the above steps are repeated until the target value increment meets the set conditions, at which point the adjusted image is determined as the compressed labeled image.
[0071] The technical solution of this disclosure involves compressing an original image to obtain a first compressed image; acquiring target adjustment probability information for the current image; wherein the current image is the original image or an image after at least one adjustment of the original image; adjusting the current image based on the adjustment probability information and compressing the adjusted image to obtain a second compressed image; determining the target value increment of the first compressed image and the second compressed image; if the target value increment meets a set condition, determining the adjusted image as a compressed labeled image for training a set image processing model based on the compressed labeled image; if the target value increment does not meet the set condition, using the adjusted image as the new current image and returning to the operation of acquiring the target adjustment probability information of the current image until the target value increment meets the set condition. Training a set neural network model based on the automatically labeled dataset obtained by the above method can not only improve the model's generalization and preprocessing effect but also improve the image compression effect of the set compression method.
[0072] Figure 4 This is a schematic diagram of the structure of a device for acquiring compressed labeled images provided in an embodiment of this disclosure, as shown below. Figure 4 As shown, the device includes:
[0073] The first compressed image acquisition module 210 is used to compress the original image to obtain the first compressed image;
[0074] The target adjustment probability information acquisition module 220 is used to acquire the target adjustment probability information of the current image; wherein, the current image is the original image or an image after at least one adjustment of the original image;
[0075] The image adjustment module 230 is used to adjust the current image based on the target adjustment probability information, and to compress the adjusted image to obtain a second compressed image;
[0076] The target value increment determination module 240 is used to determine the target value increment of the first compressed image and the second compressed image;
[0077] The compressed annotation image determination module 250 is used to determine the adjusted image as a compressed annotation image if the target value increment meets the set conditions, so as to train the set image processing model based on the compressed annotation image.
[0078] The execution module 260 is used to take the adjusted image as the new current image if the target value increment does not meet the set conditions, and return to execute the operation of obtaining the target adjustment probability information of the current image until the target value increment meets the set conditions.
[0079] Optionally, the target adjustment probability information acquisition module 220 is also used for:
[0080] The current image is input into a multi-task neural network, which outputs the initial value increment and initial adjustment probability information.
[0081] The initial adjustment probability information is optimized based on the set optimization strategy to obtain the target adjustment probability information.
[0082] Optionally, it also includes: a multi-task neural network training module, used for:
[0083] Determine the first difference information between the initial value increment and the target value increment;
[0084] Determine the second difference information between the initial adjustment probability information and the target adjustment probability information;
[0085] The loss function is determined based on the first and second difference information.
[0086] Training a multi-task neural network based on a loss function.
[0087] Optionally, returning to execution module 260 is also used for:
[0088] Return to the current image input into the trained multi-task neural network.
[0089] Optionally, the multi-task neural network includes a feature extraction sub-network, a value sub-network, and a policy sub-network; the feature extraction sub-network is connected to the value sub-network and the policy sub-network, respectively; the current image is input into the multi-task neural network, and the outputs initial value increment and initial adjustment probability information, including:
[0090] Input the current image into the feature extraction subnetwork, and output a feature map;
[0091] The feature maps are input into the value subnetwork and the policy subnetwork respectively, and the initial value increment and initial adjustment probability information are output.
[0092] Optionally, the optimization strategy can be set to any of the following: time difference strategy, dynamic programming strategy, and Monte Carlo tree search.
[0093] Optionally, the target adjustment probability information includes the probability that each pixel in the current image will be adjusted according to a set method; the image adjustment module 230 is also used for:
[0094] Obtain the pixel point corresponding to the maximum probability in the target adjustment probability information, and use it as the target pixel point;
[0095] Adjust the pixel value of the target pixel according to the set method; the set method includes increasing or decreasing the pixel value by a set value.
[0096] Optionally, the target value increment determination module 240 is also used for:
[0097] Determine the peak signal-to-noise ratio (PSNR) difference information between the first compressed image and the second compressed image;
[0098] Determine the compression ratio difference information between the first compressed image and the second compressed image;
[0099] The target value increment is obtained by weighted summation of PSNR difference information and compression ratio difference information.
[0100] The above-described apparatus can execute the methods provided in all the foregoing embodiments of this disclosure, and has the corresponding functional modules and beneficial effects for executing the above methods. Technical details not described in detail in this embodiment can be found in the methods provided in all the foregoing embodiments of this disclosure.
[0101] The following is for reference. Figure 5 The diagram illustrates a structural schematic of an electronic device 300 suitable for implementing embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs, desktop computers, or various forms of servers, such as standalone servers or server clusters. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0102] like Figure 5 As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a memory device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0103] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0104] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing a method of word recommendation. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 309, or installed from a storage device 308, or installed from a ROM 302. When the computer program is executed by a processing device 301, it performs the functions defined in the methods of embodiments of this disclosure.
[0105] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0106] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0107] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0108] The aforementioned computer-readable medium carries one or more programs. When the electronic device executes the aforementioned one or more programs, the electronic device causes the following: compresses the original image to obtain a first compressed image; acquires target adjustment probability information of the current image; wherein the current image is the original image or an image after at least one adjustment of the original image; adjusts the current image based on the target adjustment probability information and compresses the adjusted image to obtain a second compressed image; determines the target value increment of the first compressed image and the second compressed image; if the target value increment meets a set condition, the adjusted image is determined as a compressed labeled image for training a set image processing model based on the compressed labeled image; if the target value increment does not meet the set condition, the adjusted image is used as the new current image, and the operation of acquiring the target adjustment probability information of the current image is returned until the target value increment meets the set condition.
[0109] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0110] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0111] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units are not, in some cases, intended to limit the specific unit.
[0112] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0113] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0114] According to one or more embodiments of the present disclosure, the present disclosure discloses a method for obtaining compressed labeled images, including:
[0115] The original image is compressed to obtain the first compressed image;
[0116] Obtain target adjustment probability information for the current image; wherein, the current image is the original image or an image after at least one adjustment of the original image;
[0117] The current image is adjusted based on the target adjustment probability information, and the adjusted image is compressed to obtain a second compressed image;
[0118] Determine the target value increment of the first compressed image and the second compressed image;
[0119] If the target value increment meets the set conditions, the adjusted image is determined as a compressed labeled image, and the set image processing model is trained based on the compressed labeled image;
[0120] If the target value increment does not meet the set condition, the adjusted image is taken as the new current image, and the operation of obtaining the target adjustment probability information of the current image is returned until the target value increment meets the set condition.
[0121] Furthermore, the target adjustment probability information of the current image is obtained, including:
[0122] The current image is input into a multi-task neural network, which outputs the initial value increment and initial adjustment probability information.
[0123] The initial adjustment probability information is optimized based on the set optimization strategy to obtain the target adjustment probability information.
[0124] Furthermore, after determining the target value increments of the first compressed image and the second compressed image, the method further includes:
[0125] Determine the first difference information between the initial value increment and the target value increment;
[0126] Determine the second difference information between the initial adjustment probability information and the target adjustment probability information;
[0127] The loss function is determined based on the first difference information and the second difference information;
[0128] The multi-task neural network is trained based on the loss function.
[0129] Furthermore, the operation of obtaining the target adjustment probability information of the current image is returned, including:
[0130] Return to the execution and input the current image into the trained multi-task neural network.
[0131] Further, the multi-task neural network includes a feature extraction sub-network, a value sub-network, and a policy sub-network; the feature extraction sub-network is connected to the value sub-network and the policy sub-network, respectively; the current image is input into the multi-task neural network, and the initial value increment and initial adjustment probability information are output, including:
[0132] The current image is input into the feature extraction subnetwork, and a feature map is output.
[0133] The feature maps are input into the value subnetwork and the policy subnetwork respectively, and the initial value increment and initial adjustment probability information are output.
[0134] Furthermore, the optimization strategy can be any of the following: time difference strategy, dynamic programming strategy, and Monte Carlo tree search.
[0135] Further, the target adjustment probability information includes the probability that each pixel in the current image will be adjusted according to a set method; adjusting the current image based on the target adjustment probability information includes:
[0136] Obtain the pixel point corresponding to the maximum probability in the target adjustment probability information, and use it as the target pixel point;
[0137] The pixel value of the target pixel is adjusted according to the setting method; wherein, the setting method includes increasing or decreasing the pixel value by a set value.
[0138] Further, determining the target value increment of the first compressed image and the second compressed image includes:
[0139] Determine the peak signal-to-noise ratio (PSNR) difference information between the first compressed image and the second compressed image;
[0140] Determine the compression ratio difference information between the first compressed image and the second compressed image;
[0141] The target value increment is obtained by weighted summation of the PSNR difference information and the compression ratio difference information.
[0142] Note that the above description is merely a preferred embodiment and the technical principles employed in this disclosure. Those skilled in the art will understand that this disclosure is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of this disclosure. Therefore, although this disclosure has been described in detail through the above embodiments, it is not limited to the above embodiments. Many other equivalent embodiments may be included without departing from the concept of this disclosure, and the scope of this disclosure is determined by the scope of the appended claims.
Claims
1. A method for obtaining compressed labeled images, characterized in that, include: The original image is compressed to obtain the first compressed image; Obtain target adjustment probability information for the current image; wherein, the current image is the original image or an image after at least one adjustment of the original image; The current image is adjusted based on the target adjustment probability information, and the adjusted image is compressed to obtain a second compressed image; Determine the target value increment of the first compressed image and the second compressed image; If the target value increment meets the set conditions, the adjusted image is determined as a compressed labeled image, and the set image processing model is trained based on the compressed labeled image; If the target value increment does not meet the set condition, the adjusted image is taken as the new current image, and the operation of obtaining the target adjustment probability information of the current image is returned until the target value increment meets the set condition. Among them, obtaining the target adjustment probability information of the current image includes: The current image is input into a multi-task neural network, which outputs the initial value increment and initial adjustment probability information. The initial adjustment probability information is optimized based on the set optimization strategy to obtain the target adjustment probability information.
2. The method according to claim 1, characterized in that, After determining the target value increments of the first compressed image and the second compressed image, the method further includes: Determine the first difference information between the initial value increment and the target value increment; Determine the second difference information between the initial adjustment probability information and the target adjustment probability information; The loss function is determined based on the first difference information and the second difference information; The multi-task neural network is trained based on the loss function.
3. The method according to claim 2, characterized in that, Returning to the operation of obtaining target adjustment probability information for the current image, including: Return to the execution and input the current image into the trained multi-task neural network.
4. The method according to claim 1, characterized in that, The multi-task neural network includes a feature extraction subnetwork, a value subnetwork, and a policy subnetwork; the feature extraction subnetwork is connected to the value subnetwork and the policy subnetwork, respectively. The current image is input into a multi-task neural network, which outputs initial value increment and initial adjustment probability information, including: The current image is input into the feature extraction subnetwork, and a feature map is output. The feature maps are input into the value subnetwork and the policy subnetwork respectively, and the initial value increment and initial adjustment probability information are output.
5. The method according to claim 1, characterized in that, The optimization strategy can be any one of the following: time difference strategy, dynamic programming strategy, and Monte Carlo tree search.
6. The method according to claim 1, characterized in that, The target adjustment probability information includes the probability that each pixel in the current image will be adjusted according to a set method; adjusting the current image based on the target adjustment probability information includes: Obtain the pixel corresponding to the maximum probability in the target adjustment probability information, and use it as the target pixel; The pixel value of the target pixel is adjusted according to the setting method; wherein, the setting method includes increasing or decreasing the pixel value by a set value.
7. The method according to claim 1, characterized in that, Determining the target value increment of the first compressed image and the second compressed image includes: Determine the peak signal-to-noise ratio (PSNR) difference information between the first compressed image and the second compressed image; Determine the compression ratio difference information between the first compressed image and the second compressed image; The target value increment is obtained by weighted summation of the PSNR difference information and the compression ratio difference information.
8. A device for acquiring compressed labeled images, characterized in that, A method for acquiring compressed labeled images according to any one of claims 1-7, comprising: The first compressed image acquisition module is used to compress the original image to obtain the first compressed image; The target adjustment probability information acquisition module is used to acquire target adjustment probability information of the current image; wherein, the current image is the original image or an image after the original image has been adjusted at least once; An image adjustment module is used to adjust the current image based on the target adjustment probability information, and to compress the adjusted image to obtain a second compressed image; The target value increment determination module is used to determine the target value increment of the first compressed image and the second compressed image; The compressed labeled image determination module is used to determine the adjusted image as a compressed labeled image if the target value increment meets the set conditions, so as to train the set image processing model based on the compressed labeled image; The return execution module is configured to, if the target value increment does not meet the set conditions, use the adjusted image as the new current image and return to execute the operation of obtaining the target adjustment probability information of the current image until the target value increment meets the set conditions.
9. An electronic device, characterized in that, The electronic device includes: One or more processing devices; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the method for acquiring compressed labeled images as described in any one of claims 1-7.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the method for acquiring compressed labeled images as described in any one of claims 1-7.