Image pooling method, device, terminal device and computer readable storage medium

By performing weighted pooling on the image, the problem of loss of detail in existing pooling layer methods is solved, achieving image processing effects that are more in line with human vision.

CN114792289BActive Publication Date: 2026-05-29WUHAN TCL CORP RES CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN TCL CORP RES CO LTD
Filing Date
2021-01-26
Publication Date
2026-05-29

Smart Images

  • Figure CN114792289B_ABST
    Figure CN114792289B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an image pooling method and device, terminal equipment and computer readable storage medium. The method comprises: obtaining a to-be-processed image; performing first pooling processing on the to-be-processed image to obtain a reference image; assigning corresponding weights to the pixel points of the to-be-processed image according to the reference image; and performing second pooling processing on the to-be-processed image and the reference image according to the weights to obtain a target image. According to the embodiments of the present application, weights are assigned to pixel points at different positions of the to-be-processed image according to a weight formula, and the to-be-processed image and the reference image are subjected to second pooling processing according to the weights to obtain a final target image, thereby retaining more image details and being more consistent with the subjective perception of human vision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of image processing technology, and in particular relates to an image pooling method, apparatus, terminal device and computer-readable storage medium. Background Technology

[0002] With the continuous development of AI technology, image-based neural network technology has also entered a stage of rapid development. In today's world, image resolution is increasingly higher, and pooling layers have become a common and necessary network layer used to reduce resolution during neural network training. It has significant advantages such as increasing the receptive field, maintaining translational integrity, and reducing the difficulty of neural network optimization, making it an indispensable part of existing neural networks and of great importance.

[0003] Existing pooling layers include average pooling, max pooling, and random pooling. Although these pooling methods are common and easy to use, and each has its own characteristics in neural networks, they blur the edges and lose a lot of details, resulting in poor visual subjective perception and failing to meet the human visual needs for semantic information and edge details. Summary of the Invention

[0004] This application provides an image pooling method, apparatus, terminal device, and computer-readable storage medium to solve the problem of insufficient image detail in existing pooling techniques.

[0005] In a first aspect, embodiments of this application provide an image pooling method, including:

[0006] Obtain the image to be processed;

[0007] The image to be processed is subjected to the first pooling process to obtain the reference image;

[0008] The corresponding weights are assigned to the pixels of the image to be processed based on the reference image;

[0009] Based on the weights, a second pooling process is performed on the image to be processed and the reference image to obtain the target image.

[0010] Secondly, embodiments of this application provide an image processing apparatus, comprising:

[0011] The acquisition module is used to acquire the image to be processed;

[0012] The first pooling module is used to perform first pooling processing on the image to be processed to obtain a reference image;

[0013] The weighting module is used to assign corresponding weights to the pixels of the image to be processed based on the reference image.

[0014] The second pooling module is used to perform a second pooling process on the image to be processed and the reference image according to the weights to obtain the target image.

[0015] Thirdly, embodiments of this application provide a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the image pooling method described in any of the first aspects above.

[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the image pooling method described in any of the first aspects above.

[0017] Fifthly, embodiments of this application provide a computer program product that, when run on a terminal device, causes the terminal device to execute the steps of the image pooling method described in any of the first aspects above.

[0018] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.

[0019] The beneficial effects of this application embodiment compared with the prior art are as follows: This application embodiment provides an image pooling method, apparatus, terminal device, and computer-readable storage medium. The method includes: acquiring an image to be processed; performing a first pooling process on the image to be processed to obtain a reference image; assigning corresponding weights to the pixels of the image to be processed according to the reference image; and performing a second pooling process on the image to be processed and the reference image according to the weights to obtain a target image. This application embodiment obtains the final target image by assigning weights to pixels at different positions in the image to be processed according to a weight formula, and then performing a second pooling process on the image to be processed and the reference image according to the weights, thereby preserving more image details and better conforming to the subjective perception of human vision. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart illustrating the image pooling method provided in an embodiment of this application;

[0022] Figure 2 This is a schematic diagram of the image pooling apparatus provided in the embodiments of this application;

[0023] Figure 3 This is a schematic diagram of the structure of the terminal device provided in the embodiments of this application. Detailed Implementation

[0024] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0025] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0026] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0027] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determination" or "if the described condition or event is detected" may be interpreted, depending on the context, as "once determination," "in response to determination," "once the described condition or event is detected," or "in response to the detection of the described condition or event."

[0028] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0029] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0030] The technical solutions provided in the embodiments of this application will be described in detail below.

[0031] See Figure 1 The diagram illustrates a flowchart of an image pooling method. This method can be applied to terminal devices, such as mobile phones, tablets, or computers; the type of terminal device is not limited here. The method may include the following steps:

[0032] Step S101: Obtain the image to be processed, I.

[0033] Step S102: Perform a first pooling process on the image to be processed I to obtain a reference image. The image to be processed can be a face image, landscape image, or car image, etc., and its format can be JPEG, TIFF, or RAW, etc. This application does not impose any limitations on this. There are two methods for obtaining the reference image:

[0034] The first method: First, downsample the image I to be processed to obtain the downsampled image I. B .

[0035] A box filter can be used to downsample the image I to be processed. The size of the downsampled image is determined by the ratio between the target image and the image to be processed, for example, it is 1 / 2 or 1 / 4 of the length and width of the image to be processed. Alternatively, nearest neighbor interpolation, bilinear interpolation, or trilinear interpolation can also be used to downsample the image I to obtain the downsampled image. The target image is the image we ultimately want to obtain.

[0036] This application embodiment reduces high-frequency components in the image to be processed by downsampling, thereby obtaining low-frequency components, which is beneficial for obtaining global information of the image to be processed.

[0037] Then, the downsampled image I is processed using a filter template. B After smoothing, the reference image I is obtained.

[0038] The filter template can be an approximate Gaussian filter template, a bilateral filter template, or a guided filter template. When using an approximate Gaussian filter template, a 3x3 or 5x5 matrix can be selected. For example, when selecting a 3x3 approximate Gaussian filter template, the formula for generating the reference image is as follows:

[0039]

[0040] The second method: Based on the learnable parameters and the image to be processed, a reference image is obtained, where the formula for obtaining the reference image is:

[0041]

[0042] Where F(q) represents the neighborhood Ω p The learnable parameters of a neural network.

[0043] In one embodiment, the image to be processed is input into a neural network to generate a reference image, wherein the learnable parameters of the neural network are F(q), and the image to be processed is processed by the neural network to obtain a neighborhood Ω of the original image. p The reference image is generated by mapping the reference image to a single pixel.

[0044] Ω p Ω represents a neighborhood in the image to be processed. The size of the neighborhood is typically 2*2 or 3*3, and the neighborhood can contain multiple pixels. p Let p represent a pixel in the reference image corresponding to the pixel in the neighborhood, i.e., the neighborhood corresponds to pixel p in the reference image, and F(q) represent the neighborhood Ω. p The learnable parameters of a neural network.

[0045] Before obtaining the reference image based on the learnable parameters and the image to be processed, the process also includes training the neural network to obtain its learnable parameters and / or trainable parameters. This mainly includes two stages:

[0046] Forward propagation phase: The phase in which data is propagated from lower to higher levels, i.e., the forward propagation phase.

[0047] Backpropagation Phase: When the results obtained from the forward propagation do not match the expectations, the error is propagated from higher to lower levels for training; this is the backpropagation phase. The specific training process is as follows:

[0048] Step 1: Initialize the weights of the neural network. All parameters of F(q) are initialized to 0, α and λ are initialized to 1, and then a zero-mean Gaussian perturbation is added to all parameters to break the symmetry.

[0049] Step 2: The input data is propagated forward through convolutional layers, pooling layers, and fully connected layers to obtain the output value. The input data includes the image to be processed, as well as arbitrary images and feature maps.

[0050] Step 3: Calculate the error between the output value of the neural network and the target value, where the target value is set by the user according to their needs.

[0051] Step 4: When the error exceeds our expected value, the error is propagated back into the network, and training proceeds from higher to lower layers. This backpropagation phase sequentially calculates the errors of the fully connected layers, pooling layers, and convolutional layers. The error of each layer can be understood as how much of the total error of the neural network it should bear. When the error is less than or equal to our expected value, the weights are updated based on the error, and training ends.

[0052] Step S103: Assign corresponding weights to the pixels of the image to be processed based on the reference image.

[0053] The neighborhood Ω on the image to be processed p Different weights ω are assigned to pixels q at different positions within the pixel. α,λ [p, q], the greater the difference between the pixel value of pixel q and the pixel value of pixel p, the greater the weight of pixel q. Proportional to weight The larger the difference, the greater the weight. Here, q represents a neighborhood Ω on the image to be processed. p A pixel within the reference image, where p represents the corresponding neighboring pixel in the reference image, ω α,λ [p,q] represents the weight of pixel q.

[0054] For example, the formula for the weights is as follows:

[0055]

[0056] In this formula, I[q] represents the pixel value of pixel q in the image to be processed. Let ε be the pixel value of pixel P in the reference image, α and λ be fixed constants, and n be a natural number. ε is a very small fixed constant whose specific value can be determined based on the actual situation. α and λ are both trainable parameters that are adaptively adjusted during neural network training. α and λ are optimized to ensure they are both positive numbers.

[0057] Step S104: According to the weights, perform a second pooling process on the image to be processed and the reference image to obtain the target image, which is the pooled image of the image to be processed.

[0058] This embodiment of the application obtains the final target image by performing a second pooling process on the image to be processed and the reference image according to the weights assigned to different pixel points at different positions in the image to be processed in step S103. For example, weights ω are assigned to different pixel points q at different positions in the image to be processed based on their pixel values. α,λ [p,q], based on weight ω α,λ[p,q] performs a second pooling process on the image to be processed and the reference image to obtain the final target image output O(I)[p], where O(I)[p] represents the pixel value of the pixel corresponding to pixel position P in the target image. The larger the difference between the pixel value of pixel q and the pixel value of pixel p, the larger the weight corresponding to pixel q, i.e. Proportional to weight The greater the difference in pixel values, the greater the weight. The greater the difference in pixel values, the richer the colors in the image detail areas. In this embodiment, by assigning greater weight to the detail areas, the image detail areas have a greater impact and contribution on the final output target image, thereby preserving more image details and better conforming to the subjective perception of human vision.

[0059] In one embodiment, a pooling formula is further determined based on the weights; according to the pooling formula, a second pooling process is performed on the image to be processed and the reference image to obtain the target image. For example, the reference image and the image to be processed are input into a neural network, and a second pooling process is performed through a detail-preserving pooling filter to obtain the target image.

[0060] The optional pooling formula is as follows:

[0061]

[0062] Furthermore, the embodiments of this application preserve more image details through the second pooling filter that retains details, which is more in line with the subjective perception of human vision and is also more conducive to subsequent image processing, such as image segmentation, image classification, image differencing, or image compression.

[0063] Although the noise will be over-amplified as a result, the existence of trainable parameters α and λ precisely limits the over-amplification of the noise.

[0064] It should be noted that other sorting schemes that can be easily conceived by those skilled in the art within the technical scope disclosed in this invention should also be within the protection scope of this invention, and will not be elaborated here.

[0065] See Figure 2 This is a schematic diagram of an image pooling apparatus provided in an embodiment of this application. For ease of explanation, only the parts related to the embodiments of the present invention are shown, including:

[0066] Acquisition module 21 is used to acquire the image to be processed;

[0067] The first pooling module 22 is used to perform first pooling processing on the image to be processed to obtain a reference image;

[0068] The weighting module 23 is used to assign corresponding weights to the pixels of the image to be processed based on the reference image;

[0069] The second pooling module 24 is used to perform a second pooling process on the image to be processed and the reference image according to the weights to obtain the target image, which is the pooled image of the image to be processed.

[0070] Weight module 23 is also used for neighborhood Ω on the image to be processed. p Different weights are assigned to pixels q at different positions within the image. The greater the difference between the pixel value of pixel q and the pixel value of pixel p, the greater the weight of pixel q. Here, q represents a neighborhood Ω in the image to be processed. p A pixel within the reference image, where p represents a pixel in the neighborhood of the reference image.

[0071] Optionally, the formula for the weights is:

[0072] Where q represents a neighborhood Ω on the image to be processed. p A pixel within the reference image, where p represents the corresponding neighboring pixel in the reference image, ω α,λ [p,q] represents the weight of pixel q, and I[q] represents the pixel value of pixel q in the image to be processed. Let ε be the pixel value of pixel P in the reference image, α and λ be a fixed constant, and n be a natural number.

[0073] The first pooling module 22 is also used to obtain a reference image based on learnable parameters and the image to be processed, wherein the formula for obtaining the reference image is:

[0074]

[0075] Where F(q) represents the neighborhood Ω p The learnable parameters of a neural network.

[0076] The second pooling module 24 is also used to determine the pooling formula according to the weights; and to perform second pooling processing on the image to be processed and the reference image according to the pooling formula to obtain the target image.

[0077] The optional pooling formula is as follows:

[0078]

[0079] Where O(I)[p] represents the pixel value of the pixel corresponding to the pixel position P in the target image.

[0080] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional modules is merely an example. In practical applications, the above functions can be assigned to different functional units or modules as needed, that is, the internal structure of the mobile terminal can be divided into different functional units or modules to complete all or part of the functions described above. The functional modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the modules in the mobile terminal can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0081] Figure 3 This is a schematic diagram of the terminal device provided in an embodiment of this application. For example... Figure 3 As shown, the terminal device 3 in this embodiment includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30. When the processor 30 executes the computer program 32, it implements the steps of the image pooling method described above, for example... Figure 1 Steps 101 to 104 are shown. Alternatively, when processor 30 executes computer program 32, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 3 The functions of modules 21 to 24 are shown.

[0082] For example, computer program 32 can be divided into one or more modules / units, one or more of which are stored in memory 31 and executed by processor 30 to complete the present invention. One or more modules / units can be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of computer program 32 in terminal device 3.

[0083] Terminal device 3 can be a desktop computer, laptop, handheld computer, or other computing device. The terminal device may include, but is not limited to, processor 30 and memory 31. Those skilled in the art will understand that... Figure 3 This is merely an example of terminal device 3 and does not constitute a limitation on terminal device 3. It may include more or fewer components than shown, or combine certain components, or different components. For example, terminal device may also include input / output devices, network access devices, buses, etc.

[0084] The processor 30 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0085] The memory 31 can be an internal storage unit of the terminal device 3, such as a hard disk or RAM of the terminal device 3. The memory 31 can also be an external storage device of the terminal device 3, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the terminal device 3. Furthermore, the memory 31 can include both internal and external storage units of the terminal device 3. The memory 31 is used to store computer programs and other programs and data required by the terminal device. The memory 31 can also be used to temporarily store data that has been output or will be output.

[0086] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps described in the various method embodiments above.

[0087] This application provides a computer program product that, when run on a mobile terminal, enables the mobile terminal to implement the steps described in the above-described method embodiments.

[0088] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0089] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0090] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0091] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / terminal devices and methods can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0092] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0093] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0094] If an integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0095] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. An image pooling method, characterized in that, include: Obtain the image to be processed; The image to be processed is subjected to a first pooling process to obtain a reference image; The pixels of the image to be processed are assigned corresponding weights based on the reference image; According to the weights, a second pooling process is performed on the image to be processed and the reference image to obtain the target image; The step of assigning corresponding weights to the pixels of the image to be processed based on the reference image includes: For the neighborhood of the image to be processed A weight is assigned to a pixel q within the image to be processed. The greater the difference between the pixel value of pixel q and the pixel value of pixel p, the greater the weight of pixel q. Here, q represents a neighborhood in the image to be processed. A pixel within the reference image, where p represents a pixel in the reference image corresponding to the neighborhood. The step of performing a second pooling process on the image to be processed and the reference image according to the weights to obtain a target image includes: inputting the reference image and the image to be processed into a neural network, and performing a second pooling process through a pooling filter that preserves details to obtain the target image.

2. The method as described in claim 1, characterized in that, Assigning corresponding weights to the pixels of the image to be processed based on the reference image includes: Where q represents a neighborhood on the image to be processed. A pixel within the reference image, where p represents a pixel in the reference image corresponding to the neighborhood. The weight representing pixel q, This represents the pixel value of pixel q in the image to be processed. The pixel value of pixel point P in the reference image. It is a fixed constant. Let n be the trainable parameters of the neural network, and n be a natural number.

3. The method as described in claim 1, characterized in that, Based on the weights, a second pooling process is performed on the image to be processed and the reference image to obtain the target image, including: The pooling formula is determined based on the weights; According to the pooling formula, a second pooling process is performed on the image to be processed and the reference image to obtain the target image; The pooling formula is as follows: in, q represents the pixel value of the pixel at position P in the target image, and q represents a neighborhood in the image to be processed. A pixel within the reference image, where p represents a pixel in the reference image corresponding to the neighborhood. The weight representing pixel q, The pixel value represents pixel q in the image to be processed.

4. The method as described in claim 1, characterized in that, The image to be processed undergoes a first pooling process to obtain a reference image, including: Based on the learnable parameters and the image to be processed, a reference image is obtained, wherein the formula for obtaining the reference image is: in, Let F(q) be the pixel value of pixel P in the reference image, where p represents the pixel in the reference image corresponding to the neighborhood, and F(q) represents the neighborhood. The learnable parameters of the neural network, where q represents a neighborhood on the image to be processed. One pixel within, The pixel value represents the pixel value of pixel q in the image to be processed.

5. The method as described in claim 4, characterized in that, Before obtaining the reference image based on the learnable parameters and the image to be processed, the process further includes: The neural network is trained to obtain the learnable parameters and / or trainable parameters of the neural network.

6. The method as described in claim 1, characterized in that, The image to be processed is subjected to a first pooling process to obtain a reference image, including: The image to be processed is downsampled to obtain a downsampled image; The downsampled image is smoothed using a filter template to obtain a reference image.

7. An image pooling apparatus, characterized in that, include: The acquisition module is used to acquire the image to be processed; The first pooling module is used to perform a first pooling process on the image to be processed to obtain a reference image; The weighting module is used to assign corresponding weights to the pixels of the image to be processed based on the reference image. The second pooling module is used to perform a second pooling process on the image to be processed and the reference image according to the weights to obtain the target image; The weighting module is also used to weight the neighborhood of the image to be processed. The weight assigned to a pixel q within the image to be processed is determined by the difference between the pixel value of pixel q and the pixel value of pixel p. The greater the difference between pixel value q and pixel value p, the greater the weight of pixel q. A pixel within the reference image, where p represents a pixel in the reference image corresponding to the neighborhood. The second pooling module is also used to input the reference image and the image to be processed into the neural network, and perform a second pooling process through a pooling filter that preserves details to obtain the target image.

8. The apparatus as claimed in claim 7, characterized in that, The weighting module is further configured to assign corresponding weights to the pixels of the image to be processed based on the reference image, wherein: Where q represents a neighborhood on the image to be processed. A pixel within the reference image, where p represents a pixel in the reference image corresponding to the neighborhood. The weight representing pixel q, This represents the pixel value of pixel q in the image to be processed. The pixel value of pixel point P in the reference image. It is a fixed constant. Let n be the trainable parameters of the neural network, and n be a natural number.

9. The apparatus as claimed in claim 7, characterized in that, The first pooling module is further configured to obtain a reference image based on learnable parameters and the image to be processed, wherein the formula for obtaining the reference image is: in, Let F(q) be the pixel value of pixel P in the reference image, where p represents the pixel in the reference image corresponding to the neighborhood, and F(q) represents the neighborhood. The learnable parameters of the neural network, where q represents a neighborhood on the image to be processed. One pixel within, The pixel value represents the pixel value of pixel q in the image to be processed.

10. The apparatus as claimed in claim 8, characterized in that, The second pooling module is further configured to determine a pooling formula based on the weights; and to perform a second pooling process on the image to be processed and the reference image according to the pooling formula to obtain a target image. The pooling formula is as follows: in, q represents the pixel value of the pixel at position P in the target image, and q represents a neighborhood in the image to be processed. A pixel within the reference image, where p represents a pixel in the reference image corresponding to the neighborhood. The weight representing pixel q, The pixel value represents the pixel value of pixel q in the image to be processed.

11. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the image pooling method as described in any one of claims 1 to 6.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the image pooling method as described in any one of claims 1 to 6.