Training Method, Image Processing Method and Apparatus of Image Processing Network
By determining a reference pixel and modeling cropping as a Markov chain, the training method optimizes network parameters and probabilities, improving the accuracy and efficiency of gaze estimation in image processing networks.
Patent Information
- Application Number
- JP2024576999
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-06-30
- Filing Date
- 2023-03-08
- Publication Date
- 2025-07-03
AI Technical Summary
Existing image processing networks for gaze estimation suffer from low accuracy due to redundant pixels in training images, which are not effectively addressed by current channel pruning methods like differentiable Markov channel pruning (DMCP).
A training method that determines a reference pixel in a labeled training image, models the cropping operation as a Markov chain, and adjusts network parameters and cropping probabilities based on output results to optimize the image processing network.
Improves the accuracy of gaze estimation by selectively retaining relevant image regions, reducing computational load and enhancing the performance of the trained image processing network.
Smart Images

Figure 2025520858000001_ABST
Abstract
Description
Technical Field
[0001] (Cross - reference to related applications) This disclosure is based on and claims priority to a Chinese patent application with application number 202210772860.1, filing date June 30, 2022, and application title "Training Method, Image Processing Method and Apparatus for Image Processing Network", and all contents of the Chinese patent application are incorporated herein by reference.
[0002] This disclosure relates to the field of computer vision technology, and in particular, to a training method, an image processing method and an apparatus for an image processing network.
Background Art
[0003] Channel pruning is widely used for accelerating and compressing models in order to deploy parameterized convolutional image processing networks (CNNs: Convolutional Neural Networks) to embedded devices or mobile devices. In related technologies, channel pruning is performed on the network by means of differentiable Markov channel pruning (DMCP). In DMCP, the channel pruning process is modeled as a Markov chain to reduce the search space. However, in gaze estimation, since the training images contain redundant pixels, the accuracy of the trained image processing network for performing gaze estimation is not high.
Summary of the Invention
[0004] In view of this, embodiments of the present disclosure at least provide a training method for an image processing network, an image processing method and an apparatus.
[0005] The technical solutions of the embodiments of the present disclosure are realized as follows.
[0006] In a first aspect, embodiments of the present disclosure provide a training method for an image processing network, the method comprising: Based on the training images labeled with true values, determining a reference pixel; Using the reference pixel as a starting point, based on the Markov chain of the training image, determining a clipping probability when processing the training image with an image processing network; Based on the output result of processing the training clipping region by the image processing network and the true value, adjusting the network parameter value and the clipping probability of the image processing network to obtain a trained image processing network, wherein the training clipping region is obtained by clipping the training image based on the clipping probability, and the method includes the above steps.
[0007] In a second aspect, embodiments of the present disclosure provide an image processing method, the method including: Obtaining a processing target image; Based on the clipping probability of the trained image processing network, performing pixel clipping on the processing target image to obtain a clipping region waiting to be processed, wherein the trained image processing network is obtained by training based on the training method of the image processing network in the first aspect above; Using the trained image processing network to process the clipping region waiting to be processed to obtain a processing result of the processing target image.
[0008] In a third aspect, embodiments of the present disclosure provide a training apparatus for an image processing network, the apparatus including: A first determination unit configured to determine a reference pixel based on a training image labeled with a true value; A second determination unit configured to determine a clipping probability when processing the training image with an image processing network based on the Markov chain of the training image using the reference pixel as a starting point; A first adjustment unit configured to adjust network parameter values of the image processing network and the cropping probability based on an output result of processing a training cropping region by the image processing network and the ground truth, so as to obtain a trained image processing network, wherein the training cropping region is obtained by cropping the training image based on the cropping probability, and the first adjustment unit.
[0009] In a fourth aspect, an embodiment of the present disclosure provides an image processing apparatus, the apparatus including: A first acquisition unit configured to acquire a processing target image; A first cropping unit configured to perform pixel cropping on the processing target image based on a cropping probability of a trained image processing network to obtain a cropping region waiting for processing, wherein the trained image processing network is obtained by training based on the above image processing network training method, and the first cropping unit; A first processing unit configured to process the cropping region waiting for processing using the trained image processing network to obtain a processing result of the processing target image, and the first processing unit.
[0010] In a fifth aspect, an embodiment of the present disclosure provides a computer device including a memory and a processor, wherein the memory stores a computer program executable on the processor, and when the processor executes the program, part or all of the steps in the above first aspect or second aspect are realized.
[0011] An embodiment of the present disclosure provides a computer-readable storage medium storing a computer program, which when executed by a processor, realizes part or all of the steps in the above first aspect or second aspect.
[0012] Embodiments of the present disclosure provide a computer program comprising computer-readable code, which, when executed on a computer device, causes a processor in the computer device to execute some or all of the steps for implementing the above first aspect or second aspect.
[0013] Embodiments of the present disclosure provide a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and when the computer program is read and executed by a computer, some or all of the steps in the above method are implemented.
[0014] In the embodiments of the present disclosure, by determining a reference pixel in a training image with a labeled ground truth, it is convenient to select an image region that needs to be retained in the subsequent training image. Then, taking the reference pixel as a starting point, the cropping operation on the training image is modeled as a Markov process, and by combining the reference pixels that are the starting points, the cropping probability of whether each pixel point in the training image is retained can be analyzed. Then, based on the output result of predicting the training image by the image processing network and the ground truth of the training image, the network parameter values and the cropping probability of each pixel in the image processing network are optimized to obtain a trained image processing network. In this way, by introducing a Markov chain, the cropping probability of the training image can be estimated, and the cropping probability and network parameter values can be adjusted according to the output result and ground truth of the training image, thereby optimizing the cropping probability and network parameter values, and further improving the performance of the trained image processing network.
[0015] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not technical solutions for limiting the present disclosure.
[0016] To make the above objects, features, and advantages of the embodiments of the present disclosure clearer, preferred embodiments will be specifically cited below and described in detail with reference to the accompanying drawings.
Brief Description of the Drawings
[0017]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8A
Figure 8B
Figure 9
Modes for Carrying Out the Invention
[0018] The drawings herein are incorporated into the specification and form a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and are used, together with the specification, to explain the technical solutions of the present disclosure.
[0019] To make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions of the present disclosure will be further described in detail below with reference to the drawings and embodiments. The described embodiments should not be regarded as limiting the present disclosure. All other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0020] As used in the following description, "some embodiments" describes a subset of all possible embodiments. However, "some embodiments" may be the same subset or a different subset of all possible embodiments, and can be combined with each other if they do not conflict.
[0021] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terms used herein are for the purpose of describing the embodiments of the present disclosure only and are not intended to limit the present disclosure.
[0022] Before further elaborating on the embodiments of the present disclosure, first, the nouns and terms referred to in the embodiments of the present disclosure are explained, and the nouns and terms referred to in the embodiments of the present disclosure shall be applied to the following interpretations.
[0023] 1) Computer vision: Using a camera and a computer instead of human eyes to identify, track and measure targets, and further perform graphic processing, so that the images processed by the computer become more suitable for human eyes to observe or transfer to device detection.
[0024] 2) Model compression: Aimed at reducing the computational amount of the model, reducing the parameter amount / volume of the model, and reducing the inference time of the model.
[0025] 3) Markov chain: A Markov chain is a set of discrete random variables. For example, if a set of random variables is given and all the values of the random variables are within a countable set, the countable set is called a Markov chain, the countable set is called the state space, and the values of the Markov chain within the state space are called states. The Markov property, also called "memorylessness", means that the random variable at step t+1 is conditionally independent of other random variables after the random variable at step t is given. In an embodiment of the present disclosure, the operation of cropping an input image is modeled as a Markov chain, and in this way, the cropping probability of a subsequent pixel depends on the cropping probability of the previous pixel.
[0026] Embodiments of the present disclosure provide a method for training an image processing network, and the method may be executed by a processor of a computer device. Here, the computer device may refer to a device with data processing capabilities such as a server, a laptop, a tablet computer, a desktop computer, a smart TV, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant, a dedicated messaging device, a portable game device), etc. FIG. 1 is a schematic flowchart of the realization of the method for training an image processing network according to an embodiment of the present disclosure. As shown in FIG. 1, the method includes the following steps S101 to S103.
[0027] In step S101, a reference pixel is determined based on a training image labeled with a ground truth.
[0028] In some embodiments, the training image may be a sample image with a ground truth labeled, and the screen scenario of the training image may be selected based on the task of the image processing network to be trained. The training image may be an image with complex image content or an image with simple image content. In some possible implementations, when the task of the image processing network is gaze estimation, the training image may be an image including eyes at each collected viewing angle, where the labeled ground truth is the gaze of the eyes. When the task of the image processing network is eye detection, the training image may be an image including the collected face, where the labeled ground truth is the eyes in the face. When the task of the image processing network is vehicle identification, the training image is an image of the collected traffic scenario, where the labeled ground truth is the vehicle in the image.
[0029] In some embodiments, the reference pixel may be a pixel selected from the training image of any frame, and the reference pixel is used to represent the center position of the image region that needs to be retained during the process of cropping the pixel. In this way, the reference pixel may be the center pixel in the training image, or a pixel at a preset position in the training image. The preset position may be a position defined by the user, or a position within a certain range in the training image, for example, an image region where a circle with a radius of one-fourth of the image width is located in the training image. During the process of network training, for the training image of each frame, one reference pixel is selected as the center pixel of the image region to be retained. In this way, by selecting the reference pixel in the training image, it may be convenient to select the image region that needs to be retained in the subsequent training image.
[0030] In step S102, taking the reference pixel as the starting point, based on the Markov chain of the training image, determine the cropping probability when processing the training image with the image processing network.
[0031] In some embodiments, for the training image of any frame, taking the reference pixel in the training image as the starting point, according to the Markov chain of the training image, the clipping probability for clipping the pixels in the training image is determined. The Markov chain of the training image may be such that the process of clipping the training image is modeled as a Markov chain. The state in the Markov chain corresponds to a pixel being retained, and the transition probability between two adjacent states corresponds to the retention probability of the next pixel when the previous pixel is retained. The retention probability of each state can be calculated by the product of the transition probabilities, and this retention probability is regarded as the importance of the pixel.
[0032] In some possible implementations, during the process of network training, taking the reference pixel in the training image as the starting point, the clipping probability of the next pixel point adjacent to the reference pixel in the training image can be determined according to the clipping probability of the reference pixel in the Markov chain, thereby determining the clipping probability of each pixel point in the training image. In this way, by modeling the clipping operation on the training image as a Markov process, starting from the reference pixel, the clipping probability of whether each pixel point in the training image is retained can be analyzed.
[0033] In step S103, based on the output result of processing the training clipping region by the image processing network and the ground truth, the network parameter value and the clipping probability of the image processing network are adjusted to obtain a trained image processing network.
[0034] In some embodiments, the output result is related to the functions implemented by the image processing network. When the function implemented by the image processing network is gaze estimation, the output result is the prediction result of estimating the gaze in the training image by the image processing network. In this way, the loss of the image processing network can be determined by the output result and the labeled ground truth, and further, the network parameter values and the pruning probability in the image processing network can be adjusted according to the loss, so as to converge the loss output by the trained image processing network after training.
[0035] In some possible implementation forms, the network parameter values of the image processing network include at least weights and may further include pruning target channels. During the network training process, at least the weights and the pruning probability are optimized to obtain a trained image processing network including the adjusted pruning probability and the adjusted network parameter values.
[0036] In some embodiments, the training pruning region is obtained by pruning the training image based on the pruning probability. During the network training process, after determining the pruning probability of pixels by a Markov chain, the pruning probability is used to prune the training image, thereby processing the pruning region using the image processing network and optimizing the image processing network. In one specific example, taking the function of the image processing network being gaze estimation as an example, the obtained output result is, that is, the gaze estimation result, and can be realized by the following steps.
[0037] In the first step, based on the pruning probability of the image processing network, the training pruning region of the training image is determined.
[0038] In some embodiments, in the training image, the training image is cropped using the cropping probability corresponding to the pixel to obtain the training cropping region of the training image. In some possible implementations, the cropping probabilities of a plurality of pixels in the training image are represented as a vector, and the vector is multiplied by the vector representing each pixel in the image. The obtained multiplication result is the training cropping region of the training image.
[0039] In the second step, the gaze estimation is performed on the training cropping region using the image processing network to obtain the output result.
[0040] In some embodiments, the image processing network is a network for performing gaze estimation. The cropped training image, that is, the training cropping region of the training image, is input into the image processing network to perform gaze estimation and obtain the output result. In this way, by applying the cropped training image to the training process, the transition probability between each pixel in the Markov chain can be optimized.
[0041] In the embodiments of the present disclosure, it is convenient to select the image region that needs to be retained in the subsequent training image by determining the reference pixel in the training image with the ground truth labeled. Then, taking the reference pixel as the starting point, the cropping operation on the training image is modeled as a Markov process, and combined with the reference pixel as the starting point, the cropping probability of whether each pixel point in the training image is retained can be analyzed. Finally, based on the output result of predicting the training image by the image processing network and the ground truth of the training image, the network parameter values in the image processing network and the cropping probability of each pixel are optimized to obtain the trained image processing network. In this way, by introducing the Markov chain, the cropping probability of the training image is estimated, and the cropping probability and the network parameter values are adjusted according to the output result and the ground truth of the training image, thereby optimizing the image processing performance of the image processing network.
[0042] In some embodiments, the image processing network may be a neural network for realizing gaze estimation. In this way, the training image is a facial image with labeled gaze information, that is, the ground truth is the labeled gaze information of the eyes in the face. In this way, the output result corresponding to the training image includes the prediction result obtained by performing gaze estimation on the training image by the image processing network. In this way, by using the training image with labeled gaze information to train the image processing network, introducing a Markov chain during the realization process to crop the pixels of the training image, and training the network with the cropped image, the accuracy of the trained image processing network in performing gaze estimation can be improved.
[0043] In some possible implementation forms, in the scenario of gaze estimation, the labeled gaze information and the gaze information in the output result corresponding to the training image include at least one of the pitch angle of the gaze, the yaw angle of the gaze, and the roll angle of the gaze. In this way, by introducing the gaze of each viewing angle into the labeled gaze information, the gaze types in the training image can be enriched, and furthermore, the trained image processing network after training can accurately predict the gaze angles at each viewing angle.
[0044] In some embodiments, by using the central pixel of the training image as the reference pixel, the cropping probability of each pixel in the training image can be determined starting from the central pixel. That is, the above step S101 can be realized by the following step S111 (not shown).
[0045] In step S111, the central pixel in the training image is determined as the reference pixel.
[0046] Here, during the training process of the image processing network, for each batch of training images input into the image processing network, the pixel located at the center position point of the training image, that is, the central pixel, is determined. During the process of image collection, since the region of interest is close to the center of the image, the central pixel is determined in the training image, and it is highly likely that the central pixel is included in the region of interest.
[0047] Here, using the central pixel as the reference pixel, thereby using the central pixel as the starting point, the cropping probability of each pixel is determined. Since the central pixel has a relatively high probability of being included in the region of interest, the cropping region finally cropped with the central pixel as the starting point has a relatively high possibility of including the region of interest.
[0048] In the above step S111, by using the central pixel in the training image as the reference pixel, the possibility that the cropping region finally cropped based on the central pixel includes the region of interest becomes higher.
[0049] After setting the reference pixel according to the above step S111, the above step S102 can be realized by the following step S112 (not shown).
[0050] In step S112, using the central pixel as the starting point, the cropping probability of each pixel in the training image is determined based on the Markov chain.
[0051] Here, in the training image, using the central pixel as the starting point, according to the cropping probability of the central pixel in the Markov chain, the cropping probability of the next pixel of the central pixel can be determined, thereby obtaining the cropping probability of each pixel in the training image.
[0052] In some possible implementation forms, according to the cropping probability of the central pixel in the Markov chain, the transition probability from the central pixel to the next pixel point is determined from the Markov chain, and the transition probabilities of all the pixel points before the next pixel point are multiplied, that is, the cropping probability of the next pixel point can be obtained.
[0053] In an embodiment of the present disclosure, by using the central pixel of the training image as the starting point and combining the clipping probability of the central pixel in the Markov chain, the clipping probabilities of other pixels can be obtained quickly and accurately.
[0054] In some embodiments, during the process of determining the clipping probability of each pixel, the transition probability from the central pixel to other pixels can be obtained by the Markov chain, and the clipping probability of each pixel can be obtained according to the transition probability. That is, step S113 described above can be realized by the following steps.
[0055] In the first step, in the Markov chain, determine the transition probability from the central pixel to the next pixel.
[0056] In some embodiments, since each state in the Markov chain represents that the pixel corresponding to the state is retained, in this way, the transition probability from the central pixel to the next retained pixel can be obtained.
[0057] In the second step, based on the transition probability of the next pixel and the transition probabilities of a plurality of pixels before the next pixel, determine the clipping probability of the next pixel.
[0058] In some embodiments, multiply the transition probability and the transition probabilities corresponding to all the pixels before the next pixel to obtain the clipping probability of the next pixel. In this way, the transition probability from the central pixel to the next retained pixel is obtained by the Markov chain, and the clipping probability of the next pixel is obtained by multiplying each transition probability, thereby clipping the training image according to the clipping probability and optimizing the training image.
[0059] In some embodiments, the clipping probability of a pixel can be determined by the following two methods.
[0060] Method 1: Based on the Markov chain in at least one direction starting from the central pixel, the clipping probability of each pixel in at least one direction starting from the central pixel is set isotropically.
[0061] Here, for at least one direction from the central pixel, the clipping probability is set isotropically based on the Markov chain. In this way, the clipping probabilities of the pixels in a plurality of directions starting from the central pixel are all the same clipping probability. The at least one direction starts from the central pixel and includes directions towards the left, right, up, or down, and can further include directions having a certain included angle with the horizontal or vertical direction. For example, with the central pixel as the origin, the at least one direction may be the positive and negative directions of the X-axis where the origin is located. In this way, in the training image, in at least one direction starting from the central pixel, for the pixels in each direction, the clipping probability of the pixel is obtained according to the transition probability between the pixels in the Markov chain, thereby accurately determining whether it is necessary to retain the pixels in each direction, and the probabilities of retaining the pixels in each direction are the same.
[0062] Method 2: Based on the Markov chain in the symmetric propagation direction starting from the central pixel, the clipping probability of the pixels in the symmetric propagation direction starting from the central pixel is set.
[0063] Here, in the training image, taking the central pixel as the symmetry point, based on the Markov chain of symmetric propagation starting from the symmetry point, the cutting probability of the pixels corresponding to the symmetric propagation is determined. That is, the cutting probability of the pixels is set based on the Markov chain of the left-right symmetric propagation probability from the central pixel. For example, it can be understood that the symmetric propagation directions starting from the central pixel are the positive and negative directions of the X-axis and the positive and negative directions of the Y-axis in the two-dimensional coordinate system with the central pixel as the origin. In this way, taking the central pixel as the symmetry point, the cutting probability for retaining the pixels in one symmetric region is determined, so that finally the cut region after cutting the training image can be positioned in the middle region of the image, and furthermore, the probability that the cut region contains the region of interest is higher.
[0064] In some embodiments, after determining the training cut region, in order to further optimize the cut region, it can be realized by the following two methods.
[0065] Method 1: In the first step, in the training image, a line including the central pixel of the training image is determined.
[0066] Here, after the central pixel is determined in the training image, a line including the pixel point corresponding to the central pixel is determined as any line in the training image passing through the point where the central pixel is located. In this way, the line may be a vertical line, a horizontal line, or a slanted line.
[0067] In the second step, the training cut region is corrected to an axisymmetric region with the line as the axis of symmetry.
[0068] Here, by correcting the training cut region according to this line, the cut region is corrected to a line-symmetric region based on the line including the central pixel. In this way, the training cut region is a line-symmetric region, and the axis of symmetry is this line.
[0069] In the above first step and second step, the training cutout region is corrected to an axisymmetric region of a line whose axis of symmetry in the training image is the line where the central pixel is located. In this way, the pixel points in the central region of the training image are included in the corrected training cutout region, and furthermore, the training cutout region becomes more advantageous for gaze estimation.
[0070] Method 2: Correct the training cutout region to a centrosymmetric region with the central pixel as the center of symmetry.
[0071] Here, by correcting the training cutout region with the central pixel point as the symmetric point, the corrected training cutout region becomes a point-symmetric region where the symmetric point is the midline pixel point. In this way, the corrected training cutout region can more surely include the central region in the training image, and furthermore, the probability that the training cutout region includes the eye region image for gaze estimation becomes higher.
[0072] In some embodiments, the clipping probability and the weights of the network are optimized by values that satisfy an objective function including the difference between the true value and the output result. That is, the above step S103 can be realized by the steps shown in FIG. 2.
[0073] In step S201, based on the output result and the true value, determine the value of the objective function.
[0074] In some embodiments, by comparing the output result in the image processing network with the true value of the training image, the difference between the output result and the true value is obtained, and based on this difference, a loss function for the image processing network to predict the output result is obtained, and this loss function is the objective function.
[0075] In some possible embodiments, the value of the objective function, i.e., the task loss of the image processing network, is obtained by comparing the output result of the image processing network predicting the entire training image with the output result corresponding to the training cut region during the training process. The objective function may be a loss function for training the values of the network parameters of the image processing network itself, such as the weights in the image processing network and the pruning probability of the pruning target channels. The weights of the image processing network and the transition probabilities of the Markov chain are updated alternately.
[0076] In step S202, based on the value of the objective function and the computational cost loss of the image processing network, the network parameter values and the cut probability are adjusted to obtain a trained image processing network.
[0077] In some embodiments, the computational cost loss of the image processing network represents the difference between the expected computational cost of the image processing network and the actual computational cost of the image processing network. For example, the difference between the expected computational cost of the image processing network and the actual computational cost of the image processing network is used as the computational cost loss. In this way, by combining the value of the objective function and the computational cost loss of the image processing network and alternately updating the network parameter values and the cut probability of the image processing network, training of the image processing network is realized, and a trained image processing network is obtained.
[0078] In the embodiments of the present disclosure, the value of the objective function is obtained from the difference between the output result of the training image and the ground truth. By combining the value of the objective function and the computational cost loss of the image processing network and optimizing the weights or cut probability, etc., of the image processing network according to the combined loss, the optimized weights and cut probability in the trained image processing network are more excellent.
[0079] In some embodiments, the value of the objective function is obtained by comparing the output result of the training cutout region with the output result of the entire training image. That is, step S201 described above can be realized by the following steps S211 to S213 (not shown).
[0080] In step S211, the output result and the ground truth are fused to obtain a first fusion result of the training cutout region.
[0081] In some embodiments, during the process of training the image processing network for gaze estimation, the output result corresponding to the training cutout region, that is, the output result of the image processing network performing gaze estimation on the training cutout region, and the ground truth of the training image corresponding to the training cutout region are multiplied element by element to obtain the first fusion result.
[0082] In step S212, the result of the image processing network performing gaze estimation on the training image and the ground truth are fused to obtain a second fusion result of the training image.
[0083] In some embodiments, during the process of training the image processing network for gaze estimation, the output result corresponding to the entire training image, that is, the output result of the image processing network performing gaze estimation on the entire training image, and the ground truth of the training image are multiplied element by element to obtain the second fusion result.
[0084] In step S213, the value of the objective function is obtained based on the ratio of the first fusion result and the second fusion result.
[0085] In some embodiments, the first fusion result and the second fusion result are serialized respectively, and the ratio between the serialized first fusion result and the second fusion result is determined. In this way, for the training image of any frame, the expected value of the training image is determined based on the ratio, and the expected value is determined as the value of the objective function.
[0086] In an embodiment of the present disclosure, by comparing the first fusion result of the training cut-out region with the second fusion result of the training image, the difference between the result of performing gaze estimation on the training cut-out region by the image processing network and the result of performing gaze estimation on the entire training image can be obtained, and this difference is represented as the value of the objective function, whereby the network parameter values such as the weights of the image processing network can be optimized by the value of the objective function.
[0087] In some embodiments, while optimizing the network parameter values of the image processing network, the pruning target channels of the image processing network are optimized. That is, while performing pixel cut-out on the training image input to the image processing network, the channels of the image processing network are cut out to obtain a trained image processing network with better performance. That is, step S202 above can be realized by the following steps S221 and S222 (not shown).
[0088] In step S221, based on the value of the objective function and the computational loss, a transition loss is obtained.
[0089] In some embodiments, during the training process of the image processing network, the value of the objective function and the computational loss are added element by element to obtain the transition loss.
[0090] In step S222, based on the value of the objective function, the pruning probability of the weights and the pruning target channels is adjusted, and based on the transition loss, the cut-out probability is adjusted to obtain the trained image processing network.
[0091] In some embodiments, the weights of the image processing network are optimized according to the value of the objective function, and at the same time, the pruning probability of the channels to be pruned is optimized. In this way, the channels in the image processing network are cut off according to the optimized pruning probability, and the image processing network is compressed, making the image processing network more portable. At the same time as cutting off the channels, the cutting probability of the pixels of the input training image is optimized, and thereby the training image is cut off according to the optimized cutting probability, improving the effectiveness of the training image.
[0092] In the embodiments of the present disclosure, by using the transition loss to optimize the cutting probability, the adjusted cutting probability reaches the optimum. The value of the objective function is used to optimize the weights and the pruning probability of the channels to be pruned, thereby performing channel cutting on the image processing network and at the same time performing pixel cutting on the input training image, and further improving the accuracy of the trained image processing network in performing gaze estimation.
[0093] The embodiments of the present disclosure provide an image processing method applicable to an electronic device. FIG. 3A is a schematic diagram of the system architecture of the road obstacle detection method according to the embodiments of the present disclosure. As shown in FIG. 3A, the system architecture includes an image acquisition device 11, a network 12, and a control terminal 13. To realize the support of one exemplary application, the image acquisition device 11 and the control terminal 13 establish a communication connection via the network 12. First, the image acquisition device 11 reports the image to be processed acquired via the network 12 to the control terminal 13, and the control terminal 13 performs pixel cutting on the image to be processed according to the cutting probability of the trained image processing network to obtain a cutting area waiting to be processed.
[0094] As an example, the image acquisition device 11 may include a vision processing device having vision information processing capabilities. The network 12 may adopt a wired or wireless connection method. Here, when the image acquisition device 11 is a vision processing device, the control terminal 13 can communicate with the vision processing device via a wired connection method, and for example, data communication can be performed via a bus.
[0095] Alternatively, in some scenarios, the image acquisition device 11 may be a vision processing device equipped with a video collection module, or may be a host equipped with a camera. At this time, the augmented reality data display method of the embodiments of the present disclosure may be executed by the image acquisition device 11, and the above system architecture may not include the network 12 and the control terminal 13.
[0096] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will further elaborate on the specific technical solutions of the invention with reference to the drawings in the embodiments of the present disclosure. The following embodiments are configured to illustrate the present disclosure, but are not intended to limit the scope of the present disclosure.
[0097] The "some embodiments" described below describe a subset of all possible embodiments, but the "some embodiments" may be the same subset or different subsets of all possible embodiments, and can be combined with each other as long as there is no conflict, which can be understood.
[0098] FIG. 3B is a schematic flowchart of the implementation of the image processing method according to the embodiments of the present disclosure, and the following description will be made with reference to the steps shown in FIG. 3B.
[0099] In step S301, the image to be processed is acquired.
[0100] Here, the image to be processed may be an image collected in any scenario, may be an image with a simple screen, or may be an image with a complex screen. The image to be processed may be an image obtained in any scenario, may be an image collected by an image acquisition device, such as a camera, or may be an image transmitted by another received device. The image to be processed may be an image for gaze estimation, for example, a facial image including eyes, or a human body image including eyes. For example, in a traffic scenario, the image for gaze estimation may be a facial image or an eye image of a driver driving a vehicle, or may be a facial image or an eye image of a passenger in the vehicle. In an application scenario of a smart device, the image for gaze estimation may be a facial image or an eye image of a person controlling the smart device. By performing gaze estimation on the image of the person, the smart device is controlled based on the direction of the gaze.
[0101] In step S302, based on the cropping probability of the trained image processing network, pixel cropping is performed on the image to be processed to obtain a cropping area waiting for processing.
[0102] Here, the trained image processing network may be obtained by training based on the image processing network training method provided in the above embodiment. The obtained image to be processed is input into the trained image processing network to obtain a cropping area waiting for processing. In the image processing network, by multiplying the vector representing the image to be processed by the optimized cropping probability, pixel cropping for the image to be processed is realized, and the obtained multiplication result is the cropping area waiting for processing.
[0103] In step S303, the trained image processing network is used to process the cropping area waiting for processing to obtain the processing result of the image to be processed.
[0104] Here, after performing pixel extraction on the image to be processed, by using the trained image processing network to process the extraction region waiting to be processed, the accuracy of the obtained processing result can be improved.
[0105] In some possible embodiments, taking the example that the trained image processing network is used to perform gaze estimation on an image, that is, after optimizing the extraction probability of the image processing network, during the gaze estimation process, according to the optimized extraction probability, extraction is performed on the obtained image, thereby obtaining the gaze estimation result of the image. Performing gaze estimation on the extraction region waiting to be processed by using the trained image processing network to obtain the processing result of the image to be processed can be realized as follows.
[0106] Here, in the scenario of gaze estimation, it is necessary to perform image extraction on the image to be processed for which gaze estimation is to be performed by using the optimized extraction probability in the trained image processing network to obtain the extraction region waiting to be processed. In this way, the screen in the extraction region waiting to be processed is, that is, the region of interest, for example, the eye region. In this way, by performing gaze estimation based on the extraction region waiting to be processed, the accuracy of gaze estimation can be further improved.
[0107] Hereinafter, the application of the image processing method provided by the embodiments of the present disclosure in an actual scenario will be described, and the model lightweighting for the gaze estimation model will be described as an example.
[0108] In related technologies, channel pruning is applied to model acceleration and compression, and a parameterized convolutional neural network is deployed to an embedded device or a mobile device. In some tasks, the computational cost can be reduced by cropping a sub-region of each input image. For example, in appearance-based gaze estimation, usually each input image is normalized so that the center of the eye or the face is located at the center of the image. In this case, it can be assumed that the important pixels are distributed near the center of the image, and the sub-region of the input image can be cropped centered to reduce the computational cost. However, in many cases, since it is impossible to determine in which region the important pixels are distributed, an image of the entire face is used as a gaze estimator. However, it is impossible to determine whether it is optimal to use an image of the entire face under a given computational constraint.
[0109] Based on this, embodiments of the present disclosure introduce a new model acceleration concept, i.e., pixel pruning, to find a common optimal region in an input image, and propose a new differentiable pixel pruning method called Differentiable Markov Pixel Pruning (DMPP). The redundant pixels of the input image are searched based on the task loss and the computational constraint. In DMPP, the pixel pruning process is modeled as a Markov chain, and the redundant pixels obtained by searching can be pruned by a simple cropping operation during inference. In gaze estimation, pruning channels and pixels together is more effective than pruning only channels.
[0110] In the training method of the image processing network provided by the embodiments of the present disclosure, a plurality of Markov chains are used to reformulate the pruning process, where the pixel pruning of the plurality of Markov chains can be realized by the following process.
[0111] Represent the input image as TIFF2025520858000002.tif5170, where TIFF2025520858000003.tif4170 represents the number of image channels, half of the height, and half of the width respectively. In an embodiment of the present disclosure, four sets of random variables TIFF2025520858000004.tif5170 are provided, and the state spaces corresponding to these random variables TIFF2025520858000005.tif5170 can be represented by the following equations (1) to (4). TIFF2025520858000006.tif6170TIFF2025520858000007.tif6170TIFF2025520858000008.tif6170TIFF2025520858000009.tif6170
[0112] Here, TIFF2025520858000010.tif7170 represents that the pixel TIFF2025520858000011.tif6170, TIFF2025520858000012.tif7170, TIFF2025520858000013.tif6170 and TIFF2025520858000014.tif6170 are respectively retained. As shown in FIG. 4, pixel pruning is modeled as a plurality of Markov chains, TIFF2025520858000015.tif7170 represents that k pixels obtained by counting in the direction of TIFF2025520858000016.tif4170 from the central pixel are retained.
[0113] In FIG. 4, the Markov chain on the X-axis 401 includes TIFF2025520858000017.tif4170 and TIFF2025520858000018.tif4170 states, and the state of TIFF2025520858000019.tif4170 is including TIFF2025520858000020.tif7170, the state of TIFF2025520858000021.tif4170 is including TIFF2025520858000022.tif7170.
[0114] The Markov chain on the Y-axis 402 is including the states of TIFF2025520858000023.tif4170 and TIFF2025520858000024.tif4170, the state of TIFF2025520858000025.tif4170 is including TIFF2025520858000026.tif7170, the state of TIFF2025520858000027.tif4170 is including TIFF2025520858000028.tif7170.
[0115] In the embodiments of the present disclosure, a rectangular frame is used to label the optimal cropping region, and a set of random variables can represent one cropped rectangle by TIFF2025520858000029.tif5170. For example, TIFF2025520858000030.tif6170, TIFF2025520858000031.tif5170, TIFF2025520858000032.tif6170, when TIFF2025520858000033.tif7170 is satisfied, the cropped rectangular frame can be represented as TIFF2025520858000034.tif5170, where TIFF2025520858000035.tif5170 and TIFF2025520858000036.tif5170 represent the lower left vertex coordinates and the upper right vertex coordinates of the rectangular frame, respectively.
[0116] To learn the optimal region for cropping, the transition probability is parameterized as shown in equations (5) to (8). TIFF2025520858000037.tif11170TIFF2025520858000038.tif12170TIFF2025520858000039.tif11170TIFF2025520858000040.tif12170
[0117] Here,[[]]END]] TIFF2025520858000041.tif4170 represents the sigmoid function, TIFF2025520858000042.tif5170 represents a set of learnable parameters. Therefore, the edge probability of retaining a pixel TIFF2025520858000043.tif6170, TIFF2025520858000044.tif7170, TIFF2025520858000045.tif5170 and TIFF2025520858000046.tif6170 are represented as TIFF2025520858000047.tif5170 respectively as in equations (9) to (12). TIFF2025520858000048.tif10170TIFF2025520858000049.tif10170TIFF2025520858000050.tif10170TIFF2025520858000051.tif10170
[0118] In the embodiments of the present disclosure, by multiplying the edge probability of retaining each pixel for each input image, optimizing the transition probability between states in an end-to-end manner is realized. The generated soft-pruned image is differentiable in terms of the transition probability, and the learnable transition probability can be optimized in an end-to-end manner. When the image I is given, the soft-pruned image TIFF2025520858000052.tif5170 can be expressed as equations (13) and (14). TIFF2025520858000053.tif6170TIFF2025520858000054.tif29170
[0119] Here, TIFF2025520858000055.tif5170 represents the product between elements. As shown in FIG. 5, based on the Markov chains 51 corresponding to the positive and negative X - axes and the Markov chains 52 corresponding to the positive and negative Y - axes, the cropping probability of the input image is determined, and the image is cropped according to the cropping probability to obtain the cropping region 501. That is, the cropping region 501 represents the image obtained after cropping, and the edge probability of retaining each pixel for each input image is multiplied. The generated image after soft pruning (i.e., the cropping region 501) is differentiable in terms of the transition probability. The image obtained after soft - pruning the pixels (i.e., the cropping region in the above - mentioned embodiment) is as shown in FIG. 6. Multiply the original image by the learned edge probability for each element of the retained pixels to obtain the images 601 - 608 in FIG. 6. Here, the target FLOP of images 601 - 604 is 0.5G, and that of images 605 - 608 is 0.5G. Image 601 has the same screen content as image 605 but is the target FLOP, image 602 has the same screen content as image 606 but is the target FLOP, image 603 has the same screen content as image 607 but is the target FLOP, and image 604 has the same screen content as image 608 but is the target FLOP.
[0120] In the embodiments of the present disclosure, budget regularization is performed on the Markov chain, and based on differentiable Markov channel pruning, the set FLOPs are used as the target of budget regularization. For the given image TIFF2025520858000056.tif4170, the expected image size after pixel pruning TIFF2025520858000057.tif5170 is calculated as follows. TIFF2025520858000058.tif42170
[0121] Similarly, the following equation (16) can be obtained. TIFF2025520858000059.tif10170
[0122] Here, the expected image size TIFF2025520858000060.tif5170 is used to calculate the expected FLOP of the entire network. Let FLOPsexp(Φ) be the expected FLOPs of the entire network and FLOPtgt be the target FLOPs. Here, TIFF2025520858000061.tif7170. The budget regularization loss is as shown in equation (17). TIFF2025520858000062.tif10170
[0123] Here, the budget regularization loss may be the computational amount loss in the above embodiments.
[0124] In the embodiments of the present disclosure, during the training process, the weights of the network and the transition probabilities of the Markov chain are alternately updated. The loss function of the weights is the task loss, that is, the gaze angular loss in the embodiments of the present disclosure, and the formula is as follows. TIFF2025520858000063.tif10170TIFF2025520858000064.tif7170
[0125] Here, I represents an image, g represents the true gaze direction, TIFF2025520858000065.tif4170 represents a gaze estimator parameterized by θ, TIFF2025520858000066.tif4170 is TIFF2025520858000067.tif4170 represents a candidate DMPP cropper parameterized by. Here, the gaze angular loss corresponds to the value of the objective function in the above embodiments. During implementation, to update the weights, the gradient is passed through P to the transformation parameter It is not propagated to TIFF2025520858000068.tif4170 because it uses almost unpruned images sampled from the Markov chain only when updating the weights. Transition probability The equation of the loss function of TIFF2025520858000069.tif4170 (corresponding to the transition loss in the above embodiment) is as follows. TIFF2025520858000070.tif7170
[0126] As can be seen from the above, the pixel pruning process can be described as a plurality of Markov chains parameterized by learnable parameters and optimized in an end-to-end manner.
[0127] In the embodiments of the present disclosure, redundant pixels of the input image are searched based on the task loss and computational constraints. In the differentiable Markov chain pixel pruning process, the pixel pruning process is modeled as a plurality of Markov chains, and the searched redundant pixels can be pruned by a simple cropping operation during inference.
[0128] In some embodiments, by performing channel pruning while performing pixel pruning, a better trade-off between the utilization of spatial information and the complexity of the network can be obtained during the realization of gaze estimation. As shown in FIG. 7, FIG. 7 is a schematic diagram of another application scenario of the training method of the image processing network according to the embodiment of the present disclosure. As can be seen from FIG. 7, in the image processing network after channel pruning, in a three-dimensional space, channel pruning is performed on the image processing network based on the Markov chain in the axial direction 701, and pixel pruning is performed on the image input to the image processing network based on the Markov chains in the axial directions 702, 703, 704, and 705 to obtain a pruning region 706. In the embodiment of the present disclosure, by combining the channel pruning of the network and the pixel pruning of the input image to train the image processing network, the accuracy of the image processing network for performing image pruning is higher, and the efficiency of gaze estimation is improved.
[0129] Those skilled in the art can understand that in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order and does not constitute any limitation to the implementation process, and the specific execution order of each step should be determined by its function and possible internal logic.
[0130] Based on the same inventive concept, the embodiments of the present disclosure further provide an apparatus corresponding to the method. Since the principle of solving the problems of the apparatus in the embodiments of the present disclosure is the same as that of the above method in the embodiments of the present disclosure, the implementation of the apparatus can refer to the implementation of the method.
[0131] Based on the foregoing embodiments, embodiments of the present disclosure provide an image processing network training device, which includes each unit included and each module included in each unit, and can be implemented by a processor in a computer device, and of course, can also be implemented by a specific logic circuit. In the implementation process, the processor may be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0132] FIG. 8A is a structural schematic diagram of the configuration of a training device for an image processing network according to an embodiment of the present disclosure. As shown in FIG. 8A, the training device 800 for the image processing network includes a first determination unit 801 configured to determine a reference pixel based on a training image with a true value labeled; a second determination unit 802 configured to determine a cropping probability when processing the training image with the image processing network based on a Markov chain of the training image with the reference pixel as a starting point; a first adjustment unit 803 configured to adjust the network parameter value and the cropping probability of the image processing network based on an output result of processing a training cropping region with the image processing network and the true value to obtain a trained image processing network, where the training cropping region is obtained by cropping the training image based on the cropping probability.
[0133] In some embodiments, the training image includes a face image, and the true value is the labeled gaze information in the face image.
[0134] In some embodiments, the labeled gaze information includes at least one of a pitch angle of the gaze, a yaw angle of the gaze, and a roll angle of the gaze.
[0135] In some embodiments, the first determination unit 801 includes a first determination sub-unit configured to determine a central pixel in the training image as the reference pixel. In some embodiments, the first determination unit 801 includes a first determination sub-unit configured to determine a central pixel in the training image as the reference pixel. Further, the second determination unit 802 is configured to determine the cutting probability of each pixel in the training image based on the Markov chain, starting from the central pixel.
[0136] In some embodiments, the second determination unit 802 includes a second determination sub-unit configured to determine the transition probability from the central pixel to the next pixel in the Markov chain, and a third determination sub-unit configured to determine the cutting probability of the next pixel based on the transition probability of the next pixel and the transition probabilities of a plurality of pixels before the next pixel.
[0137] In some embodiments, the second determination unit 802 further is configured to isotropically set the cutting probability of each pixel in at least one direction starting from the central pixel based on the Markov chain in at least one direction starting from the central pixel.
[0138] In some embodiments, the second determination unit 802 further is configured to set the cutting probability of the pixels in the symmetric propagation direction starting from the central pixel based on the Markov chain in the symmetric propagation direction starting from the central pixel.
[0139] In some embodiments, the apparatus further includes a third determination unit configured to determine a line including the central pixel of the training image in the training image, and a first correction unit configured to, after cutting the training image based on the cutting probability to obtain the training cut region, correct the training cut region to an axisymmetric region with the line as the axis of symmetry.
[0140] In some embodiments, the apparatus further includes a second modification unit configured to modify the training cutout region into a centrosymmetric region centered on the central pixel after obtaining the training cutout region by cutting out the training image based on the cutout probability.
[0141] In some embodiments, the first adjustment unit 803 includes a fourth determination sub-unit configured to determine a value of an objective function based on the output result and the ground truth, and a first adjustment sub-unit configured to adjust the network parameter value and the cutout probability based on the value of the objective function and the computational loss of the image processing network to obtain a trained image processing network.
[0142] In some embodiments, the fifth determination sub-unit includes a first fusion unit configured to fuse the output result and the ground truth to obtain a first fusion result of the training cutout region, a second fusion unit configured to fuse the result of processing the training image by the image processing network and the ground truth to obtain a second fusion result of the training image, and a first comparison unit configured to obtain the value of the objective function based on a ratio value between the first fusion result and the second fusion result.
[0143] In some embodiments, the network parameter value of the image processing network includes weights and pruning probabilities of pruning target channels, and the first adjustment sub-unit includes a first determination unit configured to obtain a transition loss based on the value of the objective function and the computational loss, and a first adjustment unit configured to adjust the weights and the pruning probabilities of the pruning target channels based on the value of the objective function, and adjust the cutout probability based on the transition loss to obtain the trained image processing network.
[0144] Embodiments of the present disclosure provide an image processing apparatus. FIG. 8B is a structural schematic diagram of the configuration of an image processing apparatus according to an embodiment of the present disclosure. As shown in FIG. 8B, the image processing apparatus 820 includes a first acquisition unit 821 configured to acquire an image to be processed, a first clipping unit 822 configured to perform pixel clipping on the image to be processed based on the clipping probability of the trained image processing network to obtain a clipping area waiting to be processed, where the trained image processing network is obtained by training based on the above-described method for training an image processing network, and a first processing unit 823 configured to process the clipping area waiting to be processed using the trained image processing network to obtain a processing result of the image to be processed.
[0145] In some embodiments, the trained image processing network is used to perform gaze estimation on an image, and the first processing module 823 is further configured to perform gaze estimation on the clipping area waiting to be processed using the trained image processing network to obtain a processing result of the image to be processed.
[0146] The description of the embodiments of the above apparatus is similar to the description of the embodiments of the above method and has the same beneficial effects as the embodiments of the method. In some embodiments, the functions or modules included in the apparatus provided by the embodiments of the present disclosure can be used to execute the methods described in the embodiments of the above method. For technical details not disclosed in the embodiments of the apparatus of the present disclosure, reference may be made to the description of the embodiments of the method of the present disclosure for understanding.
[0147] It should be noted that in the embodiments of the present disclosure, when the above-described method for training an image processing network is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present disclosure may essentially or the part contributing to the prior art may be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, a network device, etc.) to execute all or part of the methods described in the embodiments of the present disclosure. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk. In this way, the embodiments of the present disclosure are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.
[0148] The embodiments of the present disclosure provide a computer device including a memory and a processor. The memory stores a computer program executable on the processor. When the processor executes the program, it realizes some or all of the steps in the above method.
[0149] The embodiments of the present disclosure provide a computer-readable storage medium storing a computer program which, when executed by a processor, realizes some or all of the steps in the above method. The computer-readable storage medium may be temporary or non-temporary.
[0150] The embodiments of the present disclosure provide a computer program including computer-readable code. When the computer-readable code is executed on a computer device, the processor in the computer device executes some or all of the steps for realizing the above method.
[0151] Embodiments of the present disclosure provide a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be specifically implemented by hardware, software, or a combination thereof. In some embodiments, the computer program product is embodied as a computer storage medium. In some other embodiments, the computer program product is embodied as a software product such as a software development kit (SDK).
[0152] Here, it should be pointed out that the above descriptions of the embodiments tend to emphasize the differences between the embodiments, and the same points or similar points can be referred to each other. The descriptions of the embodiments of the above device, storage medium, computer program, and computer program product are similar to the descriptions of the embodiments of the above method and have the same beneficial effects as the embodiments of the method. For the technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of the present disclosure, reference can be made to the descriptions of the embodiments of the method of the present disclosure for understanding.
[0153] It should be noted that FIG. 9 is a schematic diagram of the hardware entity of a computer device according to an embodiment of the present disclosure. As shown in FIG. 9, the hardware entity of the computer device 900 includes a processor 901, a communication interface 902, and a memory 903.
[0154] The processor 901 generally controls the overall operation of the computer device 900.
[0155] The communication interface 902 enables the computer device to communicate with other terminals or servers via a network.
[0156] The memory 903 is configured to store instructions and applications executable by the processor 901, and can also cache data (e.g., image data, audio data, voice communication data, and video communication data) processed or to be processed by the processor 901 and each module in the computer device 900, and can be realized by a flash memory (FLASH) or a random access memory (RAM: Random Access Memory). Data transmission can be performed between the processor 901, the communication interface 902, and the memory 903 via the bus 904.
[0157] It should be understood that the reference throughout the specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic associated with the embodiment is included in at least one embodiment of the present disclosure. Thus, the appearances of "in one embodiment" or "in an embodiment" in various places throughout the specification are not necessarily referring to the same embodiment. Also, these particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In various embodiments of the present disclosure, the magnitude of the numbers of the above steps / processes does not mean the order of execution before and after, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure. The numbers of the embodiments of the present disclosure above are for illustrative purposes only and do not represent the superiority or inferiority of the embodiments.
[0158] In this specification, it should be explained that the terms "including", "comprising" or any variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element limited by the phrase "including one..." does not exclude the presence of another same element in the process, method, article, or apparatus that includes the element.
[0159] In some embodiments provided by the present disclosure, it should be understood that the disclosed devices and methods may be implemented in other ways. The embodiments of the devices described above are only illustrative. For example, the division of the units is only a logical functional division, and there may be other division methods when actually implemented. For example, a plurality of units or components may be combined, or integrated into other systems, or some features may be ignored or not executed. Also, the mutual connection, direct connection, or communication connection between each component shown or discussed may be an indirect connection or communication connection through some interfaces, devices or units, and may be in electrical, mechanical or other forms.
[0160] The unit described as a separation member may or may not be physically separated. The member shown as a unit may or may not be a physical unit, that is, it may be located in one place, or distributed in a plurality of network units. According to actual needs, some or all of the units can be selected to achieve the purpose of the solution of this embodiment.
[0161] Also, each functional unit in each embodiment of the present disclosure may all be integrated into one processing unit, each individual unit may be a unit alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware, or in the form of a functional unit combining hardware and software.
[0162] Those skilled in the art can understand that all or part of the steps for implementing the embodiments of the above method may be completed by a program instructing relevant hardware. The aforementioned program may be stored in a computer-readable storage medium. When the program is executed, the steps of the embodiments of the above method are executed. The aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROMs), magnetic disks, or optical disks.
[0163] Alternatively, in the present disclosure, when the above integrated unit is implemented in the form of a software function module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present disclosure may essentially or the part that contributes to the prior art may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present disclosure. The aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0164] What is described above is only the embodiments of the present disclosure. However, the protection scope of the present disclosure is not limited thereto. Any changes or replacements that can be easily conceived by those skilled in the art within the technical scope disclosed in the present disclosure should all be included within the protection scope of the present disclosure.
Industrial Applicability
[0165] Embodiments of the present disclosure provide a training method, an image processing method, and an apparatus for an image processing network. The method includes determining a reference pixel based on a training image labeled with a ground truth, determining a cropping probability when processing the training image with the image processing network based on a Markov chain of the training image with the reference pixel as a starting point, and adjusting a network parameter value and the cropping probability of the image processing network based on an output result of processing a training cropping region with the image processing network and the ground truth to obtain a trained image processing network, where the training cropping region is obtained by cropping the training image based on the cropping probability.
Claims
1. A method for training an image processing network, executed by an electronic device, comprising: determining a reference pixel based on a training image labeled with a ground truth; using the reference pixel as a starting point, determining a cropping probability when processing the training image with the image processing network based on a Markov chain of the training image; adjusting a network parameter value and the cropping probability of the image processing network based on an output result of processing a training cropping region with the image processing network and the ground truth to obtain a trained image processing network, wherein the training cropping region is obtained by cropping the training image based on the cropping probability.
2. The training image includes a face image, and the ground truth is labeled gaze information in the face image, characterized in that the method for training an image processing network according to Claim 1.
3. The labeled gaze information includes at least one of a pitch angle of the gaze, a yaw angle of the gaze, and a roll angle of the gaze, characterized in that the method for training an image processing network according to Claim 2.
4. Determining a reference pixel based on a training image labeled with a ground truth includes determining a central pixel in the training image as the reference pixel, and using the reference pixel as a starting point, determining a cropping probability when processing the training image with the image processing network based on a Markov chain of the training image includes determining a cropping probability of each pixel in the training image based on the Markov chain with the central pixel as a starting point, characterized in that the method for training an image processing network according to Claim 2 or 3.
5. Determining a cropping probability of each pixel in the training image based on the Markov chain with the central pixel as a starting point includes determining a transition probability from the central pixel to the next pixel in the Markov chain, and determining a cropping probability of the next pixel based on the transition probability of the next pixel and transition probabilities of a plurality of pixels before the next pixel, characterized in that the method for training an image processing network according to Claim 4.
6. Determining the cutting probability of each pixel in the training image based on the Markov chain with the central pixel as the starting point includes isotropically setting the cutting probability of each pixel in at least one direction starting from the central pixel based on the Markov chain in at least one direction starting from the central pixel, characterized in that The method for training an image processing network according to claim 4 or 5.
7. Determining the cutting probability of each pixel in the training image based on the Markov chain with the central pixel as the starting point includes setting the cutting probability of the pixel in the symmetric propagation direction starting from the central pixel based on the Markov chain in the symmetric propagation direction starting from the central pixel, characterized in that The method for training an image processing network according to claim 4 or 5.
8. After obtaining the training cut region by cutting the training image based on the cutting probability, the method for training the image processing network further includes determining a line including the central pixel of the training image in the training image; and modifying the training cut region into an axisymmetric region with the line as the axis of symmetry, characterized in that The method for training an image processing network according to any one of claims 1 to 7.
9. After obtaining the training cut region by cutting the training image based on the cutting probability, the method for training the image processing network further includes modifying the training cut region into a centrosymmetric region with the central pixel as the center of symmetry, characterized in that The method for training an image processing network according to any one of claims 1 to 7.
10. Based on the output result of processing the training cut region by the image processing network and the ground truth, adjusting the network parameter value and the cutting probability of the image processing network to obtain a trained image processing network includes determining the value of the objective function based on the output result and the ground truth; and adjusting the network parameter value and the cutting probability based on the value of the objective function and the computational loss of the image processing network to obtain the trained image processing network, characterized in that The method for training an image processing network according to any one of claims 1 to 9.
11. Determining the value of the objective function based on the output result and the true value includes: Fusing the output result and the true value to obtain a first fusion result of the training cut region; Fusing the result of processing the training image by the image processing network and the true value to obtain a second fusion result of the training image; Obtaining the value of the objective function based on the ratio of the first fusion result and the second fusion result. The training method of the image processing network according to claim 10 is characterized by the above. The training method of the image processing network according to claim 10.
12. The network parameter values of the image processing network include weights and pruning probabilities of pruning target channels. Adjusting the network parameter values and the cut probability based on the value of the objective function and the computational loss of the image processing network to obtain a trained image processing network includes: Obtaining a transition loss based on the value of the objective function and the computational loss; Adjusting the weights and the pruning probabilities of the pruning target channels based on the value of the objective function, and adjusting the cut probability based on the transition loss to obtain a trained image processing network. The training method of the image processing network according to claim 10 or 11 is characterized by the above. The training method of the image processing network according to claim 10 or 11.
13. An image processing method, comprising: Obtaining a processing target image; Performing pixel cutting on the processing target image based on the cut probability of the trained image processing network to obtain a cut region waiting for processing. The trained image processing network is obtained by training based on the training method of the image processing network according to any one of claims 1 to 12; Processing the cut region waiting for processing using the trained image processing network to obtain a processing result of the processing target image. The image processing method is characterized by the above.
14. The trained image processing network is used for performing gaze estimation on an image. Processing the cut region waiting for processing using the trained image processing network to obtain a processing result of the processing target image includes: Performing gaze estimation on the cut region waiting for processing using the trained image processing network to obtain a processing result of the processing target image. The image processing method according to claim 13 is characterized by the above. The image processing method according to claim 13.
15. An apparatus for training an image processing network, A first determination unit configured to determine a reference pixel based on a training image labeled with a true value; A second determination unit configured to determine a cutting probability when processing the training image with an image processing network based on a Markov chain of the training image, with the reference pixel as a starting point; A first adjustment unit configured to adjust a network parameter value and the cutting probability of the image processing network based on an output result of processing a training cut region with the image processing network and the true value, to obtain a trained image processing network, wherein the training cut region is obtained by cutting the training image based on the cutting probability. The training apparatus for an image processing network includes the first adjustment unit.
16. An image processing apparatus, comprising: A first acquisition unit configured to acquire a processing target image; A first cutting unit configured to perform pixel cutting on the processing target image based on a cutting probability of a trained image processing network to obtain a cutting region waiting for processing, wherein the trained image processing network is obtained by training based on the training method of the image processing network according to any one of Claims 1 to 12. A first processing unit configured to process the cutting region waiting for processing using the trained image processing network to obtain a processing result of the processing target image. The image processing apparatus includes the first processing unit.
17. A computer device including a memory and a processor, wherein the memory stores a computer program executable on the processor, and when the processor executes the program, the computer device realizes the training method of the image processing network according to any one of Claims 1 to 12, or when the processor executes the program, the computer device realizes the image processing method according to Claim 13 or 14.
18. A computer-readable storage medium storing a computer program that realizes the training method of the image processing network according to any one of Claims 1 to 12 when executed by a processor, or realizes the image processing method according to Claim 13 or 14 when executed by a processor.
19. A computer program or a computer program product including a computer program or instructions for causing an electronic device to execute, when executed on the electronic device, the training method of the image processing network according to any one of claims 1 to 12, or causing the electronic device to execute the image processing method according to claim 13 or 14.