An image processing method, apparatus, electronic device, and storage medium

By obtaining multiple scale information and regions of the image in the convolutional neural network and adjusting the convolution kernel weight based on this information, the problem of insufficient scale robustness of the convolutional neural network is solved, and the image processing performance is improved.

CN113919476BActive Publication Date: 2025-06-10ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010650418.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-08
Publication Date
2025-06-10
Estimated Expiration
2040-07-08

AI Technical Summary

Technical Problem

The scale robustness of convolutional neural networks in image processing is poor, resulting in poor image processing performance in tasks that are sensitive to image scale information.

Method used

By obtaining multiple scale information and image areas of the image to be processed, the convolution kernel weight coefficients of different image areas are obtained according to the specified scale information, and convolution processing is performed to improve the scale robustness of the convolution neural network.

Benefits of technology

The scale robustness of convolutional neural networks in image processing is improved, thereby improving the performance of image processing, especially in tasks such as crowd counting and object detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113919476B_ABST
    Figure CN113919476B_ABST
Patent Text Reader

Abstract

The present application discloses an image processing method, including: obtaining multiple scale information corresponding to an image to be processed, and obtaining multiple image regions in the image to be processed, where different image regions in the multiple image regions respectively correspond to specified scale information in the multiple scale information; respectively obtaining convolution kernel weight coefficients corresponding to different image regions according to the specified scale information, where the convolution kernel weight coefficients corresponding to different image regions are the weight coefficients of the convolution kernels in a target convolutional neural network for different image regions; respectively performing convolution processing on different image regions according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernels in the target convolutional neural network to obtain convolution results for different image regions. The image processing method provided by the present application improves the scale robustness of the convolutional neural network in the process of image processing and the image processing performance of the convolutional neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly relates to an image processing method. This application also relates to an image processing apparatus, an electronic device, and a storage medium. Background Art

[0002] A convolutional neural network is a feedforward neural network that contains convolutional calculations and has a deep structure, and has now been widely used in fields such as image processing and natural language processing. The general idea when applying a convolutional neural network to image processing is as follows: Use a specified convolutional neural network that has been trained and has a specific purpose to perform image feature extraction on the image to be processed, obtain useful image feature information in the image to be processed, and then obtain the result of image processing based on this useful image feature information and output it.

[0003] When applying a convolutional neural network to image processing, since the convolutional neural network performs unified feature extraction on the image to be processed, the scale robustness of the convolutional neural network during the image processing process will be poor. In this way, when applying the convolutional neural network to image processing tasks that are sensitive to image scale information, such as: crowd counting, object detection and other image processing tasks, there will be a problem of poor image processing performance. Summary of the Invention

[0004] This application provides an image processing method, apparatus, electronic device, and storage medium to improve the scale robustness of a convolutional neural network during the image processing process, thereby improving the performance of image processing.

[0005] This application provides an image processing method, including:

[0006] Obtain multiple scale information corresponding to the image to be processed, and obtain multiple image regions in the image to be processed, where different image regions in the multiple image regions respectively correspond to specified scale information in the multiple scale information;

[0007] Respectively obtain the convolutional kernel weight coefficients corresponding to different image regions according to the specified scale information, where the convolutional kernel weight coefficients corresponding to different image regions are the weight coefficients of the convolutional kernels in the target convolutional neural network for different image regions;

[0008] Respectively perform convolutional processing on different image regions according to the convolutional kernel weight coefficients corresponding to different image regions and the convolutional kernels in the target convolutional neural network to obtain convolutional results for different image regions.

[0009] Optionally, it further includes: obtaining a processing result for the image to be processed according to the convolutional results for different image regions.

[0010] Optionally, obtaining the convolution kernel weight coefficients corresponding to different image regions according to the specified scale information includes:

[0011] Obtaining a plurality of Gaussian filters with the same convolution kernel size as that of the Gaussian convolution kernel, where each Gaussian filter in the plurality of Gaussian filters corresponds to a different specified standard deviation;

[0012] Determining the convolution kernel weight coefficients corresponding to different image regions according to the scale information corresponding to different image regions and the plurality of Gaussian filters.

[0013] Optionally, determining the convolution kernel weight coefficients corresponding to different image regions according to the scale information corresponding to different image regions and the plurality of Gaussian filters includes:

[0014] Determining the weight values of the plurality of Gaussian filters corresponding to different image regions according to the scale information corresponding to different image regions and the plurality of Gaussian filters;

[0015] Obtaining the convolution kernel weight coefficients corresponding to different image regions according to the weight values of the plurality of Gaussian filters corresponding to different image regions and the plurality of Gaussian filters.

[0016] Optionally, determining the weight values of the plurality of Gaussian filters corresponding to different image regions according to the scale information corresponding to different image regions and the plurality of Gaussian filters includes:

[0017] Determining the degree of correlation between each Gaussian filter in the plurality of Gaussian filters and different image regions according to the scale information corresponding to different image regions and the plurality of Gaussian filters;

[0018] Determining the weight value of each Gaussian filter corresponding to different image regions according to the degree of correlation between each Gaussian filter and different image regions.

[0019] Optionally, obtaining the convolution kernel weight coefficients corresponding to different image regions according to the weight values of the plurality of Gaussian filters corresponding to different image regions and the plurality of Gaussian filters includes: performing weighted averaging on the weight values of the plurality of Gaussian filters corresponding to different image regions and the plurality of Gaussian filters according to the weight value of each Gaussian filter in the plurality of Gaussian filters corresponding to different image regions and each Gaussian filter, to obtain the convolution kernel weight coefficients corresponding to different image regions.

[0020] Optionally, determining the degree of correlation between each of the plurality of Gaussian filters and the different image regions according to the scale information corresponding to the different image regions and the plurality of Gaussian filters includes:

[0021] Obtaining a first specified parameter for calculating the degree of correlation between each Gaussian filter and the different image regions, and obtaining a second specified parameter for calculating the degree of correlation between each Gaussian filter and the different image regions;

[0022] Obtaining the number of the plurality of Gaussian filters, and obtaining the serial number of each Gaussian filter in the sorting of the plurality of Gaussian filters, where the sorting of the plurality of Gaussian filters is a sorting of the plurality of Gaussian filters in ascending order of the specified standard deviation corresponding to the plurality of Gaussian filters;

[0023] Determining the degree of correlation between each Gaussian filter and the different image regions according to the first specified parameter, the second specified parameter, the number of the plurality of Gaussian filters, the scale information corresponding to the different image regions, and the serial number of each Gaussian filter in the sorting of the plurality of Gaussian filters.

[0024] Optionally, the scale information corresponding to the different image regions includes: the visibility value information corresponding to the different image regions;

[0025] Determining the degree of correlation between each Gaussian filter and the different image regions according to the first specified parameter, the second specified parameter, the number of the plurality of Gaussian filters, the scale information corresponding to the different image regions, and the serial number of each Gaussian filter in the sorting of the plurality of Gaussian filters includes:

[0026] Obtaining a target parameter for calculating the degree of correlation between each Gaussian filter and the different image regions according to the first specified parameter, the number of the plurality of Gaussian filters, the visibility value information corresponding to the different image regions, and the serial number of each Gaussian filter in the sorting of the plurality of Gaussian filters;

[0027] Calculating the degree of correlation between each Gaussian filter and the different image regions according to the target parameter and the second specified parameter.

[0028] Optionally, determining the weight value corresponding to each Gaussian filter in the different image regions according to the degree of correlation between each Gaussian filter and the different image regions includes:

[0029] Obtain the exponential power of the correlation degree of each Gaussian filter with the different image regions according to the correlation degree of each Gaussian filter with the different image regions;

[0030] Obtain the sum of the exponential powers of the correlation degrees of the multiple Gaussian filters with the different image regions according to the exponential power of the correlation degree of each Gaussian filter with the different image regions;

[0031] Obtain the weight value corresponding to each Gaussian filter in the different image regions according to the ratio of the exponential power of the correlation degree of each Gaussian filter with the different image regions to the sum of the exponential powers of the correlation degrees of the multiple Gaussian filters with the different image regions.

[0032] Optionally, it further includes:

[0033] Obtain the convolution kernel size of the convolution kernel;

[0034] Determine the convolution kernel size of the Gaussian convolution kernel according to the convolution kernel size of the convolution kernel.

[0035] Optionally, the performing convolution processing on different image regions respectively according to the convolution kernel weight coefficients corresponding to the different image regions and the convolution kernel in the target convolutional neural network to obtain the convolution results for the different image regions includes:

[0036] Obtain the simulated weight coefficients of the elements in the convolution kernel corresponding to the different image regions respectively according to the convolution kernel weight coefficients corresponding to the different image regions and the convolution kernel in the target convolutional neural network;

[0037] Perform convolution processing on the different image regions according to the simulated weight coefficients of the elements in the convolution kernel corresponding to the different image regions and the convolution kernel in the target convolutional neural network to obtain the convolution results for the different image regions.

[0038] Optionally, the obtaining the simulated weight coefficients of the elements in the convolution kernel corresponding to the different image regions respectively according to the convolution kernel weight coefficients corresponding to the different image regions and the convolution kernel in the target convolutional neural network includes: multiplying the convolution kernel weight coefficients corresponding to the different image regions by the weight coefficients of the elements in the convolution kernel, and the simulated weight coefficients of the elements in the convolution kernel corresponding to the different image regions.

[0039] Optionally, the obtaining the multiple scale information corresponding to the image to be processed and obtaining the multiple image regions in the image to be processed includes:

[0040] Obtain the image to be processed;

[0041] Extract scale information from the image to be processed to obtain the multiple scale information;

[0042] According to the multiple scale information, obtain the specified scale information, and obtain the image region corresponding to the specified scale information in the image to be processed;

[0043] According to the image region corresponding to the specified scale information in the image to be processed, obtain the multiple image regions.

[0044] On the other hand, this application also provides an image processing device, including:

[0045] An image region obtaining unit, configured to obtain multiple scale information corresponding to an image to be processed, and obtain multiple image regions in the image to be processed, where different image regions among the multiple image regions respectively correspond to the specified scale information among the multiple scale information;

[0046] A weight coefficient obtaining unit, configured to respectively obtain convolution kernel weight coefficients corresponding to different image regions according to the specified scale information, where the convolution kernel weight coefficients corresponding to different image regions are the weight coefficients of the different image regions for the convolution kernel in a target convolutional neural network;

[0047] A convolution result obtaining unit, configured to respectively perform convolution processing on different image regions according to the convolution kernel weight coefficients corresponding to the different image regions and the convolution kernel in the target convolutional neural network, to obtain convolution results for the different image regions.

[0048] On the other hand, this application also provides an electronic device, including:

[0049] A processor; and

[0050] A memory, configured to store a program for an image processing method. After the device is powered on and runs the program of the image processing method through the processor, the following steps are executed:

[0051] Obtain multiple scale information corresponding to an image to be processed, and obtain multiple image regions in the image to be processed, where different image regions among the multiple image regions respectively correspond to the specified scale information among the multiple scale information;

[0052] Respectively obtain convolution kernel weight coefficients corresponding to different image regions according to the specified scale information, where the convolution kernel weight coefficients corresponding to different image regions are the weight coefficients of the different image regions for the convolution kernel in a target convolutional neural network;

[0053] Convolve different image regions respectively according to the convolution kernel weight coefficients corresponding to the different image regions and the convolution kernels in the target convolutional neural network to obtain convolution results for the different image regions.

[0054] On the other hand, this application also provides a storage medium storing a program for an image processing method. When the program is run by a processor, the following steps are executed:

[0055] Obtain multiple scale information corresponding to the image to be processed, and obtain multiple image regions in the image to be processed. Different image regions in the multiple image regions respectively correspond to specified scale information in the multiple scale information;

[0056] Respectively obtain convolution kernel weight coefficients corresponding to different image regions according to the specified scale information. The convolution kernel weight coefficients corresponding to the different image regions are the weight coefficients of the different image regions for the convolution kernels in the target convolutional neural network;

[0057] Convolve different image regions respectively according to the convolution kernel weight coefficients corresponding to the different image regions and the convolution kernels in the target convolutional neural network to obtain convolution results for the different image regions.

[0058] On the other hand, this application also provides an object counting method, including:

[0059] Obtain the image to be processed containing the target object;

[0060] Obtain multiple depth value information corresponding to the image to be processed, and obtain multiple image regions in the image to be processed. Different image regions in the multiple image regions respectively correspond to specified depth value information in the multiple depth value information;

[0061] Respectively obtain convolution kernel weight coefficients corresponding to different image regions according to the specified depth value information. The convolution kernel weight coefficients corresponding to the specified image regions are the weight coefficients of the different image regions for the convolution kernels in the target convolutional neural network;

[0062] Convolve different image regions respectively according to the convolution kernel weight coefficients corresponding to the different image regions and the convolution kernels in the target convolutional neural network to obtain convolution results for the different image regions;

[0063] Obtain the number of the target objects according to the convolution results for the different image regions.

[0064] On the other hand, this application also provides an object detection method, including:

[0065] Obtain the image to be processed containing the target object;

[0066] Obtain multiple bounding box information corresponding to the image to be processed, and obtain multiple image regions in the image to be processed, where different image regions in the multiple image regions respectively correspond to different bounding box information in the multiple bounding box information;

[0067] Respectively obtain the convolution kernel weight coefficients corresponding to different image regions according to the different bounding box information, where the convolution kernel weight coefficients corresponding to different image regions are the weight coefficients of the different image regions for the convolution kernels in the target convolutional neural network;

[0068] Respectively perform convolution processing on different image regions according to the convolution kernel weight coefficients corresponding to the different image regions and the convolution kernels in the target convolutional neural network, and obtain convolution results for the different image regions;

[0069] Obtain the target object according to the convolution results for the different image regions.

[0070] On the other hand, this application also provides an image processing method, including:

[0071] Obtain at least one scale information corresponding to the image to be processed, and obtain at least one image region in the image to be processed, where different image regions in the at least one image region respectively correspond to the specified scale information in the at least one scale information;

[0072] Respectively obtain the convolution kernel weight coefficients corresponding to different image regions according to the specified scale information, where the convolution kernel weight coefficients corresponding to different image regions are the weight coefficients of the different image regions for the convolution kernels in the target convolutional neural network;

[0073] Respectively perform convolution processing on different image regions according to the convolution kernel weight coefficients corresponding to the different image regions and the convolution kernels in the target convolutional neural network, and obtain convolution results for the different image regions.

[0074] On the other hand, this application also provides an image processing system, including: a platform server and a user terminal;

[0075] The platform server is configured to obtain the image to be processed provided by the user terminal; obtain multiple scale information corresponding to the image to be processed, and obtain multiple image regions in the image to be processed, where different image regions in the multiple image regions respectively correspond to specified scale information in the multiple scale information; respectively obtain convolution kernel weight coefficients corresponding to different image regions according to the specified scale information, where the convolution kernel weight coefficients corresponding to different image regions are the weight coefficients of different image regions for the convolution kernels in the target convolutional neural network; respectively perform convolution processing on different image regions according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernels in the target convolutional neural network to obtain convolution results for different image regions; obtain a processing result for the image to be processed according to the convolution results for different image regions; and provide the processing result to the user terminal.

[0076] The user terminal is configured to send the image to be processed to the platform server; and obtain the processing result sent by the platform server.

[0077] Compared with the prior art, the present application has the following advantages:

[0078] In the image processing method provided by the present application, after obtaining multiple scale information corresponding to the image to be processed and obtaining multiple image regions in the image to be processed, first, respectively obtain convolution kernel weight coefficients corresponding to different image regions according to the specified scale information, where the convolution kernel weight coefficients corresponding to different image regions are the weight coefficients of different image regions for the convolution kernels in the target convolutional neural network; then, respectively perform convolution processing on different image regions according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernels in the target convolutional neural network to obtain convolution results for different image regions. The image processing method provided by the present application can ensure that different degrees of convolution processing are performed on image regions with different scale information, thereby improving the scale robustness of the convolutional neural network in the image processing process, and further improving the image processing performance of the convolutional neural network. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1A FIG. 1 is a first schematic diagram of an application scenario of the image processing method provided in an embodiment of the present application.

[0080] Figure 1B FIG. 2 is a second schematic diagram of an application scenario of the image processing method provided in an embodiment of the present application.

[0081] Figure 2Flowchart of an image processing method provided in the first embodiment of the present application.

[0082] Figure 3 Flowchart of a correlation degree acquisition method provided in the first embodiment of the present application.

[0083] Figure 4 Schematic diagram of an image processing device provided in the second embodiment of the present application.

[0084] Figure 5 Schematic diagram of an electronic device provided in the embodiments of the present application.

[0085] Figure 6 Flowchart of an object counting method provided in the fifth embodiment of the present application.

[0086] Figure 7 Flowchart of an object counting method provided in the sixth embodiment of the present application.

[0087] Figure 8 Flowchart of an image processing method provided in the seventh embodiment of the present application.

[0088] Figure 9 Schematic diagram of an image processing system provided in the eighth embodiment of the present application. Detailed implementation manners

[0089] Many specific details are set forth in the following description in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the spirit of the present application. Therefore, the present application is not limited by the specific implementations disclosed below.

[0090] In order to more clearly show the image processing method provided by the present application, first, the application scenario of the image processing method provided by the present application is introduced. In the embodiments of the present application, the image processing method is used to perform convolution processing on an image through a target convolutional neural network to obtain the target object or the number of target objects in the image. The so-called image to be processed is an image including at least one target object. The so-called target object is an object selected in the image in advance according to requirements, generally a human object, an animal object, an object object, etc. The so-called target convolutional neural network is a convolutional neural network used to extract image feature information from the image to be processed to obtain useful image feature information in the image to be processed. In the application scenario of the image processing method provided in the embodiments of the present application, the so-called scale information is information obtained by extracting scale information from the image to be processed through a preset neural network, specifically including depth value information corresponding to the image, bounding box information corresponding to the image, etc.

[0091] The execution subject of the image processing method provided in the embodiments of the present application is a computing device installed with a program or software for the image processing method provided in the embodiments of the present application.

[0092] In the embodiments of the present application, taking the application of the image processing method provided in the present application to crowd counting as an example, the image processing method provided in the present application is described in detail. At this time, the target object is a human object, the image to be processed is an image including at least one person, and the scale information corresponding to the image to be processed is multiple depth value information corresponding to the image to be processed. When the image processing method provided in the present application is applied to crowd counting, the steps of the image processing method provided in the present application are as Figure 1A shown, which is the first schematic diagram of the application scenario of the image processing method provided in the embodiments of the present application.

[0093] The image acquisition module 101A acquires the image to be processed including at least one human object. The image scale information acquisition module 102A extracts the image scale information of the image to be processed to obtain multiple scale information corresponding to the image to be processed. The specific implementation method of extracting the image scale information of the image to be processed is generally: extracting the image depth value information of the image to be processed to obtain multiple depth value information corresponding to the image to be processed. The image region division module 103A is used to obtain multiple image regions in the image to be processed according to the multiple depth value information corresponding to the image to be processed, and different image regions in the multiple image regions respectively correspond to the specified depth value information in the multiple depth value information. The weight coefficient determination module 104A is used to obtain the convolution kernel weight coefficients corresponding to different image regions respectively according to the indefinite depth value information, and the convolution kernel weight coefficients corresponding to different image regions are the weight coefficients of the convolution kernels in the target convolutional neural network for different image regions. The weight coefficient determination module 105A is used to perform convolution processing on each different image region respectively according to the convolution kernel weight coefficient corresponding to each different image region and the convolution kernel in the target convolutional neural network to obtain the convolution result for each different image region.

[0094] For the specific implementation process when the image processing method provided in the present application is applied to crowd counting, please refer to Figure 1B which is the second schematic diagram of the application scenario of the image processing method provided in the embodiments of the present application.

[0095] First, the image to be processed is obtained. First, an initial image including at least one human object is obtained. Second, the initial image is preprocessed to obtain the image to be processed. The specific implementation method of preprocessing the initial image is: performing image quality enhancement processing such as image denoising on the initial image.

[0096] Second, obtain multiple scale information in the image to be processed. Specifically, obtain multiple scale information corresponding to the image to be processed, and obtain multiple image regions in the image to be processed. That is, first obtain the depth value image corresponding to the image to be processed, and based on the depth value image, determine multiple depth value information in the image to be processed. Then, obtain different image regions corresponding to the specified depth value information in the multiple depth value information in the image to be processed, and based on the different image regions corresponding to the specified depth value information in the multiple depth value information in the image to be processed, obtain multiple image regions in the image to be processed. Among them, the depth value image can be an image obtained in advance through monocular image depth estimation based on deep learning, or an image obtained after the image to be processed is obtained and then the image depth value is extracted from the image to be processed based on monocular image depth estimation of deep learning. Since the multiple image regions are image regions obtained according to the different image regions corresponding to the specified depth value information in the image to be processed, different image regions in the multiple image regions respectively correspond to the specified depth value information in the multiple depth value information.

[0097] In the embodiments of the present application, different image regions in the multiple image regions respectively correspond to the specified scale information in the multiple scale information. The specific implementation method is generally: each image region in the multiple image regions respectively corresponds to each scale information in the multiple scale information. It should be noted that since the visibility value information in the visibility map corresponding to the image is used to identify the number of pixels in which an object of one meter in reality is imaged in the camera, in the same image, the closer to the camera, the greater the visibility value. In addition, the depth value information in the depth value image corresponding to the image is used to identify the distance of the object in the image from the camera in reality. The farther from the camera, the greater the depth value. In summary, the multiple depth value information of the image to be processed can be replaced by the multiple visibility value information corresponding to the image to be processed. That is, the specific implementation method of obtaining multiple scale information in the image to be processed is: multiple scale information in the image to be processed.

[0098] Third, obtain the convolution kernel weight coefficients corresponding to different image regions. That is, respectively obtain the convolution kernel weight coefficients corresponding to different image regions according to the specified visibility value information. The so-called convolution kernel weight coefficients corresponding to different image regions are the weight coefficients of the convolution kernel in the target convolutional neural network for different image regions. The so-called target convolutional neural network is a convolutional neural network used to extract the density feature information of the crowd image in the image when applying the image processing method provided by the present application to crowd counting.

[0099] In the application scenario of the image processing method provided in the embodiments of the present application, the specific implementation method of obtaining the convolution kernel weight coefficients corresponding to different image regions is as follows:

[0100] First, obtain multiple Gaussian filters whose convolution kernel sizes of the Gaussian convolution kernels are the same as those of the convolution kernels, and each Gaussian filter among the multiple Gaussian filters corresponds to a different specified standard deviation. The so-called convolution kernel size refers to the height size and width size corresponding to the convolution kernel. Common convolution kernel sizes corresponding to convolution kernels are: 3*3, 5*5, 7*7, etc. In the application scenario of the image processing method provided in the embodiments of the present application, specifically taking the convolution kernel size corresponding to the convolution kernel as 7*7 as an example, the application scenario of the image processing method provided in the embodiments of the present application will be described. The so-called Gaussian filter is used to, when performing image processing on an image, perform weighted averaging on the pixel value of the target pixel point in the image and the pixel values of other pixel points in the target area corresponding to the target pixel point, and use it as the pixel value of the pixel point in the processed image. The so-called standard deviation of the Gaussian filter is used to allocate the weights of different pixel points when performing weighted averaging on the pixel value of the target pixel point in the image and the pixel values of other pixel points in the target area corresponding to the target pixel point. The smaller the standard deviation, the smaller the weight of other pixel points away from the target pixel point.

[0101] In the application scenario of the image processing method provided in the embodiments of the present application, before obtaining multiple Gaussian filters whose convolution kernel sizes of the Gaussian convolution kernels are the same as those of the convolution kernels, it is necessary to first obtain the convolution kernel size of the convolution kernel, and then determine the convolution kernel size of the Gaussian convolution kernel according to the convolution kernel size of the convolution kernel. Since 95.449974% of the area is within two specified standard deviations σ on both sides of the average. That is, the random variable is in P{μ - 2σ < x < μ + 2σ} = 95.4%, so 4σ = K. Where K is the convolution kernel size of the Gaussian convolution kernel. When the Gaussian kernel size of the Gaussian filter is determined, if the Gaussian filter with a Gaussian convolution kernel size of 7*7 is equivalent to a convolution kernel with a convolution kernel size of 7*7, it is necessary to set the specified standard deviation σ to 1.25. Since 4σ = K is 4σ = 7 at this time, so σ = 1.25. If the Gaussian filter with a Gaussian convolution kernel size of 7*7 is equivalent to a convolution kernel with a convolution kernel size of 1*1, it is necessary to set the specified standard deviation σ to 0.25. Since 4σ = K is 4σ = 1 at this time, so σ = 0.25. In the application scenario of the image processing method provided in the embodiments of the present application, 30 Gaussian filters with the convolution kernel size of the Gaussian convolution kernel being 7*7 and the specified standard deviation σ being 0.25, 0.30, 0.35... 1.70, 1.75 in sequence are set as multiple Gaussian filters whose convolution kernel sizes of the Gaussian convolution kernels are the same as those of the convolution kernels. Specifically, the multiple Gaussian filters are respectively denoted as: G 1 , G 2 …G 30 .

[0102] Second, according to the weight values corresponding to different image regions of multiple Gaussian filters and the multiple Gaussian filters, obtain the convolutional kernel weight coefficients corresponding to different image regions. That is, first, according to the visibility value information p corresponding to different image regions i and the multiple Gaussian filters, determine the weight values corresponding to different image regions of the multiple Gaussian filters. Then, according to the weight values corresponding to different image regions of the multiple Gaussian filters and the multiple Gaussian filters, obtain the convolutional kernel weight coefficients corresponding to different image regions. Wherein, i is the serial number of different image regions among the multiple image regions.

[0103] The so-called specific implementation manner of determining the weight values corresponding to different image regions of the multiple Gaussian filters according to the visibility value information p corresponding to different image regions i and the multiple Gaussian filters is as follows: First, according to the visibility value information p corresponding to different image regions i and the multiple Gaussian filters, determine the correlation degree u j between each Gaussian filter G ij among the multiple Gaussian filters and different image regions. Then, according to the correlation degree u j between each Gaussian filter G ij and different image regions, determine the weight value U j corresponding to different image regions of each Gaussian filter G i . Wherein, j is the serial number in the sorting of different Gaussian filters for the multiple Gaussian filters, and u ij is the correlation degree between the j-th Gaussian filter and the i-th image region. The so-called serial number in the sorting of different Gaussian filters for the multiple Gaussian filters is the sorting of the multiple Gaussian filters in ascending order of the specified standard deviation corresponding to the multiple Gaussian filters.

[0104] In the application scenario of the image processing method provided in the embodiment of the present application, the step of determining the correlation degree u j between each Gaussian filter G ij among the multiple Gaussian filters and different image regions is as follows: First, obtain the first specified parameter θ for calculating the correlation degree u j between each Gaussian filter G ij and different image regions, and obtain the first specified parameter θ for calculating the correlation degree u j between each Gaussian filter G ijThe second specified parameter η. The so-called first specified parameter θ and second specified parameter η are parameters obtained after training with the specified parameters. Then, the number L of multiple Gaussian filters is obtained, and the serial number j of each Gaussian filter in the sorting for the multiple Gaussian filters is obtained. In the application scenario of the image processing method provided in the embodiments of the present application, L = 30. Finally, according to the first specified parameter θ, the second specified parameter η, the number L of multiple Gaussian filters, the visibility value information p corresponding to different image regions i and the serial number j of each Gaussian filter in the sorting for the multiple Gaussian filters, determine the correlation degree u between each Gaussian filter and different image regions ij .

[0105] The so-called according to the first specified parameter θ, the second specified parameter η, the number L of multiple Gaussian filters, the visibility value information p corresponding to different image regions i and the serial number j of each Gaussian filter in the sorting for the multiple Gaussian filters, determine the correlation degree u between each Gaussian filter and different image regions ij , the specific implementation method is: according to the first specified parameter, the number of multiple Gaussian filters, the visibility value information p corresponding to different image regions i and the serial number j of each Gaussian filter in the sorting for the multiple Gaussian filters, obtain the target parameter m for calculating the correlation degree between each Gaussian filter and different image regions ij = θ || p i * L - j || 2 ; according to the target parameter m ij = θ || p i * L - j || 2 and the second specified parameter η, calculate the correlation degree u between each Gaussian filter and different image regions ij ,

[0106] It should be noted that when the scale information is visibility value information, the formula for the target parameter for calculating the correlation degree between each Gaussian filter and different image regions is only: m ij = θ || p i * L - j || 2 , and further the formula for calculating the correlation degree u between each Gaussian filter and different image regions can be specifically ij ,

[0107] When the scale information is scale information other than visibility value information, although the step of determining the degree of correlation between each Gaussian filter and the different image regions is still: "According to the first specified parameter, the second specified parameter, the number of the multiple Gaussian filters, the scale information corresponding to the different image regions, and the multiple Gaussian filters, determine the degree of correlation between each Gaussian filter and the different image regions", however, the above formula: m ij = θ||p i *L-j|| 2 And

[0108] In the application scenario of the image processing method provided in the embodiments of the present application, the so-called specific implementation manner of determining the weight value U i corresponding to each Gaussian filter G ij in different image regions according to the degree of correlation u i is as follows: First, according to the degree of correlation u i between each Gaussian filter G i and different image regions, obtain the exponential power exp(u ij ) of the degree of correlation between each Gaussian filter and different image regions. Then, according to the exponential power of the degree of correlation between each Gaussian filter and different image regions, obtain the sum ∑ ij of the exponential powers of the degrees of correlation between the multiple Gaussian filters and different image regions. Finally, according to the ratio of the exponential power exp(u j ) of the degree of correlation between each Gaussian filter and different image regions to the sum ∑ ij of the exponential powers of the degrees of correlation between the multiple Gaussian filters and different image regions, obtain the weight value corresponding to each Gaussian filter in different image regions. That is, U ij = exp(u j ) / ∑ ij exp(u i ), where U ij is the weight value corresponding to each Gaussian filter in the i-th image region. In the application scenario of the image processing method provided in the embodiments of the present application, the so-called specific implementation manner of obtaining the convolution kernel weight coefficient corresponding to different image regions according to the weight values corresponding to the multiple Gaussian filters in different image regions and the multiple Gaussian filters is: According to each Gaussian filter G j in the multiple Gaussian filters, the weight value U ij corresponding to different image regions and each Gaussian filter G i in each Gaussian filter G i corresponding to different image regions and each Gaussian filter G i and each Gaussian filter G j, the weighted average of the weight values corresponding to multiple Gaussian filters in different image regions and the multiple Gaussian filters is performed to obtain the convolution kernel weight coefficients corresponding to different image regions. That is, So-called, U ij is the weight value corresponding to the j-th Gaussian filter in the i-th image region. So-called is the convolution kernel weight coefficient corresponding to the i-th image region.

[0109] Finally, the convolution results for different image regions are obtained. That is, according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernels in the target convolutional neural network respectively, convolution processing is performed on different image regions to obtain the convolution results for different image regions. The specific implementation method is as follows: First, according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernels in the target convolutional neural network respectively, the simulated weight coefficients corresponding to the elements in the convolution kernel for different image regions are obtained. Second, according to the simulated weight coefficients corresponding to the elements in the convolution kernel for different image regions and the convolution kernels in the target convolutional neural network, convolution processing is performed on different image regions to obtain the convolution results for different image regions. So-called according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernels in the target convolutional neural network respectively, the simulated weight coefficients corresponding to the elements in the convolution kernel for different image regions are obtained, including: multiplying the convolution kernel weight coefficients corresponding to different image regions by the weights corresponding to the elements in the convolution kernel, and the simulated weight coefficients corresponding to the elements in the convolution kernel for different image regions.

[0110] So-called according to the simulated weight coefficients corresponding to the elements in the convolution kernel for different image regions and the convolution kernels in the target convolutional neural network, convolution processing is performed on different image regions to obtain the convolution results for different image regions: According to the simulated weight coefficients corresponding to the elements in the convolution kernel for different image regions and the convolution kernels in the target convolutional neural network, convolution processing is performed on different image regions to obtain the crowd density feature information for different image regions.

[0111] After obtaining the crowd density feature information for different image regions, in the application scenario of the image processing method provided in the embodiments of the present application, it is also necessary to further obtain the processing result for the image to be processed according to the convolution results for different image regions. That is, according to the crowd density feature information for different image regions, the number of human objects in the image to be processed is obtained.

[0112] In the embodiments of the present application, the application scenarios of the image processing method provided in the embodiments of the present application are not specifically limited. For example, the image processing method provided in the present application can be used in scenarios such as target object detection, etc., which will not be elaborated here one by one. The application scenarios of the above-mentioned image processing method provided in the embodiments of the present application are for facilitating the understanding of the image processing method provided in the present application, rather than for limiting the image processing method provided in the present application.

[0113] First Embodiment

[0114] The first embodiment provides an image processing method, which will be described below in conjunction with Figure 2 and Figure 3 for illustration.

[0115] Figure 2 is a flowchart of an image processing method provided in the first embodiment of the present application. Figure 2 The image processing method shown includes: step S201 to step S203.

[0116] In step S201, a plurality of scale information corresponding to the image to be processed is obtained, and a plurality of image regions in the image to be processed are obtained, and different image regions in the plurality of image regions respectively correspond to the specified scale information in the plurality of scale information.

[0117] In the first embodiment of the present application, the image to be processed refers to an image including at least one target object. The target object is an object pre-selected in the image according to requirements, generally a human object, an animal object, an object object, etc. The so-called scale information is the information obtained by extracting the scale information of the image to be processed through a preset neural network, specifically including the depth value information corresponding to the image to be processed and the bounding box information corresponding to the image, etc.

[0118] In the first embodiment of the present application, the specific implementation manner of obtaining a plurality of scale information corresponding to the image to be processed and obtaining a plurality of image regions in the image to be processed is as follows: First, obtain the image to be processed. Second, perform scale information extraction on the image to be processed to obtain the plurality of scale information. Third, according to the plurality of scale information, obtain the specified scale information, and obtain the image region corresponding to the specified scale information in the image to be processed. Fourth, according to the image region corresponding to the specified scale information in the image to be processed, obtain the plurality of image regions.

[0119] In the first embodiment of the present application, specifically taking the scale information as the bounding box information as an example, at this time, the specific operations of obtaining multiple scale information corresponding to the image to be processed and obtaining multiple image regions in the image to be processed are as follows: First, obtain the bounding box image corresponding to the image to be processed, and determine multiple bounding box information in the image to be processed according to the bounding box image. Then, obtain different image regions corresponding to different bounding box information in the multiple bounding box information in the image to be processed, and obtain multiple image regions in the image to be processed according to the different image regions corresponding to different bounding box information in the multiple bounding box information. Among them, the bounding box image can be an image obtained in advance through a bounding box extraction network, or an image obtained after performing image bounding box extraction on the image to be processed based on the bounding box extraction network after obtaining the image to be processed. Since the multiple image regions are image regions obtained according to different image regions corresponding to different bounding box information in the image to be processed, different image regions in the multiple image regions respectively correspond to the specified scale information in the multiple scale information.

[0120] In the first embodiment of the present application, different image regions in the multiple image regions respectively correspond to the specified scale information in the multiple scale information. The specific implementation method is generally: each image region in the multiple image regions respectively corresponds to each scale information in the multiple scale information.

[0121] In step S202, respectively obtain the convolution kernel weight coefficients corresponding to different image regions according to the specified scale information. The convolution kernel weight coefficients corresponding to different image regions are the weight coefficients of the convolution kernel in the target convolutional neural network for different image regions.

[0122] The so-called target convolutional neural network is a convolutional neural network used to extract image features from the image to be processed to obtain useful image feature information in the image to be processed.

[0123] In the first embodiment of the present application, the specific implementation method of respectively obtaining the convolution kernel weight coefficients corresponding to different image regions according to the specified scale information is as follows: First, obtain multiple Gaussian filters with the same convolution kernel size as the convolution kernel of the Gaussian convolution kernel, and each Gaussian filter in the multiple Gaussian filters respectively corresponds to a different specified standard deviation. Then, determine the convolution kernel weight coefficients corresponding to different image regions according to the scale information corresponding to different image regions and the multiple Gaussian filters.

[0124] The so-called convolution kernel size refers to the height size and width size corresponding to the convolution kernel. Common convolution kernel sizes corresponding to convolution kernels include: 3*3, 5*5, 7*7, etc. The so-called Gaussian filter is used to, when performing image processing on an image, perform weighted averaging on the pixel value of the target pixel point in the image and the pixel values of other pixel points in the target area corresponding to the target pixel point, and use it as the pixel value of the pixel point in the processed image. The so-called standard deviation of the Gaussian filter is used to allocate the weights of different pixel points when performing weighted averaging on the pixel value of the target pixel point in the image and the pixel values of other pixel points in the target area corresponding to the target pixel point. The smaller the standard deviation, the smaller the weights of other pixel points away from the target pixel point.

[0125] In the first embodiment of the present application, before obtaining multiple Gaussian filters with the same convolution kernel size as the convolution kernel of the Gaussian convolution kernel, it is necessary to first obtain the convolution kernel size of the convolution kernel, and then determine the convolution kernel size of the Gaussian convolution kernel according to the convolution kernel size of the convolution kernel.

[0126] In the first embodiment of the present application, the process of determining the convolution kernel weight coefficients corresponding to different image regions according to the scale information corresponding to different image regions and multiple Gaussian filters is as follows: First, determine the weight values of multiple Gaussian filters corresponding to different image regions according to the scale information corresponding to different image regions and multiple Gaussian filters. Second, obtain the convolution kernel weight coefficients corresponding to different image regions according to the weight values of multiple Gaussian filters corresponding to different image regions and multiple Gaussian filters.

[0127] The so-called implementation method of determining the weight values of multiple Gaussian filters corresponding to different image regions according to the scale information corresponding to different image regions and multiple Gaussian filters is as follows: First, determine the degree of correlation between each Gaussian filter in multiple Gaussian filters and different image regions according to the scale information corresponding to different image regions and multiple Gaussian filters. Then, determine the weight values of each Gaussian filter corresponding to different image regions according to the degree of correlation between each Gaussian filter and different image regions.

[0128] It should be noted that in the first embodiment of the present application, for the steps of determining the degree of correlation between each Gaussian filter in multiple Gaussian filters and different image regions, please refer to Figure 3 , which is a flowchart of a method for obtaining the degree of correlation provided in the first embodiment of the present application.

[0129] Step S301: Obtain the first specified parameter for calculating the degree of correlation between each Gaussian filter and different image regions, and obtain the second specified parameter for calculating the degree of correlation between each Gaussian filter and different image regions.

[0130] The so-called first specified parameter and second specified parameter are parameters obtained after training with specified parameters.

[0131] Step S302: Obtain the number of multiple Gaussian filters, and obtain the serial number of each Gaussian filter in the sorting of the multiple Gaussian filters. The sorting of the multiple Gaussian filters is a sorting of the multiple Gaussian filters in ascending order of the specified standard deviations corresponding to the multiple Gaussian filters.

[0132] Step S303: Determine the degree of correlation between each Gaussian filter and different image regions according to the first specified parameter, the second specified parameter, the number of multiple Gaussian filters, the scale information corresponding to different image regions, and the serial number of each Gaussian filter in the sorting of the multiple Gaussian filters.

[0133] Since the visibility value information in the visibility map corresponding to the image is used to identify the number of pixels in which an object of one meter in reality is imaged in the camera, in the same image, the closer to the camera, the greater the visibility value. In addition, the depth value information in the depth value image corresponding to the image is used to identify the distance of the object in the image from the camera in reality. The farther from the camera, the greater the depth value. In summary, the multiple depth value information of the image to be processed can be replaced by the multiple visibility value information corresponding to the image to be processed.

[0134] When the multiple scale information corresponding to the image to be processed is the multiple visibility value information corresponding to the image to be processed, determining the degree of correlation between each Gaussian filter and different image regions according to the first specified parameter, the second specified parameter, the number of multiple Gaussian filters, the scale information corresponding to different image regions, and the serial number of each Gaussian filter in the sorting of the multiple Gaussian filters includes: First, obtain the target parameter for calculating the degree of correlation between each Gaussian filter and different image regions according to the first specified parameter, the number of multiple Gaussian filters, the visibility value information corresponding to different image regions, and the serial number of each Gaussian filter in the sorting of the multiple Gaussian filters. Then, calculate the degree of correlation between each Gaussian filter and different image regions according to the target parameter and the second specified parameter.

[0135] In the first embodiment of the present application, the implementation manner of determining the weight value corresponding to each Gaussian filter in different image regions according to the correlation degree between each Gaussian filter and different image regions is as follows: First, according to the correlation degree between each Gaussian filter and different image regions, obtain the exponential power of the correlation degree between each Gaussian filter and different image regions. Then, according to the exponential power of the correlation degree between each Gaussian filter and different image regions, obtain the sum of the exponential powers of the correlation degrees between multiple Gaussian filters and different image regions. Finally, according to the ratio of the exponential power of the correlation degree between each Gaussian filter and different image regions to the sum of the exponential powers of the correlation degrees between multiple Gaussian filters and different image regions, obtain the weight value corresponding to each Gaussian filter in different image regions.

[0136] It should be noted that in the first embodiment of the present application, obtaining the convolution kernel weight coefficients corresponding to different image regions according to the weight values corresponding to different image regions of multiple Gaussian filters and multiple Gaussian filters includes: performing weighted averaging on the weight values corresponding to different image regions of multiple Gaussian filters and multiple Gaussian filters according to the weight value corresponding to each Gaussian filter in different image regions of multiple Gaussian filters and each Gaussian filter to obtain the convolution kernel weight coefficients corresponding to different image regions.

[0137] Please refer to Figure 2 , in step S203, perform convolution processing on different image regions respectively according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernels in the target convolutional neural network to obtain the convolution results for different image regions.

[0138] In the first embodiment of the present application, the specific implementation manner of obtaining the convolution results for different image regions is as follows: First, respectively according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernels in the target convolutional neural network, obtain the simulated weight coefficients of the elements in the convolution kernel corresponding to different image regions. Then, perform convolution processing on different image regions according to the simulated weight coefficients of the elements in the convolution kernel corresponding to different image regions and the convolution kernels in the target convolutional neural network to obtain the convolution results for different image regions. The so-called obtaining the simulated weight coefficients of the elements in the convolution kernel corresponding to different image regions respectively according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernels in the target convolutional neural network includes: multiplying the convolution kernel weight coefficients corresponding to different image regions by the weight corresponding to the elements in the convolution kernel to obtain the simulated weight coefficients of the elements in the convolution kernel corresponding to different image regions.

[0139] In the first embodiment of the present application, after obtaining the convolution results for different image regions, it further includes: obtaining the processing result for the image to be processed according to the convolution results for different image regions. Specifically, when performing convolution processing on different image regions according to the analog weight coefficients corresponding to different image regions in the convolution kernel and the convolution kernel in the target convolutional neural network to obtain the convolution results for different image regions, that is, obtaining the population density feature information for different image regions, the so-called obtaining the processing result for the image to be processed according to the convolution results for different image regions is: obtaining the number of person objects in the image to be processed according to the population density feature information for different image regions.

[0140] In the image processing method provided in the first embodiment of the present application, after obtaining the multiple scale information corresponding to the image to be processed and obtaining multiple image regions in the image to be processed; first, respectively obtain the convolution kernel weight coefficients corresponding to different image regions according to the specified scale information, and the convolution kernel weight coefficients corresponding to different image regions are the weight coefficients of different image regions for the convolution kernel in the target convolutional neural network; then, respectively according to the convolution kernel weight coefficients corresponding to different image regions, which are the weight coefficients of different image regions for the convolution kernel in the target convolutional neural network. In the image processing method provided in the first embodiment of the present application, first, respectively obtain the convolution kernel weight coefficients corresponding to different image regions according to the specified scale information, and then further perform convolution processing on different image regions according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernel in the target convolutional neural network to obtain the convolution results for different image regions, which can ensure different degrees of convolution processing for image regions with different scale information, thereby improving the scale robustness of the convolutional neural network in the image processing process, and further improving the image processing performance of the convolutional neural network.

[0141] Second Embodiment

[0142] Corresponding to the application scenario of the image processing method provided in the present application and the image processing method provided in the first embodiment, the second embodiment of the present application further provides an image processing device. Since the device embodiment is basically similar to the application scenario and the first embodiment, the description is relatively simple, and for the relevant parts, refer to the partial descriptions of the application scenario and the first embodiment. The device embodiment described below is only illustrative.

[0143] Please refer to Figure 4 , which is a schematic diagram of an image processing device provided in the second embodiment of the present application.

[0144] The image processing device provided in the second embodiment of the present application includes:

[0145] An image region obtaining unit 401, configured to obtain multiple scale information corresponding to an image to be processed, and obtain multiple image regions in the image to be processed, where different image regions in the multiple image regions respectively correspond to specified scale information in the multiple scale information;

[0146] A weight coefficient obtaining unit 402, configured to respectively obtain convolution kernel weight coefficients corresponding to different image regions according to the specified scale information, where the convolution kernel weight coefficients corresponding to different image regions are weight coefficients of the convolution kernel in a target convolutional neural network for different image regions;

[0147] A convolution result obtaining unit 403, configured to respectively perform convolution processing on different image regions according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernel in the target convolutional neural network, and obtain convolution results for different image regions.

[0148] Optionally, the image processing apparatus provided in the second embodiment of the present application further includes: an image processing result obtaining unit, configured to obtain a processing result for the image to be processed according to the convolution results for different image regions.

[0149] Optionally, the weight coefficient obtaining unit 402 is specifically configured to obtain multiple Gaussian filters whose convolution kernel sizes are the same as the convolution kernel size of the convolution kernel, where each Gaussian filter in the multiple Gaussian filters respectively corresponds to a different specified standard deviation; and determine the convolution kernel weight coefficients corresponding to different image regions according to the scale information corresponding to different image regions and the multiple Gaussian filters.

[0150] Optionally, the determining the convolution kernel weight coefficients corresponding to different image regions according to the scale information corresponding to different image regions and the multiple Gaussian filters includes:

[0151] Determine weight values of the multiple Gaussian filters corresponding to different image regions according to the scale information corresponding to different image regions and the multiple Gaussian filters;

[0152] Obtain the convolution kernel weight coefficients corresponding to different image regions according to the weight values of the multiple Gaussian filters corresponding to different image regions and the multiple Gaussian filters.

[0153] Optionally, the determining the weight values of the multiple Gaussian filters corresponding to different image regions according to the scale information corresponding to different image regions and the multiple Gaussian filters includes:

[0154] Determine the degree of correlation between each Gaussian filter in the multiple Gaussian filters and the different image regions according to the scale information corresponding to the different image regions and the multiple Gaussian filters;

[0155] Determine the weight value corresponding to each Gaussian filter in the different image regions according to the degree of correlation between each Gaussian filter and the different image regions.

[0156] Optionally, the obtaining the convolution kernel weight coefficients corresponding to the different image regions according to the weight values corresponding to the multiple Gaussian filters in the different image regions and the multiple Gaussian filters includes: performing weighted averaging on the weight values corresponding to the multiple Gaussian filters in the different image regions and the multiple Gaussian filters according to the weight value corresponding to each Gaussian filter in the multiple Gaussian filters in the different image regions and each Gaussian filter, to obtain the convolution kernel weight coefficients corresponding to the different image regions.

[0157] Optionally, the determining the degree of correlation between each Gaussian filter in the multiple Gaussian filters and the different image regions according to the scale information corresponding to the different image regions and the multiple Gaussian filters includes:

[0158] Obtain a first specified parameter for calculating the degree of correlation between each Gaussian filter and the different image regions, and obtain a second specified parameter for calculating the degree of correlation between each Gaussian filter and the different image regions;

[0159] Obtain the number of the multiple Gaussian filters, and obtain the sequence number of each Gaussian filter in the sorting of the multiple Gaussian filters, where the sorting of the multiple Gaussian filters is a sorting of the multiple Gaussian filters in ascending order of the specified standard deviation corresponding to the multiple Gaussian filters;

[0160] Determine the degree of correlation between each Gaussian filter and the different image regions according to the first specified parameter, the second specified parameter, the number of the multiple Gaussian filters, the scale information corresponding to the different image regions, and the sequence number of each Gaussian filter in the sorting of the multiple Gaussian filters.

[0161] Optionally, the scale information corresponding to the different image regions includes: the visibility value information corresponding to the different image regions;

[0162] Determining the degree of correlation between each Gaussian filter and the different image regions according to the first specified parameter, the second specified parameter, the number of the multiple Gaussian filters, the scale information corresponding to the different image regions, and the sequence number of each Gaussian filter in the sorting of the multiple Gaussian filters includes:

[0163] Obtaining a target parameter for calculating the degree of correlation between each Gaussian filter and the different image regions according to the first specified parameter, the number of the multiple Gaussian filters, the visibility value information corresponding to the different image regions, and the sequence number of each Gaussian filter in the sorting of the multiple Gaussian filters;

[0164] Calculating the degree of correlation between each Gaussian filter and the different image regions according to the target parameter and the second specified parameter.

[0165] Optionally, determining the weight value corresponding to each Gaussian filter in the different image regions according to the degree of correlation between each Gaussian filter and the different image regions includes:

[0166] Obtaining the exponential power of the degree of correlation between each Gaussian filter and the different image regions according to the degree of correlation between each Gaussian filter and the different image regions;

[0167] Obtaining the sum of the exponential powers of the degrees of correlation between the multiple Gaussian filters and the different image regions according to the exponential power of the degree of correlation between each Gaussian filter and the different image regions;

[0168] Obtaining the weight value corresponding to each Gaussian filter in the different image regions according to the ratio of the exponential power of the degree of correlation between each Gaussian filter and the different image regions to the sum of the exponential powers of the degrees of correlation between the multiple Gaussian filters and the different image regions.

[0169] Optionally, it further includes:

[0170] Obtaining the convolution kernel size of the convolution kernel;

[0171] Determining the convolution kernel size of the Gaussian convolution kernel according to the convolution kernel size of the convolution kernel.

[0172] Optionally, the convolution result obtaining unit 403 is specifically configured to obtain, according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernels in the target convolutional neural network respectively, the simulated weight coefficients corresponding to different image regions for the elements in the convolution kernels; perform convolution processing on the different image regions according to the simulated weight coefficients corresponding to different image regions for the elements in the convolution kernels and the convolution kernels in the target convolutional neural network, so as to obtain the convolution results for the different image regions.

[0173] Optionally, the step of obtaining, according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernels in the target convolutional neural network respectively, the simulated weight coefficients corresponding to different image regions for the elements in the convolution kernels includes: multiplying the convolution kernel weight coefficients corresponding to different image regions by the weights corresponding to the elements in the convolution kernels, and the simulated weight coefficients corresponding to different image regions for the elements in the convolution kernels.

[0174] Optionally, the image region obtaining unit 401 is specifically configured to obtain the image to be processed; extract scale information from the image to be processed to obtain the multiple scale information; obtain the specified scale information according to the multiple scale information, and obtain the image region corresponding to the specified scale information in the image to be processed; obtain the multiple image regions according to the image region corresponding to the specified scale information in the image to be processed.

[0175] Third Embodiment

[0176] Corresponding to the application scenario of the image processing method provided in this application and the image processing method provided in the first embodiment, the third embodiment of this application also provides an electronic device. Since the third embodiment is basically similar to the application scenario and the first embodiment, the description is relatively simple, and for the relevant parts, refer to the partial descriptions of the application scenario and the first embodiment. The following description of the third embodiment is merely illustrative.

[0177] Please refer to Figure 5 , which is a schematic diagram of an electronic device provided in an embodiment of this application.

[0178] The electronic device includes: a processor 501;

[0179] And a memory 502 for storing a program of the image processing method. After the device is powered on and runs the program of the image processing method through the processor, the following steps are executed:

[0180] Obtain multiple scale information corresponding to the image to be processed, and obtain multiple image regions in the image to be processed, where different image regions in the multiple image regions respectively correspond to the specified scale information in the multiple scale information;

[0181] Obtain the convolution kernel weight coefficients corresponding to different image regions respectively according to the specified scale information, where the convolution kernel weight coefficients corresponding to the different image regions are the weight coefficients of the different image regions for the convolution kernels in the target convolutional neural network;

[0182] Perform convolution processing on different image regions respectively according to the convolution kernel weight coefficients corresponding to the different image regions and the convolution kernels in the target convolutional neural network, and obtain the convolution results for the different image regions.

[0183] It should be noted that for the detailed description of the electronic device provided in the third embodiment of the present application, reference can be made to the application scenarios of the image processing method provided in the present application and the related descriptions of the image processing method provided in the first embodiment, which will not be elaborated here.

[0184] The Fourth Embodiment

[0185] Corresponding to the application scenario of the image processing method provided in the present application and the image processing method provided in the first embodiment, the fourth embodiment of the present application further provides a storage medium. Since the fourth embodiment is basically similar to the application scenario and the first embodiment, the description is relatively simple, and the relevant parts can refer to the partial descriptions of the application scenario and the first embodiment. The following description of the fourth embodiment is only illustrative.

[0186] The storage medium stores a computer program, and when the computer program is run by a processor, the following steps are executed:

[0187] Obtain multiple scale information corresponding to the image to be processed, and obtain multiple image regions in the image to be processed, where different image regions in the multiple image regions respectively correspond to the specified scale information in the multiple scale information;

[0188] Obtain the convolution kernel weight coefficients corresponding to different image regions respectively according to the specified scale information, where the convolution kernel weight coefficients corresponding to the different image regions are the weight coefficients of the different image regions for the convolution kernels in the target convolutional neural network;

[0189] Perform convolution processing on different image regions respectively according to the convolution kernel weight coefficients corresponding to the different image regions and the convolution kernels in the target convolutional neural network, and obtain the convolution results for the different image regions.

[0190] It should be noted that for the detailed description of the storage medium provided in the fourth embodiment of the present application, reference can be made to the application scenarios of the image processing method provided in the present application and the related descriptions of the image processing method provided in the first embodiment, which will not be elaborated here.

[0191] The Fifth Embodiment

[0192] Corresponding to the application scenario of the image processing method provided in this application and the image processing method provided in the first embodiment, the fifth embodiment of this application also provides an object counting method. Since the fifth embodiment is basically similar to the application scenario and the first embodiment, the description is relatively simple. For related parts, please refer to the partial descriptions of the application scenario and the first embodiment. The following description of the fifth embodiment is merely illustrative.

[0193] Please refer to Figure 6 , which is a flowchart of an object counting method provided in the fifth embodiment of this application.

[0194] Step S601: Obtain a to-be-processed image including a target object.

[0195] In the fifth embodiment of this application, the so-called to-be-processed image is an image including at least one target object. The so-called target object is an object pre-selected in the image according to requirements, generally a human object, an animal object, an object object, etc.

[0196] Step S602: Obtain a plurality of depth value information corresponding to the to-be-processed image, and obtain a plurality of image regions in the to-be-processed image. Different image regions among the plurality of image regions respectively correspond to specified depth value information among the plurality of depth value information.

[0197] In the fifth embodiment of this application, the steps of obtaining a plurality of depth value information corresponding to the to-be-processed image and obtaining a plurality of image regions in the to-be-processed image are as follows: First, obtain a depth value image corresponding to the to-be-processed image, and determine a plurality of depth value information in the to-be-processed image according to the depth value image. Then, obtain different image regions corresponding to the specified depth value information among the plurality of depth value information in the to-be-processed image, and obtain a plurality of image regions in the to-be-processed image according to the different image regions corresponding to the specified depth value information among the plurality of depth value information. Among them, the depth value image can be an image obtained in advance through monocular image depth estimation based on deep learning, or an image obtained by performing image depth value extraction on the to-be-processed image based on monocular image depth estimation of deep learning after obtaining the to-be-processed image.

[0198] Step S603: Respectively obtain convolution kernel weight coefficients corresponding to different image regions according to the specified depth value information. The convolution kernel weight coefficients corresponding to different image regions are the weight coefficients of the convolution kernel in the target convolutional neural network for different image regions.

[0199] In the fifth embodiment of the present application, the so-called target convolutional neural network is a convolutional neural network used to extract image features from the image to be processed, so as to obtain useful image feature information in the image to be processed. The so-called convolutional kernel size refers to the height size and width size corresponding to the convolutional kernel. Common convolutional kernel sizes corresponding to convolutional kernels are: 3*3, 5*5, 7*7, etc. The so-called Gaussian filter is used to, when performing image processing on an image, perform weighted averaging on the pixel value of the target pixel point in the image and the pixel values of other pixel points in the target area corresponding to the target pixel point, and use it as the pixel value of the pixel point in the processed image. The so-called standard deviation of the Gaussian filter is used to allocate the weights of different pixel points when performing weighted averaging on the pixel value of the target pixel point in the image and the pixel values of other pixel points in the target area corresponding to the target pixel point. The smaller the standard deviation, the smaller the weights of other pixel points away from the target pixel point.

[0200] Step S604: Respectively perform convolution processing on different image regions according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernels in the target convolutional neural network, so as to obtain convolution results for different image regions.

[0201] In the fifth embodiment of the present application, the so-called performing convolution processing on different image regions according to the simulated weight coefficients corresponding to different image regions in the convolution kernel and the convolution kernels in the target convolutional neural network to obtain convolution results for different image regions means: performing convolution processing on different image regions according to the simulated weight coefficients corresponding to different image regions in the convolution kernel and the convolution kernels in the target convolutional neural network to obtain population density feature information for different image regions.

[0202] Step S605: Obtain the number of target objects according to the convolution results for different image regions.

[0203] In the fifth embodiment of the present application, the so-called obtaining a processing result for the image to be processed according to the convolution results for different image regions means: obtaining the number of human object in the image to be processed according to the population density feature information for different image regions.

[0204] Sixth Embodiment

[0205] Corresponding to the application scenario of the image processing method provided by the present application and the image processing method provided by the first embodiment, the sixth embodiment of the present application also provides an object counting method. Since the sixth embodiment is basically similar to the application scenario and the first embodiment, the description is relatively simple. For related parts, refer to the partial descriptions of the application scenario and the first embodiment. The following description of the sixth embodiment is merely illustrative.

[0206] Please refer to Figure 7, which is a flowchart of an object counting method provided in the sixth embodiment of the present application.

[0207] Step S701: Obtain a to-be-processed image containing a target object.

[0208] In the sixth embodiment of the present application, the so-called to-be-processed image is an image including at least one target object. The so-called target object is an object pre-selected in the image according to requirements, generally a human object, an animal object, an object object, etc.

[0209] Step S702: Obtain multiple bounding box information corresponding to the to-be-processed image, and obtain multiple image regions in the to-be-processed image. Different image regions in the multiple image regions respectively correspond to different bounding box information in the multiple bounding box information.

[0210] In the sixth embodiment of the present application, the steps of obtaining multiple bounding box information corresponding to the to-be-processed image and obtaining multiple image regions in the to-be-processed image are as follows: First, obtain a bounding box image corresponding to the to-be-processed image, and determine multiple bounding box information in the to-be-processed image according to the bounding box image. Then, obtain different image regions corresponding to different bounding box information in the multiple bounding box information in the to-be-processed image, and obtain multiple image regions in the to-be-processed image according to different image regions corresponding to different bounding box information in the multiple bounding box information. Among them, the bounding box image can be an image obtained in advance through a bounding box extraction network, or an image obtained after performing image bounding box extraction on the to-be-processed image based on the bounding box extraction network after obtaining the to-be-processed image. Since the multiple image regions are image regions obtained according to different image regions corresponding to different bounding box information in the to-be-processed image, different image regions in the multiple image regions respectively correspond to specified scale information in the multiple scale information.

[0211] Step S703: Obtain convolution kernel weight coefficients corresponding to different image regions according to different bounding box information. The convolution kernel weight coefficients corresponding to different image regions are the weight coefficients of the convolution kernels in the target convolutional neural network for different image regions.

[0212] In the sixth embodiment of the present application, the so-called target convolutional neural network is a convolutional neural network used to extract image features from an image to be processed, so as to obtain useful image feature information in the image to be processed. The so-called convolutional kernel size refers to the height size and width size corresponding to the convolutional kernel. Common convolutional kernel sizes corresponding to convolutional kernels are: 3*3, 5*5, 7*7, etc. The so-called Gaussian filter is used to, when performing image processing on an image, perform weighted averaging on the pixel value of a target pixel point in the image and the pixel values of other pixel points in the target area corresponding to the target pixel point, and use it as the pixel value of the pixel point in the processed image. The so-called standard deviation of the Gaussian filter is used to allocate the weights of different pixel points when performing weighted averaging on the pixel value of a target pixel point in the image and the pixel values of other pixel points in the target area corresponding to the target pixel point. The smaller the standard deviation, the smaller the weights of other pixel points away from the target pixel point.

[0213] Step S704: Respectively perform convolution processing on different image regions according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernels in the target convolutional neural network, and obtain convolution results for different image regions.

[0214] In the sixth embodiment of the present application, the so-called performing convolution processing on different image regions according to the simulated weight coefficients corresponding to different image regions in the convolution kernel and the convolution kernels in the target convolutional neural network to obtain convolution results for different image regions is: performing convolution processing on different image regions according to the simulated weight coefficients corresponding to different image regions in the convolution kernel and the convolution kernels in the target convolutional neural network, and obtaining edge feature information for different image regions.

[0215] Step S705: Obtain a target object according to the convolution results for different image regions.

[0216] In the sixth embodiment of the present application, the so-called obtaining a processing result for the image to be processed according to the convolution results for different image regions is: obtaining the number of human objects in the image to be processed according to the edge feature information for different image regions.

[0217] Seventh Embodiment

[0218] Corresponding to the application scenario of the image processing method provided by the present application and the image processing method provided by the first embodiment, the seventh embodiment of the present application also provides another object counting method. Since the seventh embodiment is basically similar to the application scenario and the first embodiment, the description is relatively simple. For related parts, refer to the partial descriptions of the application scenario and the first embodiment. The seventh embodiment described below is only illustrative.

[0219] Please refer to Figure 8, which is a flowchart of an image processing method provided in the seventh embodiment of the present application.

[0220] Step S801: obtaining at least one scale information corresponding to the image to be processed, and obtaining at least one image region in the image to be processed, wherein different image regions in the at least one image region respectively correspond to designated scale information in the at least one scale information.

[0221] In the seventh embodiment of the present application, the at least one scale information corresponding to the image to be processed is generally two or more different scale information, or it may be only one scale information. The steps of obtaining at least one scale information corresponding to the image to be processed and obtaining at least one image area in the image to be processed are: first, obtain the image to be processed. Secondly, extract the scale information of the image to be processed to obtain at least one scale information. Thirdly, obtain the specified scale information based on the at least one scale information, and obtain the image area corresponding to the specified scale information in the image to be processed. Finally, obtain at least one image area based on the image area corresponding to the specified scale information in the image to be processed. Since the at least one image area is an image area obtained based on different image areas corresponding to different scale information in the image to be processed, different image areas in the at least one image area respectively correspond to the specified scale information in the at least one scale information.

[0222] Step S802: Obtain convolution kernel weight coefficients corresponding to different image regions according to the specified scale information, where the convolution kernel weight coefficients corresponding to different image regions are weight coefficients of the convolution kernels in the target convolution neural network for different image regions.

[0223] In the seventh embodiment of the present application, the so-called target convolutional neural network is a convolutional neural network used to extract image features of the image to be processed to obtain useful image feature information in the image to be processed. The so-called convolution kernel size is the height size and width size corresponding to the convolution kernel. The convolution kernel sizes corresponding to the common convolution kernels are: 3*3, 5*5 and 7*7, etc. The so-called Gaussian filter is used to perform weighted averaging of the pixel value of the target pixel point in the image and the pixel values ​​of other pixels in the target area corresponding to the target pixel point when processing the image, as the pixel value of the pixel point of the processed image. The standard deviation of the so-called Gaussian filter is used to assign weights of different pixels when the pixel value of the target pixel point in the image is weighted averaged with the pixel values ​​of other pixels in the target area corresponding to the target pixel point. The smaller the standard deviation, the smaller the weight of other pixels away from the target pixel point.

[0224] Step S803: performing convolution processing on different image regions according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernels in the target convolution neural network, to obtain convolution results for different image regions.

[0225] In the seventh embodiment of the present application, convolution processing is performed on different image regions according to the simulated weight coefficients corresponding to different image regions by the elements in the convolution kernel and the convolution kernel in the target convolutional neural network to obtain convolution results for different image regions: convolution processing is performed on different image regions according to the simulated weight coefficients corresponding to different image regions by the elements in the convolution kernel and the convolution kernel in the target convolutional neural network to obtain edge feature information for different image regions.

[0226] Eighth embodiment

[0227] Corresponding to the application scenario of the image processing method provided by the present application and the image processing method provided by the first embodiment, the eighth embodiment of the present application also provides another object counting method. Since the eighth embodiment is basically similar to the application scenario and the first embodiment, the description is relatively simple, and the relevant parts refer to the application scenario and the partial description of the first embodiment. The eighth embodiment described below is only illustrative.

[0228] Please refer to Figure 9 , which is a schematic diagram of an image processing system provided in the eighth embodiment of the present application.

[0229] The image processing system includes: a platform server 901 and a user terminal 902.

[0230] In the eighth embodiment of the present application, the platform server 901 refers to a computing device that provides services for the software platform or application platform installed on the user terminal 902 for executing the image processing method provided by the present application, and is generally a server or server cluster in specific implementation. The user terminal 902 refers to a computing device installed with a software platform or application platform for executing the image processing method provided by the present application, and is generally a smart phone, a tablet computer, a personal computer, etc. in specific implementation.

[0231] In the eighth embodiment of the present application, the platform server 901 is configured to obtain the image to be processed provided by the user terminal 902; obtain multiple scale information corresponding to the image to be processed, and obtain multiple image regions in the image to be processed, where different image regions in the multiple image regions respectively correspond to the specified scale information in the multiple scale information; respectively obtain the convolution kernel weight coefficients corresponding to different image regions according to the specified scale information, and the convolution kernel weight coefficients corresponding to different image regions are the weight coefficients of different image regions for the convolution kernel in the target convolutional neural network; respectively perform convolution processing on different image regions according to the convolution kernel weight coefficients corresponding to different image regions and the convolution kernel in the target convolutional neural network to obtain the convolution results for different image regions; obtain the processing result for the image to be processed according to the convolution results for different image regions; and provide the processing result to the user terminal 902. The user terminal 902 is configured to send the image to be processed to the platform server 901; and obtain the processing result sent by the platform server 901.

[0232] In addition, in the eighth embodiment of the present application, after obtaining the multiple scale information corresponding to the image to be processed, the multiple image regions in the image to be processed, the convolution kernel weight coefficients corresponding to different image regions, and the convolution results for different image regions, the platform server 901 may further provide the multiple scale information corresponding to the image to be processed, the multiple image regions in the image to be processed, the convolution kernel weight coefficients corresponding to different image regions, and the convolution results for different image regions to the user terminal 902 for display by the user terminal 902. It is also possible to store the multiple scale information corresponding to the image to be processed, the multiple image regions in the image to be processed, the convolution kernel weight coefficients corresponding to different image regions, the convolution results for different image regions, and the processing result in the memory of the platform server 901. When the platform server processes the same image to be processed again, it can obtain the multiple scale information corresponding to the image to be processed, the multiple image regions in the image to be processed, the convolution kernel weight coefficients corresponding to different image regions, the convolution results for different image regions, and the processing result stored in the memory from the memory, thereby improving the processing speed for the same image to be processed. Although the present application is disclosed above with preferred embodiments, it is not used to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present application should be defined by the scope of the claims of the present application.

[0233] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0234] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (Flash RAM). The memory is an example of computer-readable media.

[0235] 1. Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0236] 2. Those skilled in the art will appreciate that the embodiments of the present application may be provided as a method, system, or computer program product. Accordingly, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

Claims

1. An image processing method, characterized in that, comprising: obtaining a plurality of scale information corresponding to the image to be processed, and obtaining a plurality of image regions in the image to be processed, wherein different image regions in the plurality of image regions respectively correspond to specified scale information in the plurality of scale information; obtaining convolution kernel weight coefficients corresponding to different image regions according to the scale information corresponding to the different image regions and a plurality of Gaussian filters, the convolution kernel weight coefficients corresponding to the different image regions being the weight coefficients of the convolution kernels in the target convolutional neural network for the different image regions, and the convolution kernel size of the Gaussian convolution kernel of the Gaussian filter being the same as the convolution kernel size of the convolution kernel; multiplying the convolution kernel weight coefficients corresponding to the different image regions by the weights corresponding to the elements in the convolution kernel to obtain simulated weight coefficients corresponding to the elements in the convolution kernel for the different image regions; performing convolution processing on the different image regions according to the simulated weight coefficients and the convolution kernels in the target convolutional neural network to obtain convolution results for the different image regions.

2. The image processing method according to claim 1, characterized in that, further comprising: obtaining a processing result for the image to be processed according to the convolution results for the different image regions.

3. The image processing method according to claim 1, characterized in that, the determining the convolution kernel weight coefficients corresponding to the different image regions according to the scale information corresponding to the different image regions and the plurality of Gaussian filters comprises: determining weight values of the plurality of Gaussian filters corresponding to the different image regions according to the scale information corresponding to the different image regions and the plurality of Gaussian filters; obtaining the convolution kernel weight coefficients corresponding to the different image regions according to the weight values of the plurality of Gaussian filters corresponding to the different image regions and the plurality of Gaussian filters.

4. The image processing method according to claim 3, characterized in that, the determining the weight values of the plurality of Gaussian filters corresponding to the different image regions according to the scale information corresponding to the different image regions and the plurality of Gaussian filters comprises: determining the degree of correlation between each Gaussian filter in the plurality of Gaussian filters and the different image regions according to the scale information corresponding to the different image regions and the plurality of Gaussian filters; determining the weight value of each Gaussian filter corresponding to the different image regions according to the degree of correlation between each Gaussian filter and the different image regions.

5. The image processing method according to claim 4, characterized in that, Obtaining the convolution kernel weight coefficients corresponding to the different image regions according to the weight values corresponding to the different image regions and the multiple Gaussian filters of the multiple Gaussian filters includes: performing weighted averaging on the weight values corresponding to the different image regions and the multiple Gaussian filters according to the weight value corresponding to each Gaussian filter in the different image regions of the multiple Gaussian filters and each Gaussian filter, to obtain the convolution kernel weight coefficients corresponding to the different image regions.

6. The image processing method according to claim 4, wherein, each Gaussian filter of the multiple Gaussian filters corresponds to a different specified standard deviation, and determining the degree of correlation between each Gaussian filter of the multiple Gaussian filters and the different image regions according to the scale information corresponding to the different image regions and the multiple Gaussian filters includes: obtaining a first specified parameter for calculating the degree of correlation between each Gaussian filter and the different image regions, and obtaining a second specified parameter for calculating the degree of correlation between each Gaussian filter and the different image regions; obtaining the number of the multiple Gaussian filters, and obtaining the serial number of each Gaussian filter in the sorting of the multiple Gaussian filters, where the sorting of the multiple Gaussian filters is a sorting of the multiple Gaussian filters in ascending order of the corresponding specified standard deviation; determining the degree of correlation between each Gaussian filter and the different image regions according to the first specified parameter, the second specified parameter, the number of the multiple Gaussian filters, the scale information corresponding to the different image regions, and the serial number of each Gaussian filter in the sorting of the multiple Gaussian filters.

7. The image processing method according to claim 6, wherein, the scale information corresponding to the different image regions includes: visibility value information corresponding to the different image regions; determining the degree of correlation between each Gaussian filter and the different image regions according to the first specified parameter, the second specified parameter, the number of the multiple Gaussian filters, the scale information corresponding to the different image regions, and the serial number of each Gaussian filter in the sorting of the multiple Gaussian filters includes: obtaining a target parameter for calculating the degree of correlation between each Gaussian filter and the different image regions according to the first specified parameter, the number of the multiple Gaussian filters, the visibility value information corresponding to the different image regions, and the serial number of each Gaussian filter in the sorting of the multiple Gaussian filters; calculating the degree of correlation between each Gaussian filter and the different image regions according to the target parameter and the second specified parameter.

8. The image processing method according to claim 4, wherein, determining the weight value corresponding to each Gaussian filter in the different image regions according to the degree of correlation between each Gaussian filter and the different image regions includes: Obtain the exponential power of the degree of correlation between each Gaussian filter and the different image regions according to the degree of correlation between each Gaussian filter and the different image regions; Obtain the sum of the exponential powers of the degrees of correlation between the multiple Gaussian filters and the different image regions according to the exponential powers of the degrees of correlation between each Gaussian filter and the different image regions; Obtain the weight value corresponding to each Gaussian filter in the different image regions according to the ratio of the exponential power of the degree of correlation between each Gaussian filter and the different image regions to the sum of the exponential powers of the degrees of correlation between the multiple Gaussian filters and the different image regions.

9. The image processing method according to claim 1, wherein, further comprising: Obtain the convolution kernel size of the convolution kernel; Determine the convolution kernel size of the Gaussian convolution kernel according to the convolution kernel size of the convolution kernel.

10. The image processing method according to claim 1, wherein, The obtaining of the multiple scale information corresponding to the image to be processed and the obtaining of the multiple image regions in the image to be processed include: Obtain the image to be processed; Extract scale information from the image to be processed to obtain the multiple scale information; Obtain the specified scale information according to the multiple scale information, and obtain the image region corresponding to the specified scale information in the image to be processed; Obtain the multiple image regions according to the image region corresponding to the specified scale information in the image to be processed.

11. An image processing apparatus, wherein, comprising: An image region obtaining unit, configured to obtain multiple scale information corresponding to an image to be processed, and obtain multiple image regions in the image to be processed, wherein different image regions in the multiple image regions respectively correspond to the specified scale information in the multiple scale information; A weight coefficient obtaining unit, configured to obtain the convolution kernel weight coefficients corresponding to different image regions according to the scale information corresponding to the different image regions and multiple Gaussian filters, the convolution kernel weight coefficients corresponding to the different image regions being the weight coefficients of the convolution kernel in a target convolutional neural network, and the convolution kernel size of the Gaussian convolution kernel of the Gaussian filter being the same as the convolution kernel size of the convolution kernel; A convolution result obtaining unit, configured to multiply the convolution kernel weight coefficients corresponding to the different image regions by the weights corresponding to the elements in the convolution kernel to obtain the simulated weight coefficients corresponding to the elements in the convolution kernel for the different image regions; perform convolution processing on the different image regions according to the simulated weight coefficients and the convolution kernel in the target convolutional neural network to obtain the convolution results for the different image regions.

12. An electronic device, wherein, comprising: A processor; and A memory, configured to store a program for an image processing method. After the device is powered on and runs the program of the image processing method through the processor, the following steps are executed: Obtain multiple scale information corresponding to the image to be processed, and obtain multiple image regions in the image to be processed, where different image regions in the multiple image regions respectively correspond to specified scale information in the multiple scale information; According to the scale information corresponding to the different image regions and multiple Gaussian filters, obtain convolution kernel weight coefficients corresponding to the different image regions. The convolution kernel weight coefficients corresponding to the different image regions are the weight coefficients of the different image regions for the convolution kernels in the target convolutional neural network. The convolution kernel size of the Gaussian convolution kernel of the Gaussian filter is the same as the convolution kernel size of the convolution kernel; Multiply the convolution kernel weight coefficients corresponding to the different image regions by the weights corresponding to the elements in the convolution kernel to obtain simulated weight coefficients corresponding to the elements in the convolution kernel for the different image regions; According to the simulated weight coefficients and the convolution kernels in the target convolutional neural network, perform convolution processing on the different image regions to obtain convolution results for the different image regions.

13. A storage medium, characterized in that, it stores a program for an image processing method, and when this program is run by a processor, it executes the following steps: Obtain multiple scale information corresponding to the image to be processed, and obtain multiple image regions in the image to be processed, where different image regions in the multiple image regions respectively correspond to specified scale information in the multiple scale information; According to the scale information corresponding to the different image regions and multiple Gaussian filters, obtain convolution kernel weight coefficients corresponding to the different image regions. The convolution kernel weight coefficients corresponding to the different image regions are the weight coefficients of the different image regions for the convolution kernels in the target convolutional neural network. The convolution kernel size of the Gaussian convolution kernel of the Gaussian filter is the same as the convolution kernel size of the convolution kernel; Multiply the convolution kernel weight coefficients corresponding to the different image regions by the weights corresponding to the elements in the convolution kernel to obtain simulated weight coefficients corresponding to the elements in the convolution kernel for the different image regions; According to the simulated weight coefficients and the convolution kernels in the target convolutional neural network, perform convolution processing on the different image regions to obtain convolution results for the different image regions.

14. An object counting method, characterized in that, it includes: Obtain the image to be processed containing the target object; Obtain multiple depth value information corresponding to the image to be processed, and obtain multiple image regions in the image to be processed, where different image regions in the multiple image regions respectively correspond to specified depth value information in the multiple depth value information; Respectively according to the specified depth value information and multiple Gaussian filters, obtain convolution kernel weight coefficients corresponding to the different image regions. The convolution kernel weight coefficients corresponding to the different image regions are the weight coefficients of the different image regions for the convolution kernels in the target convolutional neural network. The convolution kernel size of the Gaussian convolution kernel of the Gaussian filter is the same as the convolution kernel size of the convolution kernel; Multiply the convolution kernel weight coefficients corresponding to the different image regions by the weights corresponding to the elements in the convolution kernel to obtain the simulated weight coefficients corresponding to the elements in the convolution kernel for the different image regions; Perform convolution processing on the different image regions according to the simulated weight coefficients and the convolution kernel in the target convolutional neural network to obtain the convolution results for the different image regions; Obtain the number of the target objects according to the convolution results for the different image regions.

15. An object detection method, characterized in that, it includes: Obtain a to-be-processed image including a target object; Obtain multiple bounding box information corresponding to the to-be-processed image, and obtain multiple image regions in the to-be-processed image, where different image regions in the multiple image regions respectively correspond to different bounding box information in the multiple bounding box information; Respectively obtain the convolution kernel weight coefficients corresponding to the different image regions according to the different bounding box information and multiple Gaussian filters, where the convolution kernel weight coefficients corresponding to the different image regions are the weight coefficients of the different image regions for the convolution kernel in the target convolutional neural network, and the convolution kernel size of the Gaussian convolution kernel of the Gaussian filter is the same as the convolution kernel size of the convolution kernel; Multiply the convolution kernel weight coefficients corresponding to the different image regions by the weights corresponding to the elements in the convolution kernel to obtain the simulated weight coefficients corresponding to the elements in the convolution kernel for the different image regions; Perform convolution processing on the different image regions according to the simulated weight coefficients and the convolution kernel in the target convolutional neural network to obtain the convolution results for the different image regions; Obtain the target object according to the convolution results for the different image regions.

16. An image processing method, characterized in that, it includes: Obtain at least one scale information corresponding to the to-be-processed image, and obtain at least one image region in the to-be-processed image, where different image regions in the at least one image region respectively correspond to the specified scale information in the at least one scale information; Respectively obtain the convolution kernel weight coefficients corresponding to the different image regions according to the specified scale information and multiple Gaussian filters, where the convolution kernel weight coefficients corresponding to the different image regions are the weight coefficients of the different image regions for the convolution kernel in the target convolutional neural network, and the convolution kernel size of the Gaussian convolution kernel of the Gaussian filter is the same as the convolution kernel size of the convolution kernel; Multiply the convolution kernel weight coefficients corresponding to the different image regions by the weights corresponding to the elements in the convolution kernel to obtain the simulated weight coefficients corresponding to the elements in the convolution kernel for the different image regions; Perform convolution processing on the different image regions according to the simulated weight coefficients and the convolution kernel in the target convolutional neural network to obtain the convolution results for the different image regions.

17. An image processing system, characterized in that, it includes: A platform server and a user terminal; The platform server is used to obtain the to-be-processed image provided by the user terminal; Obtain multiple scale information corresponding to the image to be processed, and obtain multiple image regions in the image to be processed, where different image regions in the multiple image regions respectively correspond to specified scale information in the multiple scale information; according to the scale information corresponding to the different image regions and multiple Gaussian filters, obtain convolution kernel weight coefficients corresponding to the different image regions, where the convolution kernel weight coefficients corresponding to the different image regions are the weight coefficients of the different image regions for the convolution kernel in the target convolutional neural network, and the convolution kernel size of the Gaussian convolution kernel of the Gaussian filter is the same as the convolution kernel size of the convolution kernel; multiply the convolution kernel weight coefficients corresponding to the different image regions by the weights corresponding to the elements in the convolution kernel to obtain the simulated weight coefficients corresponding to the elements in the convolution kernel for the different image regions; according to the simulated weight coefficients and the convolution kernel in the target convolutional neural network, perform convolution processing on the different image regions to obtain convolution results for the different image regions; according to the convolution results for the different image regions, obtain a processing result for the image to be processed; provide the processing result to the user terminal; The user terminal is used to send the image to be processed to the platform server; Obtain the processing result sent by the platform server.

Citation Information

Patent Citations

  • Deep learning-based image fuzzy region detection method and apparatus

    CN106096605A

  • Convolutional neural network multinuclear parallel computing method facing GPDSP

    CN108920413A