Image processing method and device
By keeping the resolution of the main feature layer convolution layer consistent in the image recognition model and optimizing the loss function with category label weights, the problem of feature loss and inaccurate recognition of small-sized objects in image processing is solved, and the accuracy of image semantic segmentation is improved.
Patent Information
- Application Number
- CN202410383160.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2025-07-25
AI Technical Summary
In the existing image semantic segmentation technology, small-sized objects are prone to loss of detail features and inaccurate category recognition during image processing.
By using up-down sampling in the feature extraction network of the image recognition model, feature extraction is performed on different feature layers, ensuring that the resolution of the convolutional layer in the main feature layer is the same, and training and optimization is performed using a loss function with weight values corresponding to different category tags, and combining the output results of different feature layers for fusion processing.
It effectively avoids the loss of features of small-size targets during up-down sampling, improves the recognition accuracy of small-size targets, and improves the recognition performance of multiple categories.
Smart Images

Figure CN120374968A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to an image processing method and apparatus. Background Art
[0002] Image semantic segmentation technology has currently been extended to multiple common fields such as image object detection, autonomous driving, and face recognition. However, with the continuous development of deep neural networks, the segmentation effect for large-sized objects is relatively good during image semantic segmentation, while the segmentation effect for small-sized objects is not satisfactory. The reason is that as the neural network deepens, the resolution of the feature map decreases accordingly, and the information of small-scale objects will gradually be submerged in the context background during the feature extraction process, resulting in very few features of small-scale objects in the image and making it difficult to distinguish them from the background or similar objects.
[0003] In related technologies, to address the problem that small-scale objects lose features during multi-layer feature extraction, there are usually two directions: improving the neural network structure and supplementing the dataset. For example, dilated convolution or multi-scale training methods are used to improve the loss of small-object features, but this cannot fundamentally avoid the loss of feature details; while supplementing the dataset requires time-consuming and laborious manual collection and annotation of the dataset, and it cannot improve the segmentation performance for small objects of multiple categories.
[0004] Therefore, in the current image semantic segmentation technology, there are problems that small-sized objects are prone to losing detailed features and inaccurate category recognition during the image processing process. Summary of the Invention
[0005] In view of this, this application provides an image processing method and apparatus, mainly aiming to improve the problems that small-sized objects are prone to losing detailed features and inaccurate category recognition during the current image semantic segmentation technology.
[0006] In a first aspect, this application provides an image processing method, including:
[0007] Obtain an image to be processed as the input of an image recognition model;
[0008] Use the feature extraction network of the image recognition model to perform feature extraction on the image to be processed in different feature layers by means of upsampling and downsampling, and obtain the output results of different feature layers; the convolutional layers in the main feature layer of the feature extraction network have the same resolution, and the image recognition model is trained and optimized using a loss function with weight values corresponding to different category labels set.
[0009] Perform fusion processing based on the output results of different feature layers to obtain an output image.
[0010] Optionally, the training steps of the image recognition model include: obtaining an image data set labeled with class labels; the number and types of the class labels are preset in advance, and the number of the class labels is greater than or equal to two; dividing the image data set into a training set and a test set according to a preset ratio, and training and testing the image recognition model; wherein, the images in the training set and the test set do not overlap with each other.
[0011] Optionally, the training steps of the image recognition model further include: performing pixel category statistics on the images in the image data set to obtain the number of pixels corresponding to each class label; setting weight values corresponding to each class label for the loss function of the image recognition model according to the number of pixels corresponding to each class label; and training and optimizing the image recognition model by using the loss function with weight values corresponding to different class labels set.
[0012] Optionally, performing pixel category statistics on the image data set to obtain the number of pixels corresponding to each class label includes: based on the class label, performing class annotation on each pixel in all images in terms of pixels to obtain the number of pixels corresponding to each class label.
[0013] Optionally, setting weight values corresponding to each class label for the loss function of the image recognition model according to the number of pixels corresponding to each class label includes: calculating the weight value corresponding to each class label according to the proportion of the number of pixels corresponding to each class label in the sum of the total pixels of all images; and setting the loss function by using the weight value corresponding to each class label.
[0014] Optionally, calculating the weight value corresponding to each class label according to the proportion of the number of pixels corresponding to each class label in the sum of the total pixels of all images includes: performing a reciprocal processing on the proportion of the number of pixels corresponding to the class label of the target class in the sum of the total pixels of all images to obtain the weight value corresponding to the class label of the target class.
[0015] In a second aspect, the present application provides an image processing apparatus, including:
[0016] An acquisition unit configured to acquire a to-be-processed image as an input of an image recognition model;
[0017] An extraction unit configured to perform feature extraction on the to-be-processed image in different feature layers by using an upsampling and downsampling method by using a feature extraction network of the image recognition model to obtain output results of different feature layers; the resolution of the convolutional layers in the main feature layer of the feature extraction network is the same, and the image recognition model is trained and optimized by using a loss function with weight values corresponding to different class labels set.
[0018] A processing unit configured to perform a fusion process based on output results of the different feature layers to obtain an output image.
[0019] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the image processing method described in the first aspect is implemented.
[0020] In a fourth aspect, the present application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and when the processor executes the computer program, the image processing method described in the first aspect is implemented.
[0021] In a fifth aspect, the present application provides a chip, including one or more interface circuits and one or more processors; the interface circuit is configured to receive a signal from a memory of an electronic device and send the signal to the processor, and the signal includes computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device is caused to execute the image processing method described in the first aspect.
[0022] By means of the above technical solutions, an image processing method and apparatus provided by the present application first obtain an image to be processed as an input of an image recognition model, and then use a feature extraction network of the image recognition model to perform feature extraction on the image to be processed through upsampling and downsampling methods at different feature layers to obtain output results of different feature layers. Among them, the resolution of the convolutional layers in the main feature layer of the feature extraction network is the same, and the image recognition model is trained and optimized by using a loss function with weight values corresponding to different category labels already set. Finally, a fusion process is performed based on the output results of different feature layers to obtain an output image. Compared with the related art, on the one hand, by adopting the same resolution in the main feature layer, the problem of feature loss of small-size targets in the image during the upsampling and downsampling processes is avoided; on the other hand, by setting corresponding weight values for different category targets for model optimization, the problem of uneven samples caused by the small proportion of small-size targets in the training set is avoided. Thereby, more features of small-size targets are retained, and the recognition accuracy of different categories is improved.
[0023] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and together with the specification are used to explain the principles of the present application.
[0025] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0026] Figure 1 The flowchart shows a method for image processing provided by an embodiment of the present application;
[0027] Figure 2 The schematic diagram shows a network structure for extracting image features provided by an embodiment of the present application;
[0028] Figure 3 The flowchart shows the steps for training an image recognition model provided by an embodiment of the present application;
[0029] Figure 4 The schematic diagram shows the structure of an image processing device provided by an embodiment of the present application. Detailed implementation manners
[0030] Here, some embodiments of the present disclosure will be described in detail, and their examples are shown in the accompanying drawings. When the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. Various changes, deformations, and equivalents of the methods, devices, and / or systems described herein will become apparent after understanding the present disclosure. For example, the order of the operations described herein is only an example and is not limited to the orders set forth herein. Instead, except for the operations that must be performed in a specific order, changes can be made as will be apparent after understanding the present disclosure. Additionally, for the sake of clarity and conciseness, the description of features known in the art may be omitted.
[0031] The implementation manners described in some embodiments of the present disclosure below do not represent all implementation manners consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0032] Image semantic segmentation is to assign a semantic label to each pixel in an image, which is a fundamental task in image processing. Traditional semantic segmentation algorithms are mainly based on machine learning classifiers. However, the rise of deep learning, especially the development of convolutional neural networks, has greatly improved the accuracy of semantic segmentation algorithms. Image semantic segmentation can be applied in many real-life fields such as autonomous driving, image semantic enhancement, augmented reality, security, and remote sensing measurement and control. Currently, semantic segmentation methods based on deep neural networks have achieved obvious performance advantages in the field of image semantic segmentation. However, most image processing solutions perform well in segmenting large-scale objects, but the segmentation effect of small-scale objects is not satisfactory. As the neural network deepens, the resolution of the feature map decreases accordingly, and the information of small-scale objects will gradually be submerged in the context background during the feature extraction process, resulting in very few features of small-scale objects in the image and making it difficult to distinguish them from the background or similar targets. However, in scenarios such as facial feature segmentation and autonomous driving, the performance requirements for small-scale objects are very strict.
[0033] In related technologies, two aspects of improving the neural network structure or supplementing the data set are proposed to improve the segmentation accuracy of small targets. However, it does not fundamentally avoid the loss of feature details, and there are also problems of inaccurate recognition for small targets of different categories.
[0034] Therefore, in order to improve the problems of easy loss of detailed features and inaccurate category recognition of small-sized objects in the current image semantic segmentation technology during the image processing process. This embodiment provides an image processing method, as Figure 1 shown, the method includes:
[0035] S101. Obtain the image to be processed as the input of the image recognition model.
[0036] The image to be processed is also the image that needs to be subjected to image recognition and feature extraction, and it is used as the input of the image recognition model. Here, the image recognition model refers to the image recognition model that has been trained according to the training method proposed in this embodiment and can directly input the image to be processed and output the result.
[0037] S102. Use the feature extraction network of the image recognition model to perform feature extraction on the image to be processed in different feature layers by means of upsampling and downsampling, and obtain the output results of different feature layers.
[0038] Among them, the convolutional layers in the main feature layer of the feature extraction network have the same resolution, and the image recognition model is trained and optimized using a loss function with weight values corresponding to different category labels. The feature extraction network usually includes multiple feature layers, and each feature layer includes one or more convolutional layers. In this embodiment, the convolutional layers in the main feature layer have the same resolution and all maintain a high resolution. Fundamentally, it avoids the loss of detailed information, thereby retaining more detailed information about small targets.
[0039] Meanwhile, during the training stage of this model, a loss function with weight values corresponding to different category labels is used for optimization training. By statistically analyzing the proportion of pixels of different target categories in the semantic segmentation dataset and weighting the weights of different categories in the training loss function, the overall improvement of the segmentation performance of multiple types of small targets is achieved.
[0040] The upsampling and downsampling method is also to extract features by performing upsampling and downsampling between each feature layer, and finally the concat function of different feature layers outputs semantic information and detailed information of the target together. Specifically, as Figure 2 shown, an example of an image feature extraction network structure is given, including four feature layers. The main feature layer includes 4 convolutional layers, and each convolutional layer uses the same resolution, so that feature extraction can be performed using the same resolution in the main feature layer, avoiding the loss of detailed information. At the same time, upsampling and downsampling are combined with other feature layers to obtain the final output result.
[0041] S103. Perform fusion processing based on the output results of different feature layers to obtain an output image.
[0042] In this embodiment, first, the image to be processed is obtained as the input of the image recognition model, and then the feature extraction network of the image recognition model is used to extract features from the image to be processed by the upsampling and downsampling method in different feature layers, obtaining the output results of different feature layers. Among them, the convolutional layers in the main feature layer of the feature extraction network have the same resolution, and the image recognition model is obtained by training and optimizing using a loss function with weight values corresponding to different category labels. Finally, fusion processing is performed based on the output results of different feature layers to obtain an output image. Compared with the related technology, on the one hand, in this embodiment, by using the same resolution in the main feature layer, the problem of feature loss of small-sized targets in the image during the upsampling and downsampling process is avoided; on the other hand, by setting corresponding weight values for different category targets for model optimization, the problem of uneven samples caused by the small proportion of small-sized targets in the training set is avoided. Thus, more features of small-sized targets are retained, and the recognition accuracy of different categories is improved.
[0043] Optionally, the training steps of the image recognition model include: obtaining an image data set labeled with class labels; dividing the image data set into a training set and a test set according to a preset ratio, and training and testing the image recognition model; wherein, the images in the training set and the test set do not overlap with each other.
[0044] In this embodiment, an image data set labeled with class labels is obtained. The number and types of the class labels are preset in advance, and the number of class labels is greater than or equal to two. The class labels are mainly set for small-sized objects, such as people, mouths, human faces, green plants, eyes, the ground, the sky, eyebrows, human skin, hair, and other categories. These are usually relatively small or difficult-to-distinguish contents in the picture and are prone to loss during the feature extraction process due to different resolutions of the feature layers. Then, the image data set is divided into a training set and a test set according to a preset ratio, such as 4:1, and the image recognition model is trained, so that the image recognition model can recognize small-sized objects.
[0045] Optionally, the training steps of the image recognition model further include: performing pixel category statistics on the images in the image data set to obtain the pixel quantity corresponding to each class label; setting the weight value corresponding to each class label for the loss function of the image recognition model according to the pixel quantity corresponding to each class label; and training and optimizing the image recognition model by using the loss function with the weight values corresponding to different class labels set.
[0046] In this embodiment, the training process of the image recognition model is also optimized. In the above-mentioned embodiment, small target objects such as green plants may have fewer samples in the training set, so there will be a problem of inaccurate recognition during the feature extraction process. Therefore, first, pixel category statistics are performed on the images in the image data set to obtain the pixel quantity corresponding to each class label. Then, the corresponding weight values are set according to the pixel quantities corresponding to different class labels, and the model is optimized through the loss function, so that small-sized objects with fewer samples can also have a better recognition effect.
[0047] Optionally, performing pixel category statistics on the image data set to obtain the pixel quantity corresponding to each class label includes: based on the class label, performing class annotation on each pixel in all images in units of pixels to obtain the pixel quantity corresponding to each class label.
[0048] In this embodiment, to obtain the pixel quantity corresponding to each class label, a script can be written to traverse each pixel in all images in the image data set, count the class label corresponding to each pixel point, and thus obtain the pixel quantity corresponding to each class label of the entire data set.
[0049] Optionally, according to the number of pixels corresponding to each category label, set the weight value corresponding to each category label for the loss function of the image recognition model, including: calculating the weight value corresponding to each category label according to the proportion of the number of pixels corresponding to each category label in the sum of the total pixels of all images; setting the loss function using the weight value corresponding to each category label.
[0050] In this embodiment, after obtaining the number of pixels corresponding to each category label, it is then necessary to calculate the weight value according to the proportion of the number of pixels corresponding to each category label in the sum of the total pixels of all images. Specifically, if the proportion of eyebrow samples is small, the weight value corresponding to the eyebrows should be set higher, so that small-sized objects with a small number of samples can also have a good recognition effect.
[0051] Optionally, calculate the weight value corresponding to each category label according to the proportion of the number of pixels corresponding to each category label in the sum of the total pixels of all images, including: taking the reciprocal of the proportion of the number of pixels corresponding to the category label of the target category in the sum of the total pixels of all images to obtain the weight value corresponding to the category label of the target category.
[0052] In this embodiment, the loss function can be set by taking the inverse of the proportion of the number of pixels corresponding to the category label of the target category in the sum of the total pixels of all images as a feasible setting method. And in this way, different categories in the loss function are given different weight values, so that small-sized objects with a small number of samples can also have a good recognition effect.
[0053] The difference between this embodiment and the image semantic segmentation method in the related art lies in the semantic feature information extraction and neural network training stages. The trained semantic segmentation neural network model can be used in the digital image processing process of mobile phone photography and industrial camera imaging. And the solution proposed in this embodiment can significantly improve the target semantic segmentation performance, including the average precision of all categories as a whole and the precision of small targets. Table 1 shows the performance of this solution on the test set, intuitively showing the comparison of the segmentation precision between the solution in the related art (benchmark method) and the image recognition method in this embodiment (this embodiment). In terms of settings, the difference is only whether to use the above solution during training, and other settings are exactly the same.
[0054] Baseline method This embodiment mIoU 84.83% 85.43% mIoU (small objects) 77.18% 78.70%
[0055] Table 1 Comparison of target segmentation performance on the test set
[0056] Furthermore, in order to better reflect the process of model training in the technical solution proposed in this embodiment, Figure 3 The flowchart showing the training steps of the image recognition model in the above image processing method includes:
[0057] S301. Label several training images according to preset category labels to obtain an image dataset.
[0058] Here, it is illustrated by way of example. The preset category labels are set to 10 categories, including people, mouths, faces, green plants, eyes, the ground, the sky, eyebrows, human skin, and hair. 30,000 images are labeled to form an image dataset. And 5,000 of the 30,000 images are taken as the test set, and the test set and the training set do not overlap with each other.
[0059] S302. Based on the category labels, count each pixel in all images in terms of pixels to obtain the number of pixels corresponding to each category label.
[0060] Furthermore, count the category labels corresponding to each pixel point, so as to obtain the number of pixel points included in each category of the entire dataset. Among them, an image includes multiple pixel points, and the total number of pixel points is taken as 100,000 for example.
[0061] S303. Calculate the weight value corresponding to each category label according to the proportion of the number of pixels corresponding to each category label in the sum of all image total pixels.
[0062] In this step, further calculate the category weights, obtain the proportion of the number of pixels of each category in the 10 categories in the total pixel value sum (100,000), and take the reciprocal. Then, the smaller the proportion of the number of pixels, the larger the obtained weight value.
[0063] S304. Use the weight value corresponding to each category label to set the loss function.
[0064] S305. Train the image recognition model based on the image dataset and the loss function with weight values corresponding to different category labels already set.
[0065] In this embodiment, by counting the pixel ratios of different category targets in the training dataset, the training loss function is improved, different training weights are given to different category targets during training, and the problem of sample imbalance caused by the too small proportion of small targets in the training set is avoided, so that the neural network focuses more on the training of small targets during training and improves the segmentation performance of small targets in multiple categories.
[0066] In addition, in the feature extraction stage, the structure of the semantic segmentation feature extraction network is redesigned, and the same high resolution is maintained in the main feature layer to avoid information loss caused by the upsampling and downsampling processes in the existing solutions, so as to retain more feature information of small targets and improve the segmentation performance of small targets.
[0067] Further, as Figures 1 to 3For the specific implementation of the method shown, this embodiment provides an image processing device, as Figure 4 shown, the device includes: an acquisition unit 41, an extraction unit 42, and a processing unit 43.
[0068] The acquisition unit 41 is configured to acquire an image to be processed as the input of an image recognition model;
[0069] The extraction unit 42 is configured to use the feature extraction network of the image recognition model to perform feature extraction on the image to be processed in different feature layers by means of upsampling and downsampling, and obtain the output results of different feature layers; the resolution of the convolutional layers in the main feature layer of the feature extraction network is the same, and the image recognition model is trained and optimized using a loss function with weight values corresponding to different category labels set;
[0070] The processing unit 43 is configured to perform fusion processing based on the output results of different feature layers to obtain an output image.
[0071] In a specific application scenario, the acquisition unit 41 is specifically configured to acquire an image data set labeled with category labels; the number and types of the category labels are preset in advance, and the number of the category labels is greater than or equal to two; the image data set is divided into a training set and a test set according to a preset ratio, and the image recognition model is trained and tested; wherein, the images in the training set and the test set do not overlap with each other.
[0072] In a specific application scenario, the acquisition unit 41 is further specifically configured to perform pixel category statistics on the images in the image data set to obtain the number of pixels corresponding to each category label; according to the number of pixels corresponding to each category label, set weight values corresponding to each category label for the loss function of the image recognition model; use the loss function with weight values corresponding to different category labels set to train and optimize the image recognition model.
[0073] In a specific application scenario, the acquisition unit 41 is further specifically configured to, based on the category labels, perform category annotation on each pixel in all images in units of pixels to obtain the number of pixels corresponding to each category label.
[0074] In a specific application scenario, the acquisition unit 41 is further specifically configured to calculate the weight value corresponding to each category label according to the proportion of the number of pixels corresponding to each category label in the sum of the total pixels of all images; use the weight value corresponding to each category label to set the loss function.
[0075] In a specific application scenario, the obtaining unit 41 is further specifically configured to take the reciprocal of the ratio of the number of pixels corresponding to the category label of the target category to the sum of all the pixels of the image, so as to obtain the weight value corresponding to the category label of the target category.
[0076] It should be noted that for other corresponding descriptions of each functional unit involved in the image processing method provided in this embodiment, reference can be made to Figures 1 to 3 the corresponding description in, which will not be elaborated here.
[0077] Based on the method as described above in Figures 1 to 3 Accordingly, this embodiment further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method as described above in Figures 1 to 3 is implemented.
[0078] Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of various implementation scenarios of the present application.
[0079] Based on the method as described above in Figures 1 to 3 and the virtual device embodiment as described above in Figure 4 In order to achieve the above object, this embodiment of the present application further provides an electronic device, such as intelligent terminals such as smart phones, tablet computers, drones, and intelligent robots. The device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the method as described above in Figures 1 to 3 is implemented.
[0080] Optionally, the above-mentioned physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.
[0081] Those skilled in the art can understand that the above-mentioned physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or arrange different components.
[0082] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the above-mentioned physical devices, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0083] Based on the method as Figures 1 to 3 described above, and Figure 4 the virtual device embodiment as Figures 1 to 3 described above, this embodiment also provides a chip, including one or more interface circuits and one or more processors; the interface circuit is used to receive a signal from the memory of the electronic device and send the signal to the processor, and the signal includes computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device is caused to execute the method as
[0084] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware. By applying the solution of this embodiment, compared with the current existing technologies, on the one hand, by adopting the same resolution in the main feature layer, the problem of feature loss of small-size targets in the image during the upsampling and downsampling processes is avoided; on the other hand, by setting corresponding weight values for different categories of targets for model optimization, the problem of uneven samples caused by the small proportion of small-size targets in the training set is avoided. Thus, more features of small-size targets are retained, and the recognition accuracy of different categories is improved.
[0085] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element.
[0086] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments described herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. An image processing method, characterized in that, Including: Obtain the image to be processed as the input of the image recognition model; Using the feature extraction network of the image recognition model, perform feature extraction on the image to be processed in different feature layers by means of upsampling and downsampling to obtain the output results of different feature layers; the convolutional layers in the main feature layer of the feature extraction network have the same resolution, and the image recognition model is trained and optimized using a loss function with weight values corresponding to different category labels set; Based on the output results of different feature layers, perform fusion processing to obtain the output image.
2. The method according to claim 1, characterized in that, The training steps of the image recognition model include: Obtain an image dataset labeled with category labels; the number and types of the category labels are preset in advance, and the number of the category labels is greater than or equal to two; Divide the image dataset into a training set and a test set according to a preset ratio, and train and test the image recognition model; Wherein, the images in the training set and the test set do not overlap with each other.
3. The method according to claim 2, wherein The training steps of the image recognition model further include: Perform pixel category statistics on the images in the image dataset to obtain the number of pixels corresponding to each category label; According to the number of pixels corresponding to each category label, set the weight value corresponding to each category label for the loss function of the image recognition model; Use the loss function with weight values corresponding to different category labels set to train and optimize the image recognition model.
4. The method according to claim 3, characterized in that The performing pixel category statistics on the image dataset to obtain the number of pixels corresponding to each category label includes: Based on the category label, perform category annotation on each pixel in all images in units of pixels to obtain the number of pixels corresponding to each category label.
5. The method according to claim 3 or 4, characterized in that, The setting the weight value corresponding to each category label for the loss function of the image recognition model according to the number of pixels corresponding to each category label includes: Calculate the weight value corresponding to each category label according to the proportion of the number of pixels corresponding to each category label in the sum of the total pixels of all images; Use the weight value corresponding to each category label to set the loss function.
6. The method according to claim 5, characterized in that, The calculating the weight value corresponding to each category label according to the proportion of the number of pixels corresponding to each category label in the sum of the total pixels of all images includes: Perform a reciprocal processing on the proportion of the number of pixels corresponding to the category label of the target category in the sum of the total pixels of all images to obtain the weight value corresponding to the category label of the target category.
7. An image processing apparatus, characterized in that, Including: An acquisition unit configured to obtain the image to be processed as the input of the image recognition model; An extraction unit configured to use the feature extraction network of the image recognition model to perform feature extraction on the image to be processed in different feature layers by means of upsampling and downsampling to obtain the output results of different feature layers; the convolutional layers in the main feature layer of the feature extraction network have the same resolution, and the image recognition model is trained and optimized using a loss function with weight values corresponding to different category labels set; A processing unit configured to perform fusion processing based on the output results of different feature layers to obtain the output image.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
9. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
10. A chip, characterized in that, Comprising one or more interface circuits and one or more processors; the interface circuit is configured to receive a signal from a memory of an electronic device and send the signal to the processor, the signal including computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device is caused to execute the method according to any one of claims 1 to 6.