Image processing method and device, electronic equipment and storage medium

By processing the image set using a target object segmentation model and employing a deep learning network for feature extraction and bounding box recognition, the problem of difficult segmentation of small target objects is solved, achieving fast and high-quality intelligent annotation.

CN117274588BActive Publication Date: 2026-01-20WEICHAI POWER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311185538.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-14
Publication Date
2026-01-20
Estimated Expiration
2043-09-14

AI Technical Summary

Technical Problem

Existing target recognition technologies struggle to accurately segment small objects, especially when there are many objects and few pixels, resulting in low recognition accuracy and impacting user experience.

Method used

The image set is processed using a target object segmentation model, including a region generation sub-model and a target object selection sub-model. Feature extraction and bounding box recognition are performed through a deep learning network to generate high-quality recognition results.

Benefits of technology

It enables fast, high-quality intelligent annotation of small target objects, improving annotation speed and quality, and solving the problem of difficult target object segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117274588B_ABST
    Figure CN117274588B_ABST
Patent Text Reader

Abstract

An image processing method and device, electronic equipment and storage medium are disclosed. The method comprises: obtaining a to-be-processed image comprising a plurality of to-be-identified objects, and constructing a model input image set based on the to-be-processed image, wherein the to-be-identified objects are small target objects; processing the model input image set based on a target object segmentation model to determine an identification result corresponding to the to-be-processed image, wherein the target object segmentation model comprises at least one target sub-model, and the at least one target sub-model comprises a region generation sub-model and a target object screening sub-model; and labeling the to-be-identified objects in the to-be-processed image based on the identification result. The technical solution of the embodiment realizes fast and high-quality intelligent labeling of small target objects, greatly speeds up the labeling speed, and improves the labeling quality.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, and in particular to an image processing method and device, electronic equipment and a storage medium. BACKGROUND

[0002] At present, related researches on image target recognition using artificial intelligence (AI) have gradually been carried out, and through this way, the user's demand for target object recognition is met.

[0003] The existing target recognition technology has the problems of only segmenting the target, only positioning the target, or the single target object to be detected occupies few pixels, and the number of target objects is large, and the target objects may be clustered, which makes it difficult to accurately segment the target objects, and thus the target recognition accuracy is low, and the recognition result is different from the user's expected recognition result, affecting the user experience. SUMMARY

[0004] The present application provides an image processing method, device, electronic equipment and storage medium to realize fast and high-quality intelligent labeling of small target objects, greatly speeding up the labeling speed and improving the labeling quality.

[0005] According to an aspect of the present application, an image processing method is provided, which comprises:

[0006] An image processing method is provided, which comprises:

[0007] Processing the model input image set based on a target object segmentation model to determine a recognition result corresponding to the to-be-processed image, wherein the target object segmentation model comprises at least one target sub-model, and the at least one target sub-model comprises a region generation sub-model and a target object screening sub-model;

[0008] Labeling the to-be-recognized objects in the to-be-processed image based on the recognition result.

[0009] According to another aspect of the present application, an image processing device is provided, which comprises:

[0010] An image acquisition module is configured to acquire a to-be-processed image comprising a plurality of to-be-recognized objects, and to construct a model input image set based on the to-be-processed image, wherein the to-be-recognized objects are small target objects;

[0011] The recognition result determination module is configured to determine a recognition result corresponding to the to-be-recognized object by processing the model input image set based on a target object segmentation model, wherein the target object segmentation model comprises at least one target sub-model, and the at least one target sub-model comprises a region generation sub-model and a target object screening sub-model.

[0012] The object labeling module is configured to label the to-be-recognized object in the to-be-processed image based on the recognition result.

[0013] According to another aspect of the present application, an electronic device is provided, which comprises:

[0014] at least one processor; and

[0015] a memory in communication with the at least one processor; wherein

[0016] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the image processing method according to any one of the embodiments of the present application.

[0017] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to execute the image processing method according to any one of the embodiments of the present application.

[0018] The technical solution of the embodiments of the present application comprises the following steps: obtaining a to-be-processed image comprising a plurality of to-be-recognized objects, constructing a model input image set based on the to-be-processed image, further processing the model input image set based on a target object segmentation model to determine a recognition result corresponding to the to-be-processed image, and finally labeling the to-be-recognized object in the to-be-processed image based on the recognition result. The embodiments of the present application solve the problem that it is difficult to accurately segment a target object when the number of target objects is large and the number of pixels occupied by a single target object is small in related technologies, and achieve fast and high-quality intelligent labeling of small target objects, greatly accelerating the labeling speed and improving the labeling quality.

[0019] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0020] In order to make the technical solution in the embodiments of the present application clearer, the accompanying drawings needed in the embodiments will be briefly introduced. Obviously, the accompanying drawings in the following description only represent some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort.

[0021] Figure 1 is a flow chart of an image processing method according to the first embodiment of the present application;

[0022] Figure 2 is a flow chart of an image processing method according to the second embodiment of the present application;

[0023] Figure 3 is a flow chart of an image processing method according to the third embodiment of the present application;

[0024] Figure 4 is a structural schematic diagram of an image processing device according to the fourth embodiment of the present application;

[0025] Figure 5 is a structural schematic diagram of an electronic device implementing the image processing method according to the present application. DETAILED DESCRIPTION

[0026] In order to make the technical solution in the embodiments of the present application clearer, the accompanying drawings needed in the embodiments will be briefly introduced. Obviously, the accompanying drawings in the following description only represent some embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort.

[0027] It should be noted that the terms "first", "second", and the like in the description and claims of the present application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not necessarily have to include only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.

[0028] Embodiment One

[0029] Figure 1is a flowchart of an image processing method provided by an embodiment of the present application. The embodiment can be applied to the case of segmenting and labeling small target objects included in an image. The method can be executed by an image processing device, which can be implemented in the form of hardware and / or software, and can be configured in a terminal and / or a server. As shown in FIG. 1, Figure 1 The method comprises the following steps.

[0030] In S110, a to-be-processed image including a to-be-identified object is acquired, and a model input image set is constructed based on the to-be-processed image.

[0031] The to-be-processed image can be an image received by a server or a client and obtained by a user in real time through a camera, or an image stored in a related database and called by the server or the client. The to-be-processed image can include one or more to-be-identified objects. The object in the image is the to-be-identified object. Based on the neural network model of the embodiment, the identification result corresponding to the to-be-identified object can be determined. The to-be-identified object can be a small target object. The small target object can be an object with a small number of pixels in the image, or an object with a small actual volume. The to-be-identified object can be any object. Optionally, the to-be-identified object can be a corn kernel. Generally, when the number of to-be-identified objects included in the to-be-processed image is large, the pixel overlap rate between any two to-be-identified objects is high. The model input image set can be an image set including multiple images, and one of the images can be the to-be-processed image.

[0032] It should be noted that in a specific application scenario, the to-be-processed image can be acquired in real time or periodically. Alternatively, when an image uploaded by a user is detected, the image can be acquired as the to-be-processed image. The embodiment does not make a specific limitation in this regard.

[0033] For example, when a user takes a picture of certain objects and uploads the picture to the server or the client, the picture is the to-be-processed picture. Meanwhile, the server or the client can identify the to-be-processed picture based on a related algorithm, so as to determine that the objects in the to-be-processed picture are the to-be-identified objects. The number of to-be-identified objects in the to-be-processed picture can be one or more, which is not limited in the embodiment. It should be noted that, when the number of to-be-identified objects is more than one, all the objects in the picture can be regarded as to-be-identified objects; when the number of to-be-identified objects is one, the to-be-identified object in the picture can be labeled in advance, and the labeled picture can be uploaded to the server, so that the server stores the characteristic attributes of the to-be-identified object. After a plurality of to-be-processed pictures are obtained, when it is detected that at least one to-be-processed picture includes the characteristic attributes of the to-be-identified object, the object in the current to-be-processed picture can be determined as the to-be-identified object, and the to-be-processed picture can be processed.

[0034] Further, after the to-be-processed picture is obtained, the to-be-processed picture can be processed. Further, a model input picture set including the to-be-processed picture can be constructed.

[0035] Optionally, the model input picture set is constructed based on the to-be-processed picture, including: performing image enlargement processing on the to-be-processed picture based on a first preset ratio to obtain a first reference picture; performing image reduction processing on the to-be-processed picture based on a second preset ratio to obtain a second reference picture; and constructing the model input picture set based on the to-be-processed picture, the first reference picture and the second reference picture.

[0036] The first preset ratio can be understood as a ratio of image enlargement processing determined in advance. The first preset ratio can be any value. In actual application, the first preset ratio can be determined based on image processing requirements. The first reference picture can be an image obtained by enlarging the image size of the to-be-processed picture based on the first preset ratio. The second preset ratio can be understood as a ratio of image reduction processing determined in advance. The second preset ratio can be any value. In actual application, the second preset ratio can be determined based on image processing requirements. The second reference picture can be an image obtained by reducing the image size of the to-be-processed picture based on the second preset ratio.

[0037] In the embodiment, the first preset ratio and the second preset ratio can be obtained in multiple ways. The multiple ways of obtaining the first preset ratio and the second preset ratio will be described in detail below.

[0038] Optionally, the preset ratio can be obtained by calling the preset ratio from a related database in the case of obtaining the image to be processed. In actual application, a plurality of categories of objects to be recognized and image processing requirements of each category of object to be recognized can be determined in advance. Then, the image magnification ratio for image magnification and the image reduction ratio for image reduction can be set according to the image processing requirements of each category of object to be recognized, and a plurality of image magnification ratios and a plurality of image reduction ratios can be obtained. After that, the identification of the object to be recognized can be associated with the corresponding image magnification ratio and image reduction ratio and stored in the related database. Further, after obtaining the image to be processed, the object to be recognized included in the image to be recognized can be determined. Then, the associated image magnification ratio and image reduction ratio can be called from the related database according to the object identification of the determined object to be recognized. The called image magnification ratio can be used as the first preset ratio, and the called image reduction ratio can be used as the second preset ratio.

[0039] Optionally, the preset ratio can also be obtained by determining the first preset ratio and the second preset ratio in response to an editing trigger operation on the preset ratio editing item. In actual application, the preset ratio editing item can be set in advance in the display interface of the related software. Then, when it is detected that the user inputs a click operation on the preset ratio editing item through the input device or the touch point, or when it is detected that the user pauses on the preset ratio editing item for a preset time through the input device or the touch point, it can be determined that the editing start trigger operation on the preset ratio editing item is detected. Then, the trigger operation can be responded to by adjusting the preset ratio editing item to an editable state. After that, when it is detected that the user inputs an editing completion trigger operation on the preset ratio editing item through the input device, the operation can be responded to, and the value displayed in the preset ratio editing item can be used as the first preset ratio and the second preset ratio.

[0040] In actual application, after obtaining the image to be processed, the first preset ratio and the second preset ratio corresponding to the image to be processed can be obtained. Then, the image to be processed can be magnified according to the first preset ratio, and the magnified image to be processed can be used as the first reference image. After that, the image to be processed can be reduced according to the second preset ratio, and the reduced image to be processed can be used as the second reference image. Further, the image to be processed, the first reference image and the second reference image can be combined together to obtain a model input image set including three images.

[0041] S120, processing the model input image set based on the target object segmentation model to determine a recognition result corresponding to the object to be recognized.

[0042] In the embodiment, after the server or the client determines the model input image set, the model input image set can be input into the pre-trained target object segmentation model to process each image included in the model input image set based on the target image segmentation model. The target object segmentation model can include a deep learning network model of multiple sub-models. The target object segmentation model can include at least one target sub-model. The at least one target sub-model includes a region generation sub-model and a target object screening sub-model.

[0043] The region generation sub-model can be a sub-model including a region generation network and a first convolutional network and a second convolutional network connected to the region generation network respectively. The region generation network (Regions of Interest Net, ROI-Net) can be used to solve the problem of region mismatch caused by twice quantization in the ROI Pooling operation. The region generation network is a neural network model that realizes feature extraction through the ROI-Align operation. The region generation network can also be used to generate a large number of candidate boxes including image features based on the model input image. For example, the region generation network can be a small shallow fully convolutional network, for example, a fully convolutional network including 4 convolutional modules, and the convolution kernel size of each convolutional module is 3x3. The first convolutional network can be a detection head connected to the region generation network, which can be used to classify the category of the object to be recognized in the image. The second convolutional network can be another detection head connected to the region generation network, which can be used to determine the bounding box information of the candidate bounding box.

[0044] The target object screening sub-model can be a sub-model including a backbone feature extraction network, a first fully connected network connected to the backbone feature extraction network, a second fully connected network, a feature fusion network, and a third convolutional network connected to the feature fusion network. The backbone feature extraction network can be a neural network for feature extraction of an image. For example, the backbone feature extraction network can be a DarkNet-53 network, which can include five convolutional modules for five times of down-sampling. The first fully connected network can be a detection head connected to the backbone feature extraction network, which is used to perform a class classification task of the object to be recognized. The second fully connected network can be a detection head connected to the backbone feature extraction network, which is used to perform a task of determining the boundary box information of the target boundary box. The feature fusion network can be a network for up-sampling the feature image and performing feature fusion with the corresponding feature after down-sampling. It should be noted that the feature fusion network can be a network corresponding to the backbone feature network, i.e., the network structure of the feature fusion network matches the network structure of the backbone feature network. For example, when the backbone feature network includes five convolutional modules for down-sampling, the feature fusion network includes five convolutional modules for up-sampling. The third convolutional network can be a detection head connected to the feature fusion network, which can be used to perform a mask prediction task. For example, the third convolutional network can be a convolutional network with a 1x1 convolution kernel.

[0045] In the embodiment, after the model input image is input into the target object segmentation model for processing, the model can output the recognition result corresponding to the boundary box included in the image to be processed. The recognition result can be information representing the attributes of the object to be recognized included in the boundary box, and can be information judging the object class of the object to be recognized. For example, when the object to be recognized in the input image to be processed is a corn kernel, the model can output the current class (complete, incomplete, not included, or unrecognizable) of the corn kernel. Not included can be understood as that the boundary box does not include a corn kernel.

[0046] In actual application, since the target object segmentation model includes multiple target sub-models, based on the target object segmentation model, the model input image set can be processed by the multiple target sub-models in the model in sequence. Thus, the recognition result corresponding to the object to be recognized included in the image to be processed can be output.

[0047] Optionally, based on the target object segmentation model, the model input image set is processed to determine the recognition result corresponding to the object to be recognized, including: based on the region generation sub-model and the target object screening sub-model in the target object segmentation model, the model input image set is processed to obtain the recognition result.

[0048] S130, label the to-be-recognized object in the to-be-processed image based on the recognition result.

[0049] In this embodiment, after obtaining the recognition result, the to-be-recognized object in the to-be-processed image can be labeled according to the recognition result. The recognition result output by the model can include mask data of the to-be-recognized object, an object category, and boundary box information corresponding to the to-be-recognized object.

[0050] In actual application, after obtaining the recognition result output by the target object segmentation model, the obtained recognition result can be input into image processing software. Further, the mask data of the to-be-recognized object in the recognition result can be processed according to the image processing software to obtain the coordinate information of the mask contour of the to-be-recognized object. Further, the obtained mask contour coordinate information, the object category of the to-be-recognized object, and the boundary box information corresponding to the to-be-recognized object can be written into a file in a preset format to obtain an object label file. Then, the determined object label file can be read and parsed by using a preset labeling software. Thus, the to-be-recognized object included in the to-be-processed image can be labeled according to the parsed object label file. The image processing software can be any software, which can be OpenCV optionally. The preset format can be any format, which can be JSON optionally. The preset labeling software can be any software, which can be Labelme optionally.

[0051] The technical scheme of the embodiment of the present application solves the problem that in the related art, when a single target object needs to be detected, the number of target objects is large, and the number of pixels occupied by the single target object is small, which leads to difficulty in accurately segmenting the target object, and achieves fast and high-quality intelligent labeling of small target objects, greatly speeds up the labeling speed, and improves the labeling quality.

[0052] Embodiment two

[0053] Figure 2 is a flowchart of an image processing method provided by the second embodiment of the present application, which is further refined to S120 on the basis of the foregoing embodiment. The same or corresponding technical terms as in the foregoing embodiment are not described here again.

[0054] As shown in Figure 2 , the method comprises:

[0055] S210, obtain a to-be-processed image including a plurality of to-be-recognized objects, and construct a model input image set based on the to-be-processed image.

[0056] S220, performing feature extraction on the model input image set based on the region generation sub-model to determine at least one candidate recognition result set.

[0057] The candidate recognition result set can be a set including at least one candidate recognition result. The number of candidate recognition results included in each candidate recognition result set is consistent with the number of images included in each model input image set. For example, when the number of images included in the model input image set is 3, each candidate recognition result set obtained after the region generation sub-model processes the model input image set can include 3 candidate recognition results. The candidate recognition result can include a candidate region feature, an identification category corresponding to the candidate region feature, and bounding box information. The region generation sub-model includes a region generation network, and a first convolutional network and a second convolutional network connected to the region generation network, respectively. The region generation network includes at least one first convolutional module.

[0058] In actual application, after obtaining the model input image set, the model input image set can be input into the target object segmentation model, so as to perform candidate region feature extraction on the images in the model input image set based on the region generation sub-model in the target object segmentation model. Furthermore, for each image in the model input image set, candidate regions can be extracted from the image based on the neural network included in the region generation sub-model, so as to determine specific feature information corresponding to the to-be-recognized object included in the image. Thus, at least one candidate recognition result can be obtained. When the processing of each image in the model input image set based on the region generation sub-model is completed, at least one candidate recognition result set can be obtained.

[0059] Optionally, performing feature extraction on the model input image set based on the region generation sub-model to determine at least one candidate recognition result set includes: for each model input image in the model input image set, performing processing on the model input image based on the at least one first convolutional module according to a preset sliding window size corresponding to the model input object, to obtain at least one candidate region feature corresponding to the model input image; mapping the at least one candidate region feature corresponding to the first reference image and the at least one candidate region feature corresponding to the second reference image to the to-be-processed image, to obtain at least one candidate region feature corresponding to the model input image set; performing processing on each candidate region feature based on the first convolutional network, to obtain an identification category corresponding to the candidate region feature; performing processing on each candidate region feature based on the second convolutional network, to obtain bounding box information corresponding to the candidate region feature; and taking the at least one candidate region feature, the identification category corresponding to each candidate region feature, and the bounding box information corresponding to each candidate region feature as at least one candidate recognition result corresponding to the model input image set.

[0060] It should be noted that the first convolutional module in the region generation sub-model can be one or more, and the convolution kernel of each first convolutional module can be any value, and the present embodiment does not make specific limitation on this. The first convolutional module can be a module constructed based on a convolutional neural network. The convolutional neural network is a kind of feedforward neural network containing convolution calculation and having a deep structure, and is one of the representative algorithms of deep learning. At the same time, the convolutional neural network has a representation learning ability and can perform translation invariant classification on input information according to its hierarchical structure, and the present embodiment will not make specific description here.

[0061] In actual application, the model input image is input into the region generation sub-model, and a series of convolutional neural networks are used to perform down-sampling processing on the model input image. Further, a sliding window is randomly run in each pixel point in the feature image output by the last layer. For each sliding window, the center of the current sliding window can be taken as an anchor point, and a box corresponding to a set of scale and length ratios. Therefore, each sliding window can correspond to K anchor points, and accordingly, for a feature image with a resolution of HxW, HxWxK anchor points can be corresponded. Since the feature image is obtained by down-sampling processing of the model input image based on different first convolutional modules, the feature image and the model input image have a certain mapping relationship. And since the resolutions of the feature image and the model input image are different, one anchor point in the feature image corresponds to one candidate bounding box in the model input image. After mapping all the anchor points determined in the feature image to the model input image, there will be a part of the candidate bounding boxes that do not meet the candidate region condition, for example, the box exceeds the image boundary or the box size is small, etc. Some post-processing is performed on the candidate bounding boxes in the model input image, and the candidate bounding boxes that do not meet the condition are deleted, so that more reliable candidate bounding boxes can be obtained, and the image determined based on the candidate bounding boxes after screening can be taken as the candidate region feature.

[0062] In the embodiment, the preset sliding window size can be understood as a size of a sliding window preset to slide in the image. The preset sliding window size can be associated with the image size of the model input image, that is, for model input images of different image sizes, the corresponding preset sliding window sizes are also different. It should be noted that the preset sliding window size corresponding to the model input image corresponds to the ratio of the image size between the model input image and the to-be-processed image. Specifically, in the case that the preset sliding window size corresponding to the to-be-processed image is a first sliding window size, the preset sliding window size corresponding to the first reference image obtained after the to-be-processed image is enlarged based on the first preset ratio can be the product between the first sliding window size and the first preset ratio. For example, in the case that the image size of the to-be-processed image is 544x544, the corresponding preset sliding window size can be 16x16. Further, assuming that the first preset ratio is 1.2, the image size of the first reference image obtained after the to-be-processed image is enlarged based on the first preset ratio is 652.8x652.8, and the corresponding preset sliding window size is 19.2x19.2.

[0063] In actual application, for each model input image in the model input image set, the model input image can be processed according to the preset sliding window size corresponding to the model input image by at least one first convolutional module in the region generation sub-model, that is, a sliding window based on the preset sliding window size corresponding to the model input image slides in the model input image. Further, in the case that the model input image set includes model input images other than the to-be-processed image, the at least one candidate region feature corresponding to the model input image in the model input image set other than the to-be-processed image can be mapped to the to-be-processed image, that is, all the at least one candidate region features corresponding to all the model input images in the model input image set are displayed in the to-be-processed image, and the to-be-processed image displaying sliding windows of different preset sliding window sizes can be obtained. Further, all the candidate region features displayed in the to-be-processed image can be taken as the at least one candidate region feature corresponding to the model input image set.

[0064] Further, in order to preliminarily classify and identify the to-be-identified object corresponding to the candidate region feature and preliminarily regress and position the candidate bounding box, after obtaining the at least one candidate region feature, the candidate region feature can be respectively input into a first convolutional network and a second convolutional network connected with the region generation network. Then, the candidate region feature can be processed based on the first convolutional network to determine the identification category of the to-be-identified object corresponding to the candidate region feature, and the candidate region feature can be processed based on the second convolutional network to determine the bounding box information of the candidate bounding box corresponding to the candidate region feature. The identification category of the to-be-identified object can be the display state category of the to-be-identified object in the corresponding model input image. For example, when the to-be-identified object is a corn kernel, the corresponding identification category can be complete, defective, impurity, not included, or unidentifiable, etc. The bounding box information can be understood as the predicted bounding box information obtained after positioning the display position of the to-be-identified object in the corresponding model input image. The bounding box information includes the center point coordinates of the bounding box and the width and height information of the bounding box. For example, when the to-be-identified object in the model input image is a corn kernel, the output bounding box information is the information corresponding to the bounding box used to mark the corn kernel in the model input image.

[0065] Further, the at least one candidate region feature, the identification category corresponding to each candidate region feature, and the bounding box information corresponding to each candidate region feature can be used as at least one candidate identification result corresponding to the model input image set.

[0066] S230, performing region screening and merging processing on the at least one candidate region feature to determine at least one to-be-processed region feature.

[0067] In this embodiment, after obtaining the at least one candidate region feature corresponding to the model input image set, for all candidate region features displayed in the to-be-processed image, there can be a high region feature coincidence rate between some candidate region features. Therefore, the candidate region features can be subjected to region screening and merging processing. Thus, at least one to-be-processed region feature can be obtained.

[0068] In actual application, the obtained all candidate region features can be subjected to region screening and merging processing according to a preset region screening and merging algorithm deployed in the target object segmentation model. Then, the candidate region features with a region coincidence rate reaching a preset value can be merged together. Then, the region features obtained after screening and merging can be used as to-be-processed region features. Thus, at least one to-be-processed region feature can be obtained. The preset region screening and merging algorithm can be any algorithm that can realize region screening and merging. Optionally, the preset region screening and merging algorithm can be a Non-Maximum Suppression (NMS) algorithm.

[0069] S240, processing each to-be-processed region feature based on a preset size to obtain at least one to-be-input region feature.

[0070] The preset size can be understood as a size predefined for limiting the size of the input image input into the target object screening sub-model. It should be noted that the preset size can be greater than any preset sliding window size corresponding to a model input image. For example, in the case where the preset sliding window size corresponding to the to-be-processed image is 16x16, the preset size can be 32x32.

[0071] In actual application, the obtained to-be-processed region feature is a sliding window including image features generated based on the to-be-processed image, and in the case where the to-be-processed region feature includes a to-be-recognized object, the to-be-recognized object included therein is a small target object with small image size. Therefore, in order to perform more accurate object recognition and classification on the to-be-processed region feature, each to-be-processed region feature can be processed in size according to the preset size. Further, the to-be-processed region feature processed in size can be used as a to-be-input region feature.

[0072] S250, inputting the at least one to-be-input region feature into the target object screening sub-model to obtain a recognition result corresponding to the to-be-processed image.

[0073] In this embodiment, after obtaining the at least one to-be-input region feature, the at least one to-be-input region feature can be input into the target object screening sub-model to process the at least one to-be-input region feature based on the target object screening sub-model. Thus, a recognition result corresponding to the to-be-processed image can be obtained.

[0074] In this embodiment, the target object screening sub-model can include a sub-model of a backbone feature extraction network, a first fully connected network connected to the backbone feature extraction network, a second fully connected network, a feature fusion network, and a third convolutional network connected to the feature fusion network. In actual application, after inputting the at least one to-be-input region feature into the target object screening sub-model, the at least one to-be-input region feature is processed based on the backbone feature extraction network, and the feature image obtained after processing is input into the first fully connected network, the second fully connected network, and the feature fusion network, respectively. Thus, a recognition result corresponding to the to-be-processed image can be obtained.

[0075] Optionally, the at least one to-be-input region feature is input into the target object screening sub-model to obtain an identification result corresponding to the to-be-processed image, including: performing feature extraction processing on the at least one to-be-input region feature based on at least one first convolution module included in the backbone feature extraction network to obtain at least one first feature; processing the at least one first feature based on a first full connection network to obtain an identification category corresponding to each first feature; processing the at least one first feature based on a second full connection network to obtain boundary box information corresponding to each first feature; processing the at least one first feature and the model output of each first convolution module in the backbone feature extraction network based on at least one second convolution module included in the feature fusion network to obtain at least one second feature; processing the at least one second feature based on a third convolution network to obtain a mask image corresponding to each second feature; and determining the identification result based on the identification category, the boundary box information, and the mask image.

[0076] The backbone feature extraction network can be a network for performing feature information extraction, and is a network for performing down-sampling processing on model input. The backbone feature extraction network can include a plurality of first convolution modules with different convolution kernels. The backbone feature extraction network can be any feature extraction network. Optionally, the backbone feature extraction network can be Darknet-53. The first full connection network can be understood as a detection head connected to the backbone feature extraction network, and the detection head is used to perform the task of identifying the object category of the to-be-classified object. The second full connection network can be understood as another detection head connected to the backbone feature extraction network, and the detection head is used to perform the task of determining the boundary box information of the feature image. The feature fusion network can be understood as a network for performing up-sampling processing on model input. The feature fusion network can include a plurality of second convolution modules with different convolution kernels. It should be noted that the number of second convolution modules included in the feature fusion network is consistent with the number of first convolution modules included in the backbone feature extraction network. It should also be noted that, in the case that the arrangement direction of the second convolution modules included in the feature fusion network is consistent with the data output direction of the network, and the arrangement direction of the first convolution modules included in the backbone feature extraction network is consistent with the data output direction of the network, the convolution kernel of the last first convolution module is the same as the convolution kernel of the first second convolution module, the convolution kernel of the second last first convolution module is the same as the convolution kernel of the second second convolution module, the convolution kernel of the third last first convolution module is the same as the convolution kernel of the third second convolution module, and so on, and the convolution kernel of the first first convolution module is the same as the convolution kernel of the last second convolution module. The third convolution network can be understood as a detection head connected to the feature fusion network, and the detection head is used to perform the task of generating a mask image.

[0077] In actual application, after obtaining the at least one to-be-input region feature, the at least one to-be-input region feature can be input into the target object screening sub-model. Then, the at least one to-be-input region feature can be sequentially subjected to feature extraction based on at least one first convolution module included in a backbone feature extraction network in the target object screening sub-model, and the to-be-input region feature after feature information extraction can be taken as a first feature. Then, the at least one first feature can be obtained. After that, the at least one first feature can be input into a first full connection network, a second full connection network and a feature fusion network respectively. Then, the at least one first feature can be processed based on the first full connection network to determine an identification category of the to-be-identified object included in each first feature, and the identification category corresponding to each first feature can be obtained. Meanwhile, the at least one first feature can be processed based on the second full connection network to determine the bounding box information corresponding to each first feature, and the bounding box information corresponding to each first feature can be obtained. After that, the at least one first feature and the model output of each first convolution module in the backbone feature extraction network can be processed based on at least one second convolution module included in the feature fusion network. Thus, at least one second feature can be obtained.

[0078] Optionally, the processing of the at least one first feature and the model output of each first convolution module in the backbone feature extraction network based on at least one second convolution module included in the feature fusion network to obtain the at least one second feature includes: taking the at least one first feature as the model output of the last first convolution module, and taking the model output as the model input of the first second convolution module; for the second convolution modules other than the first second convolution module in the feature fusion network, adding the model output of the last convolution module of the current second convolution module and the model output of the first convolution module corresponding to the current second convolution module to serve as the model input of the current second convolution module, until the current second convolution module is the last second convolution module, and taking the model output of the last second convolution module as the at least one second feature.

[0079] The first convolution module corresponding to the current second convolution module is a first convolution module in the backbone feature extraction network satisfying a preset standard. The preset standard can be understood as a standard preset for screening the first convolution module in the backbone feature extraction network. Optionally, the preset standard can be a first convolution module with the same convolution kernel as the current second convolution module in the backbone feature extraction network.

[0080] In practical applications, for the feature fusion network, at least one first feature can be input into the first second convolutional module, and the at least one first feature is up-sampled based on the second convolutional module, and the first feature after the up-sampling processing can be taken as the model output of the first second convolutional module. Then, for other second convolutional modules in the feature fusion network except the first second convolutional module, the model output of the first second convolutional module and the model output of the second last first convolutional module in the backbone feature extraction network are added, and the result after the addition is taken as the model input of the second second convolutional module. Further, the model output of the second second convolutional module and the model output of the third last first convolutional module in the backbone feature extraction network are added, and the result after the addition is taken as the model input of the third second convolutional module, and so on. The model output of the second last second convolutional module and the model output of the first first convolutional module in the backbone feature extraction network are added, and the result after the addition is taken as the model input of the last second convolutional module. Then, the model output of the last second convolutional module can be taken as at least one second feature.

[0081] It should be noted that the backbone feature extraction network and the feature fusion network can be understood as a combination of networks that perform down-sampling processing on feature information and then perform up-sampling processing on the down-sampled feature information. Moreover, the dimension of the down-sampling processing on the feature information by each first convolutional module in the backbone feature extraction network is the same as the dimension of the up-sampling processing on the feature information by each second convolutional module in the feature fusion network. Therefore, the dimension of the second feature is the same as the dimension of the region feature to be input.

[0082] Further, at least one second feature can be input into a third convolutional network connected with the feature fusion network, so as to process each second feature based on the third convolutional network. Thus, a mask map corresponding to each second feature can be obtained. Each mask map can be an image in which the pixel value of a pixel point corresponding to the object to be recognized in the second feature is adjusted to a first pixel value, and the pixel value of other pixel points in the second feature is adjusted to a second pixel value. The first pixel value can be any value, which can be 0 or 1. The second pixel value can be any value, which can be 0 or 1.

[0083] Finally, the obtained mask map, recognition category and bounding box information can be taken as the recognition result corresponding to the image to be processed.

[0084] S260, labeling the object to be recognized in the image to be processed based on the recognition result.

[0085] The technical scheme of the embodiment of the present application, by acquiring a to-be-processed image including a plurality of to-be-identified objects, and constructing a model input image set based on the to-be-processed image, further, processing the model input image set based on a target object segmentation model, determining the identification result corresponding to the to-be-processed image, and finally, labeling the to-be-identified objects in the to-be-processed image based on the identification result, solves the problem that in the related art, when the number of target objects is large and the number of pixels occupied by a single target object is small, it is difficult to accurately segment the target object, and realizes fast and high-quality intelligent labeling of small target objects, greatly speeding up the labeling speed and improving the labeling quality.

[0086] Embodiment three

[0087] Figure 3 is a flowchart of an image processing method provided by the third embodiment of the present application, based on the foregoing embodiments, the target object segmentation model can also be pre-trained to identify and process the model input image set based on the target object segmentation model, so as to determine the identification result corresponding to the to-be-processed image. Wherein, the same or corresponding technical terms as the above embodiments are not repeated here.

[0088] As shown in Figure 3 , the method comprises:

[0089] S310, acquiring a plurality of training samples.

[0090] Among them, the training sample includes: a training sample image, the training sample image includes a to-be-identified object, a theoretical result corresponding to the training sample image, the theoretical result includes a theoretical category, theoretical boundary box information and a theoretical mask image.

[0091] Among them, the training sample image can be an image captured by a camera device, or an image reconstructed by an image reconstruction model, or an image pre-stored from a storage space. At the same time, the image includes one or more objects, which can be used as to-be-identified objects. The theoretical mask image can be a mask image obtained by pre-processing the training sample image with a binary mask. The theoretical category can be the category to which each object in the training sample image belongs. The theoretical boundary box information can be the real boundary box information corresponding to the to-be-identified object in the training sample image.

[0092] Specifically, before training the object segmentation model to be trained, a plurality of training samples need to be obtained first to train the model based on the training samples. In order to improve the accuracy of the model, as many and rich training samples as possible are obtained, a plurality of training sample images including the object to be recognized are obtained, and further, the training sample images are processed to obtain the theoretical mask image, the theoretical category corresponding to the training sample image, and the theoretical bounding box information corresponding to the object to be recognized, so as to construct rich training samples based on the above-mentioned manner.

[0093] S320, input the training sample into the object segmentation model to be trained to obtain an actual output result.

[0094] It should be noted that for each training sample, the training method of S320 can be used. Thus, the target object segmentation model can be obtained.

[0095] The model parameters in the object segmentation model to be trained are default values. The model parameters in the object segmentation model to be trained are corrected by the training sample to obtain the target object segmentation model. The actual output result includes an actual mask image, an actual category, and actual bounding box information. The actual mask image is the mask image output after the training sample image is input into the object segmentation model to be trained; the actual category is the recognition category output after the training sample image is input into the object segmentation model to be trained; and the actual bounding box information is the bounding box information corresponding to the object to be recognized output after the training sample image is input into the object segmentation model to be trained.

[0096] In this embodiment, the object segmentation model to be trained can include a region generation sub-model to be trained and an object screening sub-model to be trained. In actual application, after the training sample image is input into the object segmentation model to be trained, the training sample image is processed by the region generation sub-model to be trained to obtain at least one actual candidate recognition result. Further, the at least one actual candidate recognition result is processed according to the object screening sub-model to be trained. Thus, the actual output result corresponding to the training sample image can be obtained.

[0097] S330, according to the first loss function corresponding to the region generation sub-model to be trained and the second loss function corresponding to the object screening sub-model to be trained, loss processing is performed on the theoretical result and the actual output result.

[0098] The to-be-trained region generation sub-model and the to-be-trained object screening sub-model are models with initial parameters or default parameters. The first loss function includes a category loss function corresponding to the actual category and the theoretical category, and a bounding box loss function corresponding to the actual bounding box information and the theoretical bounding box information. The second loss function includes a mask loss function corresponding to the actual mask graph and the theoretical mask graph, a theoretical loss function corresponding to the actual category and the theoretical category, and a bounding box loss function corresponding to the actual bounding box information and the theoretical bounding box information. The category loss function can be any loss function, which can be a cross-entropy loss function. The bounding box loss function can be any loss function, which can be a Euclidean loss function. The mask loss function can be any loss function, which can be a cross-entropy loss function. Specifically, the first loss function can be obtained by adding the category loss function and the bounding box loss function. The second loss function can be obtained by adding the category loss function, the bounding box loss function, and the mask loss function.

[0099] S340, based on the loss value, the model parameters of the to-be-trained object segmentation model are corrected to obtain a target object segmentation model.

[0100] Generally, the model parameters of the to-be-trained object segmentation model are initial parameters or default parameters. When the to-be-trained object segmentation model is trained, each model parameter in the model can be corrected based on the output result of the to-be-trained object segmentation model, that is, the loss value of the to-be-trained object segmentation model is corrected to obtain a target object segmentation model. The loss value is the difference value between the actual output image and the theoretical output image.

[0101] Specifically, when the model parameters in the to-be-trained object segmentation model are corrected using the loss value, the convergence of the loss function can be used as the training target, such as whether the training error is less than the preset error, or whether the error change tends to be stable, or whether the current iteration number is equal to the preset number. If it is detected that the convergence condition is reached, such as the training error of the loss function is less than the preset error, or the error change tends to be stable, it indicates that the to-be-trained object segmentation model is trained, and at this time the iteration training can be stopped. If it is detected that the current convergence condition is not reached, other training samples can be further obtained to continue training the to-be-trained object segmentation model until the training error of the loss function is within the preset range. When the training error of the loss function converges, the to-be-trained object segmentation model that is trained is used as the target object segmentation model, that is, when the image including the to-be-recognized object is input into the target object segmentation model, the recognition result corresponding to the to-be-processed image can be accurately obtained.

[0102] S350, a to-be-processed image including a plurality of to-be-recognized objects is obtained, and a model input image set is constructed based on the to-be-processed image.

[0103] S360, processing the model input image set based on the target object segmentation model to determine the recognition result corresponding to the to-be-processed image.

[0104] S370, labeling the to-be-identified object in the to-be-processed image based on the recognition result.

[0105] The technical scheme of the embodiment of the present application, by acquiring a to-be-processed image including a plurality of to-be-identified objects, and constructing a model input image set based on the to-be-processed image, further processing the model input image set based on a target object segmentation model to determine the recognition result corresponding to the to-be-processed image, and finally labeling the to-be-identified object in the to-be-processed image based on the recognition result, solves the problem in related art that when the number of target objects to be detected is large and the number of pixels occupied by a single target object is small, it is difficult to accurately segment the target object, and realizes fast and high-quality intelligent labeling of small target objects, greatly speeding up the labeling speed and improving the labeling quality.

[0106] Embodiment Four

[0107] Figure 4 is a structural schematic diagram of an image processing device provided by Embodiment Four of the present application. As shown in the figure, the device comprises an image acquisition module 410, a recognition result determination module 420, and an object labeling module 430. Figure 4

[0108] The image acquisition module 410 is configured to acquire a to-be-processed image including a plurality of to-be-identified objects, and construct a model input image set based on the to-be-processed image, wherein the to-be-identified objects are small target objects. The recognition result determination module 420 is configured to process the model input image set based on a target object segmentation model to determine the recognition result corresponding to the to-be-identified objects, wherein the target object segmentation model includes at least one target sub-model, and the at least one target sub-model includes a region generation sub-model and a target object screening sub-model. The object labeling module 430 is configured to label the to-be-identified objects in the to-be-processed image based on the recognition result.

[0109] ​The technical scheme of the embodiment of the present application comprises the following steps: obtaining a to-be-processed image comprising a plurality of to-be-identified objects, constructing a model input image set based on the to-be-processed image, further processing the model input image set based on a target object segmentation model to determine an identification result corresponding to the to-be-processed image, and finally labeling the to-be-identified objects in the to-be-processed image based on the identification result. The problems in the related art, such as difficulty in accurately segmenting target objects when the number of target objects is large and the number of pixels occupied by a single target object to be detected is small, are solved, and fast and high-quality intelligent labeling of small target objects is achieved, which greatly accelerates the labeling speed and improves the labeling quality.

[0110] Optionally, the image acquisition module 410 comprises an image enlargement unit, an image reduction unit, and an image set construction unit.

[0111] The image enlargement unit is configured to perform image enlargement processing on the to-be-processed image based on a first preset ratio to obtain a first reference image.

[0112] The image reduction unit is configured to perform image reduction processing on the to-be-processed image based on a second preset ratio to obtain a second reference image.

[0113] The image set construction unit is configured to take the to-be-processed image, the first reference image, and the second reference image as model input images and construct the model input image set.

[0114] Optionally, the identification result determination module 420 comprises a candidate identification result determination unit, a region feature screening and merging processing unit, a region feature enlargement unit, and an identification result determination unit.

[0115] The candidate identification result determination unit is configured to process each model input image included in the model input image set based on the region generation sub-model to determine at least one candidate identification result, wherein the candidate identification result comprises a candidate region feature, an identification category corresponding to the candidate region feature, and boundary box information.

[0116] The region feature screening and merging processing unit is configured to perform region screening and merging processing on the at least one candidate region feature to determine at least one to-be-processed region feature.

[0117] The region feature enlargement unit is configured to perform enlargement processing on each to-be-processed region feature based on a preset size to obtain at least one to-be-input region feature.

[0118] The identification result determination unit is configured to input the at least one to-be-input region feature into the target object screening sub-model to obtain an identification result corresponding to the to-be-processed image.

[0119] Optionally, the region generation sub-model comprises a region generation network and first and second convolutional networks connected to the region generation network respectively, and the region generation network comprises at least one first convolutional module.

[0120] Correspondingly, the candidate recognition result determination unit comprises a candidate region feature determination subunit, a candidate region feature mapping unit, an identification category determination subunit, a bounding box information determination subunit, and a candidate recognition result determination subunit.

[0121] The candidate region feature determination subunit is configured to, for each model input image in the model input image set, process the model input image based on the at least one first convolutional module according to a preset sliding window size corresponding to the model input object, to obtain at least one candidate region feature corresponding to the model input image.

[0122] The candidate region feature mapping unit is configured to map the at least one candidate region feature corresponding to the model input image other than the to-be-processed image in the model input image set to the to-be-processed image, to obtain at least one candidate region feature corresponding to the model input image set.

[0123] The identification category determination subunit is configured to process each candidate region feature based on the first convolutional network, to obtain an identification category corresponding to the candidate region feature.

[0124] The bounding box information determination subunit is configured to process each candidate region feature based on the second convolutional network, to obtain bounding box information corresponding to the candidate region feature.

[0125] The candidate recognition result determination subunit is configured to take the at least one candidate region feature, the identification category corresponding to each candidate region feature, and the bounding box information corresponding to each candidate region feature as at least one candidate recognition result corresponding to the model input image set.

[0126] Optionally, the target object screening sub-model comprises a backbone feature extraction network, a first fully connected network and a second fully connected network connected to the backbone feature extraction network, a feature fusion network, and a third convolutional network connected to the feature fusion network.

[0127] Correspondingly, the recognition result determination unit comprises a feature extraction processing subunit, an identification category determination subunit, a bounding box information determination subunit, an up-sampling processing subunit, a mask map determination subunit, and a recognition result determination subunit.

[0128] The feature extraction processing subunit is configured to perform feature extraction processing on the at least one to-be-input region feature based on at least one first convolution module included in the backbone feature extraction network, to obtain at least one first feature.

[0129] The recognition category determination subunit is configured to process the at least one first feature based on the first full connection network, to obtain a recognition category corresponding to each of the first features.

[0130] The bounding box information determination subunit is configured to process the at least one first feature based on the second full connection network, to obtain bounding box information corresponding to each of the first features.

[0131] The up-sampling processing subunit is configured to perform up-sampling processing on the at least one first feature and model outputs of each of the first convolution modules in the backbone feature extraction network based on at least one second convolution module included in the feature fusion network, to obtain at least one second feature.

[0132] The mask map determination subunit is configured to process the at least one second feature based on the third convolution network, to obtain a mask map corresponding to each of the second features.

[0133] The recognition result determination subunit is configured to determine the recognition result based on the recognition category, the bounding box information, and the mask map.

[0134] The number of the first convolution modules included in the backbone feature extraction network is consistent with the number of the second convolution modules included in the feature fusion network.

[0135] Optionally, the up-sampling processing subunit is specifically configured to take the at least one first feature as a model output of a last first convolution module, and take the model output as a model input of a first second convolution module; for other second convolution modules in the feature fusion network except the first second convolution module, add a model output of a previous convolution module of a current second convolution module and a model output of a first convolution module corresponding to the current second convolution module, to obtain a model input of the current second convolution module, until the current second convolution module is a last second convolution module, and take a model output of the last second convolution module as the at least one second feature; wherein the first convolution module corresponding to the current second convolution module is a first convolution module in the backbone feature extraction network that meets a preset standard corresponding to the current second convolution module.

[0136] Optionally, the device further includes a training sample acquisition module, a training sample input module, a loss processing module, and a model parameter correction module:

[0137] a training sample acquisition module, configured to acquire a plurality of training samples, wherein the training samples comprise training sample images, objects to be recognized in the training sample images, theoretical results corresponding to the training sample images, the theoretical results comprising theoretical categories, theoretical bounding box information, and theoretical mask images;

[0138] a training sample input module, configured to input the training samples into the object segmentation model to be trained to obtain actual output results, wherein the actual output results comprise actual mask images, actual categories, and actual bounding box information;

[0139] a loss processing module, configured to perform loss processing on the theoretical results and the actual output results according to a first loss function corresponding to the region generation sub-model to be trained and a second loss function corresponding to the object screening sub-model to be trained;

[0140] a model parameter correction module, configured to correct model parameters of the object segmentation model to be trained based on the loss value to obtain the target object segmentation model;

[0141] The first loss function comprises a category loss function corresponding to the actual categories and the theoretical categories, and a bounding box loss function corresponding to the actual bounding box information and the theoretical bounding box information; and the second loss function comprises a mask loss function corresponding to the actual mask images and the theoretical mask images, a theoretical loss function corresponding to the actual categories and the theoretical categories, and a bounding box loss function corresponding to the actual bounding box information and the theoretical bounding box information.

[0142] The image processing apparatus provided in the embodiments of the present application can execute the image processing method provided in any of the embodiments of the present application, and has the function modules and beneficial effects corresponding to the execution method.

[0143] Embodiment five

[0144] Figure 5 A structural schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.

[0145] As Figure 5As shown, the electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., connected to the at least one processor 11 in communication. The memory stores computer programs executable by the at least one processor 11, and the processor 11 can perform various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 12 or loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0146] Various components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc., an output unit 17, such as various types of displays, a speaker, etc., a storage unit 18, such as a magnetic disk, an optical disk, etc., and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0147] The processor 11 can be various general and / or special-purpose processing components having processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the image processing method.

[0148] In some embodiments, the image processing method can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the image processing method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the image processing method by any other appropriate means, such as by means of firmware.

[0149] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0150] Computer programs used to implement the processes of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program, when executed, can cause instructions defined in the flow charts and / or block diagrams to be implemented. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package and partially on a remote machine or entirely on a remote machine or server.

[0151] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store computer programs for use by or in connection with an instruction execution system, apparatus, or device. Computer-readable storage media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0152] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0153] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0154] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0155] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in a different order, as long as the desired results of the present disclosure are achieved, and the present disclosure is not limited herein.

[0156] The above detailed description does not limit the scope of the present disclosure. It is understood that various modifications, combinations, sub-combinations, and alternatives can be made to the detailed disclosure without departing from the spirit and principles of the present disclosure. Any modifications, equivalent substitutions, improvements, and the like that are made within the spirit and principles of the present disclosure are included in the scope of the present disclosure.

Claims

1. An image processing method, characterized by, The method comprises the following steps: acquiring a to-be-processed image comprising a plurality of to-be-identified objects, and constructing a model input image set based on the to-be-processed image, wherein the to-be-identified objects are small target objects; processing the model input image set based on a target object segmentation model to determine an identification result corresponding to the to-be-processed image, wherein the target object segmentation model comprises at least one target sub-model, the at least one target sub-model comprises a region generation sub-model and a target object screening sub-model; the region generation sub-model comprises a region generation network, and a first convolutional network and a second convolutional network connected to the region generation network, respectively; the region generation network comprises at least one first convolutional module; the target object screening sub-model comprises a backbone feature extraction network, a first fully connected network, a second fully connected network, and a feature fusion network connected to the backbone feature extraction network, and a third convolutional network connected to the feature fusion network; annotating the to-be-identified objects in the to-be-processed image based on the identification result; the method of processing the model input image set based on the target object segmentation model to determine an identification result corresponding to the to-be-identified objects comprises the following steps: processing each model input image included in the model input image set based on the region generation sub-model to determine at least one candidate identification result, wherein the candidate identification result comprises a candidate region feature, an identification category corresponding to the candidate region feature, and boundary box information; performing region screening and merging processing on the at least one candidate region feature to determine at least one to-be-processed region feature; processing each to-be-processed region feature based on a preset size to obtain at least one to-be-input region feature; inputting the at least one to-be-input region feature into the target object screening sub-model to obtain an identification result corresponding to the to-be-processed image. The generating sub-model based on the region processes each model input image included in the model input image set, determines at least one candidate recognition result, including: for each model input image in the model input image set, based on the at least one first convolutional module, the model input image is processed according to the preset sliding window size corresponding to the model input object, at least one candidate region feature corresponding to the model input image is obtained; the at least one candidate region feature corresponding to the model input image in the model input image set except the to-be-processed image is mapped to the to-be-processed image, at least one candidate region feature corresponding to the model input image set is obtained; based on the first convolutional network, each candidate region feature is processed respectively, to obtain the recognition category corresponding to the candidate region feature; based on the second convolutional network, each candidate region feature is processed respectively, to obtain the bounding box information corresponding to the candidate region feature; the at least one candidate region feature, the recognition category corresponding to each candidate region feature, and the bounding box information corresponding to each candidate region feature are taken as at least one candidate recognition result corresponding to the model input image set.

2. The method of claim 1, wherein, The model input image set is constructed based on the to-be-processed image, including: The to-be-processed image is image enlarged based on a first preset ratio, to obtain a first reference image; The to-be-processed image is image reduced based on a second preset ratio, to obtain a second reference image; The to-be-processed image, the first reference image and the second reference image are taken as model input images, and the model input image set is constructed.

3. The method of claim 1, wherein, The at least one to-be-input region feature is input into the target object screening sub-model, to obtain a recognition result corresponding to the to-be-processed image, including: The at least one to-be-input region feature is feature extracted based on at least one first convolutional module included in the backbone feature extraction network, to obtain at least one first feature; The at least one first feature is processed based on the first fully connected network, to obtain a recognition category corresponding to each first feature; The at least one first feature is processed based on the second fully connected network, to obtain bounding box information corresponding to each first feature; The at least one first feature and the model output of each first convolutional module in the backbone feature extraction network are processed based on at least one second convolutional module included in the feature fusion network, to obtain at least one second feature; The at least one second feature is processed based on the third convolutional network, to obtain a mask map corresponding to each second feature; The recognition result is determined based on the recognition category, the bounding box information and the mask map; The number of the first convolutional modules included in the backbone feature extraction network is consistent with the number of the second convolutional modules included in the feature fusion network.

4. The method of claim 3, wherein, The at least one second feature is obtained by processing the at least one first feature and the model output of each first convolutional module in the backbone feature extraction network based on at least one second convolutional module included in the feature fusion network, including: The at least one first feature is taken as the model output of the last first convolutional module, and the model output is taken as the model input of the first second convolutional module; For other second convolutional modules in the feature fusion network except the first second convolutional module, the model output of the last convolutional module of the current second convolutional module and the model output of the first convolutional module corresponding to the current second convolutional module are added to be taken as the model input of the current second convolutional module, until the current second convolutional module is the last second convolutional module, and the model output of the last second convolutional module is taken as the at least one second feature; Wherein, the first convolutional module corresponding to the current second convolutional module is the first convolutional module in the backbone feature extraction network that meets the preset standard corresponding to the current second convolutional module.

5. The method of claim 1, wherein, The target object segmentation model is obtained by the following steps: Obtain a plurality of training samples, wherein the training samples include: a training sample image, the training sample image includes an object to be identified, a theoretical result corresponding to the training sample image, the theoretical result includes a theoretical category, theoretical bounding box information and a theoretical mask image; Input the training sample into a to-be-trained object segmentation model to obtain an actual output result, wherein the actual output result includes an actual mask image, an actual category and actual bounding box information; According to the first loss function corresponding to the to-be-trained region generation sub-model and the second loss function corresponding to the to-be-trained object screening sub-model, the theoretical result and the actual output result are loss processed; Based on the loss value, the model parameters of the to-be-trained object segmentation model are corrected to obtain the target object segmentation model; Wherein, the first loss function includes a category loss function corresponding to the actual category and the theoretical category, and a bounding box loss function corresponding to the actual bounding box information and the theoretical bounding box information; the second loss function includes a mask loss function corresponding to the actual mask image and the theoretical mask image, a theoretical loss function corresponding to the actual category and the theoretical category, and a bounding box loss function corresponding to the actual bounding box information and the theoretical bounding box information.

6. An image processing apparatus characterized by comprising: Including: An image acquisition module is configured to acquire a to-be-processed image including a plurality of to-be-identified objects, and construct a model input image set based on the to-be-processed image, wherein the to-be-identified objects are small target objects; The recognition result determination module is configured to determine a recognition result corresponding to the to-be-recognized object by processing the model input image set based on a target object segmentation model, wherein the target object segmentation model comprises at least one target sub-model, and the at least one target sub-model comprises a region generation sub-model and a target object screening sub-model; the region generation sub-model comprises a region generation network, a first convolutional network and a second convolutional network connected to the region generation network respectively, and the region generation network comprises at least one first convolutional module; the target object screening sub-model comprises a backbone feature extraction network, a first full connection network, a second full connection network and a feature fusion network connected to the backbone feature extraction network, and a third convolutional network connected to the feature fusion network; The object labeling module is configured to label the to-be-recognized object in the to-be-processed image based on the recognition result. The recognition result determination module comprises a candidate recognition result determination unit, a region feature screening and merging processing unit, a region feature amplification unit and a recognition result determination unit; wherein the candidate recognition result determination unit is configured to process each model input image included in the model input image set based on the region generation sub-model, and determine at least one candidate recognition result, wherein the candidate recognition result comprises a candidate region feature, an identification category corresponding to the candidate region feature and boundary box information; the region feature screening and merging processing unit is configured to perform region screening and merging processing on the at least one candidate region feature, and determine at least one to-be-processed region feature; the region feature amplification unit is configured to perform amplification processing on each to-be-processed region feature based on a preset size, and obtain at least one to-be-input region feature; and the recognition result determination unit is configured to input the at least one to-be-input region feature into the target object screening sub-model, and obtain a recognition result corresponding to the to-be-processed image. The candidate recognition result determination unit comprises a candidate region feature determination subunit, a candidate region feature mapping unit, an identification category determination subunit, a bounding box information determination subunit, and a candidate recognition result determination subunit. The candidate region feature determination subunit is configured to, for each model input image in the model input image set, process the model input image based on the at least one first convolution module according to the preset sliding window size corresponding to the model input object, to obtain at least one candidate region feature corresponding to the model input image. The candidate region feature mapping unit is configured to map the at least one candidate region feature corresponding to the model input image other than the to-be-processed image in the model input image set to the to-be-processed image, to obtain at least one candidate region feature corresponding to the model input image set. The identification category determination subunit is configured to process each candidate region feature based on the first convolution network, to obtain an identification category corresponding to the candidate region feature. The bounding box information determination subunit is configured to process each candidate region feature based on the second convolution network, to obtain bounding box information corresponding to the candidate region feature. The candidate recognition result determination subunit is configured to take the at least one candidate region feature, the identification category corresponding to each candidate region feature, and the bounding box information corresponding to each candidate region feature as at least one candidate recognition result corresponding to the model input image set.

7. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the image processing method of any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing the processor to execute when the computer instructions are executed, thereby realizing the image processing method of any one of claims 1-5. The computer readable storage medium stores computer instructions for causing the processor to execute when the computer instructions are executed, thereby realizing the image processing method of any one of claims 1-5.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and storage medium

    CN112036516A

  • Vehicle indicator light identification method and device, computer equipment and storage medium

    CN114170587A