Image processing method and device
By using the target window technology in image processing, image blocks are extracted from the target image and mask image for segmentation processing, the problem of low edge segmentation accuracy in traditional methods is solved, and the segmentation accuracy of the target object is significantly improved.
Patent Information
- Application Number
- CN202510134149.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-06-20
AI Technical Summary
The accuracy of the edge segmentation result of the traditional image processing method at the target object is lower than that of the internal area segmentation result, affecting the accuracy of the overall segmentation result.
By obtaining the first segmentation result of the target image, the target window corresponding to the edge area of the target mask image is determined, and the window size is adjusted according to the size of the target object, the image blocks corresponding to the target window are extracted from the target image and the mask image for finer segmentation processing, and a second segmentation result with higher segmentation accuracy is obtained.
It significantly improves the segmentation accuracy of the target object in the target image, and achieves more accurate and detailed segmentation of the target object.
Smart Images

Figure CN120182591A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to an image processing method and apparatus. Background Art
[0002] In the field of image processing, accurately segmenting target objects in an image is a key technology. However, the accuracy of the edge segmentation result of the target object by traditional segmentation methods is lower than that of the internal region segmentation result, thus affecting the accuracy of the overall segmentation result. Summary of the Invention
[0003] In view of this, at least one image processing method and apparatus are provided in the embodiments of this application.
[0004] The technical solutions of the embodiments of this application are implemented as follows:
[0005] On the one hand, an image processing method is provided in the embodiments of this application. The method includes: obtaining a first segmentation result corresponding to a target image, where the first segmentation result includes a target mask image corresponding to a target object in the target image; determining a target window corresponding to an edge region of the target mask image, where the size of the target window is determined according to the target object; obtaining a target image block and a mask image block, where the target image block is an image block in the target image corresponding to the target window, and the mask image block is an image block in the target mask image corresponding to the target window; and obtaining a second segmentation result of the target image according to the target image block and the mask image block, where the segmentation accuracy of the second segmentation result is greater than that of the first segmentation result.
[0006] On the other hand, an image processing apparatus is provided in the embodiments of this application. The apparatus includes: a first obtaining module, configured to obtain a first segmentation result corresponding to a target image, where the first segmentation result includes a target mask image corresponding to a target object in the target image; a determining module, configured to determine a target window corresponding to an edge region of the target mask image, where the size of the target window is determined according to the target object; a second obtaining module, configured to obtain a target image block and a mask image block, where the target image block is an image block in the target image corresponding to the target window, and the mask image block is an image block in the target mask image corresponding to the target window; and a third obtaining module, configured to obtain a second segmentation result of the target image according to the target image block and the mask image block, where the segmentation accuracy of the second segmentation result is greater than that of the first segmentation result.
[0007] In the embodiments of the present application, by obtaining the first segmentation result corresponding to the target image, the approximate position and shape of the target object in the target image can be initially identified. At the same time, by determining the target window corresponding to the edge region of the target mask image and adjusting the window size according to the size of the target object, the key region of the target object can be accurately located. Image patches corresponding to the target window, namely the target image patches and the mask image patches, are respectively extracted from the target image and the target mask image, and a more refined segmentation process is performed using the target image patches and the mask image patches, so as to obtain a second segmentation result with higher segmentation accuracy. In this way, the present application can make full use of the information in the target image and the mask image, thereby realizing more accurate and detailed segmentation of the target object. Based on the embodiments provided by the present application, the segmentation accuracy of the target object in the target image can be significantly improved.
[0008] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the technical solution of the present application. Description of the Drawings
[0009] The drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to explain the technical solution of the present application.
[0010] Figure 1 Schematic diagram of the implementation process of an image processing method provided by an embodiment of the present application Figure 1 ;
[0011] Figure 2 Schematic diagram of the implementation process of an image processing method provided by an embodiment of the present application Figure 2 ;
[0012] Figure 3 Schematic diagram of the implementation process of an image processing method provided by an embodiment of the present application Figure 3 ;
[0013] Figure 4 Schematic diagram of the implementation process of an image processing method provided by an embodiment of the present application Figure 4 ;
[0014] Figure 5 Schematic diagram of the implementation process of an image processing method provided by an embodiment of the present application Figure 5 ;
[0015] Figure 6 Schematic data flow diagram of an image processing method provided by an embodiment of the present application;
[0016] Figure 7 Schematic diagram of the composition structure of an image processing device provided by an embodiment of the present application;
[0017] Figure 8 A schematic diagram of the hardware entity of a computer device provided by an embodiment of the present application. Specific implementation manners
[0018] In order to make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be further elaborated in detail below in conjunction with the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0019] In the following description, "some embodiments" are involved, which describe subsets of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict. The terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing this application and are not intended to limit this application.
[0021] An embodiment of the present application provides an image processing method, which can be executed by a processor of a computer device. Among them, the computer device can refer to devices with data processing capabilities such as servers, laptop computers, tablet computers, desktop computers, smart TVs, set-top boxes, mobile devices (such as mobile phones, portable video players, personal digital assistants, dedicated messaging devices, portable game devices), etc.
[0022] Figure 1 A schematic diagram of the implementation process of an image processing method provided by an embodiment of the present application Figure 1 , as Figure 1 shown, the method includes the following steps S101 to step S103:
[0023] Step S101, obtain a first segmentation result corresponding to the target image, where the first segmentation result includes a target mask image corresponding to a target object in the target image.
[0024] The target image is an image to be processed or analyzed, which may be a photo, video frame, or other forms of image data. Accordingly, the target object is a specific object to be segmented from the target image, and the target object may be any entity in the image, such as a person, animal, plant, building, vehicle, etc.
[0025] In some embodiments, the target image may be preliminarily segmented by an image segmentation algorithm to obtain a target mask image corresponding to the target object in the target image as the first segmentation result. The first segmentation result is used to distinguish the area where the target object is located from other areas in the target image.
[0026] Exemplarily, the above-mentioned image segmentation algorithm may be, but is not limited to, a threshold segmentation algorithm, a region growing algorithm, a clustering segmentation algorithm, a segmentation algorithm based on a deep learning model, and the like.
[0027] In some embodiments, the target mask image is a binary image, which is used to represent the position and shape of the target object in the target image. In the target mask image, the area where the target object is located is assigned a specific value (usually white or 1), while other areas are assigned another different value (usually black or 0).
[0028] Step S102: determining a target window corresponding to an edge region of the target mask image, wherein a size of the target window is determined according to the target object.
[0029] The edge region of the target mask image is a boundary region between the target object and the background or other objects in the target mask image, and the edge region includes contour information of the target object.
[0030] The target window is a window set in the edge area, and is used to extract contour information of the target object in the target mask image.
[0031] In some embodiments, the number of the target windows is multiple, and the multiple target windows include all the contour information of the target object, that is, each target window includes part of the contour information of the target object. The contour information between different target windows may overlap or may not overlap.
[0032] In some embodiments, the size of the target window can be determined according to the size of the target object. Exemplarily, the size of the target window can be proportional to the size of the target object. The larger the size of the target object, the larger the size of the target window. In other embodiments, the size of the target window can be determined according to the type of the target object. Exemplarily, multiple categories and window sizes corresponding to each category can be preset, and the size of the target window can be determined by detecting the category of the target object.
[0033] In some embodiments, the edge region in the target mask image may be recognized first; one or more target windows may be set according to the position and size information of the edge region, wherein the set multiple target windows can completely cover the contour information of the target object. During the process of setting the multiple target windows, the overlapping degree between the target windows may be controlled by a preset precision parameter to control the integrity and accuracy of the contour information. Among them, the higher the precision parameter is set, the higher the overlapping degree between the target windows is, and the lower the precision parameter is set, the lower the overlapping degree between the target windows is. When the precision parameter is set to the minimum value, there is no overlapping contour information between adjacent target windows, and the contour information between adjacent target windows is continuous.
[0034] In some embodiments, the window center of the above target window may fall on the contour of the target object in the target mask image, so that the target window can more accurately cover the contour information of the target object.
[0035] Step S103, obtain a target image block and a mask image block, where the target image block is the image block in the target image corresponding to the target window, and the mask image block is the image block in the target mask image corresponding to the target window.
[0036] In some embodiments, for each target window obtained in step S102, the target image may be intercepted based on the range of the target window in the target image to obtain the above target image block; correspondingly, the target mask image may be intercepted based on the range of the target window in the target mask image, and then the above mask image block may be obtained.
[0037] It should be noted that the target image block and the mask image block are related to each other, and the mask image block is used to distinguish the area where the target object is located in the target image block from other areas.
[0038] Step S104, obtain a second segmentation result of the target image according to the target image block and the mask image block, and the segmentation accuracy of the second segmentation result is greater than the segmentation accuracy of the first segmentation result.
[0039] In some embodiments, the target image block and the mask image block may be input into a trained target segmentation model to obtain a sub-segmentation result corresponding to the target image block; by combining the sub-segmentation results corresponding to each target image block, a segmentation result of the target object in the target image, that is, the above second segmentation result, may be obtained.
[0040] In some embodiments, the input of the trained target segmentation model is a target image patch and the preliminary segmentation result corresponding to the target image patch (i.e., a part of the segmentation result obtained based on the target window in the target mask image), and the output of the trained target segmentation model is the sub-segmentation result of the target image patch.
[0041] In the embodiments of the present application, by obtaining the first segmentation result corresponding to the target image, the approximate position and shape of the target object in the target image can be initially identified. At the same time, by determining the target window corresponding to the edge region of the target mask image and adjusting the window size according to the size of the target object, the key region of the target object can be accurately located; image patches corresponding to the target window are respectively extracted from the target image and the target mask image, that is, the target image patch and the mask image patch, and more refined segmentation processing is performed using the target image patch and the mask image patch, and a second segmentation result with higher segmentation accuracy can be obtained. In this way, the present application can make full use of the information in the target image and the mask image, so as to achieve more accurate and detailed segmentation of the target object. Based on the embodiments provided in the present application, the segmentation accuracy of the target object in the target image can be significantly improved.
[0042] Figure 2 is a schematic implementation process of an image processing method provided by an embodiment of the present application Figure 2 , and this method can be executed by the processor of a computer device. Based on Figure 1 , the method further includes steps S201 to S202, which will be described in combination with Figure 2 the steps shown.
[0043] Step S201: Obtain the size information of the target object.
[0044] In some embodiments, the size information of the target object is used to characterize the physical size (such as length, width, height) of the target object in the real world; at this time, the size information of the target object can be obtained in the following ways: based on an image recognition algorithm, the category of the target object is recognized to obtain the category of the target object, and the physical size of the target object in the real world is determined based on the category of the target object; or, receive a segmentation request from the user for the target object, and the segmentation request may carry the above size information, and parse the segmentation request to obtain the physical size of the target object in the real world.
[0045] Exemplarily, taking the vehicle segmentation scenario as an example, the physical size (such as length, width, and height) of the vehicle can be queried by identifying the model and manufacturer information of the vehicle in the image; or, the segmentation request for the vehicle carries the size information for the vehicle, and the semantic information carried by the segmentation request can be "segment the sedan with a vehicle length of 4 meters", so that the physical size of the vehicle can also be obtained.
[0046] In some other embodiments, the size information of the target object is used to characterize the pixel size of the target object in the target image (such as the number of pixels corresponding to the width and height); at this time, the size information of the target object can be obtained in the following manner: using a target detection algorithm (such as a convolutional neural network, edge detection, etc.) to identify the position of the target object in the image; according to the position information of the target object, calculating the number of pixels corresponding to the width and height of the target object by pixel counting.
[0047] In some embodiments, the above-mentioned obtaining of the size information of the target object can be achieved through step S2011.
[0048] Step S2011: Determine the size information of the target object according to the target mask image.
[0049] Wherein, the target mask image is used to represent the position and shape of the target object in the target image, and thus, the size information of the target object in the target mask image can be determined.
[0050] In some embodiments, the target mask image can be analyzed to determine the bounding box or the minimum circumscribed rectangle of the target object in the target mask image, and further determine the size parameters such as the length and width of the target object.
[0051] In the embodiments of the present application, determining the size information of the target object according to the target mask image can efficiently extract the precise boundary of the target object from the binary mask image, thereby improving the accuracy of the size information.
[0052] Step S202: Determine the size of the target window according to the size information of the target object.
[0053] In some embodiments, a mapping relationship between the size information of the target object and the size of the target window can be preset. After obtaining the size information of the target object, based on this mapping relationship, the size of the target window required currently can be obtained.
[0054] In some embodiments, the size information may include the size parameters of the target object in at least one direction. Exemplarily, the at least one direction may be perpendicular to each other, such as the size parameters in the three dimensions of length, width, and height.
[0055] In some embodiments, the size parameters in each direction can be judged separately. When the size parameters in each direction all indicate that the target object belongs to the target size level, the window size corresponding to the target size level is determined as the size of the target window. Exemplarily, there is a target object, and the size information of the target object includes length, width, and height. Among them, if the length, width, and height all correspond to the window size of the low level, the window size of the low level is determined as the size of the target window.
[0056] In other embodiments, the corresponding target size levels are determined respectively based on the size parameters in each direction, and the window size corresponding to the largest target size level is determined as the size of the target window. Continuing with the above example, if the length and width correspond to the window size of the low level, but the height corresponds to the window size of the high level, the window size of the high level is determined as the size of the target window.
[0057] In other embodiments, the corresponding target size levels are determined respectively based on the size parameters in each direction, and the window size corresponding to the smallest target size level is determined as the size of the target window. Continuing with the above example, if the length and width correspond to the window size of the low level, but the height corresponds to the window size of the high level, the window size of the low level is determined as the size of the target window.
[0058] In other embodiments, a comprehensive size parameter is determined based on the size parameters in each direction, and the window size corresponding to the comprehensive size parameter is determined as the size of the target window. Continuing with the above example, first, a comprehensive size parameter (such as taking the average value) is determined based on the length, width, and height, and the window size corresponding to the level of the comprehensive size parameter is determined as the size of the target window.
[0059] In the embodiments of the present application, by obtaining the size information of the target object, the actual size of the target object in the image or scene can be accurately understood, providing an accurate reference basis for the subsequent selection of the target window. At the same time, determining the size of the target window according to the size information of the target object can ensure that the target window neither misses the contour information of the target object nor contains too much background information, thereby improving the efficiency and accuracy of the subsequent second segmentation.
[0060] Figure 3 is a schematic implementation process of an image processing method provided by the embodiments of the present application Figure 3 , which can be executed by the processor of the computer device. Based on Figure 1 , the method further includes steps S301 to step S302, which will be described in combination with Figure 3 the steps shown.
[0061] Step S301: Obtain the smoothness information corresponding to the contour of the target object.
[0062] Among them, the smoothness information is used to characterize the smoothness of the contour of the target object in the target image.
[0063] In some embodiments, the edge detection of the target object can be performed on the target image first to obtain the edge information of the target object in the target image; the smoothness calculation is performed on the edge information to obtain the smoothness information.
[0064] In some embodiments, the above edge detection can be implemented by at least one of the following methods: Sobel algorithm, Prewitt algorithm, Laplacian algorithm, and Canny algorithm. Among them, the Sobel algorithm detects edges by calculating the gradient of the image, and uses two convolution kernels to detect edges in the horizontal and vertical directions respectively; the Prewitt algorithm is similar to the Sobel algorithm, but uses different convolution kernels to detect edges; the Laplacian algorithm detects edges by calculating the second derivative of the image; the Canny algorithm calculates edges through steps such as Gaussian filtering, gradient calculation, and non-maximum suppression.
[0065] In some embodiments, the above process of calculating the smoothness of the edge information can be implemented by the following method: Calculate the perimeter of the edge of the target object based on the edge information of the target object (determined by accumulating the Euclidean distances between adjacent points on the contour); calculate the area of the target object based on the edge information; use the area and the perimeter to determine the smoothness information; for example, the reciprocal of the ratio of the square of the perimeter of the contour to four times the area times π can be used to calculate the smoothness information, and the larger this value is, the smoother the contour is. Or, the method of curvature analysis is used to calculate the curvature values of each point on the contour; statistical analysis is performed on the curvature values of all points on the contour, for example, calculating the standard deviation or coefficient of variation of the curvature; the smoothness information is determined based on the value of the statistic, for example, the smaller the standard deviation or the lower the coefficient of variation, the smoother the contour is.
[0066] In some embodiments, the above-mentioned obtaining of the smoothness information corresponding to the contour of the target object can be achieved through steps S3011 to S3012.
[0067] Step S3011: Obtain the edge information of the target mask image.
[0068] In some embodiments, edge detection of the target object can be performed on the target mask image to obtain the edge information of the target object in the target mask image. Among them, since the target mask image is a binary image, the edge of the target object is actually the position where the pixel value in the target mask image changes from 1 (representing the target object) to 0 (representing the background) or from 0 to 1.
[0069] Step S3012: Determine the smoothness information according to the edge information.
[0070] In some embodiments, the above edge information may include the coordinate information of each edge pixel point among multiple edge pixel points in the target mask image. Furthermore, the smoothness of the edge can be evaluated by calculating parameters such as the continuity, change rate, or curvature of the edge pixel points, and then the smoothness information can be obtained.
[0071] Step S302: Determine the size of the target window according to the smoothness information.
[0072] In some embodiments, when the smoothness information characterizes a higher smoothness level of the contour of the target object, that is, when it characterizes that the contour of the target object is smoother, the size of the target window is larger; when the smoothness information characterizes a lower smoothness level of the contour of the target object, that is, when it characterizes that the contour of the target object is rougher, the size of the target window is smaller.
[0073] It can be understood that when the contour of the target object is smooth, it means that in the first segmentation result, the shape of the target object changes relatively evenly without too many details or mutations. At this time, using a larger target window can more effectively capture the overall features of the target object in the local area. At the same time, since a wider area is covered, the attention to details in the subsequent segmentation process can be reduced, thereby improving the stability and accuracy of the segmentation. A larger window can also reduce the computational amount and improve the processing efficiency. On the contrary, when the contour of the target object is rough, it means that in the first segmentation result, the shape of the target object contains more detailed features. At this time, using a small window can more precisely focus on these detailed features. At the same time, the small window can more carefully analyze the edge changes and improve the accuracy of the subsequent segmentation process.
[0074] In the embodiments of the present application, the size of the target window can be dynamically adjusted according to the contour smoothness of the target object, so as to improve the processing efficiency while ensuring the segmentation accuracy.
[0075] In some other embodiments, different sizes of the target window may also be set for different contour regions of the target object. The contour region is different contour segments of the contour of the target object in the target image. In the embodiments of the present application, the size of the corresponding target window may be set for each contour segment. Based on Figure 1 , the method further includes: segmenting the contour of the target object to obtain at least two contour segments; obtaining the sub-smoothness information corresponding to each contour segment; and determining the size of the target window corresponding to each contour segment based on the sub-smoothness information corresponding to each contour segment.
[0076] The target window corresponding to a contour segment refers to the window used to extract the contour information within the contour segment of the target object in the target mask image.
[0077] Here, the manner of determining the sub-smoothness information corresponding to each contour segment is the same as the manner of determining the smoothness information corresponding to the contour of the target object in step S301 above, and the implementation manner in step S301 may be referred to during implementation; correspondingly, the manner of determining the size of the target window corresponding to the contour segment based on the sub-smoothness information corresponding to the contour segment is the same as the manner of determining the size of the target window according to the smoothness information in step S302 above, and the implementation manner in step S302 may be referred to during implementation.
[0078] In the embodiments of the present application, fine processing of the contour of the target object can be realized, and the size of the target window can be dynamically adjusted according to the characteristics of different contour segments. When some contour segments are relatively smooth, a larger target window can more effectively capture the overall features and improve the processing efficiency; while when the contour segments are relatively rough, a smaller target window can more accurately focus on the detail features and improve the segmentation accuracy.
[0079] In some other embodiments, the first size of the target window may also be determined based on the size information of the target object; the second size of the target window is determined based on the smoothness information corresponding to the contour of the target object; and the size of the target window is determined based on the first size and the second size.
[0080] In the embodiments of the present application, by considering the first size and the second size together, the size of the target window can be optimized by integrating the overall size and edge features of the target object. Furthermore, the size of the target window can be more flexibly determined according to the actual features of the target object, effectively capturing the edge details while ensuring that the main part of the target object is covered, thereby improving the accuracy and efficiency of subsequent segmentation processing.
[0081] Figure 4 is the implementation process schematic of an image processing method provided by the embodiments of the present application Figure 4, this method can be executed by the processor of a computer device. Based on Figure 1 , the target window includes a first window and a second window, the first window is smaller than the second window, and the first window is located within the second window. The method further includes step S401, which will be described in conjunction with Figure 4 the steps shown.
[0082] Step S401: Determine the size of the first window and the size of the second window according to the target object.
[0083] Among them, the number of the target windows is multiple, and the multiple target windows include all the contour information of the target object. For each target window, it can actually include at least one sub-window, and the sizes of the sub-windows are different; in some embodiments, the at least one sub-windows are in a nested relationship in sequence, that is, the largest window among the N sub-windows includes the remaining N - 1 windows, the second largest window includes the remaining N - 2 windows, and so on, and the smallest window does not include any window.
[0084] Exemplarily, the at least one sub-window can include a first window and a second window, and the sizes of the first window and the second window are different. For example, the first window is smaller than the second window, and the first window is located within the second window. For the convenience of understanding this solution, the following embodiments will be described with the first window and the second window, and it is not limited to the number of windows included in the target window.
[0085] In some embodiments, the method of determining the size of the first window based on the target object and the method of determining the size of the second window based on the target object are the same as the method of determining the size of the target window based on the size information of the target object and / or the smoothness information corresponding to the contour of the target object in the above embodiments, and the method provided in the above embodiments can be referred to during implementation.
[0086] In the embodiments of the present application, during the process of re-determining the second segmentation result for the target object within each target window, due to the setting of the first window and the second window with different sizes, thus, the smaller first window can focus more on the detailed features, and the larger second window can capture extensive information inside the target object and its surrounding environment. Furthermore, different ranges of image information can be fully utilized to improve the accuracy of the subsequent segmentation process.
[0087] In some embodiments, the above-mentioned determining the size of the first window and the size of the second window according to the target object can be implemented through step S4011.
[0088] Step S4011: Determine the size ratio relationship between the first window and the second window according to the target object.
[0089] In some embodiments, the size ratio relationship is the relative size relationship in terms of size between the first window and the second window. Exemplarily, the size ratio relationship can be represented by a proportionality coefficient or a multiple. For the convenience of describing the embodiments of the present application, the following size ratio relationships are all described by taking the proportionality coefficient obtained by dividing the size of the large window by the size of the small window as an example.
[0090] In some embodiments, the size ratio relationship can be determined according to the size of the target object. Exemplarily, the size ratio relationship can be directly proportional to the size of the target object. The larger the size of the target object, the larger the size ratio relationship, that is, the greater the size gap between the first window and the second window; or, the size ratio relationship can be inversely proportional to the size of the target object. The larger the size of the target object, the smaller the size ratio relationship, that is, the smaller the size gap between the first window and the second window. In other embodiments, the size of the target window can be determined according to the type of the target object. Exemplarily, multiple categories and the corresponding size ratio relationships for each category can be preset. By detecting the category of the target object, the size ratio relationship can be determined accordingly.
[0091] To further improve the accuracy of the subsequent segmentation process, the mapping relationship between the size of the target object and the size ratio relationship, and / or the mapping relationship between the category of the target object and the size ratio relationship can be learned through deep learning methods.
[0092] Taking the example of learning the mapping relationship between the size of the target object and the size ratio relationship through deep learning methods, the first prediction model for predicting the size ratio relationship based on the object size can be trained through the following scheme:
[0093] (1) Collect image data containing target objects of different sizes and categories, and label the size and segmentation result (as the standard result) of each target object. Segment the image data to obtain the first segmentation result corresponding to the image data, that is, obtain the target mask image corresponding to the image data.
[0094] (2) Obtain the deep learning model to be trained, the input of which is the size of the target object (which can be the width, height, length or a combination of at least two), and the output is the predicted size ratio relationship. Among them, the model can select a convolutional neural network (CNN) or other neural network structures suitable for processing image data.
[0095] (3) Use the prepared image data and annotation information to train the deep learning model. During the training process, the size of the target object can be output to the deep learning model to be trained, and the predicted size ratio relationship output by the deep learning model to be trained can be obtained. Based on this predicted size ratio relationship, the sizes of the first window and the second window can be calculated; based on the sizes of the first window and the second window, the subsequent segmentation process (which may include the above steps S103 and S104) can be performed to obtain the corresponding second segmentation result; compare the second segmentation result with the annotation information, calculate the loss function (such as cross-entropy loss, IoU loss, etc.), and adjust the model parameters of the deep learning model to be trained through the backpropagation algorithm. This process can be iterated until the model converges, and the trained deep learning model is output as the first prediction model.
[0096] Of course, the second prediction model for predicting the size ratio relationship based on the object category can be trained using an implementation process similar to the above solution, which will not be repeated here.
[0097] Step S4012: Determine the sizes of the first window and the second window according to the size ratio relationship.
[0098] In some embodiments, based on the preset minimum window size, and based on this minimum window size and the above size ratio relationship, the sizes of each of the at least one sub-window can be obtained in sequence. That is to say, the size of the smallest sub-window is the minimum window size. After that, based on this size ratio relationship, the smallest sub-window is enlarged in sequence, and thus each sub-window can be obtained. Taking the first window and the second window as an example, the size of the first window can be set as the minimum window size, and then, based on this size ratio relationship, the size of the first window is enlarged to obtain the size of the second window.
[0099] In other embodiments, based on the preset maximum window size, and based on this maximum window size and the above size ratio relationship, the sizes of each of the at least one sub-window can be obtained in sequence. That is to say, the size of the largest sub-window is the maximum window size. After that, based on this size ratio relationship, the largest sub-window is reduced in sequence, and thus each sub-window can be obtained. Taking the first window and the second window as an example, the size of the second window can be set as the maximum window size, and then, based on this size ratio relationship, the size of the second window is reduced to obtain the size of the first window.
[0100] In the embodiment of the present application, by determining the size ratio relationship between the first window and the second window according to the size or type of the target object, the window size can be flexibly adjusted to meet the needs of different target objects; at the same time, by learning the mapping relationship between the size or type of the target object and the size ratio relationship through a deep learning method, the accuracy and adaptability of the window size determination can be further improved. Based on the embodiment provided by the present application, accurate segmentation of different target objects can be achieved, and the efficiency and accuracy of the segmentation process can be improved.
[0101] Figure 5 This is a schematic diagram of an implementation process of an image processing method provided in an embodiment of the present application. Figure 5 The method can be executed by a processor of a computer device. The target image block includes a first target image block obtained from the target image through the first window and a second target image block obtained through the second window; the mask image block includes a first mask image block obtained from the target mask image through the first window and a second mask image block obtained through the second window. The above step S104 can be updated to steps S501 to S505, which will be combined with Figure 5 The steps shown are explained.
[0102] Step S501: Perform image enhancement processing on the first target image block and the first mask image block to obtain a first enhanced target image block and a first enhanced mask image block, wherein the resolution of the first enhanced target image block and the first enhanced mask image block is the same as that of the second target image block.
[0103] In some embodiments, the target image block is an image block in the target image corresponding to the target window. For each target window, since the target window includes at least one sub-window, the image block corresponding to the target window includes the image block corresponding to each sub-window. It can be understood that since the sizes of the sub-windows are different, the resolutions of the image blocks corresponding to the sub-windows are also different; illustratively, taking at least one sub-window including a first window and a second window as an example, the image block corresponding to the target window includes a first target image block obtained from the target image through the first window, and a second target image block obtained from the target image through the second window. Similarly, the mask image block includes a first mask image block obtained from the target mask image through the first window, and a second mask image block obtained from the target mask image through the second window.
[0104] In some embodiments, since the resolution of the first target image block is less than that of the second target image block, and the resolution of the first mask image block is also less than that of the second target image block, in order to facilitate the subsequent fusion of image information, it is necessary to align the resolutions of the first target image block and the first mask image block with that of the second target image block. Therefore, it is necessary to perform image enhancement processing on the first target image block and the first mask image block, and this image enhancement processing is used to increase the resolutions of the first target image block and the first mask image block and increase them to be the same as that of the second target image block.
[0105] In some embodiments, the above-mentioned image enhancement processing may be, but is not limited to, image interpolation, super-resolution reconstruction and other solutions. Among them, image interpolation is a solution to calculate the values of new pixels through interpolation algorithms (such as nearest neighbor interpolation, bilinear interpolation, and bicubic interpolation) according to the information of existing pixels, so as to increase the pixel density of the image; super-resolution reconstruction is a solution to extract more useful information from low-resolution images and generate higher-resolution images by using algorithms such as deep learning.
[0106] In some embodiments, the performing image enhancement processing on the first target image block and the first mask image block includes: performing image enhancement processing on the first target image block and the first mask image block when the target object meets the target conditions.
[0107] Among them, the target conditions include at least one of the following: the size of the target object is less than the target size, and the smoothness of the contour of the target object is less than the target smoothness.
[0108] It can be understood that when the size of the target object is less than the target size, it means that the target object may be relatively small in the target image (or relatively small in the real scene), and the detail information may not be prominent enough. Therefore, it is necessary to magnify this detail information through image enhancement processing to facilitate the provision of detail information in the subsequent second segmentation process.
[0109] Of course, when the smoothness of the contour of the target object is less than the target smoothness, it means that the contour of the target object may be relatively rough or irregular, and it also shows that the contour at the current position contains more detail information. Therefore, it is also necessary to generate more detail information through image enhancement processing to facilitate the provision of detail information in the subsequent second segmentation process.
[0110] In the embodiments of the present application, when there is more detailed information in the contour of the target object, image enhancement processing can be performed on the original first target image block and the first mask image block, thereby generating more detailed information, providing a richer and more accurate reference for the subsequent second segmentation process. In this way, not only the accuracy of segmentation is improved, but also the adaptability of the algorithm to target objects of different sizes and shapes is enhanced.
[0111] Step S502: Obtain first fusion image information according to the first enhanced target image block and the first enhanced mask image block.
[0112] Step S503: Obtain second fusion image information according to the second target image block and the second mask image block.
[0113] In some embodiments, the above steps S502 and S503 are used to integrate the original image information and mask information within a window to obtain a feature representation containing comprehensive information. For example, integration can be performed by means such as channel merging, feature fusion (such as weighted summation, etc.).
[0114] Among them, in step S502, the enhanced first target image block and the first mask image block are subjected to channel merging or feature fusion to generate first fusion image information; in step S503, the second target image block and the second mask image block are subjected to channel merging or feature fusion to generate second fusion image information.
[0115] Step S504: Input the first fusion image information and the second fusion image information into the target network to obtain a corrected mask image block output by the target network.
[0116] In some embodiments, the corrected mask image block is a binary image, which is used to represent the position and shape of the target object in the second target image block. In the corrected mask image block, the area where the target object is located is assigned a specific value (usually white or 1), while other areas are assigned another different value (usually black or 0).
[0117] In some embodiments, the target network is a deep learning model for object segmentation. The input of the target network is the first fusion image information and the second fusion image information, and its output is the segmented mask image corresponding to the second target image block, that is, the corrected mask image block.
[0118] In the embodiments of the present application, the first fused image information and the second fused image information respectively integrate the information of the target image patches and the mask image patches under different resolutions and different sub-windows. By fusing this information, the target network can obtain more comprehensive context and detail features, which helps to more accurately identify the boundaries and shapes of the target objects. Additionally, since the input information carries the mask image patches at different scales, in this way, the target network can better locate the target objects during the segmentation process and reduce the possibility of mis-segmentation.
[0119] Step S505: Obtain the second segmentation result according to the corrected mask image patch.
[0120] In some embodiments, for each target window, step S504 can obtain the corrected mask image patch corresponding to each target window. Different corrected mask image patches include the contours of different positions of the target object. Furthermore, all the obtained corrected mask image patches can be combined to obtain the complete contour of the target object; then, based on the complete contour of the target object, the second segmentation result can be obtained.
[0121] In the embodiments of the present application, the second segmentation result is obtained through the obtained corrected mask image patch. In this way, the complete contour of the target object can be reconstructed, thereby solving the problem of inaccurate contours that may exist in the preliminary segmentation result.
[0122] Please refer to Figure 6 , which shows the data flow diagram of the image processing method provided by the present application. As Figure 6 shown, for the second segmentation result at a position (i.e., at a target window) in the target image 611, it can be obtained through the following scheme:
[0123] (1) Obtain the target image 611, and the first segmentation result corresponding to the target image 611 is the target mask image 612;
[0124] (2) Determine the sizes of the first window 621 and the second window 622 based on the size determination scheme of the target window provided in the above embodiments;
[0125] (3) The first target image patch 631 can be determined by using the first window 621 and the target image 611, and the second target image patch 641 can be determined by using the second window 622 and the target image 611. Correspondingly, the first mask image patch 632 can be determined by using the first window 621 and the target mask image 612, and the second mask image patch 642 can be determined by using the second window 622 and the target mask image 612;
[0126] (3) Image enhancement processing is respectively performed on the first target image block 631 and the first mask image block 632, and a first enhanced target image block 651 and a first enhanced mask image block 652 can be obtained. Here, it can be seen that compared with the first target image block 631, the overall content range included in the first enhanced target image block 651 does not change, but only the image resolution increases. At the same time, compared with the second target image block 641, although the image resolutions of the first enhanced target image block 651 and the second target image block 641 are the same, the image contents included are different. Similarly, the above relationship also holds among the first enhanced mask image block 652, the second mask image block 642, and the first mask image block 632.
[0127] (4) The first fusion image information 662 is obtained according to the first enhanced target image block 651 and the first enhanced mask image block 652;
[0128] (5) The second fusion image information 661 is obtained according to the second target image block 641 and the second mask image block 642;
[0129] (6) The first fusion image information 662 and the second fusion image information 661 are input into the target network 67, and a corrected mask image block 68 output by the target network 67 is obtained.
[0130] Based on the embodiments provided in this application, accurate segmentation of specific positions in the target image can be achieved, and the detail performance and quality of the segmentation result can be improved.
[0131] Based on the foregoing embodiments, an image processing device is provided in an embodiment of this application. The device includes each unit included and each module included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; during implementation, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.
[0132] Figure 7 As shown in the composition structure diagram of an image processing device provided in an embodiment of this application, Figure 7 as shown, the image processing device 700 includes: a first obtaining module 710, a determining module 720, a second obtaining module 730, and a third obtaining module 740, where:
[0133] A first acquisition module 710, configured to acquire a first segmentation result corresponding to a target image, where the first segmentation result includes a target mask image corresponding to a target object in the target image;
[0134] A determination module 720, configured to determine a target window corresponding to an edge region of the target mask image, where a size of the target window is determined according to the target object;
[0135] A second acquisition module 730, configured to acquire a target image block and a mask image block, where the target image block is an image block corresponding to the target window in the target image, and the mask image block is an image block corresponding to the target window in the target mask image;
[0136] A third acquisition module 740, configured to acquire a second segmentation result of the target image according to the target image block and the mask image block, where a segmentation accuracy of the second segmentation result is greater than a segmentation accuracy of the first segmentation result.
[0137] In some embodiments, the apparatus further includes a size determination module; the size determination module is configured to acquire size information of the target object; and determine the size of the target window according to the size information of the target object.
[0138] In some embodiments, the size determination module is further configured to determine the size information of the target object according to the target mask image.
[0139] In some embodiments, the apparatus further includes a size determination module; the size determination module is configured to acquire smoothness information corresponding to a contour of the target object; and determine the size of the target window according to the smoothness information.
[0140] In some embodiments, the size determination module is further configured to acquire edge information of the target mask image; and determine the smoothness information according to the edge information.
[0141] In some embodiments, the apparatus further includes a size determination module; the target window includes a first window and a second window, the first window is smaller than the second window, and the first window is located within the second window; the size determination module is configured to determine the size of the first window and the size of the second window according to the target object.
[0142] In some embodiments, the size determination module is further configured to determine a size ratio relationship between the first window and the second window according to the target object; and determine the size of the first window and the size of the second window according to the size ratio relationship.
[0143] In some embodiments, the target image block includes a first target image block obtained from the target image through the first window and a second target image block obtained through the second window; the mask image block includes a first mask image block obtained from the target mask image through the first window and a second mask image block obtained through the second window. The third obtaining module is further configured to perform image enhancement processing on the first target image block and the first mask image block to obtain a first enhanced target image block and a first enhanced mask image block, where the resolutions of the first enhanced target image block and the first enhanced mask image block are the same as those of the second target image block; obtain first fusion image information according to the first enhanced target image block and the first enhanced mask image block; obtain second fusion image information according to the second target image block and the second mask image block; input the first fusion image information and the second fusion image information into a target network to obtain a corrected mask image block output by the target network; and obtain the second segmentation result according to the corrected mask image block.
[0144] In some embodiments, the third obtaining module is further configured to perform image enhancement processing on the first target image block and the first mask image block when the target object meets a target condition, where the target condition includes at least one of the following: the size of the target object is smaller than a target size, and the smoothness of the contour of the target object is less than a target smoothness.
[0145] The description of the above device embodiments is similar to that of the above method embodiments and has similar beneficial effects to those of the method embodiments. In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the methods described in the above method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0146] It should be noted that in the embodiments of the present application, if the above image processing method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related art can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. In this way, the embodiments of the present application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.
[0147] An embodiment of the present application provides a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above method.
[0148] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements some or all of the steps in the above method. The computer-readable storage medium can be transient or non-transient.
[0149] An embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code runs in a computer device, the processor in the computer device executes to implement some or all of the steps in the above method.
[0150] An embodiment of the present application provides a computer program product. The computer program product includes a non-transient computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above method. The computer program product can be specifically implemented by means of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0151] It should be noted here that the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities can be referred to each other. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects to the method embodiments. For the technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.
[0152] Figure 8 The following is a schematic diagram of the hardware entity of a computer device provided by an embodiment of the present application. As Figure 8 shown, the hardware entity of the computer device 800 includes: a processor 801 and a memory 802. Among them, the memory 802 stores a computer program that can run on the processor 801. When the processor 801 executes the program, it implements the steps in the method of any of the above embodiments.
[0153] The memory 802 stores a computer program that can run on the processor. The memory 802 is configured to store instructions and applications executable by the processor 801, and can also cache data to be processed or already processed by the processor 801 and each module in the computer device 800 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM).
[0154] When the processor 801 executes the program, it implements the steps of the image processing method in any one of the above. The processor 801 generally controls the overall operation of the computer device 800.
[0155] An embodiment of the present application provides a computer storage medium. The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the image processing method in any one of the above embodiments.
[0156] It should be noted here that the descriptions of the above storage medium and device embodiments are similar to the descriptions of the above method embodiments, and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the storage medium and device embodiments of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.
[0157] The above processor can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that other electronic devices for implementing the functions of the above processor are also possible, and the embodiments of the present application do not make specific limitations.
[0158] The above computer storage medium / memory can be a read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it can be various terminals including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0159] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the order numbers of the above steps / processes do not mean the order of execution. The order of execution of each step / process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0160] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, the element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the element.
[0161] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.
[0162] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0163] In addition, in each embodiment of the present application, all the functional units can be integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit. The above integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional units. Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage medium includes various media that can store program codes, such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs.
[0164] Alternatively, if the above integrated units of the present application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The foregoing storage medium includes various media that can store program codes, such as removable storage devices, ROM, magnetic disks, or optical discs.
[0165] As described above, this is only the implementation mode of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.
Claims
1. An image processing method, comprising: Obtaining a first segmentation result corresponding to the target image, wherein the first segmentation result includes a target mask image corresponding to the target object in the target image; Determine a target window corresponding to an edge area of the target mask image, wherein a size of the target window is determined according to the target object; Obtaining a target image block and a mask image block, wherein the target image block is an image block in the target image corresponding to the target window, and the mask image block is an image block in the target mask image corresponding to the target window; A second segmentation result of the target image is obtained according to the target image block and the mask image block, and the segmentation accuracy of the second segmentation result is greater than the segmentation accuracy of the first segmentation result.
2. The method according to claim 1, further comprising: Obtaining size information of the target object; The size of the target window is determined according to the size information of the target object.
3. The method according to claim 2, wherein obtaining the size information of the target object comprises: The size information of the target object is determined according to the target mask image.
4. The method according to any one of claims 1 to 3, further comprising: Obtaining smoothness information corresponding to the contour of the target object; The size of the target window is determined according to the smoothness information.
5. The method according to claim 4, wherein obtaining the smoothness information corresponding to the contour of the target object comprises: Obtaining edge information of the target mask image; The smoothness information is determined according to the edge information.
6. The method according to any one of claims 1 to 3, wherein the target window comprises a first window and a second window, the first window is smaller than the second window, and the first window is located within the second window, and the method further comprises: The size of the first window and the size of the second window are determined according to the target object.
7. The method according to claim 6, wherein determining the size of the first window and the size of the second window according to the target object comprises: Determining a size ratio relationship between the first window and the second window according to the target object; The size of the first window and the size of the second window are determined according to the size ratio relationship.
8. The method according to claim 6, wherein the target image block comprises a first target image block obtained from the target image through the first window and a second target image block obtained through the second window; the mask image block comprises a first mask image block obtained from the target mask image through the first window and a second mask image block obtained through the second window, and obtaining a second segmentation result of the target image according to the target image block and the mask image block comprises: Performing image enhancement processing on the first target image block and the first mask image block to obtain a first enhanced target image block and a first enhanced mask image block, wherein the resolution of the first enhanced target image block and the first enhanced mask image block is the same as that of the second target image block; Obtaining first fused image information according to the first enhanced target image block and the first enhanced mask image block; Obtaining second fused image information according to the second target image block and the second mask image block; Inputting the first fused image information and the second fused image information into a target network to obtain a corrected mask image block output by the target network; The second segmentation result is obtained according to the modified mask image block.
9. The method according to claim 8, wherein the performing image enhancement processing on the first target image block and the first mask image block comprises: When the target object meets the target condition, the first target image block and the first mask image block are subjected to image enhancement processing, and the target condition includes at least one of the following: the size of the target object is smaller than the target size, and the smoothness of the target object contour is smaller than the target smoothness.
10. An image processing device, comprising: A first obtaining module, configured to obtain a first segmentation result corresponding to a target image, wherein the first segmentation result includes a target mask image corresponding to a target object in the target image; A determination module, used to determine a target window corresponding to an edge area of the target mask image, wherein a size of the target window is determined according to the target object; a second obtaining module, configured to obtain a target image block and a mask image block, wherein the target image block is an image block in the target image corresponding to the target window, and the mask image block is an image block in the target mask image corresponding to the target window; The third obtaining module is used to obtain a second segmentation result of the target image according to the target image block and the mask image block, wherein the segmentation accuracy of the second segmentation result is greater than the segmentation accuracy of the first segmentation result.