Image segmentation method and device and image segmentation model
By determining the motion state information of the image frame in image segmentation and selecting a matching image segmentation model, the problem of reducing image segmentation accuracy caused by moving objects is solved, and more efficient and accurate image segmentation is achieved.
Patent Information
- Application Number
- CN202510186571.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-30
AI Technical Summary
During the image segmentation process, especially in the case of moving objects, the object movement offset in adjacent frame images is too large, resulting in a decrease in the accuracy of image segmentation.
By obtaining the current frame image group, determining its motion state information, and selecting a target image segmentation model that matches the motion state information, the current frame image group is processed. This model contains different number of fusion neural network modules for the fusion of the previous frame image information and the current frame image information.
Improve the accuracy and efficiency of image segmentation, especially in the case of moving objects, reduce information loss at the edge, enhance model performance and reduce power consumption.
Smart Images

Figure CN120070472A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and particularly to an image segmentation method, apparatus, and image segmentation model. Background Art
[0002] Currently, image segmentation technology has been widely applied in scenarios such as background blurring, virtual background, video portrait matting, driverless, security monitoring, etc. To meet the high-precision and high-real-time image segmentation requirements in these application scenarios, open-source image segmentation algorithms are usually adopted to quickly identify the object segmentation regions in the image to be processed, so as to implement the image processing tasks in the corresponding scenarios.
[0003] However, in the process of image segmentation in application scenarios such as the motion analysis of athletes on the sports field, high-speed traffic flow monitoring, or defect detection of moving items on the conveyor belt, it is very easy for the object to be segmented to be in a moving state, resulting in too large a movement offset of the object in adjacent frame images, reducing the accuracy of image segmentation. Summary of the Invention
[0004] In view of the above problems, this application provides the following solutions:
[0005] In a first aspect of this application, an image segmentation method is provided, and the method includes:
[0006] Obtain the current frame image group to be processed;
[0007] Process the current frame image group to determine the motion state information of the shooting object in the current frame image group;
[0008] Select a target image segmentation model that matches the motion state information, and process the current frame image group to obtain the current frame image segmentation result; the number of fusion neural network modules in the working state in the target image segmentation models that match different motion state information is different, and the fusion neural network module is used for fusing the information of the previous frame image and the current frame image.
[0009] In a possible implementation, the step of selecting a target image segmentation model that matches the motion state information and processing the current frame image group to obtain the current frame image segmentation result includes:
[0010] According to the motion state information, select a target image segmentation model that matches the motion state information from multiple trained candidate image segmentation models;
[0011] Input the current frame image group into the target image segmentation model for processing to obtain the current frame image segmentation result.
[0012] In a possible implementation, selecting a target image segmentation model that matches the motion state information and processing the current frame image group to obtain a current frame image segmentation result includes:
[0013] According to the motion state information, select a target decoding module that matches the motion state information from multiple candidate decoding modules included in the target image segmentation model;
[0014] Based on the encoding module and the target decoding module in the target image segmentation model, process the current frame image group to obtain a current frame image segmentation result;
[0015] Among them, the number of fusion neural network modules included in different candidate decoding modules is different.
[0016] In a possible implementation, selecting a target image segmentation model that matches the motion state information and processing the current frame image group to obtain a current frame image segmentation result includes:
[0017] According to the motion state information, select a target fusion neural network module that matches the motion state information from multiple fusion neural network modules included in the decoding module of the target image segmentation model, and use the other fusion neural network modules among the multiple fusion neural network modules as non-target fusion neural network modules;
[0018] Input the current frame image group into the target image segmentation model, and obtain a current frame image segmentation result under the condition of controlling the non-target fusion neural network modules to be in a non-working state.
[0019] In a possible implementation, where:
[0020] The larger the motion offset range of the shooting object represented by different motion state information between adjacent frame image groups, the fewer the number of fusion neural network modules in the target image segmentation model that match the motion state information and are in a working state;
[0021] The direction of decreasing the number of fusion neural network modules in the working state starts from the bottom layer of the target image segmentation model.
[0022] In a possible implementation, obtaining the current frame image group to be processed includes:
[0023] Obtain the current frame original image in the image sequence;
[0024] Perform downsampling on the current frame original image to obtain a current frame image to be segmented corresponding to the current frame original image; the image size of the current frame image to be segmented is smaller than the image size of the current frame original image.
[0025] Among them, the current frame original image and the current frame image to be segmented form a current frame image group.
[0026] In a possible implementation, selecting a target image segmentation model that matches the motion state information and processing the current frame image group to obtain a current frame image segmentation result includes:
[0027] Selecting an encoding module in the target image segmentation model that matches the motion state information to extract features from the current frame image to be segmented in the current frame image group, and obtaining current frame image features of multiple different sizes;
[0028] Selecting a decoding module in the target image segmentation model that matches the motion state information to decode the current frame image features of multiple different sizes, and obtaining current frame image decoding information;
[0029] According to the upsampling module in the target image segmentation model, processing the current frame image decoding information and the current frame original image in the current frame image group to obtain a current frame image segmentation result.
[0030] In a possible implementation, processing the current frame image group to determine the motion state information of the shooting object in the current frame image group includes any of the following implementation manners:
[0031] Performing motion detection on the current frame image to be segmented included in the current frame image group to obtain the motion state information of the shooting object in the current frame image to be segmented;
[0032] Obtaining the action features of the shooting object in the current frame image to be segmented included in the current frame image group, and determining the motion state information of the shooting object in the current frame image to be segmented according to the action features;
[0033] Obtaining the motion offset between the current frame image to be segmented included in the current frame image group and the previous frame image to be segmented included in the previous frame image group, and determining the motion offset as the motion state information of the shooting object in the current frame image group;
[0034] Among them, target image segmentation models containing different numbers of fusion neural network modules in the working state respectively correspond to different motion offset ranges.
[0035] In a possible implementation, selecting a target image segmentation model that matches the motion state information and processing the current frame image group to obtain a current frame image segmentation result includes:
[0036] Select a target image segmentation model that matches the motion state information, process an image group of continuously specified number of frames, and obtain the corresponding frame image segmentation results, so as to, after obtaining the next image group of the specified number of frames, perform processing on the next image group to determine the motion state information of the shooting object in the next image group;
[0037] Among them, the continuously specified number of frames is multiple consecutive frames starting from the current frame;
[0038] The fewer the number of fusion neural network modules in the working state in the target image segmentation model, the larger the continuously specified number of frames.
[0039] The second aspect of the present application provides an image segmentation model, including: an encoding module and a decoding module, where:
[0040] At least one layer in the decoding module does not include a fusion neural network module, or at least one layer includes a fusion neural network module in a non-working state.
[0041] In a possible implementation, the number of decoding modules is multiple, and among the multiple layers included in each of the multiple decoding modules, the number of layers including the fusion neural network module is different.
[0042] The third aspect of the present application provides an image segmentation device, including:
[0043] A current frame image group acquisition module, configured to acquire a current frame image group to be processed;
[0044] A motion state information determination module, configured to process the current frame image group to determine the motion state information of the shooting object in the current frame image group;
[0045] A selection module, configured to select a target image segmentation model that matches the motion state information, process the current frame image group, and obtain a current frame image segmentation result;
[0046] Among them, the number of fusion neural network modules in the working state in the target image segmentation models that match different motion state information is different, and the fusion neural network module is used for fusing the information of the previous frame image and the current frame image. Description of the Drawings
[0047] Combined with the drawings and referring to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn to scale.
[0048] Figure 1Schematic flowchart of the image segmentation method proposed in the first embodiment of the present application;
[0049] Figure 2 Schematic flowchart of the image segmentation method proposed in the second embodiment of the present application;
[0050] Figure 3 Schematic diagram of the network structure of a traditional image segmentation model;
[0051] Figure 4 Schematic diagram of an image segmentation model structure and its operation process applicable to the image segmentation method proposed in the present application;
[0052] Figure 5 Schematic diagram of another image segmentation model structure and its operation process applicable to the image segmentation method proposed in the present application;
[0053] Figure 6 In the image segmentation method proposed in the present application, an optional flowchart for obtaining the motion state information of the photographed object;
[0054] Figure 7 Schematic flowchart of the image segmentation method proposed in the third embodiment of the present application;
[0055] Figure 8 Schematic diagram of the network structure of an image segmentation model applicable to the image segmentation method proposed in the present application;
[0056] Figure 9 Schematic diagram of the network structure of another image segmentation model applicable to the image segmentation method proposed in the present application;
[0057] Figure 10 Schematic diagram of yet another image segmentation model structure and its operation process applicable to the image segmentation method proposed in the present application;
[0058] Figure 11 Schematic diagram of the structure of an image segmentation device provided in the embodiments of the present application;
[0059] Figure 12 Schematic diagram of the hardware structure of an electronic device applicable to the image segmentation method proposed in the present application. Detailed implementation manners
[0060] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the implementation part of the present application are only for explaining the specific embodiments of the present application, rather than aiming to limit the present application. The embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art know that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are equally applicable to similar technical problems.
[0061] In the description and claims of this application and the above-mentioned drawings, terms such as "first" and "second" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing embodiments of this application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device comprising a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.
[0062] To solve the above problems, an embodiment of this application provides an image segmentation method. The image segmentation method of the embodiment of this application will be introduced in detail below with reference to the drawings.
[0063] Refer to Figure 1 , which is a schematic flowchart of the image segmentation method proposed in the first embodiment of this application. This method can be applied to an electronic device, which can be a server or a terminal device with image processing capabilities, or can be implemented by a server in cooperation with a terminal device. The product form of the electronic device can be determined according to actual needs.
[0064] As Figure 1 shown, the image segmentation method proposed in this embodiment may include but is not limited to the following steps:
[0065] Step S11, obtain the current frame image group to be processed;
[0066] In the embodiments of the present application, the current frame image group may be based on the original image currently captured by the image acquisition device in the image shooting mode or the video shooting mode, or may be based on the current frame original image to be processed in the image sequence (such as video stream or image stream) from the server, or may be based on the current frame original image in the image sequence from the storage device or other data sources, etc. Therefore, the present application may determine the current frame original image obtained in this way but not limited to this way as the current frame image group, or may, after obtaining the current frame original image, process the current frame original image (such as downsampling or cropping, etc.) to form the current frame image group with the processed image and the current frame original image, that is to say, the current frame image group obtained in step S11 may only contain one original image (i.e., the current frame original image), or may contain one original image (i.e., the current frame original image) and one processed image (i.e., the image obtained by processing the current frame originally). The present application does not limit the content included in the current frame image group and its obtaining method. The current frame original image used to obtain the current frame image group may come from image sequences such as online or offline video / image streams, etc. Its data source may be a storage device integrated or external to the electronic device, or may be an image acquisition device integrated or external to the electronic device, etc. The present application does not limit the data source of each frame of the original image and may be determined according to the situation.
[0067] Exemplarily, in order to implement the action analysis of athletes on the sports field, the athletes on the sports field may be photographed, and the photographed image containing the athletes may be processed by the image segmentation method proposed by the present application to obtain the segmented area where the athletes are located in the image and identify the current actions of the athletes. Based on this, the corresponding frame image group may be obtained based on the original image captured in real time on the sports field, or the image sequence or video stream captured by the image acquisition device (such as a camera or a monitoring device, etc.) may be transmitted to the electronic device for caching, so that when the electronic device executes the image segmentation method proposed by the present application, the current frame original image cached may be read, and the current frame image group to be processed currently may be obtained accordingly.
[0068] Optionally, the electronic device may also determine the data source providing the image data of the sports field and establish a communication connection with the data source during the process of responding to the action analysis task of the athletes specified on the sports field, or in response to the athlete image segmentation subtask obtained by splitting the task, and obtain the current frame image group to be processed according to the method described above but not limited to it.
[0069] Similarly, in other application scenarios such as high-speed traffic flow monitoring or defect detection of items moving on a conveyor belt, when performing image segmentation tasks corresponding to the scenarios (such as vehicle recognition and segmentation, item defect recognition and segmentation, etc.), the electronic device can, in response to the image segmentation task, based on the current frame original image in the corresponding captured high-speed image sequence or conveyor belt image sequence, etc., obtain the current frame image group to be processed, so as to realize the corresponding image processing requirements by executing the method proposed in this application, such as determining whether there are defects in the high-speed traffic flow or each item on the conveyor belt, and if so, determining the defect positions in the items. The implementation process is not elaborated one by one in this application.
[0070] It should be understood that in application scenarios including but not limited to those listed above, each frame image group to be processed can be consecutive frames or non-consecutive frames. Generally, it is required that the corresponding original image contains the object to be segmented, such as athletes, vehicles or items, etc. Optionally, the images taken at the same time can be multiple images taken from multiple perspectives simultaneously, which can correspond to the same frame number or a group of consecutive frame numbers to obtain the image group of the corresponding frame. Then, according to the image segmentation method proposed in this application, the corresponding segmentation result can be accurately and reliably obtained. For example, based on the multi-perspective image segmentation result of an item to improve the accuracy of item defect detection; based on the multi-perspective image segmentation result of an athlete, improve the accuracy and reliability of athlete motion analysis and avoid misjudgment, etc. The image segmentation process for each frame image group is similar and is not elaborated one by one in this application.
[0071] Step S12: Process the current frame image group to determine the motion state information of the shooting object in the current frame image group;
[0072] In the application scenarios listed above, the shooting object in the shooting scene may be in a static state or a moving state. However, when in a moving state, especially in a high-speed moving state, such as sprinters / speed skaters / skiers in a sports stadium, due to the image acquisition device being unable to reach the same motion state, the relative position relationship between the shooting object and the image acquisition device (which can be represented by shooting parameters such as shooting distance or shooting angle) will constantly change, which will cause a large offset range of the same shooting object in adjacent frame images captured by the image acquisition device, and even cause the shooting object to be located in the edge area of the image and be blurred and distorted, that is, the information loss at the image edge is large, reducing the image segmentation accuracy and efficiency.
[0073] To address the above issues and improve the accuracy and efficiency of image segmentation, embodiments of this application propose dynamically adjusting the image segmentation steps for the current frame image group in combination with the actual motion state of the captured object, so as to solve the problem of low image segmentation accuracy and efficiency caused by using an image segmentation model with a fixed network structure to perform segmentation processing on the original images of captured objects in various motion states through the same image segmentation steps.
[0074] Therefore, after the currently acquired current frame image group to be processed, it is possible to identify the motion state information of the captured object in the current frame image group to determine whether the captured object is currently in a stationary state or a motion state, and in which motion state level such as low speed / medium speed / high speed, etc., and then perform subsequent processing steps based on this motion state information. Among them, in combination with the above description of the current frame image group, in the case where the current frame image group includes the current frame original image and the image after processing the current frame original image, this application will identify the motion state information of the captured object for the processed image. Compared with the implementation method of directly identifying the running state information of the captured object for the current frame original image, since the amount of data contained in the processed image is greatly reduced, the computational amount of motion state information recognition is significantly reduced, the recognition efficiency is improved, and since the motion state of the captured object in the images before and after processing is the same, it is ensured that the motion state information recognized from the processed image is accurate and reliable, thus ensuring the accuracy and reliability of the subsequent model selection result.
[0075] It should be understood that in the case where the current frame image group only includes the current frame original image, this application can directly identify the motion state information of the captured object for the current frame original image to reliably execute the subsequent model selection process. This application does not limit the implementation method of how to identify the motion state information of the captured object from the image (such as the original image or its processed image).
[0076] In some embodiments, this application can identify and track and detect the captured object in the application scenario based on deep learning methods to obtain the motion state information of the captured object. This process can be implemented through a motion detection model trained based on deep learning algorithms. This application does not elaborate on the training and its inference implementation process of this motion detection model. Optionally, this application can also use the optical flow method (such as the Horn-Schunck algorithm (a global optical flow estimation algorithm) or the Lucas-Kanade algorithm (an optical flow estimation algorithm for two-frame difference), etc.) to implement the running state detection of the captured object, that is, based on the pixel intensity change in the current frame image group, obtain the motion state information. This application does not elaborate on the implementation process of how to implement step S12 based on the optical flow method. In practical applications, especially in complex dynamic scenarios, it is possible to preferentially use the optical flow method and deep learning methods to implement step S12.
[0077] In a possible implementation, after estimating the motion parameters of the photographed object by the optical flow method, algorithms such as the Kalman filter can also be used to predict and track the motion trajectory of the photographed object, that is, combining motion estimation and target tracking calculations to determine the actual motion state of the photographed object, improving the reliability and accuracy of the motion state information. In addition, the present application can also implement step S12 by the inter-frame difference method, that is, by comparing the pixel gray value differences of two consecutive frames or multiple groups of frames to determine the current motion state information of the photographed object. Or step S12 can be implemented by the template matching method, that is, by a sparse motion estimation technique to find similar templates in different groups of frame images to estimate the motion state of the photographed object.
[0078] In addition, the present application can also detect the key points and corresponding feature descriptors in the current group of frame images by the feature point matching method, find the most matching key points in another group of frame images (such as the previous group of frame images), and infer the motion state of the photographed object through the matching relationship. The present application does not limit the method for detecting the motion state of the photographed object for implementing step S12, which can include but is not limited to one or more combinations listed above, and can also be determined according to the type and application requirements of the actual application scenario, flexibly selecting the method for detecting the motion state of the photographed object, and the implementation process will not be described in detail one by one in the present application.
[0079] It should be noted that since the same motion state of the photographed object usually exists for a period of time, making the motion states of the photographed object in multiple consecutive groups of frame images the same, therefore, when the photographed object switches to the current motion state, the first group of frame images of it is processed to obtain the motion state information of the photographed object. After obtaining the next group of frame images to be processed, the motion state information can be directly determined as the motion state information of the photographed object in the next group of frame images, without spending a large amount of time and computing resources to process the next group of frame images. Optionally, the present application can pre-configure the number of frames for detecting the motion state of the photographed object according to experience or the speed level of the current motion state, etc. After determining the motion state information of the photographed object in the first group of frame images in a certain motion state (which can be implemented according to the detection method described above), the photographed objects in multiple groups of frame images with the number of frames in this interval all have this motion state information. After passing through this number of frames, step S12 is implemented again according to the detection method described above.
[0080] Step S13: Select a target image segmentation model that matches the motion state information, and process the current group of frame images to obtain the current frame image segmentation result; the number of working fusion neural network modules in the target image segmentation models that match different motion state information is different, and the fusion neural network module is used for fusing the information of the previous frame image and the current frame image.
[0081] Following the above analysis, when determining the motion state information of the shooting object in the current frame image group and learning the motion state of the current shooting object at what speed level, select an image segmentation method that matches the currently obtained motion state information from the image segmentation methods configured for different motion state information or different speed levels of motion states, and process the current frame image group to quickly obtain a high-precision image segmentation result of the current frame.
[0082] Among them, the different image segmentation methods described above can be distinguished by different numbers of fusion neural network modules. Since the role of this fusion neural network module is to fuse the information of the previous frame image and the current frame image to improve the image segmentation accuracy, the more the number of fusion neural network modules corresponding to the actually selected image segmentation method in the process of processing the current frame image group, the more times the fusion step is executed. This will not only increase the image segmentation duration, but also, if the shooting object is in a motion state, the greater the loss of information at the edge, and the lower the accuracy of the finally obtained image segmentation result, which may not meet the subsequent image processing tasks.
[0083] It can be seen that when the shooting object is in a motion state, only by selecting an appropriate number of fusion neural network modules to participate in the processing of the current frame image group can the image segmentation accuracy be guaranteed, that is, the model performance is improved and the power consumption is reduced. For example, the higher the speed level of the current motion state of the shooting object, the fewer the number of fusion neural network modules that can be selected to participate in the processing of the current frame image group to avoid reducing the segmentation accuracy due to fusing too much historical frame image information. Exemplarily, if the image segmentation model used in this application is the RVM (Robust High Resolution Video Matting with Temporal Guidance) model, this application can dynamically select the number of ConvGRU (Convolutional Gated Recurrent Unit) operators included in the decoding module of the RVM model to improve the image segmentation accuracy and image segmentation efficiency.
[0084] Based on the above analysis, in some embodiments, this application can pre-train different image segmentation models corresponding to different motion state information, and the different image segmentation models contain different numbers of fusion neural network modules. In this way, an image segmentation model corresponding to the motion state information of the shooting object in the current frame image group can be directly selected as the target image segmentation model to implement the processing of the current frame image group. In this case, each fusion neural network module included in this image segmentation model is in a working state and participates in the processing of the current frame image group.
[0085] In a possible implementation method, the present application can also pre-determine the number of fusion neural network modules in the working state corresponding to different motion state information. In this way, an image segmentation model can be pre-trained. After obtaining the motion state information of the shooting object in the current frame image group, it is determined which of the multiple fusion neural network modules located in different layers in this image segmentation model need to participate in the processing of the current frame image group. That is, starting from the fusion neural network module in the first layer of the image segmentation model, according to the number of fusion neural network modules matching the motion state information, the fusion neural network modules in the working state are sequentially selected to participate in the processing of the current frame. For the fusion neural network modules in other layers, they are switched to the non-working state and no longer participate in the processing of the current frame image.
[0086] Based on this, the target image segmentation model matching the motion state information in step S13 can be the image segmentation model after selecting the fusion neural network modules in the working state according to the method described above. In this way, when the shooting object is in different motion states, the difference between the target image segmentation models selected by the present application lies in the different numbers of fusion neural network modules in the working state, which can be achieved by controlling the same image segmentation model. Compared with the implementation method of training multiple image segmentation models corresponding to different motion state information respectively, the model training time and resource consumption are greatly reduced.
[0087] It should be noted that for the two implementation methods of implementing step S13 described above, the present application does not limit the network structure of the trained image segmentation model. A combined structure of an encoding module and a decoding module can be adopted. In addition to a certain number of fusion neural network modules, the decoding module can determine other module types and structures according to the algorithm type for training the image segmentation model. It can be determined with reference to the network structure of traditional image segmentation models such as the RVM model, and the present application does not elaborate on this one by one here.
[0088] In addition, for the implementation process of training an image segmentation model and adaptively selecting the number of fusion neural network modules in the working state corresponding to the currently obtained running state information, it can be achieved by selecting a decoding module with the corresponding number of fusion neural network modules from multiple candidate decoding modules and combining it with the encoding module to form the target image segmentation model, or by controlling the fusion neural network modules with this number of fusion neural network modules in a decoding module to enter the working state, etc. It can be determined according to the deployment information such as the number and structure of the image segmentation models actually deployed in the electronic device. The present application does not limit the implementation method of step S13.
[0089] Among them, processing the current frame image group according to the target image segmentation model, the obtained current frame image segmentation result may include one or more of the segmentation contour of the shooting object in the current frame image, the bounding box, the object category, and the confidence score of the shooting object belonging to the object category, etc., which can be determined according to the image segmentation requirements of the actual application scenario, so as to complete subsequent image processing task requirements such as object recognition, scene understanding or target tracking, such as athlete tracking detection and motion analysis, defective product detection, vehicle flow statistics, etc. Therefore, in the face of different application scenarios and image processing task requirements, the forms and contents of the image segmentation results of each frame obtained in this application may be different, and this application does not make restrictions.
[0090] In summary, in the embodiment of this application, after obtaining the current frame image group to be processed, compared with directly processing the current frame image group using an image segmentation model with a pre-trained fixed network structure, this application takes into account the accuracy difference in the image segmentation processing of shooting objects in different motion states by this image segmentation model, and proposes to first determine the motion state information of the shooting object in the current frame image group, and then select a target image segmentation model that matches the motion state information to process the current frame image group, so as to quickly obtain a current frame image segmentation result with higher accuracy. It can be seen that when the motion state information of the shooting object is different, the number of fusion neural network modules in the working state in the target image segmentation model for processing the current frame image group is different. This application adaptively selects the number of fusion neural network modules participating in the current frame image processing according to this motion state information, avoiding reducing the image segmentation accuracy due to too many or too few fusion neural network modules in the working state, and ensuring that the obtained image segmentation result can reliably meet the requirements of subsequent image processing tasks.
[0091] In some embodiments, the above-mentioned target image segmentation model and each candidate image segmentation model can be trained based on common image segmentation models in actual application scenarios. For example, by training based on the RVM algorithm, an RVM model with a new network structure (different from the network structure of the traditional RVM model) is obtained. In this case, each fusion neural network module in the above-mentioned image segmentation model refers to the ConvGRU neural network module in the decoding module of the RVM model. The operation principle of ConvGRU in the RVM model is not elaborated in detail in this application. It should be noted that the image segmentation models mentioned in the context of this application include but are not limited to the RVM model. For fusion neural network modules that capture the relationship between the current frame and the previous frame image information in the image sequence by fusing adjacent frame image information, enabling the image segmentation model to better understand the continuity and changes between frames, facilitating the generation of more coherent and accurate foreground segmentation results (i.e., the image segmentation results of the captured object), and improving the generalization ability of the image segmentation model, it can include but is not limited to the ConvGRU neural network module.
[0092] In a possible implementation, the fusion neural network module can be a Dynamic Feature Fusion (DFF) module, which adaptively fuses multi-scale local features through global information. It receives two input feature maps (for example, the current frame feature map of the current frame image group and the feature map of the previous frame image group, i.e., the previous frame feature map), and generates a fused feature map through operations such as channel concatenation, global pooling, attention mechanism, and channel dimensionality reduction to complete subsequent captured object recognition and segmentation. It can be seen that in this embodiment, the importance of each feature is adjusted through dynamic learning and they are fused together to obtain more accurate and powerful results, which is more suitable for tasks such as image recognition and video analysis that require dynamic adjustment of feature importance. It can be understood that in this embodiment, the target image segmentation model selected in this application can be other types of image segmentation models with a DFF module and different from the RVM model. In a possible implementation, the fusion neural network module in the image segmentation model can also be a FreqFusion feature fusion module, that is, a plug-and-play feature fusion module, which can be used to fuse high-resolution features and low-resolution features. It fuses features from different scales through a specific fusion strategy (such as addition or concatenation) to improve the accuracy and reliability of the captured object in the obtained feature map, so as to better complete the subsequent captured object recognition and segmentation steps.
[0093] In a possible implementation, the fusion neural network module in the image segmentation model can also be a multi-modal fusion module. Especially in some multi-modal image segmentation tasks, the multi-modal fusion module can fuse image features of different modalities (such as optical images and radar images). This module usually combines a convolutional neural network (CNN) and a recurrent neural network (RNN) to extract and fuse features, thereby generating more accurate image segmentation results. The operation principle of this multi-modal fusion module is not described in detail in this application. Based on this, in this embodiment, the target image segmentation model selected in this application and each candidate image segmentation model can be a multi-modal image segmentation model with a multi-modal fusion module, such as M4oE (Medical Multimodal Mixture of Experts, a model for multi-modal medical image segmentation), FusionSAM (Segment Anything Model, a multi-modal image segmentation model for processing multi-modal fusion tasks in natural images), Sa2VA (Dense Video Multi-Modal Large Model, a multi-modal video large model), or a multi-modal segmentation model based on feature fusion such as the MFNet (Multi-spectral Fusion Network) model and the AFNet (Attention Fusion Network) model.
[0094] It should be noted that for a certain number of fusion neural network modules in the image segmentation model proposed in this application, they are preferably configured as the same type of fusion neural network module, including but not limited to the several types of fusion neural network modules listed above. However, in some embodiments, according to actual processing requirements, it can also be a combination of one or more of the various fusion neural network modules listed above to construct an image segmentation model, etc. This application does not limit the number and type of fusion neural network modules in the image segmentation model and can be determined according to the situation. The following embodiments of this application only take the ConvGRU neural network module included in the RVM model as an example for illustration. For other fusion neural network modules with the above corresponding image processing functions, the implementation process of participating in the image segmentation method proposed in this application is similar, and this application does not give detailed examples one by one.
[0095] Refer to Figure 2 , which is a schematic flowchart of the image segmentation method proposed in the second embodiment of this application. This embodiment can refine the description of the acquisition process of each frame image group to achieve the purpose of reducing the processing duration of the image segmentation model and improving the overall segmentation efficiency. As Figure 2 shown, the image segmentation method proposed in this embodiment may include:
[0096] Step S21: Obtain the original image of the current frame in the image sequence;
[0097] Step S22: Downsample the original image of the current frame to obtain the image to be segmented of the current frame corresponding to the original image of the current frame; the image size of the image to be segmented of the current frame is smaller than the image size of the original image of the current frame.
[0098] In the inference application of an image segmentation model, in order to increase the computational amount of image segmentation and improve the segmentation efficiency, before the image segmentation model extracts features from the original image, usually the high-resolution and large-size original image is first downsampled to reduce the performance pressure of the input original image on the subsequent encoding and decoding processes. Taking Figure 3 the traditional image segmentation model shown as an example, in the network structure of this image segmentation model, as Figure 3 shown, a downsampling module is configured before its encoding module. After the original image of the current frame (i.e., the high-resolution image) is input into this image segmentation model, it is necessary to first perform downsampling processing on the high-resolution original image through the downsampling module, and then input the processed low-resolution image into the encoding module for subsequent processing.
[0099] Exemplarily, if the original image input into Figure 3 the shown image segmentation model is a 4K resolution large-size image, the neural network processor needs to consume a long time to perform 4-fold downsampling (i.e., downsampling, which can also be called decimation) on it to obtain a low-resolution and small-size image suitable for subsequent processing. The time-consuming of this downsampling process occupies 40% of the entire image segmentation processing process, bringing great performance pressure to the processing performance of the entire image segmentation model and resulting in a relatively high model power consumption.
[0100] To improve the above problems, this application proposes to improve the traditional image segmentation model as Figure 3 shown, remove the downsampling module of this image segmentation model, simplify the calculation steps of the image segmentation model, and still be able to use the image segmentation principle of this image segmentation model. This application proposes that before using the image segmentation model, the high-resolution and large-size original image of the current frame is downsampled in advance to obtain the low-resolution and small-size image to be segmented of the current frame required by the encoding module. Then, the original image of the current frame and the image to be segmented of the current frame corresponding to the original image of the current frame form the image group of the current frame, and then the target image segmentation model matching the motion state information of the shooting object is called / started, so that the target image segmentation model no longer needs to spend a long time performing downsampling processing on the input image, that is, the online processing duration of the model is reduced and the model power consumption is lowered.
[0101] Based on this, still taking the original current frame image with a large size of 4K resolution as an example for illustration. After obtaining the original current frame image with 4K resolution from the image sequence in the embodiments of the present application, a suitable downsampling operator or image processing software can be flexibly selected to process the original current frame image, and a current frame image to be segmented with a small size of 1K resolution is obtained. The current frame image group composed of the original current frame image and the current frame image to be segmented is used as the model input. Compared with the traditional image segmentation model shown in Figure 3 which only takes the original current frame image as the model input, the present application directly provides a small-size image to be segmented of the same frame content for the image segmentation model, so that the subsequent model can directly perform encoding processing on the current frame image to be segmented without having to perform downsampling processing on the input image, saving the consumption of the model downsampling processing.
[0102] Among them, since the downsampling process of the original current frame image in the present application is no longer limited within the image segmentation model, the present application can flexibly select a more suitable downsampling operator or image processing software according to the actual application scenario to implement the preprocessing process of the image sequence, that is, step S21 and step S22, improving the diversity and accuracy of the original image downsampling method and helping to improve the model performance.
[0103] In a possible implementation, if the image sequence is an offline video source, each frame of the original image to be processed in the offline video source can be extracted based on ffmpeg (an open-source computer program that can be used to record, convert digital audio and video, and convert them into streams), and downsampled to obtain video data containing each frame of the image to be segmented. It should be understood that in this preprocessing process, other high-precision downsampling operators can also be used to process each frame of the original image in the offline video source, which will not have an adverse impact on the performance and functions of the image segmentation method in the actual application scenario.
[0104] In a possible implementation, if the image sequence is an online video source, such as the image data collected in real time by the camera of an electronic device, etc., the resource consumption of the electronic device for executing the downsampling operator needs to be considered. When selecting the downsampling operator, its impact on the performance and power consumption of the electronic device needs to be considered, but there are still multiple downsampling operator candidates and it will not be limited to one downsampling operator. In this scenario, the CPU or GPU (Graphics Processing Unit) of the electronic device and other processors can pre-downsample the original image by several times (such as reducing it by 2 times or 4 times, etc., to reduce the size and data volume of the image, and the downsampling multiple can be determined according to actual needs, and the present application does not limit this), to obtain the corresponding frame image to be segmented with a small size. The implementation process is not described in detail in the present application.
[0105] Based on the above analysis, in the implementation process of step S22, the Nearest Neighbor Interpolation method can be adopted, that is, by selecting the original pixel in the original image of the current frame that is closest to the target pixel to generate the downsampled image (i.e., the image to be segmented in the current frame), and the downsampling can be quickly completed.
[0106] Optionally, the present application can also adopt the Bilinear Interpolation method, that is, by considering the four nearest neighbor pixels around the target pixel in the original image of the current frame and calculating the value of the target pixel by weighted averaging according to their distances to obtain the image to be segmented in the current frame. Or adopt the Bicubic Interpolation method, that is, by considering the 16 nearest neighbor pixels around the target pixel in the original image of the current frame and using a cubic polynomial function for weighted averaging to obtain the image to be segmented in the current frame. These two downsampling algorithms can obtain a higher-quality image to be segmented, but the computational amount is relatively large, and they are more suitable for offline processing scenarios to avoid affecting the performance and power consumption of the electronic device in the online scenario. Of course, if the performance and power consumption of the electronic device are sufficient, these two downsampling operators can still be selected in the online scenario.
[0107] In a possible implementation, the present application can also adopt the Gaussian Downsampling method, that is, by using a Gaussian filter to smooth the original image of the current frame and then performing downsampling to obtain the image to be segmented in the current frame. This method can reduce the noise and details of the image and can maintain the overall structure and texture of the image. Or adopt the Median Downsampling method, that is, by dividing the original image of the current frame into small blocks and calculating the median of each block to generate the downsampled image, that is, the image to be segmented in the current frame. This method can also reduce the noise and details of the image and maintain the edges and structure of the image.
[0108] In practical applications, in addition to the several downsampling methods listed above, if it is necessary to retain the bright part details of the original image, the present application can also choose to use the Max Downsampling method, that is, by dividing the image into small blocks and calculating the maximum value of each block to generate the image to be segmented in the current frame. If it is necessary to retain the dark part details of the original image, the present application can also choose to use the Min Downsampling method, that is, by dividing the image into small blocks and calculating the minimum value of each block to generate the image to be segmented in the current frame, etc. The present application can flexibly select the downsampling operator to implement step S22 according to the image processing requirements of the actual application scenario, and the implementation process will not be elaborated in the present application. It can be seen that the present application is relative to Figure 3The traditional image segmentation model shown can only adopt the implementation method of using a fixed downsampling operator to downsample the input original image, which improves the implementation diversity and flexibility of downsampling the original image and better meets the image processing requirements of actual application scenarios.
[0109] Step S23: Perform motion detection on the image to be segmented in the current frame to obtain the motion state information of the object being photographed in the image to be segmented in the current frame.
[0110] Combined with the detection method of the motion state of the object being photographed described above, compared with the implementation method of processing the original image in the current frame to obtain the motion state information of the object being photographed, the embodiment of the present application proposes to process the image to be segmented in the current frame to obtain the motion state information of the same object being photographed. Since the image size of the image to be segmented in the current frame is smaller than the image size of the original image in the current frame, but the downsampling process in step S21 retains the motion state information of the object being photographed, the computational amount of processing the image to be segmented in the current frame is smaller than that of processing the original image in the current frame. On the basis of ensuring the obtained motion state information of the object being photographed, the motion state detection efficiency can be greatly improved.
[0111] In a possible implementation, the present application can perform motion detection on the image to be segmented in the current frame based on a pre-trained motion detection model to obtain the motion state information of the current object being photographed. Among them, the motion state detection model can be obtained by training an initial motion state detection model such as the YOLO (You Only Look Once) series model or the Mediapipe model based on deep learning on multiple frame sample images of the object being photographed in different running states. It can analyze the motion state of the object being photographed by calculating the joint angles through real-time monitoring of the body key points of the object being photographed. The implementation process is not described in detail in the present application.
[0112] In a possible implementation, the present application can also implement step S23 based on the motion state detection method of traditional machine learning. The motion state of the object being photographed in the image to be segmented in the current frame can be detected through machine learning algorithms such as decision trees, neural networks, or classifiers for motion state recognition. Or based on the feature screening method of multi-information fusion, a small number of feature values are extracted and combined with multiple classifiers to achieve high-precision motion state recognition. In addition, the spatio-temporal location and classification of the actions of the object being photographed in the image to be segmented in the current frame can also be performed through algorithms such as Spatio-Temporal Action Detection (STAD) to obtain the motion state information of the object being photographed, etc.
[0113] It can be seen that in the implementation process of processing the current frame of the image to be segmented to obtain the motion state information of the shooting object in the current frame of the image to be segmented, in addition to the above several implementation methods described and included in step S23, the present application can also obtain the action characteristics of the shooting object in the current frame of the image to be segmented, and determine the motion state information of the shooting object in the current frame of the image to be segmented based on the action characteristics. This implementation process can be based on but not limited to the STAD model, or after using a feature extraction model to obtain the action characteristics of the shooting object in the current frame of the image to be segmented, analyze or classify the extracted action characteristics, etc., to determine the motion state category of the shooting object at present, so as to obtain the motion state information of the shooting object in the current frame of the image to be segmented, etc.
[0114] It should be noted that for the implementation method of processing the current frame of the image to be segmented proposed in the present application to obtain the motion state information of the shooting object in the current frame of the image to be segmented, it includes but is not limited to the several implementation methods described above. According to the requirements of the actual application scenario, such as in the application scenarios of real-time and high-precision, or in the application scenarios with limited resources / requiring quick implementation, the respective matching motion state detection methods can be different.
[0115] Step S24, select the encoding module in the target image segmentation model that matches the motion state information, and perform feature extraction on the current frame of the image to be segmented to obtain current frame image features of multiple different sizes;
[0116] In the embodiment of the present application, the image segmentation model includes an encoding module, a decoding module, and an upsampling module. After determining the motion state information of the current shooting object, in order to quickly obtain a high-precision current frame image segmentation result, if multiple image segmentation models corresponding to different motion state information are pre-trained, the target image segmentation model that matches the motion state information of the shooting object in the current frame of the image to be segmented will be selected, and the pre-processed current frame of the image to be segmented and the current frame of the original image will be used as model inputs and input into the target image segmentation model for processing. In this case, the encoding modules of different image segmentation models can be the same.
[0117] Optionally, as described above for the image segmentation model, a target image segmentation model can be pre-trained, and the number of fusion neural network modules in the working state that match the actual motion state information of the shooting object can be selected to complete the image segmentation process. In this case, the encoding module with a fixed structure of the image segmentation model can be used to perform feature extraction on the current frame of the image to be segmented.
[0118] Among them, during the process of the target image segmentation model processing the input image, the present application does not require the target image segmentation model to perform downsampling on the input image online. Instead, it can directly input the current frame of the image to be segmented into the encoding model for feature extraction. For example, by using convolutional layers of different sizes through convolutional operations, features such as the color, texture, and shape of the current frame of the image to be segmented are gradually extracted to obtain multiple current frame image features of different sizes, that is, feature maps of different resolutions, for subsequent feature fusion and segmentation tasks. The present application does not limit the network structure of the encoding module, which can be determined according to the type of the target image segmentation model.
[0119] In a possible implementation, if the target image segmentation model belongs to the RVM model, its encoding module can use mobileNetV3-large (a lightweight neural network, which is used as a feature extractor for complex computer vision tasks) as the backbone (main network), and then connect the LR-ASPP (Lite R-ASPP, Lite Reduced Atrous Spatial Pyramid Pooling, a lightweight semantic segmentation network suitable for mobile devices) module to provide efficient segmentation performance. Under this network structure, the encoding module can use feature extraction network layers (such as convolutional layers) of scales of 1 / 2, 1 / 4, 1 / 8, 1 / 16, etc. to sequentially extract features from the current frame of the image to be segmented, and correspondingly obtain the current frame image features of the corresponding sizes. However, it is not limited to the size described in this embodiment and can be flexibly adjusted according to the image processing requirements of the actual application scenario.
[0120] It should be understood that, in addition to the several multi-modal image segmentation models listed above, the above-mentioned target image segmentation model of the present application can also be a SegNet model (A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation), an FCN (Fully Convolutional Networks) model, a U-Net (Convolutional Networks for Biomedical Image Segmentation) model, a DeepLab (Deep Learning for Semantic Segmentation) model, a Mask R-CNN (Mask Region-based Convolutional Neural Network) model, or a PSPNet (Pyramid Scene Parseing Network) model, etc. The encoding modules in different types of target image segmentation models can be different, which can be determined according to the design principles of the corresponding models, and the present application will not give detailed examples one by one here.
[0121] Step S25: Select the decoding module in the target image segmentation model that matches the motion state information, and decode the current frame image features of the multiple different sizes to obtain the current frame image decoding information; the number of fusion neural network modules in the working state in the decoding modules that match different motion state information is different;
[0122] Following the above analysis, for the current frame image features of multiple different sizes of the current frame image to be segmented output by the encoding module, they can be input into different network layers of the decoding module in the target image segmentation model that matches the motion state information of the current shooting object. After fusing the output of the previous network layer with the current frame feature map of the corresponding size by using this network layer, combined with the fusion neural network module in the working state (such as the ConvGRU neural network module in the decoding module of the RVM model, or the multi-modal fusion module in the multi-modal image segmentation model, or the DFF module in the D-Net (Dynamic Neural Network) model, etc.), decoding is performed to obtain the decoding information corresponding to the size of this network layer, and the current frame image decoding information of the current frame image to be segmented is obtained layer by layer.
[0123] In a possible implementation, if an image segmentation model (i.e., the target image segmentation model) is pre-trained in this application and the image segmentation model includes an encoding module and a decoding module, as Figure 4 shown, step S24 can specifically extract features from the current frame image to be segmented according to the encoding module in the target image segmentation model, that is, the selection operation of the encoding module does not need to be performed during the encoding stage. However, during the decoding stage, it is necessary to select the number of working fusion neural network modules that match the motion state information of the current shooting object, control the corresponding number of fusion neural network modules in the decoding module to enter the working state, and other fusion neural network modules are in the non-working state. By obtaining a decoding module network structure that matches the current obtained motion state information, decode the current frame image features of multiple different sizes output by the encoding module to obtain the current frame image decoding information.
[0124] Of course, if the pre-trained image segmentation model includes multiple decoding modules corresponding to different motion state information respectively, as Figure 5 shown, during the implementation of step S25, a decoding module that matches the motion state information of the currently obtained shooting object can be directly selected as the target decoding module, and the current frame image features of multiple different sizes output by the decoding module are decoded. The decoding implementation process is not described in detail in this application and can be determined according to the decoding principle of the type of image segmentation model selected for use.
[0125] Exemplarily, when the target image segmentation model is trained based on the RVM algorithm, its decoding module can be a cyclic decoder, and multiple sizes of fusion neural network modules (such as ConvGRU) in the working state can be used to aggregate time information, that is, fuse the previous frame image information with the current frame image information to reduce the amount of calculation. The implementation process can be determined according to the operation principle of the cyclic decoder and is not described in detail in this application. Among them, the cyclic decoder includes a bottleneck block (the bottleneck layer, that is, the number of channels of multiple network layers is like a bottleneck, the input channel decreases from large to small and then increases from small to large), multiple upsampling blocks (upsampling layers), and an output block (output layer).
[0126] Combined with the network structure of the encoding module described in the above embodiments, the bottleneck block module can be connected to the LR-ASRR module of the encoding module to process the current frame image features of the 1 / 16 size. The fusion neural network module (such as ConvGRU) in the bottleneck block module operates only on half of the channels through segmentation, and the output result is then spliced and fused with the channels of the other half. The obtained image features are input into an upsampling block to upsample the current frame image features of the 1 / 8 size output by the encoding module. The obtained 1 / 4 size image features are input into the next upsampling block and continue to be upsampled with the current frame image features of the 1 / 4 size output by the encoding module. The obtained 1 / 2 size image features are input into the next next upsampling block and continue to be upsampled with the current frame image features of the 1 / 2 size output by the encoding module. The obtained image features are input into the output layer for processing.
[0127] Among them, during the upsampling process of each upsampling block module, upsampling can be achieved through bilinear interpolation. After splicing the bilinear interpolated image features with the current frame image features of the corresponding size output by the encoding module, it is used as the input of the next upsampling block module. After being processed by the average pooling layer, a convolutional layer (Conv), a batch normalization layer (i.e., Batch Normalization, BN layer), and a rectified linear layer (such as RectifiedLinear Unit, ReLU, a non-linear activation function) are used to achieve the merging of image features and the reduction of channels, and then the same ConvGRU operation as the bottleneck block module is performed. The implementation process is not described in detail in this application.
[0128] After that, the output after bilinear interpolation and the image features of the current frame image to be processed can be spliced and fused. Here, there is no need to perform an average pooling operation on the current frame image features. Two repeated stacks composed of a convolutional layer (Conv), a BN layer, and a ReLU layer can be directly used to generate the hidden features of the current frame image to be segmented. These hidden features are mapped (i.e., the last convolutional layer Conv) to become the output, that is, the current frame image decoding information composed of hidden features, an alpha prediction map (which is the Alpha channel prediction map in the foreground segmentation task and can be a grayscale image, where the value of each pixel ranges from 0 to 1, indicating the probability that the pixel belongs to the foreground. A value of 1 indicates completely belonging to the foreground, a value of 0 indicates completely belonging to the background, and a value between 0 and 1 indicates that the pixel may be at the junction of the foreground and the background) and foreground prediction (i.e., object prediction).
[0129] Step S26: According to the upsampling module in the target image segmentation model, process the decoded information of the current frame image and the original image of the current frame to obtain the segmentation result of the current frame image.
[0130] As can be seen from the above analysis, the decoded information of the current frame image is obtained by sequentially encoding and decoding the current frame image to be segmented, that is, the low-resolution and small-size image, by the encoding module and the decoding module, in the hope of restoring the image features of the low-resolution current frame image to be segmented to a high-resolution segmentation result. However, some image features will be lost in the process of obtaining the current frame image to be segmented, affecting the accuracy of segmentation prediction. Therefore, the target image segmentation model will use the upsampling module to combine the spatial information in the large-size and high-resolution original image of the current frame to process the decoded information of the current frame image and generate a clearer and more accurate segmentation result of the current frame image.
[0131] Still taking Figure 4 and Figure 5 the network structure of the target image segmentation model shown as an example for illustration, its upsampling module can be a Deep Guided Filter (DGF). At this time, the decoded information of the current frame image output by the decoding module can include hidden features, a foreground prediction map (a rough prediction map of the captured object), and an alpha prediction map. Input these decoded information of the current frame image and the original image of the current frame into the DGF module to output high-quality segmentation results of the current frame image such as a clear and accurate alpha prediction map and a prediction map of the captured object. The operation principle of the DGF module is not described in detail in this application.
[0132] In summary, in the embodiment of this application, in order to reduce the online processing time and resource consumption of the image segmentation model, after obtaining the original image of the current frame to be processed in the image sequence, it will not be directly input into the target image segmentation model. Instead, the large-size and high-resolution original image of the current frame is first downsampled to obtain a small-size and low-resolution current frame image to be segmented, and then the current frame image group composed of the original image of the current frame and the current frame image to be segmented is used to replace the original image of the current frame and the upsampling operator in the traditional image segmentation model as the input of the image segmentation model, so that the image segmentation model does not need to downsample the input image again, reducing the power consumption of the image segmentation model.
[0133] For the image segmentation model used in this application, instead of using an image segmentation model with a fixed network structure, the motion state information of the current captured object is considered, and a target image segmentation model that matches it is selected to process the current frame image group, so as to solve the problem that in a shooting scene in a high-speed motion state, using a fixed image segmentation model still causes excessive loss at the edge and reduces the image segmentation accuracy.
[0134] During the above image segmentation process, the encoding module in the target image segmentation model selected by the present application directly extracts multi-scale features from the current frame image to be segmented, so as to capture details and overall features at different scales, enhance the image understanding ability of the image segmentation model, improve the accuracy and robustness of image segmentation. Then, a decoding module that matches the motion state information of the current shooting object is selected to decode the current frame image features of multiple different scales. Since the number of fusion neural network modules in the decoding module that are in the working state at this time matches the motion state information, it is ensured that the decoded information of the current frame image obtained is the optimal decoded result in the current motion scenario, avoiding the decrease in image segmentation accuracy caused by excessive fusion of the previous frame image information. Thus, the upsampling model processes the decoded information of the current frame image and the original image of the current frame, and can quickly obtain a high-precision current frame image segmentation result.
[0135] Among them, for the acquisition of the motion state information of the current shooting object, the present application processes the small-scale current frame image to be segmented, which reduces the amount of image processing calculation and improves the efficiency of obtaining the motion state of the shooting object compared with the acquisition method of processing the large-scale original image of the current frame.
[0136] In some other embodiments proposed by the present application, in the image segmentation method described in the above embodiments, in order to obtain the motion state information of the current shooting object, the present application can also obtain the motion offset between the current frame image to be segmented included in the current frame image group and the previous frame image to be segmented included in the previous frame image group, and determine the motion offset as the motion state information of the shooting object in the current frame image group. It should be noted that in practical applications, target image segmentation models with different numbers of fusion neural network modules in the working state respectively correspond to different motion offset ranges. In this way, the number of fusion neural network modules corresponding to the motion offset range can be determined according to the motion offset range to which the motion offset between adjacent frame images to be segmented belongs, so as to implement the selection of the target image segmentation model that matches the motion state information of the current shooting object.
[0137] In a possible implementation, the above motion offset can be obtained by the frame difference method. Based on this, as Figure 6 shown, the implementation method of processing the current frame image group to determine the motion state information of the shooting object in the current frame image group may include:
[0138] Step S61, obtain a first sub-image of the current frame image to be segmented included in the current frame image group, and a second sub-image of the previous frame image to be segmented included in the previous frame image group;
[0139] Step S62: Perform an absolute difference operation on the first sub - image and the second sub - image to obtain the motion offset map of the adjacent - frame images to be segmented.
[0140] Step S63: Based on the motion offset map, determine the number of offset pixel points between the adjacent - frame images to be segmented as the motion offset amount, so as to determine the motion state information of the object being photographed in the current - frame image group.
[0141] In the embodiments of the present application, each original image in the image sequence can be an RGB image. The current - frame image to be segmented corresponding to the current - frame original image can be converted into a grayscale image, denoted as the first sub - image. Read the second sub - image (grayscale image) of the previous - frame image to be segmented that has been obtained during the process of obtaining the previous - frame image segmentation result, so as to reduce the data volume and simplify the subsequent calculation process. The implementation method of grayscale conversion of the image to be segmented in the present application will not be elaborated in detail.
[0142] After that, a per - pixel absolute - difference calculation can be performed on the first sub - image and the second sub - image to obtain the absolute value of the grayscale - value difference of the corresponding pixels between the adjacent - frame images to be segmented as the pixel value of this pixel, thus constituting the motion offset map of the adjacent - frame images to be segmented. Among them, if the object being photographed is currently in a high - speed motion state, the absolute difference of the pixels corresponding to the object being photographed in the motion offset map is very large; if it is in a low - speed motion state, the absolute difference of the pixels corresponding to the object being photographed in the motion offset map is relatively small (compared with the absolute difference in the high - speed motion state); if it is in a relatively static state, the absolute difference of the pixels corresponding to the object being photographed in the motion offset map is very small.
[0143] It can be seen that the higher the speed level of the motion state of the object being photographed, the greater the absolute difference of the pixels corresponding to the object being photographed in the motion offset map. The present application can perform threshold processing on the motion offset map to determine which speed - level motion state the current object being photographed is in, that is, which shooting scene the object being photographed in the current - frame image group is in, such as high - / medium - / low - speed motion scenes or static scenes, etc.
[0144] Based on the above analysis, the number of offset pixel points in the motion offset map can be counted, that is, the number of offset pixel points between adjacent images to be segmented. The number of offset pixel points is determined as the motion offset amount of the object being photographed between adjacent - frame images to be segmented, which can be expressed as offset_num, and directly used as the motion state information of the object being photographed in the current - frame image to be segmented. Subsequently, the motion state of the object being photographed can be determined through threshold processing. Optionally, according to actual needs, the present application can also, after determining the motion state of the object being photographed according to the above - mentioned threshold - processing method, determine the identifier representing this motion state as the motion state information, etc., so as to facilitate directly selecting the most suitable target image - segmentation model for processing the current - frame image group subsequently.
[0145] In a possible implementation, the present application can configure an appropriate number of different offset thresholds for the number of types of different pre-deployed motion states (i.e., the number of different motion state information), and form a motion offset range by adjacent offset thresholds to determine a candidate image segmentation model corresponding to each different motion offset range (where the number of fusion neural network modules in different candidate image segmentation models is different), or the decoding module in the same image segmentation model (where the multiple candidate decoding modules included in this image segmentation model (such as Figure 5 the decoding module 1, decoding module 2, ……, decoding module n shown, and the present application does not limit the size of the integer n) each contain a different number of fusion neural network modules), or the number of fusion neural network modules in the working state included in the same decoding module in the same image segmentation model (such as Figure 4 the processing method shown).
[0146] In this way, in practical applications, the motion offset determined above can be compared with different offset thresholds to determine the motion offset range to which the motion offset belongs, and the image segmentation model with the fusion neural network module in the working state corresponding to it (which can be implemented by but not limited to the three methods listed above) is determined as the target image segmentation model to process the current frame image group and obtain the current frame image segmentation result. Among them, as analyzed above, the larger the motion offset range of the shooting object corresponding to different motion state information between adjacent frame image groups, the fewer the number of fusion neural network modules in the working state in the target image segmentation model matching the motion state information.
[0147] Moreover, during the image segmentation process, during the decoding process of the image segmentation model, it is necessary to gradually restore large-size and high-resolution image features, and it is necessary to expand high-dimensional image information and semantic information to low dimensions. Less spatial information is considered in this process. In order to require more spatial information for alignment in the segmentation accuracy, but the less spatial information is provided by the actual decoding process closer to the output layer, resulting in a larger offset, and even misalignment, which leads to a lower segmentation accuracy of the shooting object. Therefore, the present application proposes that the direction of reducing the number of fusion neural network modules in the working state starts from the bottom layer of the target image segmentation model, such as starting from the ConvGRU in the last upsampling layer included in the recursive decoder of the RVM model shown above, to minimize the adverse impact on the model performance. The present application does not limit the size of the number of fusion neural network modules in the working state corresponding to different motion state information, which can be determined based on experience or experiments or during the model training process.
[0148] In some embodiments, after obtaining the running state information of the current shooting object according to the method described above, based on the motion state information, a target image segmentation model matching the motion state information can be selected from multiple trained candidate image segmentation models, so as to input the current frame image group into the target image segmentation model for processing to obtain the current frame image segmentation result. Combining the above detailed description of the processing process of the current frame image group, here, the encoding module of the target image segmentation model can extract features from the current frame image to be segmented, input the obtained current frame image features of multiple different sizes into the decoding module of the target image segmentation model for decoding processing, input the obtained current frame image decoding information and the current frame original image into the upsampling module of the target image segmentation model for processing to obtain the current frame image segmentation result. The implementation process can refer to the description of the corresponding part of the above embodiments, and this embodiment will not be elaborated here.
[0149] Among them, the number of fusion neural network modules included in each of the multiple candidate image segmentation models is different, and they can be respectively applied to the high-precision segmentation processing of shooting objects in the running states of corresponding speed levels. These candidate image segmentation models can all be trained based on the RVM algorithm (or other machine learning / deep learning algorithms). Sample images with corresponding motion state information can be selected to train the initial image segmentation model with the corresponding number of fusion neural network modules, or the general initial image segmentation model can be trained, and the number of required fusion neural network modules can be continuously adjusted during the training process to obtain the image segmentation model with the best segmentation accuracy and performance in the current scenario as the candidate image segmentation model. The training process of each candidate image segmentation model in this application will not be elaborated.
[0150] Exemplarily, this application takes the example of pre-configuring multiple candidate image segmentation models with different numbers of fusion neural network modules, corresponding to scenarios with different motion state information. These three candidate image segmentation models are respectively denoted as the first image segmentation model, the second image segmentation model, and the third image segmentation model, and the number of fusion neural network modules included in each of these three image segmentation models decreases layer by layer. They respectively correspond to the relatively static scenario, the low-speed motion scenario, and the high-speed motion scenario. In this application scenario, the motion speed of the shooting object increases sequentially, and the specific motion speed value is not limited.
[0151] According to the above scenario, referring to Figure 7 the flowchart of the image segmentation method proposed in Embodiment 3 of this application shown, the image segmentation method may include but is not limited to:
[0152] Step S71, obtaining the current frame original image in the image sequence;
[0153] Step S72: Downsample the original image of the current frame to obtain the image to be segmented of the current frame corresponding to the original image of the current frame; the image size of the image to be segmented of the current frame is smaller than the image size of the original image of the current frame.
[0154] Step S73: Obtain the motion offset between the image to be segmented of the current frame and the image to be segmented of the previous frame.
[0155] Regarding the implementation processes of Steps S71 - S73, reference can be made to the descriptions of the corresponding parts in the above embodiments, and details are not elaborated in this embodiment.
[0156] Step S74: Determine whether the motion offset is greater than the first offset threshold. If so, proceed to Step S75; if not, execute Step S78.
[0157] Step S75: Determine whether the motion offset is greater than the second offset threshold. If so, proceed to Step S76; if not, execute Step S77.
[0158] Step S76: Select the third image segmentation model as the target image segmentation model.
[0159] Step S77: Select the second image segmentation model as the target image segmentation model.
[0160] Step S78: Select the first image segmentation model as the target image segmentation model.
[0161] Combined with the above - mentioned related descriptions of the process of obtaining the motion offset of the photographed object in adjacent - frame images, since the greater the motion speed of the photographed object, the greater the obtained motion offset tends to be. The first image segmentation model pre - trained in this application is more suitable for high - precision segmentation processing of photographed objects in relatively static scenes, the second image segmentation model is more suitable for high - precision segmentation processing of photographed objects in low - speed motion scenes, and the third image segmentation model is more suitable for high - precision segmentation processing of photographed objects in high - speed motion scenes. In this way, the critical values of the motion offset between these three scenarios can be determined as the offset thresholds, that is, the boundary values of the motion offset ranges corresponding to each scenario. The offset critical value used to divide the motion offset ranges of the relatively static scene and the low - speed motion scene is denoted as the first offset threshold, and the offset critical value used to divide the motion offset ranges of the low - speed motion scene and the high - speed motion scene is denoted as the second offset threshold. It can be seen that the first offset threshold is less than the second offset threshold. The application does not limit the numerical values of each offset threshold, which can be obtained based on experience or training.
[0162] According to the motion state information acquisition method described above, after determining the motion offset of the shooting object between the current frame of the image to be segmented and the previous frame of the image to be segmented, that is, the motion offset between the current frame of the original image and the previous frame of the original image, the motion offset can be compared with each offset threshold. If the motion offset is less than or equal to the first offset threshold, it can be considered that the shooting object is in a relatively static state, and the application scenario where the shooting object is located is a relatively static scenario. To improve the image segmentation accuracy, a relatively large number of fusion neural network modules can be selected to participate in the image segmentation process. At this time, from three candidate image segmentation models, the first image segmentation model with the largest number of fusion neural network modules can be selected and determined as the target image segmentation model.
[0163] If the currently obtained motion offset is greater than the first offset threshold and less than or equal to the second offset threshold, it can be considered that the shooting object is in a low-speed motion state, and the application scenario where the shooting object is located is a low-speed motion scenario. At this time, there may be a certain degree of blurred distortion in the edge area of the current frame of the image to be segmented. If the first image segmentation model is still used for processing, the segmentation accuracy may decrease due to information loss at the edge. To solve this problem, the present application can select the second image segmentation model in which there are no fusion neural network modules in the bottom layer of the decoding module or several adjacent layers as the target image segmentation model, reduce the number of times of fusing the information of the previous frame of image, reduce the computational complexity of the model, and in order to minimize the adverse impact on the model performance, the present application will start from the bottom layer to reduce the fusion neural network modules, and reduce the time series information captured by the model by reducing the minimum number of such fusion neural network modules to ensure better image segmentation accuracy.
[0164] Optionally, in the case where it is determined that the motion offset is greater than the first offset threshold, the second image segmentation model or the third image segmentation model can also be directly selected as the target image segmentation model without further judging the second offset threshold. To further improve the image segmentation accuracy and ensure the model performance, the judgment method of multiple offset thresholds proposed in the embodiments of the present application is preferably used.
[0165] If the currently obtained motion offset is greater than the second offset threshold, it can be considered that the shooting object is in a high-speed motion state, and the application scenario where the shooting object is located is a high-speed motion scenario. At this time, there will usually be a serious blurred distortion problem in the edge area of the shooting object in the current frame of the image to be segmented. If the decoding module still uses a relatively large number of fusion neural network modules to participate in the image segmentation, the image segmentation accuracy will be greatly reduced due to large information loss at the edge, such as the incomplete segmentation area of the shooting object, etc. To improve this situation, the present application will select the third image segmentation model that contains a relatively small number or even no fusion neural network modules as the target image segmentation model to process the current frame of the image group.
[0166] Exemplarily, if the candidate image segmentation models are all trained based on the RVM algorithm, the second image segmentation model can be a network structure as shown in Figure 8 . For the last two upsampling modules in its decoding module, compared with the earlier upsampling modules, the fusion neural network module ConvGRU is no longer deployed. The first image segmentation model can be such that ConvGRU can still be deployed in both of the last two upsampling modules in the decoding module, or ConvGRU is no longer deployed in the last upsampling module, etc. The third image segmentation model can be a network structure as shown in Figure 9 . In its decoding module, ConvGRU may no longer be deployed, but it is not limited to the network structures shown in Figure 8 and Figure 9 . Through testing, it is known that the downsampling processing duration of the original image by the downsampling module in the traditional RVM model accounts for 30%, and the processing duration of all ConvGRUs in the decoding module accounts for 27%. According to the method described above, this application pre-downsamples the original image in advance. After obtaining the small-sized image to be segmented and then inputting it into the model, and adaptively selects the target image segmentation model that matches the motion state information of the current shooting object to implement the processing of the image to be segmented and the original image. Since the time-consuming ratio of this selection process is less than 1%, this application can improve the overall processing duration of the model by 20%-40%, improve the image segmentation efficiency, and ensure the image segmentation accuracy.
[0167] Combined with the above analysis, this application can also pre-determine the corresponding motion offset range of each candidate image segmentation model itself. After obtaining the motion offset of the current shooting object, the candidate image segmentation model corresponding to the motion offset range where the motion offset is located is determined as the target image segmentation model. This application does not limit the size of the motion offset range corresponding to each candidate image segmentation model and its determination method. It can be flexibly configured or adjusted according to the actual situation.
[0168] It should be understood that the number of pre-trained candidate image segmentation models includes but is not limited to the three candidate image segmentation models listed above that are respectively applicable to relatively static scenes, low-speed motion scenes, and high-speed motion scenes. Optionally, this application can also make a finer-grained division of the motion offset to obtain more motion offset ranges, corresponding to training a larger number of candidate image segmentation models. The implementation process is similar, and this application does not give detailed examples one by one.
[0169] Step S79: Input the current frame original image and the current frame image to be segmented into the target image segmentation model for processing to obtain the current frame image segmentation result.
[0170] According to the method described above, a target image segmentation model that matches the running offset of the current shooting object is selected. The encoding module of the target image segmentation model is used to extract image features from the current frame of the image to be segmented. The obtained current frame image features of multiple different sizes are input into the decoding module of the target image segmentation model for decoding processing. After obtaining the current frame image decoding information, the upsampling module of the target image segmentation model is used to process the current frame image decoding information and the current frame original image to obtain a high-precision current frame image segmentation result. The implementation process is not described in detail in this application.
[0171] It can be seen that the embodiment of this application needs to perform shooting object segmentation processing on the original shooting image in the current application scenario to meet the downstream image processing tasks. After obtaining the current frame original image, in order to reduce the online processing time of the image segmentation model and improve the flexibility and accuracy of image downsampling processing, this application can flexibly select an appropriate downsampling operator to pre-downsample the current frame original image to obtain a current frame of the image to be segmented with a smaller size to complete subsequent processing, so that the subsequent used image segmentation model does not need to perform the downsampling step on the input image, solving the impact of this downsampling step on the model performance and power consumption.
[0172] Moreover, for the image segmentation model to be used, this application will adaptively select the most suitable image segmentation model in combination with the current motion state of the shooting object to avoid reducing the segmentation accuracy due to too many fusion neural network modules in the image segmentation model. For this, in this embodiment, by obtaining the motion offset between the current frame of the image to be segmented and the previous frame of the image to be segmented, and comparing it with different pre-configured offset thresholds, the scene type of the current shooting object is determined, such as the speed level of the motion state, so as to select a target image segmentation model that matches it to complete the processing of the current frame image group and quickly obtain a high-precision image segmentation result to reliably meet the requirements of downstream image processing tasks.
[0173] In some other embodiments proposed in this application, as Figure 5 shown, an image segmentation model is pre-trained, but this image segmentation model includes an encoding module and multiple candidate decoding modules. The number of fusion neural network modules included in each of these multiple candidate decoding modules is different, and the corresponding motion state information is also different. According to but not limited to the method described above, after obtaining the motion state information of the shooting object in the current frame image group, the target decoding module that matches the motion state information can be selected from the multiple candidate decoding modules included in the target image segmentation model based on this motion state information; thus, based on the encoding module and the target decoding module in the target image segmentation model, the current frame image group is processed to obtain the current frame image segmentation result.
[0174] Still taking the corresponding scenario of Embodiment 3 above as an example for illustration, if multiple candidate decoding modules include a first candidate decoding module, a second candidate decoding module, and a third candidate decoding module that sequentially correspond to a relatively static scenario, a low-speed motion scenario, and a high-speed motion scenario, when it is determined that the currently obtained motion offset is less than or equal to the first offset threshold, the first candidate decoding module can be selected as the target decoding module; when it is determined that the motion offset is greater than the first offset threshold and less than or equal to the second offset threshold, the second candidate decoding module can be selected as the target decoding module; when it is determined that the motion offset is greater than the second offset threshold, the third candidate decoding module can be selected as the target decoding module.
[0175] In this regard, a decoding control signal representing the target decoding module can be generated, so that after the encoding module of the target image segmentation model completes image feature extraction, in response to the decoding control signal, the target decoding module is controlled to be in a working state, and the obtained current frame image features of multiple different sizes are input into the target decoding module for decoding; optionally, the decoding control signal can also be input when the current frame image group is input into the target image segmentation model. At this time, the corresponding target decoding module is directly controlled to be in a working state, and other candidate decoding modules are in a non-working state, so as to complete the processing of the input current frame image to be segmented through the encoding module and the target decoding module, and obtain a high-precision image segmentation result.
[0176] In some other embodiments proposed in this application, as Figure 4 shown, an image segmentation model is pre-trained. The image segmentation model includes an encoding module and a decoding module, but the working states of the respective fusion neural network modules included in the decoding module are controllable, and the number of fusion neural network modules that match different motion state information can be determined in advance. According to but not limited to the method described above, after obtaining the motion state information of the shooting object in the current frame image group, the target fusion neural network module that matches the motion state information can be selected from the multiple fusion neural network modules included in the decoding module of the target image segmentation model (in this application, the fusion neural network modules deployed from the top layer of the decoding module can be selected layer by layer), and the other fusion neural network modules among the multiple fusion neural network modules are used as non-target fusion neural network modules.
[0177] After that, the current frame image group can be input into the target image segmentation model. When the non-target fusion neural network module is controlled to be in a non-operating state, the current frame image segmentation result can be obtained. It can be seen that in this case, the target image segmentation model only controls the selected target fusion neural network module that matches the motion state information to be in an operating state and participates in the segmentation process of the current frame of the image to be processed. For other non-target fusion neural network modules, they do not participate in this segmentation process, avoiding excessive fusion of the information of the previous frame image, so that the image segmentation result is too affected by the information loss at the edge and the image segmentation accuracy of the current frame is reduced.
[0178] In a possible implementation, for the switching control between the operating state and the non-operating state of the fusion neural network module, it can be achieved by configuring corresponding control circuits for each fusion neural network module to directly shield the non-target fusion neural network module or freeze the parameters of the non-target fusion neural network module. This way can retain the position of the non-target fusion neural network module but does not participate in the operation. In the case where the motion state information of the photographed object changes and the non-target fusion neural network module is selected as the target fusion neural network module, it can still participate in the operation. The present application does not limit the control circuit structure for implementing this control, such as a switch circuit connected in parallel with the fusion neural network module.
[0179] Based on the image segmentation methods described in the above embodiments, in practical applications, the motion state of the photographed object is often not an instantaneous state. Usually, there will be a situation where the motion state information of the photographed object in consecutive multiple frames of the original image / to-be-segmented image is the same. For this consecutive multiple-frame image group, the network structure of the target image segmentation model selected according to the method described above is the same, and the selection operation can be not repeated to reduce the calculation amount and improve the image segmentation efficiency of the entire image sequence. Therefore, the present application proposes that after the target image segmentation model is first selected, the subsequent consecutive m-frame image groups can directly use this target image segmentation model for direct processing without repeating the model selection process. m can be an integer greater than or equal to 2 and less than or equal to 5, and can be a fixed value determined by means of experience or experiment or training, etc., and is not limited to the value range of [2, 5]. Of course, m can also be dynamically determined according to the motion offset of the photographed object in adjacent frames. For example, the larger the motion offset, the larger the value of m. The present application does not limit the value of m.
[0180] Based on this, in the process of implementing the above-mentioned method of selecting a target image segmentation model that matches the motion state information and performing segmentation processing on the image to be segmented to obtain the image segmentation result of the current frame, the target image segmentation model that matches the motion state information can be selected according to, but not limited to, the method described above, and an image group with a continuous specified number of frames (such as m mentioned above, which can be multiple consecutive frames starting from the current frame. Generally, the fewer the number of fusion neural network modules in the working state in the target image segmentation model, the larger the continuous specified number of frames) can be processed to obtain the image segmentation results of the corresponding frames. After obtaining the next image group with the specified number of frames, continue to process the next image group according to the method described above, determine the motion state information of the shooting object in the next image group, and then process multiple consecutive image groups according to the method proposed in this embodiment.
[0181] Combined with the relevant descriptions of the image segmentation methods proposed in the above embodiments, the embodiments of the present application also propose an image segmentation model. Figure 4 、 Figure 8 and Figure 9 Shown in the image segmentation model network structure, the image segmentation model proposed in the embodiments of the present application includes an encoding module and a decoding module, where: at least one layer in the decoding module does not include a fusion neural network module, or at least one layer includes a fusion neural network module in a non-working state.
[0182] It can be seen that when the image segmentation model proposed in the present application includes a decoding module, starting from the bottom layer of the decoding module, one or more layers (which do not include the output module in the decoding module) can be selected not to deploy a fusion neural network module (such as ConvGRU). In this way, as Figure 10 Shown, multiple different image segmentation models with different numbers of fusion neural network modules can be deployed as candidate image segmentation models for the electronic device to flexibly select the target image segmentation model that is most suitable for the actual application scenario in different application scenarios and complete the image segmentation processing for that application scenario. For example, in a motion scenario, using this image segmentation model (such as Figure 8 or Figure 9 Shown network structure) to process the original image captured according to the image segmentation model proposed in the present application can quickly obtain a high-precision image segmentation result, avoid excessive information loss at the edge and reduce the segmentation accuracy, improve the model performance, and reduce the power consumption of the model and improve the image segmentation efficiency due to reducing the number of fusions.
[0183] Optionally, as Figure 4As shown, the present application can deploy the fusion neural network module in each layer of the decoding module (excluding the output module in the decoding module), but its working state needs to be controllable. In this way, when the shooting object is in a motion state at different speed levels, according to the image segmentation method proposed in the present application, at least one layer containing the fusion neural network module in the decoding module can be adaptively selected to be in a non-working state, reducing the number of fusion neural network modules participating in the operation to ensure the image segmentation accuracy. The implementation process can refer to the description of the corresponding part of the method embodiment above, and this embodiment will not be elaborated here.
[0184] In some embodiments, when the image segmentation model includes an encoding module and multiple decoding modules, as Figure 5 shown, among the multiple layers included in each of these multiple decoding modules, the number of layers containing the fusion neural network module is different, that is, the number of fusion neural network modules included in different decoding modules is different, so that they are respectively matched with different motion state information. In practical applications, according to the motion state information of the shooting object in the actual application scenario, a decoding module that matches it is selected to participate in the model operation to ensure the accuracy of the obtained image segmentation result.
[0185] Among them, regarding the selection process of the candidate image segmentation model or candidate decoding module or target fusion neural network module described above, after it is completed by the processor of the electronic device, the corresponding control signal can be sent to the corresponding target image segmentation model for implementation. The implementation process can refer to the description of the corresponding method embodiment above, and this embodiment will not be elaborated here.
[0186] For each image segmentation model described in the above embodiments, RGB images and their corresponding portrait masks (masks) in public datasets such as videomatte240k, iamgematte, and coco can be used to form a training image set. The RGB images and their corresponding portrait mask images are respectively downsampled by the same multiple to obtain corresponding small-size low-resolution training image groups and their training labels (that is, the downsampled portrait mask images). Then, according to the image segmentation method proposed in the present application, the initial image segmentation model deployed in any of the above ways can be trained to obtain multiple candidate image segmentation models applicable to different application scenarios, that is, matching different motion state information, or multiple candidate decoding modules of an image segmentation model, or different numbers of target fusion neural network modules in a decoding model of an image segmentation model, etc., to obtain different image segmentation network structures respectively matching different motion state information.
[0187] Among them, in the above model training process, the loss value between the predicted face segmentation result output by the model and the corresponding training label can be minimized to adjust the model parameters and / or the number of fusion neural network modules in the working state, so as to obtain an image segmentation model of an image segmentation network structure that matches the corresponding motion state information. The training process will not be elaborated in this application. It should be noted that for different image processing tasks applicable to different application scenarios, the source of the training image set used to train the applicable image segmentation model may include, but is not limited to, several publicly available data sets listed above, and can also be obtained from a dedicated database for implementing the corresponding image processing task in this application scenario or other data sources, depending on the situation, and this application will not give examples one by one.
[0188] The above introduces an image segmentation method and an image segmentation model provided by an embodiment of this application. The following will introduce the device for executing the above image segmentation method.
[0189] Refer to Figure 11 , which is a schematic structural diagram of an image segmentation device provided by an embodiment of this application. As Figure 11 shown, the image segmentation device may include:
[0190] A current frame image group acquisition module 111, configured to acquire a current frame image group to be processed;
[0191] A motion state information determination module 112, configured to process the current frame image group to determine the motion state information of the shooting object in the current frame image group;
[0192] A selection module 113, configured to select a target image segmentation model that matches the motion state information, process the current frame image group, and obtain a current frame image segmentation result;
[0193] Among them, the number of fusion neural network modules in the working state in the target image segmentation models that match different motion state information is different, and the fusion neural network module is used to perform the fusion of the previous frame image information and the current frame image information.
[0194] In a possible implementation, the above selection module 113 may include:
[0195] A first selection unit, configured to select a target image segmentation model that matches the motion state information from multiple trained candidate image segmentation models according to the motion state information;
[0196] A first processing unit, configured to input the current frame image group into the target image segmentation model for processing to obtain a current frame image segmentation result.
[0197] In a possible implementation, the above selection module 113 may include:
[0198] A second selection unit, configured to select a target decoding module that matches the motion state information from multiple candidate decoding modules included in the target image segmentation model according to the motion state information;
[0199] A second processing unit, configured to process the current frame image group based on the encoding module and the target decoding module in the target image segmentation model to obtain a current frame image segmentation result;
[0200] Wherein, the number of fusion neural network modules included in different candidate decoding modules is different.
[0201] In a possible implementation, the above selection module 113 may include:
[0202] A third selection unit, configured to select a target fusion neural network module that matches the motion state information from multiple fusion neural network modules included in the decoding module of the target image segmentation model according to the motion state information, and use the other fusion neural network modules among the multiple fusion neural network modules as non-target fusion neural network modules;
[0203] A third processing unit, configured to input the current frame image group into the target image segmentation model, and obtain a current frame image segmentation result in a case where the non-target fusion neural network modules are controlled to be in a non-working state.
[0204] In the embodiments of the present application, the larger the motion offset range of the shooting object represented by different motion state information between adjacent frame image groups, the fewer the number of fusion neural network modules in the target image segmentation model that are in a working state and match the motion state information; the reduction direction of the number of fusion neural network modules in the working state starts from the bottom layer of the target image segmentation model.
[0205] In some embodiments, the current frame image group acquisition module 111 may include:
[0206] A current frame original image acquisition unit, configured to acquire a current frame original image in the image sequence;
[0207] A downsampling unit, configured to perform downsampling on the current frame original image to obtain a current frame image to be segmented corresponding to the current frame original image; the image size of the current frame image to be segmented is smaller than the image size of the current frame original image;
[0208] Wherein, the current frame original image and the image to be segmented form a current frame image group.
[0209] Based on this, the above selection module 113 may include:
[0210] A fourth selection unit, configured to select an encoding module in a target image segmentation model that matches the motion state information, perform feature extraction on a current frame image to be segmented in the current frame image group, and obtain current frame image features of multiple different sizes;
[0211] A fifth selection unit, configured to select a decoding module in the target image segmentation model that matches the motion state information, decode the current frame image features of multiple different sizes, and obtain current frame image decoding information;
[0212] A fourth processing unit, configured to process the current frame image decoding information and the current frame original image in the current frame image group according to an upsampling module in the target image segmentation model, and obtain a current frame image segmentation result.
[0213] In some embodiments, the above selection module 113 may include any of the following units:
[0214] A motion detection unit, configured to perform motion detection on a current frame image to be segmented included in the current frame image group, and obtain motion state information of a shooting object in the current frame image to be segmented;
[0215] A motion state information determination unit, configured to obtain an action feature of a shooting object in a current frame image to be segmented included in the current frame image group, and determine the motion state information of the shooting object in the current frame image to be segmented according to the action feature;
[0216] A motion offset amount acquisition unit, configured to acquire a motion offset amount between a current frame image to be segmented included in the current frame image group and a previous frame image to be segmented included in the previous frame image group, and determine the motion offset amount as the motion state information of the shooting object in the current frame image group;
[0217] Wherein, target image segmentation models containing different numbers of fusion neural network modules in a working state respectively correspond to different motion offset ranges.
[0218] In some embodiments, the above selection module 113 may include:
[0219] A selection processing unit, configured to select a target image segmentation model that matches the motion state information, process an image group of continuously specified number of frames, and obtain corresponding frame image segmentation results, so as to, after obtaining the next frame image group of the specified number of frames, perform processing on the next frame image group and determine the motion state information of the shooting object in the next frame image group;
[0220] Among them, the continuous specified number of frames is multiple consecutive frames starting from the current frame; the fewer the number of fusion neural network modules in the target image segmentation model in the working state, the larger the continuous specified number of frames.
[0221] An embodiment of the present application also provides a computer-readable storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any image segmentation method provided by the embodiment of the present application.
[0222] Among them, the computer-readable storage medium can be any available medium that the electronic device can store, or a data storage device such as a training device or a data center that integrates one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state disk (SSD)), etc.
[0223] An embodiment of the present application also provides a computer program product, including one or more computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device can implement any image segmentation method provided by the embodiment of the present application.
[0224] Among them, when the computer-readable instructions are loaded and executed on the electronic device, the processes or functions described in the embodiment of the present application are fully or partially generated. The electronic device can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer-readable instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer-readable instructions can be transmitted from a website, a computer, a training device, or a data center to another website, a computer, a training device, or a data center in a wired (for example, coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.) manner, which can be determined according to the actual application scenario.
[0225] Refer to Figure 12 , which is a schematic hardware structure diagram of an electronic device applicable to the image segmentation method proposed in the present application. The electronic device can include, but is not limited to, one or more terminal devices such as a mobile phone, a tablet computer, a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, a netbook, a robot, or a business terminal, etc. The image segmentation method proposed in the present application can also be implemented by configuring a server. As Figure 12As shown, the electronic device may include, but is not limited to: at least one communication component 121, at least one memory 122, and at least one processor 123, where:
[0226] Communication may occur between at least one communication component 121, at least one memory 122, and at least one processor 123 via a bus. The bus may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 12 only a single bidirectional line is shown herein, but this does not mean there is only one bus or one type of bus.
[0227] The communication component 121 may be used to receive an image sequence, such as receiving each frame of the original image to be processed sent by an image acquisition device, or reading an image sequence or the current frame of the original image to be processed from a database, etc. The source of the image sequence is not limited in this application. It may also be used to implement data or instruction transmission between internal components of the electronic device.
[0228] Based on this, in the embodiments of this application, the communication component 121 may include communication components corresponding to wireless communication methods such as Wi-Fi, Bluetooth, 5G / 6G, etc., so that the electronic device can achieve data transmission with other devices through these communication components. It may also include one or more interfaces supporting wired communication methods, such as General-Purpose Input / Output (GPIO) interfaces, USB interfaces, Universal Asynchronous Receiver / Transmitter (UART) interfaces, etc., to achieve data transmission between internal components of the electronic device. The composition structure of the communication component 121 for implementing this function and its corresponding communication transmission mechanism are not limited in this application and may be determined according to the situation.
[0229] The memory 122 may be used to store multiple computer instructions for implementing the image segmentation method proposed in the embodiments of this application; the processor 123 may load and execute the multiple computer instructions stored in the memory 122 to implement each step of the image segmentation method proposed in the embodiments of this application. The implementation process may refer to the description of the corresponding part in the method embodiments above.
[0230] In the embodiments of the present application, the memory 122 may include storage media such as floppy disks, USB flash drives, external hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs. The processor 123 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), a digital signal processor (DSP), an application specific integrated circuit (ASIC), or a field-programmable gate array (FPGA).
[0231] It should be understood that Figure 12 the structure of the electronic device shown does not constitute a limitation on the electronic device in the embodiments of the present application. In practical applications, the electronic device may include more or fewer components than Figure 12 those shown, or combine certain components. For example, when the electronic device is a terminal device, it may further include an image collector, a pick-up, a speaker, a display screen, various sensors, an antenna, a power module, a radio frequency component, an external port, and other input / output components, which can be determined according to the processing function requirements, and the present application will not give detailed examples one by one.
[0232] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.
[0233] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. That is to say, in the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product in the form of a software product can be stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, etc., including several instructions for causing a computer device (which can be a personal computer, training device, or network device, etc.) to execute the image segmentation method described in each embodiment of the present application.
Claims
1. An image segmentation method, the method comprising: Get the current frame image group to be processed; Processing the current frame image group to determine motion state information of the photographed object in the current frame image group; A target image segmentation model matching the motion state information is selected, and the current frame image group is processed to obtain a current frame image segmentation result; the number of fusion neural network modules in working state is different in the target image segmentation models matching different motion state information, and the fusion neural network module is used to fuse the previous frame image information with the current frame image information.
2. According to the method of claim 1, the step of selecting a target image segmentation model that matches the motion state information and processing the current frame image group to obtain a current frame image segmentation result comprises: According to the motion state information, selecting a target image segmentation model that matches the motion state information from a plurality of trained candidate image segmentation models; The current frame image group is input into the target image segmentation model for processing to obtain a current frame image segmentation result.
3. According to the method of claim 1, the step of selecting a target image segmentation model that matches the motion state information and processing the current frame image group to obtain a current frame image segmentation result comprises: According to the motion state information, selecting a target decoding module that matches the motion state information from a plurality of candidate decoding modules included in the target image segmentation model; Based on the encoding module and the target decoding module in the target image segmentation model, the current frame image group is processed to obtain a current frame image segmentation result; Among them, different candidate decoding modules contain different numbers of fusion neural network modules.
4. The method according to claim 1, wherein the step of selecting a target image segmentation model that matches the motion state information and processing the current frame image group to obtain a current frame image segmentation result comprises: According to the motion state information, a target fusion neural network module matching the motion state information is selected from a plurality of fusion neural network modules included in a decoding module of the target image segmentation model, and other fusion neural network modules in the plurality of fusion neural network modules are used as non-target fusion neural network modules; The current frame image group is input into the target image segmentation model, and the current frame image segmentation result is obtained while controlling the non-target fusion neural network module to be in a non-working state.
5. The method according to any one of claims 1 to 4, wherein: The larger the motion offset range of the photographed object represented by the different motion state information between adjacent frame image groups, the smaller the number of fusion neural network modules in working state in the target image segmentation model matching the motion state information; The direction of reducing the number of the fusion neural network modules in working state is starting from the bottom layer of the target image segmentation model.
6. According to the method according to any one of claims 1 to 4, the step of obtaining the current frame image group to be processed comprises: Get the original image of the current frame in the image sequence; Down-sampling the original image of the current frame to obtain an image of the current frame to be segmented corresponding to the original image of the current frame; The image size of the current frame to be segmented is smaller than the image size of the current frame original image; The current frame original image and the current frame image to be segmented constitute a current frame image group.
7. The method according to claim 6, wherein the step of selecting a target image segmentation model that matches the motion state information and processing the current frame image group to obtain a current frame image segmentation result comprises: Selecting an encoding module in a target image segmentation model that matches the motion state information, performing feature extraction on a current frame image to be segmented in the current frame image group, and obtaining a plurality of current frame image features of different sizes; Selecting a decoding module in the target image segmentation model that matches the motion state information, decoding the current frame image features of the multiple different sizes, and obtaining current frame image decoding information; According to the up-sampling module in the target image segmentation model, the current frame image decoding information and the current frame original image in the current frame image group are processed to obtain a current frame image segmentation result.
8. The method according to claim 6, wherein the processing the current frame image group to determine the motion state information of the photographed object in the current frame image group comprises any of the following implementations: Performing motion detection on the current frame image to be segmented included in the current frame image group to obtain motion state information of the photographed object in the current frame image to be segmented; Acquire the motion features of the subject in the current frame image to be segmented contained in the current frame image group, and determine the motion state information of the subject in the current frame image to be segmented according to the motion features; Acquire a motion offset between a current frame to be segmented image included in the current frame image group and a previous frame to be segmented image included in the previous frame image group, and determine the motion offset as motion state information of the photographed object in the current frame image group; in, The target image segmentation model includes different numbers of fusion neural network modules in working state corresponding to different motion offset ranges.
9. The method according to claim 5, wherein the step of selecting a target image segmentation model that matches the motion state information and processing the current frame image group to obtain a current frame image segmentation result comprises: Selecting a target image segmentation model that matches the motion state information, processing a group of images of a specified number of consecutive frames to obtain a corresponding frame image segmentation result, so as to process the next group of images of the specified number of frames after obtaining the next group of images of the specified number of frames, and determining the motion state information of the photographed object in the next group of images; Wherein, the continuous specified number of frames is a continuous number of frames starting from the current frame; The fewer the number of fusion neural network modules in working state in the target image segmentation model, the larger the number of continuous specified frames.
10. An image segmentation model, comprising: Encoding module and decoding module, where: At least one layer in the decoding module does not include a fusion neural network module, or the fusion neural network module included in at least one layer is in a non-working state.
11. According to the model described in claim 10, there are multiple decoding modules, and each of the multiple decoding modules contains multiple layers, and the number of layers containing the fusion neural network module is different.
12. An image segmentation device, the device comprising: A current frame image group acquisition module is used to acquire the current frame image group to be processed; A motion state information determination module, used to process the current frame image group to determine the motion state information of the photographed object in the current frame image group; A selection module is used to select a target image segmentation model that matches the motion state information, process the current frame image group, and obtain a current frame image segmentation result; Among them, the number of fusion neural network modules in working state in the target image segmentation model matching different motion state information is different, and the fusion neural network module is used to fuse the image information of the previous frame with the image information of the current frame.
Citation Information
Cited By
Automatic wiring cable segmentation method based on hybrid expert model
CN121366288A