Model training method and device, image segmentation method and device, electronic equipment and chip

By merging the convolutional network of the image segmentation model, it is simplified into a second convolutional layer, which solves the problems of large size and high computing resources of the image segmentation model, and realizes lightweight and efficient image segmentation.

CN120373363APending Publication Date: 2025-07-25BEIJING X RING TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411472780.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-21
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing image segmentation model has large size and high computing resource requirements, resulting in low image segmentation efficiency.

Method used

By merging the convolutional networks in the image segmentation model, multiple subnets are merged into a second convolutional layer, simplifying the model structure.

Benefits of technology

It realizes the lightweight image segmentation model, reduces computing resource requirements and power consumption, improves image segmentation efficiency and inference speed, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373363A_ABST
    Figure CN120373363A_ABST
Patent Text Reader

Abstract

The invention relates to a model training method and device, an image segmentation method and device, electronic equipment and a chip, and belongs to the technical field of image processing. The method comprises the steps that a training sample is acquired, and the training sample comprises a sample image and a sample segmentation result of the sample image; based on the training sample, training an image segmentation model, the image segmentation model comprising N convolutional networks, each convolutional network comprising a plurality of sub-networks, and the plurality of sub-networks of each convolutional network comprising M first convolutional layers; performing parameter merging on a plurality of sub-networks in the ith convolutional network to obtain an ith second convolutional layer; and replacing the ith convolutional network with the ith second convolutional layer to update the image segmentation model, and taking the updated image segmentation model as a final image segmentation model. Therefore, the image segmentation model is lighter, has the advantages of small size, small required computing resources, low power consumption and high reasoning speed, and is high in image segmentation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a model training method, an image segmentation method, an apparatus, an electronic device, a storage medium, and a chip. Background Art

[0002] Currently, with the continuous development of artificial intelligence technologies, image segmentation models have been widely used in fields such as intelligent image matting, face recognition, pedestrian detection, traffic control, and medical imaging, and have advantages such as high efficiency and high automation. For example, an image can be input into an image segmentation model, and the segmentation result can be output by the image segmentation model. However, the volume of the image segmentation model and the required computing resources in related technologies are relatively large, resulting in a problem of low image segmentation efficiency. Summary of the Invention

[0003] The present disclosure provides a model training method, an image segmentation method, an apparatus, an electronic device, a computer-readable storage medium, and a chip to at least solve the problem of low image segmentation efficiency in related technologies. The technical solution of the present disclosure is as follows:

[0004] According to a first aspect of an embodiment of the present disclosure, a model training method is provided, including: obtaining a training sample, where the training sample includes a sample image and a sample segmentation result of the sample image; training an image segmentation model based on the training sample, where the image segmentation model includes N convolutional networks, each convolutional network includes a plurality of sub-networks, and a plurality of sub-networks of each convolutional network each include M first convolutional layers, and both N and M are positive integers greater than 1; performing parameter merging on the plurality of sub-networks in the i-th convolutional network to obtain a second convolutional layer of the i-th, where i is a positive integer not greater than N; replacing the i-th convolutional network with the second convolutional layer of the i-th to update the image segmentation model, and using the updated image segmentation model as the final image segmentation model.

[0005] According to a second aspect of an embodiment of the present disclosure, an image segmentation method is provided, including: obtaining an original image to be segmented; inputting the original image into an image segmentation model, and outputting a target segmentation result of the original image by the image segmentation model, where the image segmentation model includes N second convolutional layers, N is a positive integer greater than 1, and the image segmentation model is obtained by using the model training method described in the first aspect of the embodiment of the present disclosure.

[0006] According to a third aspect of the embodiments of the present disclosure, there is provided a model training device, including: an acquisition module configured to acquire training samples, where the training samples include sample images and sample segmentation results of the sample images; a training module configured to train an image segmentation model based on the training samples, where the image segmentation model includes N convolutional networks, each convolutional network includes a plurality of sub-networks, and each of the plurality of sub-networks of each convolutional network includes M first convolutional layers, and both N and M are positive integers greater than 1; a merging module configured to perform parameter merging on the plurality of sub-networks in the i-th convolutional network to obtain a second convolutional layer of the i-th, where i is a positive integer not greater than N; and an updating module configured to replace the i-th convolutional network with the second convolutional layer of the i-th to update the image segmentation model, and use the updated image segmentation model as the final image segmentation model.

[0007] According to a fourth aspect of the embodiments of the present disclosure, there is provided an image segmentation device, including: an acquisition module configured to acquire an original image to be segmented; a processing module configured to input the original image into an image segmentation model, and the image segmentation model outputs a target segmentation result of the original image, where the image segmentation model includes N second convolutional layers, N is a positive integer greater than 1, and the image segmentation model is obtained by using the model training method described in the first aspect of the embodiments of the present disclosure.

[0008] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including a processor; a memory for storing processor-executable instructions; where the processor is configured to implement the steps of the methods described in the first to second aspects of the embodiments of the present disclosure.

[0009] According to a sixth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the methods described in the first to second aspects of the embodiments of the present disclosure are implemented.

[0010] According to a seventh aspect of the embodiments of the present disclosure, there is provided a chip, which can execute the method described in the first or second aspect, and the chip can include an image processing chip ISP or a GPU and other chips.

[0011] The chip includes a processor, the processor is coupled to a memory, and when the processor executes computer instructions stored in the memory, the method described in the first or second aspect is implemented.

[0012] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects: Using a relatively complex convolutional network in the training process of the image segmentation model helps to improve the image segmentation accuracy. After the training is completed, the convolutional network is simplified into a second convolutional layer, making the image segmentation model more lightweight. The image segmentation model has the advantages of small volume, small required computing resources, low power consumption, and fast inference speed, improving the image segmentation efficiency, thereby enhancing the user experience in the image segmentation scenario. In addition, it can also reduce the storage space and computing resources occupied by the image segmentation model on the hardware device, facilitating the deployment of the image segmentation model in hardware devices with limited resource capacity.

[0013] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0015] Figure 1 is a flowchart of a model training method shown according to an exemplary embodiment.

[0016] Figure 2 is a schematic diagram of a model training method shown according to an exemplary embodiment.

[0017] Figure 3 is a flowchart of a model training method shown according to another exemplary embodiment.

[0018] Figure 4 is a flowchart of an image segmentation method shown according to an exemplary embodiment.

[0019] Figure 5 is a schematic diagram of an image segmentation model shown according to an exemplary embodiment.

[0020] Figure 6 is a flowchart of an image segmentation method shown according to another exemplary embodiment.

[0021] Figure 7 is a block diagram of a model training device shown according to an exemplary embodiment.

[0022] Figure 8 is a block diagram of an image segmentation device shown according to an exemplary embodiment.

[0023] Figure 9 is a block diagram of an electronic device shown according to an exemplary embodiment.

[0024] Figure 10 It is a block diagram of a chip shown according to an exemplary embodiment. Detailed implementation manners

[0025] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0027] In the technical solutions of the present disclosure, the acquisition, storage, use, processing, etc. of data all comply with the provisions of relevant laws and regulations.

[0028] Figure 1 It is a flowchart of a model training method shown according to an exemplary embodiment. As Figure 1 shown, the model training method of the embodiments of the present disclosure includes the following steps.

[0029] S101, obtain training samples, where the training samples include sample images and sample segmentation results of the sample images.

[0030] It should be noted that the execution subject of the model training method of the embodiments of the present disclosure is an electronic device, such as a mobile phone, a notebook, a desktop computer, a vehicle-mounted terminal, a smart home appliance, a wearable device, etc. Among them, the wearable device may include a wrist-worn device (such as a smart watch, a smart bracelet), a head-worn device, a foot-worn device, etc.

[0031] In the embodiments of the present disclosure, the method may be executed by a chip, such as an ISP (Image signal Processor, image processing chip) or a GPU (Graphics Processing Unit, image processing unit), etc. The chip may be integrated in a terminal device or an AI (Artificial Intelligence, artificial intelligence) server.

[0032] The model training method according to the embodiments of the present disclosure can be executed by the model training device according to the embodiments of the present disclosure. The model training device according to the embodiments of the present disclosure can be configured in any electronic device to execute the model training method according to the embodiments of the present disclosure.

[0033] It should be noted that there are no excessive restrictions on the sample images. For example, they may include RGB, HSV, HSL, YCbCr, Lab, YUV images, etc. There are no excessive restrictions on the sample segmentation results. For example, they may include the sample segmentation regions of the sample images, the sample categories of the sample pixel points in the sample images, the sample probabilities of the sample pixel points under each candidate category, etc. For example, the sample segmentation results can be obtained through manual annotation.

[0034] For example, taking the resolution of the sample image as H*W, the sample segmentation result is the target feature map. The resolution of the target feature map is H*W, and the number of channels is C. The channels of the target feature map correspond one-to-one with the candidate categories. The pixel value of the pixel point in the target feature map under the p-th channel is 0 or 1. 0 is used to represent that the category of the pixel point is not the candidate category corresponding to the p-th channel, and 1 is used to represent that the category of the pixel point is the candidate category corresponding to the p-th channel. Among them, the selected category can be used to indicate the segmentation object corresponding to the target feature map, such as a human body, the sky, etc.

[0035] For example, taking the number of channels of the target feature map as 3, the candidate category corresponding to the first channel of the target feature map is the sky, the candidate category corresponding to the second channel of the target feature map is the human body, and the candidate category corresponding to the third channel of the target feature map is the background. If the pixel values of the pixel point in the target feature map under the first to third channels are 1, 0, 0 respectively, then the category of the pixel point is the sky. If the pixel values of the pixel point in the target feature map under the first to third channels are 0, 1, 0 respectively, then the category of the pixel point is the human body. If the pixel values of the pixel point in the target feature map under the first to third channels are 0, 0, 1 respectively, then the category of the pixel point is the background.

[0036] S102, based on the training samples, train the image segmentation model, where the image segmentation model includes N convolutional networks, each convolutional network includes multiple sub-networks, and each of the multiple sub-networks of each convolutional network includes M first convolutional layers, and both N and M are positive integers greater than 1.

[0037] It should be noted that there are no excessive restrictions on the image segmentation model. For example, it may include FCN (Fully Convolutional Network), U-Net, Transformer model, etc.

[0038] It should be noted that there is no excessive limitation on the number of convolutional networks included in the image segmentation model. Different convolutional networks may be the same or different. There is no excessive limitation on the number of sub-networks included in each convolutional network. Different sub-networks in a convolutional network may be the same or different. There is no excessive limitation on the connection method between sub-networks in a convolutional network. For example, it may include serial connection, parallel connection, etc. Different first convolutional layers may be the same or different. There is no excessive limitation on the size of the convolutional kernel of the first convolutional layer. For example, it can be 1*1, 2*2, 3*3, etc.

[0039] It can be understood that in addition to the M first convolutional layers, the convolutional network may also include other sub-networks, such as a normalization layer, a pooling layer, etc. Among them, the normalization layer may include a BN (Batch Normalization) layer, and the pooling layer may include an AVG (Average Pooling) layer.

[0040] In one implementation, multiple sub-networks in each convolutional network are connected in series.

[0041] In one implementation, each convolutional network includes Q paths of convolution. Each path of convolution includes at least one first convolutional layer. The Q paths of convolution are connected in parallel, where Q is a positive integer greater than 1.

[0042] In some examples, M = 5. Each convolutional network includes four paths of convolution. The 1st first convolutional layer forms the 1st path of convolution. The 2nd to 3rd first convolutional layers are connected in series to form the 2nd path of convolution. The 4th first convolutional layer forms the 3rd path of convolution. The 5th first convolutional layer forms the 4th path of convolution. The 1st to 4th paths of convolution are connected in parallel. For example, the 2nd first convolutional layer is the input layer of the 2nd path of convolution, and the 3rd first convolutional layer is the output layer of the 2nd path of convolution.

[0043] For example, the sizes of the convolutional kernels of the 1st, 2nd, and 4th first convolutional layers are all 1*1, and the sizes of the convolutional kernels of the 3rd and 5th first convolutional layers are both K*K.

[0044] For example, as Figure 2 shown, the convolutional network includes the 1st to 5th first convolutional layers, the 1st to 6th BN layers, and the AVG layer. The 1st first convolutional layer and the 1st BN layer are connected in series to form the 1st path of convolution. The 2nd first convolutional layer, the 2nd BN layer, the 3rd first convolutional layer, and the 3rd BN layer are connected in series to form the 2nd path of convolution. The 4th first convolutional layer, the 4th BN layer, the AVG layer, and the 5th BN layer are connected in series to form the 3rd path of convolution. The 5th first convolutional layer and the 6th BN layer are connected in series to form the 4th path of convolution. The four paths of convolution are connected in parallel.

[0045] In some examples, each convolutional network further includes a fusion network, and the method further includes performing a fusion process on the feature maps of the Q-channel convolutional outputs through the fusion network to obtain the feature maps output by the convolutional network.

[0046] It should be noted that performing a fusion process on the feature maps of the Q-channel convolutional outputs through the fusion network to obtain the feature maps output by the convolutional network can be implemented by any image fusion method in the related art, and no excessive limitation is made here.

[0047] For example, performing a fusion process on the feature maps of the Q-channel convolutional outputs through the fusion network to obtain the feature maps output by the convolutional network includes obtaining the sum value of the pixel values of the p-th channel of the pixel points on the feature maps of the Q-channel convolutional outputs through the fusion network as the pixel value of the p-th channel of the pixel points on the feature maps output by the convolutional network.

[0048] For example, performing a fusion process on the feature maps of the Q-channel convolutional outputs through the fusion network to obtain the feature maps output by the convolutional network includes performing a weighted average on the pixel values of the p-th channel of the pixel points on the feature maps of the Q-channel convolutional outputs through the fusion network as the pixel value of the p-th channel of the pixel points on the feature maps output by the convolutional network.

[0049] It should be noted that training the image segmentation model based on the training samples can be implemented by any model training method in the related art, and no excessive limitation is made here.

[0050] In one implementation manner, training the image segmentation model based on the training samples includes inputting the sample images into the image segmentation model, outputting the predicted segmentation results of the sample images by the image segmentation model, and training the image segmentation model based on the predicted segmentation results and the sample segmentation results.

[0051] In some examples, training the image segmentation model based on the predicted segmentation results and the sample segmentation results includes obtaining the loss function of the image segmentation model based on the predicted segmentation results and the sample segmentation results, and training the image segmentation model based on the loss function. It should be noted that no excessive limitation is made on the loss function. For example, it may include CE (CrossEntropy), MSE (Mean-Square Error), KL (Kullback-Leibler) divergence, contrast loss function, etc.

[0052] S103. Merge the parameters of multiple sub-networks in the i-th convolutional network to obtain the i-th second convolutional layer, where i is a positive integer not greater than N.

[0053] It should be noted that merging the parameters of multiple sub-networks in a convolutional network means merging the parameters of multiple sub-networks in the convolutional network into the parameters of a second convolutional layer, that is, simplifying the convolutional network into a second convolutional layer. Different second convolutional layers may be the same or different, and there are no excessive restrictions on the convolutional kernel size of the second convolutional layer. For example, it can be 2*2, 3*3, etc. Merge the parameters of multiple sub-networks in the i-th convolutional network.

[0054] For example, continuing with Figure 2 the convolutional network shown as an example, the parameters of the first convolutional layer 1 to 5, the BN layer 1 to 6, and the AVG layer can be merged to obtain a second convolutional layer.

[0055] For example, taking the parameter merging of the serially connected convolutional layers 1 and 2 to obtain a merged convolutional layer as an example, the weight parameters and bias terms of convolutional layers 1 and 2 are as follows:

[0056] F 1 ∈R D×C×1×1 ,b 1 ∈R D

[0057] F 2 ∈R E×D×K×K ,b 2 ∈R E

[0058] Among them, F 1 is the weight parameter of convolutional layer 1, C is the number of channels of the feature map input to convolutional layer 1, D is the number of channels of the feature map output by convolutional layer 1, 1×1 is the convolutional kernel size of convolutional layer 1, b 1 is the bias term of convolutional layer 1, F 2 is the weight parameter of convolutional layer 2, D is the number of channels of the feature map input to convolutional layer 2, E is the number of channels of the feature map output by convolutional layer 2, K×K is the convolutional kernel size of convolutional layer 2, b 2 is the bias term of convolutional layer 2.

[0059] The process of obtaining the weight parameter F′ of the merged convolutional layer is as follows:

[0060] F′ = F 2 * TRANS(F 1 )

[0061] Among them, * is the convolution operator, and TRANS(·) is the transposed convolution operator.

[0062] The process of obtaining the bias term b′ of the merged convolutional layer is as follows:

[0063]

[0064] For example, taking the parameter merging of the parallel-connected convolutional layers 1 and 2 to obtain a merged convolutional layer as an example, the weight parameters and bias terms of convolutional layers 1 and 2 are as described in the above embodiments. The convolutional layer 1 can be converted into a convolutional layer X with a kernel size of K×K by padding 0 around it, and then the parameter merging of convolutional layer X and convolutional layer 2 is performed to obtain a merged convolutional layer. The weight parameter of the merged convolutional layer is the sum of the weight parameters of convolutional layer X and convolutional layer 2, and the bias term of the merged convolutional layer is the sum of the bias terms of convolutional layer X and convolutional layer 2.

[0065] For example, taking the parameter merging of convolutional layer 1 and BN layer to obtain a merged convolutional layer as an example, the weight parameters and bias terms of convolutional layer 1 are as described in the above embodiments, and the parameters of the BN layer are as follows:

[0066]

[0067] Among them, x is an original input data within the current batch, mean is the average value of all original input data within the current batch, var is the variance of all original input data within the current batch, γ is a scaling parameter, and β is an offset parameter.

[0068] The process of obtaining the weight parameter F′ of the merged convolutional layer is as follows:

[0069]

[0070] The process of obtaining the bias term b′ of the merged convolutional layer is as follows:

[0071]

[0072] Among them, REP(·) is a broadcast operator, and REP(b 1 ) is the broadcast bias term.

[0073] In one implementation, parameter merging is performed on multiple sub-networks in the i-th convolutional network to obtain the i-th second convolutional layer, including parameter merging of M first convolutional layers in the i-th convolutional network to obtain the i-th second convolutional layer.

[0074] S104, replacing the i-th convolutional network with the i-th second convolutional layer to update the image segmentation model, and using the updated image segmentation model as the final image segmentation model.

[0075] It should be noted that the final image segmentation model includes N second convolutional layers, and each second convolutional layer is obtained by parameter merging of multiple sub-networks in the convolutional network.

[0076] The model training method provided by the embodiments of the present disclosure obtains training samples, where the training samples include sample images and sample segmentation results of the sample images. Based on the training samples, an image segmentation model is trained. The image segmentation model includes N convolutional networks, each convolutional network includes multiple sub-networks, and multiple sub-networks of each convolutional network include M first convolutional layers. Both N and M are positive integers greater than 1. Parameter merging is performed on multiple sub-networks in the i-th convolutional network to obtain the i-th second convolutional layer, where i is a positive integer not greater than N. The i-th convolutional network is replaced with the i-th second convolutional layer to update the image segmentation model, and the updated image segmentation model is used as the final image segmentation model. Thus, a relatively complex convolutional network is adopted in the training process of the image segmentation model, which helps to improve the image segmentation accuracy. After training, the convolutional network is simplified to a second convolutional layer, making the image segmentation model more lightweight. The image segmentation model has the advantages of small volume, small required computing resources, low power consumption, and fast inference speed, improving the image segmentation efficiency, thereby enhancing the user experience in the image segmentation scenario. In addition, the storage space and computing resources occupied by the image segmentation model on hardware devices can be reduced, facilitating the deployment of the image segmentation model in hardware devices with limited resource capacity.

[0077] Based on any of the above embodiments, each convolutional network includes Q paths of convolution. Each path of convolution includes at least one first convolutional layer, and the Q paths of convolution are connected in parallel, where Q is a positive integer greater than 1. Each path of convolution includes multiple sub-networks.

[0078] Figure 3 is a flowchart of a model training method shown according to another exemplary embodiment. As Figure 3 shown, the model training method of the embodiments of the present disclosure includes the following steps.

[0079] S301, Obtain training samples, where the training samples include sample images and sample segmentation results of the sample images.

[0080] S302, Based on the training samples, train an image segmentation model, where the image segmentation model includes N convolutional networks, each convolutional network includes multiple sub-networks, and multiple sub-networks of each convolutional network include M first convolutional layers. Both N and M are positive integers greater than 1.

[0081] For the relevant content of steps S301 - S302, reference can be made to the above embodiments and will not be elaborated here.

[0082] S303, Perform parameter merging on multiple sub-networks of the j-th path of convolution in the i-th convolutional network to obtain the j-th merged convolutional layer, where j is a positive integer not greater than Q.

[0083] It should be noted that merging the parameters of multiple sub-networks of a certain path of convolution means merging the parameters of multiple sub-networks in a certain path of convolution into the parameters of a merged convolution layer, that is, simplifying a certain path of convolution into a merged convolution layer.

[0084] For example, continuing with Figure 2 the convolutional network shown, the kernel sizes of the first convolutional layers 1, 2, and 4 are all 1*1, and the kernel sizes of the first convolutional layers 3 and 5 are both K*K.

[0085] The parameters of the first convolutional layer 1 and the BN layer 1 can be merged to obtain a merged convolutional layer 1 with a kernel size of 1*1, that is, simplifying the first path of convolution into the merged convolutional layer 1.

[0086] The parameters of the first convolutional layer 2 and the BN layer 2 can be merged to obtain a fourth convolutional layer 1 with a kernel size of 1*1. The parameters of the first convolutional layer 3 and the BN layer 3 can be merged to obtain a fourth convolutional layer 2 with a kernel size of K*K. The parameters of the fourth convolutional layer 1 and the fourth convolutional layer 2 can be merged to obtain a merged convolutional layer 2 with a kernel size of K*K, that is, simplifying the second path of convolution into the merged convolutional layer 2.

[0087] The parameters of the first convolutional layer 4, the BN layer 4, the AVG layer, and the BN layer 5 can be merged to obtain a merged convolutional layer 3 with a kernel size of 1*1, that is, simplifying the third path of convolution into the merged convolutional layer 3.

[0088] The parameters of the first convolutional layer 5 and the BN layer 6 can be merged to obtain a merged convolutional layer 4 with a kernel size of K*K, that is, simplifying the fourth path of convolution into the merged convolutional layer 4.

[0089] S304. Merge the parameters of the Q merged convolutional layers to obtain the i-th second convolutional layer.

[0090] For example, continuing with the above-mentioned merged convolutional layers 1 to 4, the parameters of the merged convolutional layers 1 to 4 can be merged to obtain a second convolutional layer with a kernel size of K*K, that is, simplifying the convolutional network into a second convolutional layer.

[0091] In one embodiment, parameter merging is performed on Q merged convolutional layers to obtain the i-th second convolutional layer, including determining a reference convolutional layer from the Q merged convolutional layers, and taking the remaining convolutional layers other than the reference convolutional layer in the Q merged convolutional layers as the third convolutional layer, performing parameter conversion on the third convolutional layer to obtain a converted convolutional layer with the same kernel size as the reference convolutional layer to update the third convolutional layer, and performing parameter merging on the third convolutional layer and the reference convolutional layer to obtain the i-th second convolutional layer. Thus, parameter conversion can be performed on the third convolutional layer to make the kernel size of the third convolutional layer the same as that of the reference convolutional layer, facilitating subsequent parameter merging of the third convolutional layer and the reference convolutional layer with the same kernel size to obtain the second convolutional layer.

[0092] It should be noted that the Q merged convolutional layers include a reference convolutional layer and a third convolutional layer. The number of reference convolutional layers is at least one, and the sum of the number of reference convolutional layers and the number of third convolutional layers is Q.

[0093] In some examples, determining a reference convolutional layer from the Q merged convolutional layers includes taking the merged convolutional layer with the largest kernel size as the reference convolutional layer.

[0094] For example, continuing with the above-mentioned merged convolutional layers 1 to 4 as an example, both merged convolutional layers 2 and 4 can be used as the reference convolutional layers, and both merged convolutional layers 1 and 3 can be used as the third convolutional layers.

[0095] Perform parameter conversion on merged convolutional layer 1 to obtain converted convolutional layer 1 with a kernel size of 3*3, and update merged convolutional layer 1 to converted convolutional layer 1, that is, the kernel size of the updated merged convolutional layer 1 is 3*3.

[0096] Perform parameter conversion on merged convolutional layer 3 to obtain converted convolutional layer 2 with a kernel size of 3*3, and update merged convolutional layer 3 to converted convolutional layer 2, that is, the kernel size of the updated merged convolutional layer 3 is 3*3.

[0097] Perform parameter merging on merged convolutional layers 1 to 4 to obtain the second convolutional layer.

[0098] S305, Replace the i-th convolutional network with the i-th second convolutional layer to update the image segmentation model, and use the updated image segmentation model as the final image segmentation model.

[0099] For the relevant content of step S305, reference can be made to the above-mentioned embodiments, which will not be elaborated here.

[0100] The model training method provided by the embodiments of the present disclosure merges the parameters of multiple sub-networks of the j-th path convolution in the i-th convolutional network to obtain the j-th merged convolutional layer, where j is a positive integer not greater than Q. Then, the parameters of the Q merged convolutional layers are merged to obtain the i-th second convolutional layer. Thus, the parameters of each path convolution in the convolutional network can be merged separately to obtain multiple merged convolutional layers, and then the parameters of the multiple merged convolutional layers are merged to obtain the second convolutional layer, so as to simplify the convolutional network into a second convolutional layer.

[0101] Figure 4 is a flowchart of an image segmentation method shown according to an exemplary embodiment, as Figure 4 shown, the image segmentation method of the embodiments of the present disclosure includes the following steps.

[0102] S401, Obtain the original image to be segmented.

[0103] S402, Input the original image into the image segmentation model, and the image segmentation model outputs the target segmentation result of the original image, where the image segmentation model includes N second convolutional layers, and N is a positive integer greater than 1.

[0104] It should be noted that the execution subject of the image segmentation method of the embodiments of the present disclosure is an electronic device, such as a mobile phone, a notebook, a desktop computer, a vehicle-mounted terminal, a smart home appliance, a wearable device, etc. Among them, the wearable device may include a wrist-worn device (such as a smart watch, a smart bracelet), a head-mounted device, a foot-worn device, etc.

[0105] The image segmentation method of the embodiments of the present disclosure can be executed by the image segmentation device of the embodiments of the present disclosure. The image segmentation device of the embodiments of the present disclosure can be configured in any electronic device to execute the image segmentation method of the embodiments of the present disclosure.

[0106] It should be noted that the image segmentation model is obtained by using the model training method proposed by the present disclosure. For the relevant content of the original image, reference can be made to the relevant content of the sample image in the above embodiments. For the relevant content of the target segmentation result, reference can be made to the relevant content of the sample segmentation result in the above embodiments, which will not be elaborated here. The image segmentation model generates the target segmentation result, which can be implemented by any image segmentation method in the related art, and will not be limited here.

[0107] In one implementation, the method further includes performing multi-scale feature extraction on the original image through the image segmentation model to obtain multi-scale feature maps, and performing fusion processing on the multi-scale feature maps to obtain the target feature map as the target segmentation result.

[0108] The image segmentation method provided by the embodiments of the present disclosure obtains an original image to be segmented, inputs the original image into an image segmentation model, and the image segmentation model outputs a target segmentation result of the original image. Among them, the image segmentation model is obtained by using the model training method proposed by the present disclosure, which helps to improve the accuracy of image segmentation. The image segmentation model is more lightweight, and has the advantages of small volume, small required computing resources, low power consumption, and fast inference speed, improving the image segmentation efficiency, and thus enhancing the user experience in the image segmentation scenario.

[0109] Based on any of the above embodiments, the image segmentation model includes a downsampling network, an upsampling network, an ARM (Attention Refinement Module) network, and a splicing network. It should be noted that no excessive limitations are imposed on the downsampling network, the upsampling network, the ARM network, and the splicing network.

[0110] For example, the downsampling network may include multiple first convolutional layers and multiple activation layers. Among them, the activation layers may include SiLU (SigmoidLinear Unit), ReLu (Rectified Linear Unit), etc.

[0111] For example, the upsampling network may include an interpolation network, PixelShuffle, a transposed convolutional network, etc. Among them, the interpolation network may include a bilinear interpolation network, a bicubic interpolation network, a nearest neighbor interpolation network, etc. PixelShuffle is also called a sub-pixel convolutional neural network, a sub-pixel convolutional network, a pixel shuffle network, and the transposed convolutional network is also called a deconvolution network.

[0112] For example, the ARM network may include multiple first convolutional layers, multiple activation layers, and a GAP (Global AveragePooling) layer.

[0113] For example, the splicing network may include a Concat network.

[0114] For example, as Figure 5 shown, the image segmentation model includes a downsampling network, upsampling networks 1 to 4, ARM networks 1 to 4, and splicing networks 1 to 4. Among them, the downsampling network includes first convolutional layers 1 to 12.

[0115] For example, the convolutional kernels of the first convolutional layers 1 to 12 are all 3*3.

[0116] For example, the upsampling networks 1 to 4 are all transposed convolutional networks with a convolutional kernel size of 2*2.

[0117] For example, the output ends of the first convolutional layers 1 to 12 are respectively connected to the input end of an activation layer.

[0118] Figure 6 It is a flowchart of an image segmentation method shown according to another exemplary embodiment. As Figure 6 shown, the method of the embodiments of the present disclosure includes the following steps.

[0119] S601, obtain the original image to be segmented, and input the original image into the image segmentation model.

[0120] For the relevant content of step S601, reference can be made to the above embodiments, which will not be elaborated here.

[0121] S602, perform multi-scale feature extraction on the original image through the downsampling network to obtain a multi-scale first feature map.

[0122] It should be noted that there are no excessive restrictions on the scale of feature extraction. For example, taking the resolution of the original image as 512*512 as an example, the resolution of the first feature map may include 256*256, 128*128, 64*64, 32*32, 16*16.

[0123] For example, continuing with Figure 5 as an example, the resolution of the original image is 512*512, and the number of channels is 3.

[0124] Perform feature extraction on the original image through the first convolutional layer 1 to obtain a first feature with a resolution of 256*256 and a channel number of 32 Figure 1 .

[0125] Perform feature extraction on the first feature through the first convolutional layer 2 Figure 1 to obtain a first feature with a resolution of 256*256 and a channel number of 32 Figure 2 .

[0126] Perform feature extraction on the first feature through the first convolutional layer 3 Figure 2 to obtain a first feature with a resolution of 128*128 and a channel number of 64 Figure 3 .

[0127] Perform feature extraction on the first feature through the first convolutional layer 4 Figure 3 to obtain a first feature with a resolution of 128*128 and a channel number of 64 Figure 4 .

[0128] Perform feature extraction on the first feature through the first convolutional layer 5 Figure 4 to obtain a first feature with a resolution of 64*64 and a channel number of 128 Figure 5 .

[0129] The first convolution layer 6 performs feature extraction on the first feature Figure 5 to obtain a first feature map with a resolution of 64*64 and 128 channels Figure 6 .

[0130] The first convolution layer 7 performs feature extraction on the first feature Figure 6 to obtain a first feature map with a resolution of 32*32 and 256 channels Figure 7 .

[0131] The first convolution layer 8 performs feature extraction on the first feature Figure 7 to obtain a first feature map with a resolution of 32*32 and 256 channels Figure 8 .

[0132] The first convolution layer 9 performs feature extraction on the first feature Figure 8 to obtain a first feature map with a resolution of 16*16 and 256 channels Figure 9 .

[0133] The first convolution layer 10 performs feature extraction on the first feature Figure 9 to obtain a first feature map with a resolution of 16*16 and 256 channels Figure 10 .

[0134] The first convolution layer 11 performs feature extraction on the first feature Figure 10 to obtain the first feature map 11 with a resolution of 16*16 and 256 channels

[0135] The first convolution layer 12 performs feature extraction on the first feature map 11 to obtain the first feature map 12 with a resolution of 8*8 and 256 channels

[0136] S603. The upsampling network performs upsampling on the first feature map with the smallest scale to obtain the first second feature map, and performs upsampling on the multi-scale concatenated feature maps respectively to obtain the remaining second feature maps at multiple scales

[0137] For example, continuing with Figure 5 as an example, the upsampling network 1 performs upsampling on the first feature map 12 to obtain a second feature with a resolution of 16*16 and 256 channels Figure 1 .

[0138] The upsampling network 2 performs upsampling on the concatenated feature Figure 1 to obtain a second feature with a resolution of 32*32 and 256 channels Figure 2 .

[0139] The upsampling network 3 performs upsampling on the concatenated featureFigure 2 Perform upsampling processing to obtain a second feature with a resolution of 64*64 and 128 channels Figure 3 .

[0140] Perform upsampling processing on the concatenated feature through the upsampling network 4 Figure 3 to obtain a second feature with a resolution of 128*128 and 64 channels Figure 4 .

[0141] S604. Feature enhancement processing is performed on the h-th second feature map through the ARM network to obtain the h-th third feature map, where the scale of the h-th second feature map is the same as the scale of the h-th third feature map, and h is a positive integer.

[0142] For example, continuing with Figure 5 as an example, feature enhancement processing is performed on the second feature through the ARM network 1 Figure 1 to obtain a third feature with a resolution of 16*16 and 256 channels Figure 1 .

[0143] Feature enhancement processing is performed on the second feature through the ARM network 2 Figure 2 to obtain a third feature with a resolution of 32*32 and 256 channels Figure 2 .

[0144] Feature enhancement processing is performed on the second feature through the ARM network 3 Figure 3 to obtain a third feature with a resolution of 64*64 and 128 channels Figure 3 .

[0145] Feature enhancement processing is performed on the second feature through the ARM network 4 Figure 4 to obtain a third feature with a resolution of 128*128 and 64 channels Figure 4 .

[0146] In one implementation, feature enhancement processing is performed on the h-th second feature map through the ARM network to obtain the h-th third feature map, including obtaining the attention weight of each channel of the h-th second feature map through the ARM network, and performing feature enhancement processing on the h-th second feature map according to the attention weight of each channel of the h-th second feature map through the ARM network to obtain the h-th third feature map. Thus, the attention weight of each channel of the second feature map can be obtained through the ARM network to perform feature enhancement processing on the second feature map to obtain the third feature map.

[0147] It should be noted that the attention weights of different channels of the second feature map may be different.

[0148] In some examples, obtaining the attention weights of each channel of the h-th second feature map through the ARM network includes performing feature extraction on the h-th second feature map through the first convolutional layer in the ARM network to obtain the seventh feature map, performing a non-linear transformation on the seventh feature map through the activation layer in the ARM network to obtain the eighth feature map, compressing the eighth feature map into a global feature vector through the GAP layer in the ARM network, obtaining the attention weights of each channel of the h-th second feature map through the first convolutional layer in the ARM network based on the global feature vector, and performing a non-linear transformation on the attention weights of each channel of the h-th second feature map through the activation layer in the ARM network to update the attention weights of each channel of the h-th second feature map.

[0149] In some examples, performing feature enhancement processing on the h-th second feature map according to the attention weights of each channel of the h-th second feature map through the ARM network to obtain the h-th third feature map includes obtaining the product between the pixel value of a pixel point in the p-th channel of the h-th second feature map and the attention weight of the p-th channel of the h-th second feature map through the ARM network to obtain the pixel value of the pixel point in the p-th channel of the h-th third feature map.

[0150] S605, performing splicing processing on the first feature map and the third feature map of the same scale through the splicing network to obtain a multi-scale spliced feature map.

[0151] It should be noted that the splicing processing refers to splicing according to the channel dimension, and the number of channels of the spliced feature map is the sum of the number of channels of the spliced first feature map and the third feature map.

[0152] For example, continuing with Figure 5 as an example, performing splicing processing on the first feature map 11 and the third feature Figure 1 through the splicing network 1 to obtain a spliced feature map with a resolution of 16 * 16 and 512 channels Figure 1 .

[0153] Performing splicing processing on the first feature Figure 8 and the third feature Figure 2 through the splicing network 2 to obtain a spliced feature map with a resolution of 32 * 32 and 512 channels Figure 2 .

[0154] Performing splicing processing on the first feature Figure 6 and the third feature Figure 3 through the splicing network 3 to obtain a spliced feature map with a resolution of 64 * 64 and 256 channels Figure 3 .

[0155] Performing splicing processing on the first feature Figure 4 and the third featureFigure 4 Perform splicing processing to obtain a spliced feature with a resolution of 128*128 and 128 channels Figure 4 .

[0156] In one embodiment, the image segmentation model further includes a first feature extraction network, and the method further includes obtaining a third feature with a scale equal to the maximum scale Figure 1 A first feature map with the same scale is used as the fourth feature map. The fourth feature map is subjected to feature extraction through the first feature extraction network to obtain a fifth feature map for updating the fourth feature map, where the scale of the fifth feature map is the same as that of the fourth feature map. Thus, the fourth feature map is a first feature map with a larger scale and resolution, which helps to restore better spatial features, retain more spatial features, avoid loss of spatial features, and thereby improve the image segmentation accuracy

[0157] In some examples, the first feature extraction network includes 3 second convolutional layers

[0158] For example Figure 5 As shown, the image segmentation model further includes a first feature extraction network, and the first feature extraction network includes second convolutional layers 13 to 15. For example, the output ends of the first convolutional layers 13 to 15 are respectively connected to the input end of an activation layer

[0159] For example, continuing with Figure 5 as an example, the first feature Figure 4 is used as the fourth feature map, and the first feature Figure 4 is subjected to feature extraction through the first convolutional layer 13 to obtain a ninth feature with a resolution of 128*128 and 128 channels Figure 1 , and the ninth feature Figure 1 is subjected to feature extraction through the first convolutional layer 14 to obtain a ninth feature with a resolution of 128*128 and 128 channels Figure 2 , and the ninth feature Figure 2 is subjected to feature extraction through the first convolutional layer 15 to obtain a fifth feature map with a resolution of 128*128 and 64 channels, and the first feature Figure 4 is updated to the fifth feature map

[0160] In one embodiment, the method further includes using the spliced feature map with the maximum scale as the target segmentation result

[0161] The image segmentation method provided by the embodiments of the present disclosure extracts multi-scale features from the original image through a downsampling network to obtain multi-scale first feature maps. The smallest-scale first feature map is upsampled through an upsampling network to obtain the first second feature map, and the remaining second feature maps of multiple scales are obtained by performing upsampling processing on the multi-scale concatenated feature maps respectively. The h-th second feature map is subjected to feature enhancement processing through an ARM network to obtain the h-th third feature map, where the scale of the h-th second feature map is the same as that of the h-th third feature map, and h is a positive integer. The first feature map and the third feature map of the same scale are concatenated through a concatenation network to obtain multi-scale concatenated feature maps. Thus, the feature maps can be subjected to feature enhancement processing through the ARM network, that is, the attention mechanism is used to highlight important features and suppress unimportant features, improving the feature expression ability of the image segmentation model, and thus improving the image segmentation accuracy.

[0162] Based on any of the above embodiments, the image segmentation model further includes a second feature extraction network, and the method further includes extracting features from the concatenated feature map of the largest scale through the second feature extraction network to obtain a sixth feature map, where the scale of the sixth feature map is the same as that of the concatenated feature map of the largest scale, and the sixth feature map is upsampled through the upsampling network to obtain a target feature map as the target segmentation result, where the scale of the target feature map is the same as that of the original image.

[0163] For example, as Figure 5 shown, the image segmentation model further includes a second feature extraction network and an upsampling network 5, and the second feature extraction network includes second convolutional layers 16 to 17. For example, the output end of the second convolutional layer 16 is connected to the input end of an activation layer.

[0164] For example, continuing with Figure 5 as an example, the concatenated feature Figure 4 is subjected to feature extraction through the second convolutional layer 16 to obtain a tenth feature map with a resolution of 128*128 and 32 channels.

[0165] The tenth feature map is subjected to feature extraction through the second convolutional layer 17 to obtain a sixth feature map with a resolution of 128*128 and C channels.

[0166] The sixth feature map is upsampled through the upsampling network 5 to obtain a target feature map with a resolution of 512*512 and C channels.

[0167] Figure 7 is a block diagram of a model training device shown according to an exemplary embodiment. Refer to Figure 7, the model training device 100 of the embodiments of the present disclosure includes: an acquisition module 110, a training module 120, a merging module 130, and an updating module 140.

[0168] The acquisition module 110 is configured to acquire training samples, where the training samples include sample images and sample segmentation results of the sample images;

[0169] The training module 120 is configured to train an image segmentation model based on the training samples, where the image segmentation model includes N convolutional networks, each convolutional network includes a plurality of sub-networks, and the plurality of sub-networks of each convolutional network each include M first convolutional layers, and both N and M are positive integers greater than 1;

[0170] The merging module 130 is configured to perform parameter merging on the plurality of sub-networks in the i-th convolutional network to obtain the i-th second convolutional layer, where i is a positive integer not greater than N;

[0171] The updating module 140 is configured to replace the i-th convolutional network with the i-th second convolutional layer to update the image segmentation model, and use the updated image segmentation model as the final image segmentation model.

[0172] In an embodiment of the present disclosure, each convolutional network includes Q paths of convolutions, each path of convolution includes at least one first convolutional layer, and the Q paths of convolutions are connected in parallel, where Q is a positive integer greater than 1.

[0173] In an embodiment of the present disclosure, each path of convolution includes a plurality of sub-networks, and the merging module 130 is further configured to: perform parameter merging on the plurality of sub-networks of the j-th path of convolution in the i-th convolutional network to obtain the j-th merged convolutional layer, where j is a positive integer not greater than Q; perform parameter merging on the Q merged convolutional layers to obtain the i-th second convolutional layer.

[0174] In an embodiment of the present disclosure, the merging module 130 is further configured to: determine a reference convolutional layer from the Q merged convolutional layers, and use the remaining convolutional layers other than the reference convolutional layer in the Q merged convolutional layers as the third convolutional layer; perform parameter conversion on the third convolutional layer to obtain a converted convolutional layer with a convolutional kernel size consistent with that of the reference convolutional layer to update the third convolutional layer; perform parameter merging on the third convolutional layer and the reference convolutional layer to obtain the i-th second convolutional layer.

[0175] In one embodiment of the present disclosure, M = 5. Each convolutional network includes four-way convolution. The first first convolutional layer constitutes the first way of convolution. The second to third first convolutional layers are connected in series to form the second way of convolution. The fourth first convolutional layer constitutes the third way of convolution. The fifth first convolutional layer constitutes the fourth way of convolution. The first to fourth ways of convolution are connected in parallel.

[0176] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.

[0177] The model training device provided by the embodiment of the present disclosure obtains training samples, where the training samples include sample images and sample segmentation results of the sample images. Based on the training samples, an image segmentation model is trained. The image segmentation model includes N convolutional networks, each convolutional network includes multiple sub-networks, and the multiple sub-networks of each convolutional network include M first convolutional layers. Both N and M are positive integers greater than 1. The parameters of the multiple sub-networks in the i-th convolutional network are merged to obtain the i-th second convolutional layer, where i is a positive integer not greater than N. The i-th convolutional network is replaced with the i-th second convolutional layer to update the image segmentation model, and the updated image segmentation model is used as the final image segmentation model. Thus, a relatively complex convolutional network is adopted in the training process of the image segmentation model, which helps to improve the image segmentation accuracy. After training, the convolutional network is simplified to a second convolutional layer, making the image segmentation model more lightweight. The image segmentation model has the advantages of small volume, small required computing resources, low power consumption, and fast inference speed, improving the image segmentation efficiency, thereby enhancing the user experience in the image segmentation scenario. In addition, the storage space and computing resources occupied by the image segmentation model on the hardware device can be reduced, facilitating the deployment of the image segmentation model in hardware devices with limited resource capacity.

[0178] Figure 8 is a block diagram of an image segmentation device shown according to an exemplary embodiment. Referring to Figure 8 , the image segmentation device 200 of the embodiment of the present disclosure includes: an acquisition module 210 and a processing module 220.

[0179] The acquisition module 210 is configured to acquire an original image to be segmented;

[0180] The processing module 220 is configured to input the original image into the image segmentation model, and the image segmentation model outputs a target segmentation result of the original image. The image segmentation model includes N second convolutional layers, N is a positive integer greater than 1, and the image segmentation model is obtained by using the model training method proposed by the present disclosure.

[0181] In one embodiment of the present disclosure, the image segmentation model includes a downsampling network, an upsampling network, an attention refinement ARM network, and a splicing network;

[0182] The processing module 220 is further configured to: perform multi-scale feature extraction on the original image through the downsampling network to obtain multi-scale first feature maps; perform upsampling processing on the first feature map with the smallest scale through the upsampling network to obtain the first second feature map, and perform upsampling processing on the multi-scale spliced feature maps respectively to obtain the remaining second feature maps of multiple scales; perform feature enhancement processing on the h-th second feature map through the ARM network to obtain the h-th third feature map, where the scale of the h-th second feature map is the same as the scale of the h-th third feature map, and h is a positive integer; perform splicing processing on the first feature map and the third feature map with the same scale through the splicing network to obtain multi-scale spliced feature maps.

[0183] In one embodiment of the present disclosure, the processing module 220 is further configured to: obtain the attention weight of each channel of the h-th second feature map through the ARM network; perform feature enhancement processing on the h-th second feature map according to the attention weight of each channel of the h-th second feature map through the ARM network to obtain the h-th third feature map.

[0184] In one embodiment of the present disclosure, the image segmentation model further includes a first feature extraction network, and the processing module 220 is further configured to: obtain a first feature map with a scale consistent with that of the third feature with the largest scale as the fourth feature map; perform feature extraction on the fourth feature map through the first feature extraction network to obtain a fifth feature map to update the fourth feature map, where the scale of the fifth feature map is the same as the scale of the fourth feature map. Figure 1 In one embodiment of the present disclosure, the first feature extraction network includes 3 second convolutional layers.

[0185] In one embodiment of the present disclosure, the image segmentation model further includes a second feature extraction network, and the processing module 220 is further configured to: perform feature extraction on the spliced feature map with the largest scale through the second feature extraction network to obtain a sixth feature map, where the scale of the sixth feature map is the same as the scale of the spliced feature map with the largest scale; perform upsampling processing on the sixth feature map through the upsampling network to obtain a target feature map as the target segmentation result, where the scale of the target feature map is the same as the scale of the original image.

[0186] In one embodiment of the present disclosure, the image segmentation model further includes a second feature extraction network, and the processing module 220 is further configured to: perform feature extraction on the spliced feature map with the largest scale through the second feature extraction network to obtain a sixth feature map, where the scale of the sixth feature map is the same as the scale of the spliced feature map with the largest scale; perform upsampling processing on the sixth feature map through the upsampling network to obtain a target feature map as the target segmentation result, where the scale of the target feature map is the same as the scale of the original image.

[0187] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0188] The image segmentation device provided by the embodiments of the present disclosure obtains an original image to be segmented, inputs the original image into an image segmentation model, and the image segmentation model outputs a target segmentation result of the original image. Among them, the image segmentation model is obtained by using the model training method proposed by the present disclosure, which helps to improve the accuracy of image segmentation. The image segmentation model is more lightweight, and has the advantages of small volume, small required computing resources, low power consumption, and fast inference speed, improving the image segmentation efficiency, and thus enhancing the user experience in the image segmentation scenario.

[0189] Figure 9 It is a block diagram of an electronic device shown according to an exemplary embodiment.

[0190] As Figure 9 shown, the above electronic device 300 includes:

[0191] A memory 310 and a processor 320, a bus 330 connecting different components (including the memory 310 and the processor 320). The memory 310 stores a computer program, and when the processor 320 executes the program, it implements the model training method and the image segmentation method described in the embodiments of the present disclosure.

[0192] The bus 330 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any bus structure in a variety of bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0193] The electronic device 300 typically includes a variety of electronic device-readable media. These media can be any available media accessible by the electronic device 300, including volatile and non-volatile media, removable and non-removable media.

[0194] The memory 310 may further include a computer system-readable medium in the form of volatile memory, such as a random access memory (RAM) 340 and / or a cache memory 350. The electronic device 300 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 360 may be used to read and write non-removable, non-volatile magnetic media ( Figure 9 not shown, commonly referred to as a "hard disk drive"). AlthoughFigure 9 is not shown. Disk drives for reading from and writing to removable non-volatile disks (such as "floppy disks") and optical disk drives for reading from and writing to removable non-volatile optical disks (such as CD-ROMs, DVD-ROMs, or other optical media) may be provided. In these cases, each drive may be connected to the bus 330 through one or more data medium interfaces. The memory 310 may include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present disclosure.

[0195] A program / utilities 380 having a set (at least one) of program modules 370 may be stored, for example, in the memory 310. Such program modules 370 include—but are not limited to—an operating system, one or more application programs, other program modules, and program data. Implementations of a network environment may be included in each or some combination of these examples. The program modules 370 generally perform the functions and / or methods in the embodiments described in the present disclosure.

[0196] The electronic device 300 may also communicate with one or more external devices 390 (such as a keyboard, a pointing device, a display 391, etc.), and may also communicate with one or more devices that enable a user to interact with the electronic device 300, and / or communicate with any device that enables the electronic device 300 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 392. Moreover, the electronic device 300 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 393. As Figure 9 shown, the network adapter 393 communicates with other modules of the electronic device 300 through the bus 330. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 300, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0197] The processor 320 executes various functional applications and data processing by running the programs stored in the memory 310.

[0198] It should be noted that for the implementation process and technical principles of the electronic device in this embodiment, refer to the foregoing explanations of the model training method and image segmentation method of the embodiments of the present disclosure, which will not be elaborated here.

[0199] The electronic device provided by the embodiments of the present disclosure can execute the model training method described above to obtain training samples, where the training samples include sample images and sample segmentation results of the sample images. Based on the training samples, an image segmentation model is trained. The image segmentation model includes N convolutional networks, each convolutional network includes multiple sub-networks, and each of the multiple sub-networks of each convolutional network includes M first convolutional layers. Both N and M are positive integers greater than 1. Parameter merging is performed on the multiple sub-networks in the i-th convolutional network to obtain the i-th second convolutional layer, where i is a positive integer not greater than N. The i-th convolutional network is replaced with the i-th second convolutional layer to update the image segmentation model, and the updated image segmentation model is used as the final image segmentation model. Thus, a relatively complex convolutional network is adopted during the training process of the image segmentation model, which helps to improve the image segmentation accuracy. After training, the convolutional network is simplified into a second convolutional layer, making the image segmentation model more lightweight. The image segmentation model has the advantages of small volume, small required computing resources, low power consumption, and fast inference speed, improving the image segmentation efficiency, thereby enhancing the user experience in the image segmentation scenario. In addition, the storage space and computing resources occupied by the image segmentation model on the hardware device can be reduced, facilitating the deployment of the image segmentation model in hardware devices with limited resource capacity.

[0200] To implement the above embodiments, the present disclosure also proposes a computer-readable storage medium, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the model training method and the image segmentation method provided by the present disclosure are implemented.

[0201] Optionally, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0202] Figure 10 It is a block diagram of a chip shown according to an exemplary embodiment.

[0203] As Figure 10 shown, the above chip 400 includes one or more interface circuits 420 and one or more processors 410; the interface circuit 420 is used to receive signals and send signals to the processor 410. The signals include computer instructions stored in the memory. When the processor 410 executes the computer instructions, the chip 400 executes the steps of the parameter adjustment method provided by the present disclosure.

[0204] The processor 410 and the interface circuit 420 can be interconnected by a line.

[0205] The chip 400 further includes a memory 430, and all or part of the memory 430 may be outside the chip 400.

[0206] The interface circuit 420 is connected to the memory 430. The interface circuit 420 can be used to receive signals from the memory 430 or other devices, and the interface circuit 420 can be used to send signals to the processor 410 or other devices. For example, the interface circuit 420 can read the instructions stored in the memory 430 and send the instructions to the processor 410.

[0207] The interface circuit 420 can obtain data, program instructions, and / or information, etc. in the internal storage area of the chip 400; it can also obtain data, program instructions, and / or information, etc. outside the chip 400.

[0208] It should be noted that terms such as interface circuit, interface, transceiver pin, transceiver, etc. can be replaced with each other.

[0209] To implement the above embodiments, the present disclosure also provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor of an electronic device, it implements the model training method and image segmentation method as described above.

[0210] Those skilled in the art will readily think of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0211] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A model training method, characterized in that Including: Obtain training samples, where the training samples include sample images and sample segmentation results of the sample images; Based on the training samples, train an image segmentation model, where the image segmentation model includes N convolutional networks, each convolutional network includes multiple sub-networks, and multiple sub-networks of each convolutional network each include M first convolutional layers, and both N and M are positive integers greater than 1; Merge the parameters of multiple sub-networks in the i-th convolutional network to obtain the i-th second convolutional layer, where i is a positive integer not greater than N; Replace the i-th convolutional network with the i-th second convolutional layer to update the image segmentation model, and use the updated image segmentation model as the final image segmentation model.

2. The method according to claim 1, wherein Each convolutional network includes Q paths of convolutions, each path of convolution includes at least one first convolutional layer, and the Q paths of convolutions are connected in parallel, where Q is a positive integer greater than 1.

3. The method according to claim 2, wherein Each path of convolution includes multiple sub-networks. The step of merging the parameters of multiple sub-networks in the i-th convolutional network to obtain the i-th second convolutional layer includes: Merge the parameters of multiple sub-networks of the j-th path of convolution in the i-th convolutional network to obtain the j-th merged convolutional layer, where j is a positive integer not greater than Q; Merge the parameters of Q merged convolutional layers to obtain the i-th second convolutional layer.

4. The method according to claim 3, wherein The step of merging the parameters of Q merged convolutional layers to obtain the i-th second convolutional layer includes: Determine a reference convolutional layer from the Q merged convolutional layers, and use the remaining convolutional layers other than the reference convolutional layer in the Q merged convolutional layers as the third convolutional layer; Perform parameter conversion on the third convolutional layer to obtain a converted convolutional layer with a convolutional kernel size consistent with that of the reference convolutional layer to update the third convolutional layer; Merge the parameters of the third convolutional layer and the reference convolutional layer to obtain the i-th second convolutional layer.

5. The method according to claim 2, wherein M = 5, each convolutional network includes four paths of convolutions. The 1st first convolutional layer forms the 1st path of convolution, the 2nd to 3rd first convolutional layers are connected in series to form the 2nd path of convolution, the 4th first convolutional layer forms the 3rd path of convolution, the 5th first convolutional layer forms the 4th path of convolution, and the 1st to 4th paths of convolutions are connected in parallel.

6. An image segmentation method, characterized in that, Including: Obtain an original image to be segmented; Input the original image into the image segmentation model, and the image segmentation model outputs the target segmentation result of the original image. The image segmentation model includes N second convolutional layers, N is a positive integer greater than 1, and the image segmentation model is obtained by using the model training method described in any one of claims 1-5.

7. The method according to claim 6, characterized in that, The image segmentation model includes a downsampling network, an upsampling network, an attention refinement ARM network, and a splicing network; The method further includes: Perform multi-scale feature extraction on the original image through the downsampling network to obtain multi-scale first feature maps; Perform upsampling processing on the first feature map with the smallest scale through the upsampling network to obtain the 1st second feature map, and perform upsampling processing on the multi-scale spliced feature maps respectively to obtain the remaining second feature maps of multiple scales; Performing feature enhancement processing on the h-th second feature map through the ARM network to obtain the h-th third feature map, where the scale of the h-th second feature map is the same as the scale of the h-th third feature map, and h is a positive integer; Performing splicing processing on the first feature map and the third feature map of the same scale through the splicing network to obtain a multi-scale spliced feature map.

8. The method according to claim 7, characterized in that, The performing feature enhancement processing on the h-th second feature map through the ARM network to obtain the h-th third feature map includes: Obtaining the attention weight of each channel of the h-th second feature map through the ARM network; Performing feature enhancement processing on the h-th second feature map according to the attention weight of each channel of the h-th second feature map through the ARM network to obtain the h-th third feature map.

9. The method according to claim 7, wherein The image segmentation model further includes a first feature extraction network, and the method further includes: Obtaining a first feature map with a scale consistent with that of the third feature map of the largest scale as the fourth feature map; Performing feature extraction on the fourth feature map through the first feature extraction network to obtain a fifth feature map to update the fourth feature map, where the scale of the fifth feature map is the same as the scale of the fourth feature map.

10. The method according to claim 9, wherein The first feature extraction network includes 3 second convolutional layers.

11. The method according to claim 7, wherein The image segmentation model further includes a second feature extraction network, and the method further includes: Performing feature extraction on the spliced feature map of the largest scale through the second feature extraction network to obtain a sixth feature map, where the scale of the sixth feature map is the same as the scale of the spliced feature map of the largest scale; Performing upsampling processing on the sixth feature map through the upsampling network to obtain a target feature map as the target segmentation result, where the scale of the target feature map is the same as the scale of the original image.

12. A model training device, characterized in that, including: An acquisition module configured to acquire training samples, where the training samples include sample images and sample segmentation results of the sample images; A training module configured to train an image segmentation model based on the training samples, where the image segmentation model includes N convolutional networks, each convolutional network includes multiple sub-networks, and multiple sub-networks of each convolutional network include M first convolutional layers, and both N and M are positive integers greater than 1; A merging module configured to merge the parameters of multiple sub-networks in the i-th convolutional network to obtain the i-th second convolutional layer, where i is a positive integer not greater than N; An update module configured to replace the i-th convolutional network with the i-th second convolutional layer to update the image segmentation model, and use the updated image segmentation model as the final image segmentation model.

13. An image segmentation device, characterized in that, including: An acquisition module configured to acquire an original image to be segmented; A processing module configured to input the original image into an image segmentation model, and output a target segmentation result of the original image by the image segmentation model, where the image segmentation model includes N second convolutional layers, N is a positive integer greater than 1, and the image segmentation model is obtained by using the model training method according to any one of claims 1-5.

14. An electronic device, characterized in that, Comprising: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to: Implement the steps of the method according to any one of claims 1-11.

15. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the program instructions are executed by the processor, the steps of the method according to any one of claims 1-11 are implemented.

16. A chip, comprising a processor, the processor being coupled to a memory, and when the processor executes computer instructions stored in the memory, implementing the method according to any one of claims 1-11.