Construction method and device of image segmentation decoder, equipment and medium
By combining sparse upsampling layers and dynamic pruning techniques with neural predictors and multi-objective evolutionary algorithms, a high-efficiency image segmentation decoder is constructed, solving the problems of computational redundancy and low parameter efficiency in traditional decoders, and achieving efficient image segmentation in the financial and medical fields.
Patent Information
- Application Number
- CN202510991666.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-11-07
AI Technical Summary
In existing technologies, traditional decoders suffer from computational redundancy and low parameter efficiency, making it difficult to perform image segmentation with efficient decoders. Furthermore, manually designing sparse structures makes it difficult to balance accuracy and efficiency.
Feature vectors are obtained by using sparse upsampling layers and dynamic pruning techniques. Combined with a pre-defined neural predictor and a multi-objective evolutionary algorithm, a target segmentation decoder is constructed. The sparse upsampling layer and dynamic pruning reduce model parameters and computational cost. The neural predictor is used to evaluate the performance of candidate decoding architectures, and the multi-objective evolutionary algorithm is used to select the optimal decoder.
It achieves efficient image segmentation in the financial and medical fields, reduces computation and the number of parameters, improves the operating efficiency of the decoding structure, and balances accuracy and efficiency. It is suitable for accurate segmentation of invoice images and medical images.
Smart Images

Figure CN120912690A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image segmentation, which can be applied to the fields of finance and insurance, medical health, etc., and in particular to a construction method and device of an image segmentation decoder, equipment and a medium. BACKGROUND
[0002] With the rapid development of technology, image segmentation technology plays a key role in many fields. In the medical field, accurate image segmentation helps disease diagnosis, surgery planning, etc.; in the financial field, image segmentation can also be used for bill identification, chart analysis, etc. In recent years, the image segmentation method based on deep learning has made significant progress, and the encoder-decoder architecture has become the mainstream. However, the traditional decoder uses fixed up-sampling operations, which has the problems of computational redundancy and low parameter efficiency. In the optimization of the decoder, the hand-designed sparse structure depends on human experience and is difficult to balance accuracy and efficiency, or the structure is designed automatically through neural architecture search, but the optimization of decoder sparsity requires high computational cost, which leads to the problem that the image segmentation cannot be performed through an efficient decoder in the prior art. SUMMARY
[0003] The embodiments of the present application provide a construction method, device, equipment and medium of an image segmentation decoder, aiming to solve the problem that the image segmentation cannot be performed through an efficient decoder in the prior art.
[0004] In a first aspect, the embodiments of the present application provide a construction method of an image segmentation decoder, which includes: extracting features of a collected training image through a preset encoder to obtain a feature map; performing dynamic pruning on the feature map through a preset sparse up-sampling layer and a preset convolution layer to obtain a feature vector; constructing a structure of the feature vector and a preset candidate decoding architecture through a preset neural predictor to obtain an intersection-over-union score of each candidate decoding architecture; and calculating an fitness according to the intersection-over-union score through a preset multi-objective evolutionary algorithm, and determining a target segmentation decoder in the candidate decoding architecture according to the calculated fitness score.
[0005] In a second aspect, the embodiments of the present application also provide a construction device of an image segmentation decoder, which includes: a convolution unit configured to extract features of a collected training image through a preset encoder to obtain a feature map; a pruning unit configured to perform dynamic pruning on the feature map through a preset sparse up-sampling layer and a preset convolution layer to obtain a feature vector; a construction unit configured to construct a structure of the feature vector and a preset candidate decoding architecture through a preset neural predictor to obtain an intersection-over-union score of each candidate decoding architecture; and a calculation unit configured to calculate an fitness according to the intersection-over-union score through a preset multi-objective evolutionary algorithm, and determine a target segmentation decoder in the candidate decoding architecture according to the calculated fitness score.
[0006] In a third aspect, an embodiment of the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above method when executing the computer program.
[0007] In a fourth aspect, an embodiment of the present application further provides a computer readable storage medium, wherein the storage medium stores a computer program, and the computer program comprises program instructions, and the program instructions can implement the above method when being executed by a processor.
[0008] An embodiment of the present application provides a construction method, device, equipment and medium of an image segmentation decoder. The method comprises: extracting features of a collected training image through a preset encoder to obtain a feature map; performing dynamic pruning on the feature map through a preset sparse up-sampling layer and a preset convolution layer to obtain a feature vector; constructing a structure of the feature vector and a preset candidate decoding architecture through a preset neural predictor to obtain an intersection-over-union score of each candidate decoding architecture; and calculating an adaptability according to the intersection-over-union score through a preset multi-objective evolutionary algorithm, and determining a target segmentation decoder from the candidate decoding architectures according to the calculated adaptability score. The training data is extracted for features through the encoder to obtain a feature map, so as to capture key information in the image and facilitate subsequent decoder construction. The process of obtaining the feature vector through sparse up-sampling and dynamic pruning can effectively reduce the number of parameters and the amount of calculation of the model and improve the operation efficiency of the decoding structure. The feature vector and the preset candidate decoding architecture are input into the preset neural predictor to obtain the intersection-over-union score, so as to intuitively compare the advantages and disadvantages of each candidate architecture and provide a basis for subsequent selection of the optimal decoder. Finally, the multi-objective evolutionary algorithm and the intersection-over-union score are used to screen the architecture with the optimal comprehensive performance from the numerous candidate decoding architectures as the target segmentation decoder, so as to balance the conflicts between different targets, obtain a decoder that performs well in multiple aspects, and thus efficiently perform image segmentation. Whether in the field of finance for fine segmentation of bill images or in the field of medicine for accurate segmentation of images, the target segmentation decoder can play an important role and help efficient development of related businesses. BRIEF DESCRIPTION OF DRAWINGS
[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0010] Figure 1 A flowchart of the construction method of the image segmentation decoder provided by an embodiment of the present application is shown in the figure.
[0011] Figure 2 A sub-flowchart diagram of a construction method of an image segmentation decoder provided by an embodiment of the present application is shown in FIG. 4;
[0012] Figure 3 A sub-flowchart diagram of a construction method of an image segmentation decoder provided by an embodiment of the present application is shown in FIG. 4;
[0013] Figure 4 A sub-flowchart diagram of a construction method of an image segmentation decoder provided by an embodiment of the present application is shown in FIG. 4;
[0014] Figure 5 A sub-flowchart diagram of a construction method of an image segmentation decoder provided by an embodiment of the present application is shown in FIG. 4;
[0015] Figure 6 A sub-flowchart diagram of a construction method of an image segmentation decoder provided by an embodiment of the present application is shown in FIG. 4;
[0016] Figure 7 A schematic block diagram of a construction device of an image segmentation decoder provided by an embodiment of the present application is shown in FIG. 6;
[0017] Figure 8 A schematic block diagram of a computer device provided by an embodiment of the present application is shown in FIG. 7. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, but not all embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0019] It should be understood that, when used in the specification and the appended claims, the terms “comprise” and “include” indicate the presence of described features, integers, steps, operations, elements, and / or components, but do not exclude one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0020] It should also be understood that the terms used herein in the specification and the appended claims are only for the purpose of describing particular embodiments and do not intend to limit the present application. As used in the specification and the appended claims of the present application, unless otherwise clearly indicated by the context, the singular forms “a”, “an” and “the” are intended to include the plural forms as well.
[0021] It should be further understood that the term "and / or" used in the description and claims of the application herein is used to mean any combination of one or more of the associated listed items, and all possible combinations, and includes these combinations.
[0022] Referring to Figure 1 , Figure 1 A flowchart of the construction method of the image segmentation decoder provided by the embodiment of the present application is shown in the figure. The construction method of the image segmentation decoder in the embodiment can be applied to many key fields, such as accurate segmentation analysis of bill images in the financial field and accurate segmentation of medical image lesion areas in the medical field. By using the method, the sparse perception evolution search optimization structure is used, the neural predictor is used to accelerate the evaluation, the multi-objective function is used for joint optimization, the calculation amount is reduced while the accuracy is guaranteed, and thus a decoder that performs well in multiple aspects is obtained, so that the image segmentation can be efficiently performed.
[0023] Figure 1 A flowchart of the construction method of the image segmentation decoder provided by the embodiment of the present application is shown in the figure. As shown in the figure, the method includes the following steps S110-S140.
[0024] S110, the collected training image is subjected to feature extraction by a preset encoder to obtain a feature map.
[0025] In the embodiment, the training images are representative and can cover various situations and changes in the target application scenarios. For example, in medical image segmentation, if lung nodules are to be segmented, the training images collected should include nodules of different sizes, shapes, and positions, as well as lung images of different patients. If the text and numbers on the bill are to be segmented in the financial field, the bill images collected should include bills of various fonts, font sizes, layout methods, and different backgrounds. The encoder is a network structure for extracting features from input images. It should be noted that the preset encoder and the subsequently constructed decoder are applied to an innovative sparse neural architecture search framework (SparseNAS) in the embodiment, and the preset encoder extracts multi-level features using a standard convolutional network (such as ResNet). The training images collected are subjected to feature extraction by the preset encoder to obtain feature maps. Specifically, the training images are first subjected to a convolution operation of the first layer, a plurality of small convolution kernels (such as 7x7) are used to convolve the images, and low-level features such as edges and textures of the images are extracted to obtain preliminary feature maps. Then, the feature maps are subjected to down-sampling by a max-pooling layer to reduce the dimension. With the increase of the number of network layers, the subsequent convolution layers extract higher-level features, such as the shape and boundary of a tumor. After the multi-layer convolution and pooling operations, a set of feature maps with rich semantic information are finally obtained, which will be used for subsequent decoder to segment the tumor region or to segment the bill in the financial field. The feature maps obtained according to the collected images and the preset encoder provide a data basis for the construction of the subsequent decoder.
[0026] S120, dynamically pruning the feature maps through a preset sparse up-sampling layer and a preset convolution layer to obtain a feature vector.
[0027] In the embodiment, the sparse up-sampling layer is a network layer that performs dynamic pruning operation on the feature map output by the encoder while expanding the spatial dimension. The preset convolution layer is one of the core components of the convolutional neural network. In the convolution layer, the convolution kernel calculates the channels of the input feature map. The feature vector is obtained by performing dynamic pruning on the feature map through the preset sparse up-sampling layer and the preset convolution layer. Specifically, in the medical field, the learnable interpolation kernel in the sparse up-sampling layer can be adjusted according to the feature map of the brain tumor image, so that the feature map after up-sampling can more accurately restore the shape and position of the tumor region. At the same time, the dynamic channel gating mechanism evaluates the channels of the feature map and removes some channel information that is irrelevant to tumor segmentation, such as some feature channels in normal brain tissue that do not contribute to tumor segmentation. Then, the up-sampled feature map enters the preset convolution layer. The dynamic grouping convolution dynamically allocates convolution kernel groups according to the distribution of the current feature map, and only uses the convolution kernel groups related to tumor features to perform calculation. After a series of dynamic pruning operations, a feature vector is finally obtained, which contains key feature information of the tumor in the brain MRI image and provides accurate feature representation for subsequent tumor segmentation. After the dynamic pruning processing of the preset sparse up-sampling layer and the preset convolution layer, the feature map is converted into a feature vector, which contains key feature information after sparse processing, has lower dimension and higher semantic information density, and better represents important features of the input image, providing strong support for subsequent image segmentation tasks.
[0028] In an embodiment, as shown in FIG. 1, Figure 2 The step S120 includes steps S121-S122.
[0029] S121, performing interpolation operation on the feature map according to a preset activation function and a preset interpolation kernel function of the preset sparse up-sampling layer, to obtain a primary feature vector;
[0030] S122, pruning the feature channels corresponding to the primary feature vector through a preset convolution layer and inter-group sparsity constraint, to obtain the feature vector.
[0031] In the embodiment, the preset interpolation kernel function is used for up-sampling operation of the feature map in the spatial dimension. Unlike the traditional fixed interpolation kernel (such as nearest neighbor interpolation and bilinear interpolation), the interpolation kernel here is learnable, and its parameters will be automatically adjusted according to the task requirements during model training to better recover the spatial information of the feature map. The activation function is used to introduce a nonlinear factor, so that the model can learn more complex feature representations. In the sparse up-sampling layer, the activation function can perform nonlinear transformation on the interpolated feature map to enhance the expression ability of the model. The feature map is interpolated according to the preset activation function and the preset interpolation kernel function of the preset sparse up-sampling layer to obtain a primary feature vector, and the specific formula is:
[0032] U(x) = Interp(x) 0 σ(W g x + b g )
[0033] where x R H×W×C , x is the feature map, containing spatial dimension height and width H x W and channel number C, R is a real number, Interp(·) is a preset interpolation kernel function, which is a learnable interpolation kernel, supporting nearest neighbor / bilinear and other modes, W g R C ×C , b g R, W g is a gating weight, b g is a bias term, and σ(·): Sigmoid is the preset activation function, used to output the channel importance weight g i (0, 1), it should be noted that when g i <0.05 (adjustable threshold), the channel is forced to zero in back propagation, realizing gradient-aware sparsification, and U(x) is the primary feature vector. The inter-group sparsity constraint is a way to group the feature channels of the convolutional layer by introducing a regularization term or a dynamic mask, and encourage the feature channels in each group to have similar importance, while allowing differences between different groups. The primary feature vector is pruned by the preset convolutional layer and the inter-group sparsity constraint to obtain the feature vector, and the specific formula is:
[0034]
[0035] where y is the feature vector, * represents convolution operation, and 0 represents element-wise multiplication. U(x) is the primary feature vector, K is the total number of groups, which is an evolutionary variable and may change during model training. It is usually a Gaussian random value during initialization, and is a variable that needs to be updated during model training. It determines how many groups the convolution operation is divided into, and different groups can have different convolution kernels and masks.k This is the k-th convolutional kernel group. Each convolutional kernel group contains multiple convolutional kernels, where mk is a binary mask, approximately differentiable using Gumbel-Softmax. Its elements can only take values of 0 or 1, and are used to filter the features after convolution. When m... k When an element is 0, the corresponding convolution result is set to zero during summation; when it is 1, the convolution result is retained. This method enables dynamic feature selection and sparsity constraints. The inter-group sparsity constraint is... Where, ||m k ||0 represents m k The zero norm, i.e., m k The number of non-zero elements in the mask. This constraint requires that the mask m of all groups... k The total number of non-zero elements in China does not exceed Where C is the number of feature map channels. This indicates rounding down. The purpose of this constraint is to encourage the model to maintain a certain degree of sparsity during grouped convolution, reduce unnecessary computation and parameters, and improve the model's efficiency and generalization ability. The overall process is as follows: First, the input feature map x is processed by a learnable interpolation kernel and combined with gating weights and biases. Channel importance weights are generated through an activation function. When the importance weight < 0.05, the corresponding channel is set to zero during backpropagation to achieve sparsity and obtain a new feature vector. Then, a convolution variant with dynamic grouping masks is used. After convolving the feature map with K groups of convolution kernels, it is multiplied element-wise with a binary mask. Then, the feature channels are pruned according to the sparsity constraint between groups, so that different feature channels are dynamically assigned to active groups, and ineffective groups are eliminated in the evolution. This achieves sparse feature representation while maintaining the flexibility of the model, reducing the amount of computation and the number of parameters, thereby improving the performance and efficiency of the decoder model.
[0036] S130. The feature vector and the preset candidate decoding architecture are constructed using a preset neural predictor to obtain the cross-union score of each candidate decoding architecture.
[0037] In the embodiment, the feature vector is a vector with lower dimension and higher semantic information density, and the preset candidate decoding architecture is a plurality of pre-generated decoder architectures with different structures. These architectures differ in the number of layers, connection methods, and used operations (such as convolution, deconvolution, skip connection, etc.). The preset neural predictor is a neural network model that can predict the performance index, specifically the intersection over union (IoU) score, that can be achieved by decoding using the candidate decoding architecture according to the input feature vector and the information of the candidate decoding architecture. The IoU score is an index for measuring the degree of overlap between two regions and is used to evaluate the similarity between the predicted result and the true result in tasks such as target detection and image segmentation. The feature vector and the preset candidate decoding architecture are structured by the preset neural predictor to obtain the IoU score of each candidate decoding architecture. Specifically, the feature vector and each preset candidate decoding architecture are provided as input to the preset neural predictor, which performs internal calculation and processing after receiving the input. It first integrates and analyzes the input feature vector and candidate decoding architecture information, and performs feature transformation and combination through various layers in the neural network (such as fully connected layers, convolution layers, etc., depending on the design of the neural predictor). In this process, the neural predictor learns the mapping relationship between the feature vector and different candidate decoding architectures, and then predicts the IoU score between the final output result and the true result if the input feature is decoded using the candidate decoding architecture, where the IoU score reflects the performance level that each candidate decoding architecture can achieve when processing the current input feature. The IoU score of each candidate decoding architecture is obtained by structuring the preset neural predictor to understand the performance level of different candidate decoding architectures and facilitate the selection of the best decoder architecture.
[0038] In an embodiment, as shown in FIG. 13, the step S130 further includes steps S131-S132. Figure 3
[0039] S131, according to the feature vector and the node number and operation type in the preset candidate decoding architecture, the graph structure corresponding to each preset candidate decoding architecture is obtained by structuring the preset neural predictor.
[0040] S132, the graph structure is predicted by the preset model in the preset neural predictor and the neural network to obtain the corresponding IoU score.
[0041] In the embodiment, the preset candidate decoding architecture is a network structure composed of different nodes and operations. The number of nodes determines the size and complexity of the architecture, and the operation type includes convolution, deconvolution, pooling, jump connection and various operations that may be used in the decoding process. The graph structure reflects the data flow and calculation method of different architectures when processing feature vectors. The specific process of structure construction by the preset neural predictor according to the feature vector and the number of nodes and operation types in the preset candidate decoding architecture is as follows:
[0042] a = [AdjMat || OpEmb] e R N×(N+d)
[0043] Wherein, a is the graph structure, AdjMat is an N x N adjacency matrix (N is the number of nodes), OpEmb is a d-dimensional operation type embedding (up-sampling / convolution operation), and || is a concatenation operation. The subsequent operation type is determined according to the process of obtaining the feature vector, such as the specific number of convolution channels. The graph structure is predicted by the preset model in the preset neural predictor and the neural network to obtain the corresponding IoU score. Specifically, the neural predictor inputs the graph structure into the transformer model to obtain the output, which is then input into an MLP (Multilayer Perceptron) layer neural network to obtain the final output, which is the IoU score. By outputting an IoU score for each graph structure (i.e. each preset candidate decoding architecture), the similarity between the predicted result and the true result when using the corresponding candidate decoding architecture for image segmentation and other tasks is reflected, so as to select a better decoder structure.
[0044] In an embodiment, the step S132 further includes a step S1321 before the step S132.
[0045] S1321, loss optimization of the initial neural predictor by a preset loss function and a preset sample data to obtain the preset neural predictor.
[0046] In the embodiment, the preset loss function is a function used to measure the difference between the predicted result of the neural predictor and the true value. The preset sample data is data used to train and optimize the neural predictor. The initial neural predictor is a neural network model that has not been fully trained. It has a certain structure and parameters, but these parameters are randomly initialized or set by a simple method, and it cannot accurately complete the prediction task. The initial neural predictor is loss optimized by the preset loss function and the preset sample data to obtain the preset neural predictor. Specifically, the Huber loss function is used for loss optimization, and the optimization process is as follows:
[0047]
[0048] wherein, is the IoU score predicted by the neural predictor according to the sample data, y i is the real IoU score in the training sample, δ = 0.1 is the robust threshold, B represents the training batch, and i is the i-th sample data in the current batch. The predictor can achieve an IoU prediction error of ±0.03 after pre-training on several random architectures. The neural predictor is trained and optimized according to the loss function, so that it can accurately predict the IoU score according to the input feature vector and candidate decoding architecture information.
[0049] S140, fitness calculation is performed on the candidate decoding architecture according to the IoU score by using a preset multi-objective evolutionary algorithm, and a target segmentation decoder is determined in the candidate decoding architecture according to the calculated fitness score.
[0050] In this embodiment, the preset multi-objective evolutionary algorithm is an optimization algorithm simulating the biological evolution process, which is used to find the optimal solution among multiple objectives (such as segmentation accuracy, computational efficiency, etc.). The fitness score is used to reflect the numerical value of the comprehensive performance of the candidate architecture on multiple objectives. The target segmentation decoder is selected from the candidate decoding architecture, which performs best in fitness calculation and can meet the decoding requirements of specific tasks. It can be used in the final image segmentation task in medical images (such as CT, MRI images) or financial bill images. According to the IoU score, fitness calculation is performed on the candidate decoding architecture by using a preset multi-objective evolutionary algorithm, and a target segmentation decoder is determined in the candidate decoding architecture according to the calculated fitness score. Specifically, for the MLP+Transformer decoder in this embodiment, the fitness score of each candidate decoder needs to be calculated according to the IoU score by using the fitness function, and the calculation results are sorted to obtain the best decoder architecture with the highest score, which is used as the target segmentation decoder. By combining the multi-objective evolutionary algorithm and the IoU score, the target segmentation decoder best suited for specific tasks can be selected from the candidate decoding architecture, meeting the needs of different fields (such as finance and medicine) for image segmentation accuracy and efficiency.
[0051] In an embodiment, as shown in Figure 4 the step S140 includes steps S141-S143.
[0052] S141, the number of floating point operations and the sparsity of the candidate decoding architecture are calculated;
[0053] S142, the corresponding fitness score is obtained by calculating the number of floating point operations, the sparsity and the IoU score according to a preset fitness formula;
[0054] S143. The candidate decoding architecture with the highest fitness score is determined as the target segmentation decoder.
[0055] In this embodiment, the number of floating-point operations (FLOPs) is a metric for measuring the computational complexity of the model, representing the number of floating-point operations performed by the model during forward inference. A higher number of FLOPs indicates a greater computational load and higher demand for computing resources. Sparsity reflects the proportion of zero parameters or zeroed features in the model. Higher sparsity indicates that more parameters or features are removed from the model, helping to reduce storage and computational overhead. The number of FLOPs and sparsity of the candidate decoding architecture are calculated specifically by analyzing the operations of each layer (such as convolutional layers, fully connected layers, etc.) in the candidate decoding architecture. For example, for a convolutional layer, its number of FLOPs is approximately the product of the height, width, number of input channels, number of output channels, kernel height, and kernel width of the input feature map. Adding the number of FLOPs of all layers in the architecture yields the total number of FLOPs for the entire candidate decoding architecture. The sparsity calculation result is obtained by dividing the number of parameters or features set to zero in the statistical model by the total number of parameters or features. For example, in sparsely constrained convolution, some feature channels are set to zero by inter-group sparsity constraints. The number of zeroed channels divided by the total number of channels is the sparsity of that layer. The sparsity of each layer is appropriately weighted or averaged to obtain the sparsity of the entire architecture. The corresponding fitness score is calculated based on the number of floating-point operations, the sparsity, and the intersection-union ratio (IU / I) score using a preset fitness formula. Specifically, the preset fitness formula is:
[0056] F(a)=IoU(a)-λ1FLOPs(a)+λ2Sparsity(a)
[0057] Wherein, IoU is the Intersection over Union (IoU) score output by the aforementioned neural predictor, FLOPs is the number of floating-point operations, Sparsity is the sparsity, weight coefficients λ1 = 0.8, λ2 = 0.5 (configurable), a is the graph structure of the candidate decoding architecture, and F(a) is the fitness score of the candidate decoding architecture. The candidate decoding architecture with the highest fitness score is determined as the target segmentation decoder. Specifically, from all candidate decoding architectures, the architecture with the highest fitness score is selected as the target segmentation decoder. By comprehensively considering the number of floating-point operations, sparsity, and IoU score, the target segmentation decoder most suitable for a specific task can be selected from the candidate decoding architectures.
[0058] In one embodiment, such as Figure 5 As shown, step S143 further includes steps S1431-S1433.
[0059] S1431、according to the fitness score, ranking the candidate decoding architectures in descending order, performing evolutionary operation on the candidate decoding architectures beyond a preset ranking range, and obtaining new candidate decoding architectures;
[0060] S1432、recalculating the fitness of the new candidate decoding architectures, repeating the steps of ranking the candidate decoding architectures in descending order according to the fitness score, performing evolutionary operation on the candidate decoding architectures beyond a preset ranking range, and obtaining new candidate decoding architectures until a preset termination condition is reached;
[0061] S1433、determining the target segmentation decoder according to the result of the last ranking in descending order.
[0062] In the embodiment, the preset ranking range is a preset ranking range, for example, the top 20% of the architecture, the architectures within the range are reserved, and the architectures outside the range will be evolved. The evolution operation is an operation of modifying the candidate decoding architecture to generate a new architecture, thereby increasing the diversity of the architecture. The preset termination condition is a preset condition, for example, a condition of reaching a maximum number of iterations, a condition of fitness score convergence (i.e., a small change in fitness score in multiple iterations), and the like, for judging whether the evolution process is ended. The candidate decoding architectures are sorted in descending order according to the fitness scores, the candidate decoding architectures outside the preset ranking range are evolved to obtain new candidate decoding architectures. Specifically, assuming that there are 10 candidate decoding architectures initially, they are sorted from high to low according to the fitness scores calculated before. The preset ranking range is the top 3, that is, architecture 1, architecture 2, and architecture 3 are reserved, and architectures 4-10 are evolved. The evolution operation can be modifying some layers of the architecture, such as changing the convolution kernel size of a convolution layer, increasing or decreasing the number of channels, and the like. For example, for architecture 4, the convolution kernel size of a convolution layer thereof is changed from 3x3 to 5x5 to obtain a new candidate decoding architecture 4'. The fitness scores of the newly generated candidate decoding architectures (such as architecture 4' and the like) are recalculated, all candidate decoding architectures (including the reserved and newly generated) are sorted in descending order according to the fitness scores, and then the architectures outside the preset ranking range are evolved again. This process is repeatedly performed, and new candidate decoding architectures can be generated in each iteration, and their fitness is reevaluated. Assuming that the preset termination condition is reaching 10 iterations or the average change of the fitness scores being less than 0.01. During the iteration process, it is continuously checked whether the termination condition is met. When the preset termination condition is reached, a final sorting in descending order is performed. The architecture with the highest fitness score is selected as the target segmentation decoder. For example, after multiple iterations, the fitness score of architecture 1 is still the highest, and then architecture 1 is determined as the target segmentation decoder for the segmentation of financial bill images or medical images. By continuously optimizing the candidate decoding architectures by using the evolution algorithm, the target segmentation decoder with the best performance is finally selected to meet the demand of the financial and medical fields for image segmentation
[0063] In an embodiment, as shown in FIG. 14B, the step S1431 further includes steps S14311-S14312. Figure 6
[0064] S14311, topological variation, sparse variation, and crossover recombination variation are performed on the candidate decoding architecture;
[0065] S14312, determining the new candidate decoding architecture according to the variation result.
[0066] In this embodiment, the topology variation is a variation that changes the topology structure of the candidate decoding architecture, such as increasing or decreasing the number of network layers, changing the connection mode between layers, etc. The sparsity variation is to adjust the sparsity related parameters in the candidate decoding architecture, such as changing the threshold of the inter-group sparsity constraint, adjusting the generation mode of the binary mask, etc., thereby affecting the proportion of zero parameters or zeroed features in the model. The crossover recombination variation is to exchange and combine the partial structures of two or more candidate decoding architectures, thereby generating a new architecture variation. For example, according to the above examples, the evolution operation is performed on architecture 4-architecture 10, then the topology variation is performed on architecture 4, such as replacing one of the deconvolution layers with a dilated convolution layer to expand the receptive field and obtain more context information, adjusting the sparsity generation mode of architecture 5, such as changing the generation probability of the binary mask, and selecting architecture 6 and architecture 7 for crossover recombination. For example, the feature extraction module of architecture 6 has a good effect on extracting lesion features in medical images, and the decoding module of architecture 7 can better restore image details. After crossover recombination, the new architecture combines the advantages of the two, which is expected to improve the accuracy of medical image segmentation. The specific variation architecture and mode are not limited, and the evolution operation can be performed. By integrating the new architecture generated by the variation and the reserved architecture, a set of new candidate decoding architectures is formed, which is used for subsequent evaluation and selection to find a decoding architecture more suitable for medical image segmentation or the financial field.
[0067] Figure 7 is a schematic block diagram of an image segmentation decoder construction device 200 provided by an embodiment of the present application. As shown in Figure 7 corresponding to the above image segmentation decoder construction method, the present application further provides an image segmentation decoder construction device. The image segmentation decoder construction device includes a unit for executing the above image segmentation decoder construction method, and the device can be configured in a terminal such as a desktop computer, a tablet computer, a laptop computer, etc. Specifically, referring to Figure 7 , the image segmentation decoder construction device includes a convolution unit 210, a pruning unit 220, a construction unit 230, and a calculation unit 240.
[0068] The convolution unit 210 is configured to perform feature extraction on the collected training images through a preset encoder to obtain feature maps.
[0069] The pruning unit 220 is configured to perform dynamic pruning on the feature maps through a preset sparse up-sampling layer and a preset convolution layer to obtain feature vectors.
[0070] In an embodiment, the pruning unit 220 includes an interpolation unit and a constraint unit.
[0071] an interpolation unit, configured to perform interpolation operation on the feature map according to a preset activation function and a preset interpolation kernel function of the preset sparse up-sampling layer, to obtain a primary feature vector;
[0072] a constraint unit, configured to prune feature channels corresponding to the primary feature vector through a preset convolution layer and an inter-group sparsity constraint, to obtain the feature vector.
[0073] a construction unit 230, configured to perform structure construction on the feature vector and a preset candidate decoding architecture through a preset neural predictor, to obtain an intersection-over-union score of each candidate decoding architecture.
[0074] In an embodiment, the construction unit 230 includes a construction subunit and a prediction unit.
[0075] the construction subunit, configured to perform structure construction on the feature vector and a node number and an operation type in the preset candidate decoding architecture through the preset neural predictor, to obtain a graph structure corresponding to each preset candidate decoding architecture;
[0076] the prediction unit, configured to perform prediction on the graph structure through a preset model and a neural network in the preset neural predictor, to obtain the corresponding intersection-over-union score.
[0077] In an embodiment, the construction unit 230 includes an optimization unit.
[0078] the optimization unit, configured to perform loss optimization on an initial neural predictor through a preset loss function and preset sample data, to obtain the preset neural predictor.
[0079] a calculation unit 240, configured to perform fitness calculation on the intersection-over-union score through a preset multi-objective evolutionary algorithm, and determine a target segmentation decoder from the candidate decoding architectures according to the calculated fitness score. In an embodiment, the calculation unit 240 includes a first calculation unit, a second calculation unit, and a determination unit.
[0080] the first calculation unit, configured to calculate a floating-point operation number and a sparsity of the candidate decoding architecture;
[0081] the second calculation unit, configured to perform calculation through a preset fitness formula calculation according to the floating-point operation number, the sparsity, and the intersection-over-union score, to obtain a corresponding fitness score;
[0082] the determination unit, configured to determine the candidate decoding architecture with the highest fitness score as the target segmentation decoder.
[0083] It should be noted that the specific implementation process of the construction device 200 and each unit of the image segmentation decoder can be clearly understood by those skilled in the art, and can be referred to the corresponding description in the foregoing method embodiment. For the convenience and brevity of description, it will not be repeated here.
[0084] The construction device of the image segmentation decoder can be implemented in the form of a computer program, which can run on a computer device as shown in the figure. Figure 8 .
[0085] Please refer to Figure 8 , Figure 8 is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device 500 can be a terminal or a server, wherein the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, a wearable device, and the like electronic device with communication function. The server can be a stand-alone server or a server cluster composed of multiple servers.
[0086] Please refer to Figure 8 , the computer device 500 includes a processor 502, a memory and a network interface 505 connected through a system bus 501, wherein the memory can include a non-volatile storage medium 503 and an internal memory 504.
[0087] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. The computer program 5032 includes program instructions, which when executed, can cause the processor 502 to perform a composite laser printing method.
[0088] The processor 502 is configured to provide computing and control capabilities to support the operation of the entire computer device 500.
[0089] The internal memory 504 provides an environment for the running of the computer program 5032 in the non-volatile storage medium 503, which when executed by the processor 502, can cause the processor 502 to perform a construction method of an image segmentation decoder.
[0090] The network interface 505 is configured to perform network communication with other devices. Those skilled in the art can understand that Figure 8 the structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device 500 to which the scheme of the present application is applied. The specific computer device 500 can include more or less components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0091] The processor 502 is configured to run the computer program 5032 stored in the memory to implement the steps of the above method.
[0092] It should be understood that, in the embodiments of the present application, the processor 502 can be a central processing unit (CPU), and the processor 502 can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0093] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned method embodiments can be completed by a computer program instructing related hardware. The computer program includes program instructions, and the computer program can be stored in a storage medium, which is a computer-readable storage medium. The program instructions are executed by at least one processor in the computer system to implement the process steps of the above-mentioned method embodiments.
[0094] Therefore, the present application also provides a storage medium. The storage medium can be a computer-readable storage medium. The storage medium stores a computer program, wherein the computer program includes program instructions. The program instructions are executed by the processor to make the processor perform the steps of the above method.
[0095] The storage medium can be a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk, etc. Various computer-readable storage media that can store program codes.
[0096] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been described in general terms in the above description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0097] In several embodiments provided by the present application, it should be understood that the disclosed apparatus and method can be implemented in other manners. For example, the division of the apparatus embodiments is merely illustrative. For example, the division of the units can be changed, and some units can be combined or integrated into another system, or some features can be ignored or not executed. The above described embodiments are merely specific implementation, and thus, are not intended to limit the protection scope of the present application.
[0098] The steps in the method embodiments of the present application can be adjusted, combined and deleted according to actual needs. The units in the apparatus embodiments of the present application can be combined, divided and deleted according to actual needs. In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit.
[0099] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a storage medium. Based on such understanding, the technical solutions of the present application essentially or the part of the prior art that makes a contribution, or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application.
[0100] The above describes merely specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method of constructing an image segmentation decoder, characterized by, The method comprises the steps of: extracting features of the collected training images through a preset encoder to obtain feature maps; performing dynamic pruning on the feature maps through a preset sparse up-sampling layer and a preset convolution layer to obtain feature vectors; constructing a structure of the feature vectors and a preset candidate decoding architecture through a preset neural predictor to obtain an intersection-over-union score of each candidate decoding architecture; calculating the fitness of the intersection-over-union score through a preset multi-objective evolutionary algorithm, and determining a target segmentation decoder from the candidate decoding architectures according to the calculated fitness score.
2. The method of claim 1, wherein, The step of performing dynamic pruning on the feature maps through a preset sparse up-sampling layer and a preset convolution layer to obtain feature vectors comprises the steps of: performing interpolation operation on the feature maps according to a preset activation function and a preset interpolation kernel function of the preset sparse up-sampling layer to obtain primary feature vectors; pruning feature channels corresponding to the primary feature vectors through a preset convolution layer and inter-group sparsity constraint to obtain the feature vectors.
3. The method of claim 1, wherein, The step of constructing a structure of the feature vectors and a preset candidate decoding architecture through a preset neural predictor to obtain an intersection-over-union score of each candidate decoding architecture comprises the steps of: constructing a graph structure corresponding to each preset candidate decoding architecture through the preset neural predictor according to the feature vectors, the number of nodes and the operation type in the preset candidate decoding architecture; predicting the graph structure through a preset model in the preset neural predictor and a neural network to obtain the corresponding intersection-over-union score.
4. The method of claim 3, wherein, Before the step of obtaining the corresponding intersection-over-union score, the method further comprises the steps of: optimizing an initial neural predictor through a preset loss function and preset sample data to obtain the preset neural predictor.
5. The method of claim 1, wherein, The step of calculating the fitness of the intersection-over-union score through a preset multi-objective evolutionary algorithm, and determining a target segmentation decoder from the candidate decoding architectures according to the calculated fitness score comprises the steps of: calculating the number of floating-point operations and the sparsity of the candidate decoding architectures; calculating the corresponding fitness score through a preset fitness formula according to the number of floating-point operations, the sparsity and the intersection-over-union score; determining the candidate decoding architecture with the highest fitness score as the target segmentation decoder.
6. The method of claim 5, wherein, The step of determining the candidate decoding architecture with the highest fitness score as the target segmentation decoder further comprises the steps of: sorting the candidate decoding architectures in descending order according to the fitness scores, performing evolutionary operation on the candidate decoding architectures outside a preset sorting range to obtain new candidate decoding architectures; repeating the steps of sorting the candidate decoding architectures in descending order according to the fitness scores, performing evolutionary operation on the candidate decoding architectures outside a preset sorting range to obtain new candidate decoding architectures until a preset termination condition is reached; determining the target segmentation decoder according to the result of the last descending sorting.
7. The method of claim 6, wherein, The step of performing evolutionary operation on the candidate decoding architectures outside a preset sorting range to obtain new candidate decoding architectures comprises the steps of: topological variation, sparse variation and crossover recombination variation are performed on the candidate decoding architecture; a new candidate decoding architecture is determined according to the variation result.
8. An apparatus for constructing an image segmentation decoder, characterized in that comprise: a convolution unit configured to perform feature extraction on the collected training image through a preset encoder to obtain a feature map; a pruning unit configured to perform dynamic pruning on the feature map through a preset sparse up-sampling layer and a preset convolution layer to obtain a feature vector; a construction unit configured to perform structure construction on the feature vector and a preset candidate decoding architecture through a preset neural predictor to obtain an intersection-over-union score of each candidate decoding architecture; a calculation unit configured to perform fitness calculation on the intersection-over-union score through a preset multi-objective evolutionary algorithm, and determine a target segmentation decoder from the candidate decoding architectures according to the calculated fitness score.
9. A computer device, comprising: The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method of any one of claims 1-7 when executing the computer program.
10. A storage medium, characterized by The storage medium stores a computer program, the computer program comprises program instructions, and the program instructions can implement the method of any one of claims 1-7 when executed by a processor.