An adaptive neural architecture search method for multi-domain image segmentation tasks

By building a flexible and modular search space and combining particle swarm optimization algorithms in neural architecture search, the problem of difficulty in designing a neural network architecture suitable for multiple image segmentation tasks in the existing technology is solved, and efficient and flexible neural network architecture search and optimization is achieved, which significantly improves the accuracy and efficiency of image segmentation tasks.

CN119204084BActive Publication Date: 2025-05-06NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411748733.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-02
Publication Date
2025-05-06
Estimated Expiration
2044-12-02

AI Technical Summary

Technical Problem

The prior art is difficult to design an efficient neural network architecture suitable for multiple image segmentation tasks, and the traditional NAS method consumes huge computing resources, takes a long time to search process, lacks flexibility, and cannot adaptively adjust the network architecture between multiple different tasks.

Method used

By building a flexible and modular search space, combining particle swarm optimization algorithm (PSO), reward and punishment mechanism and multi-objective optimization strategy, an adaptive neural architecture search method is designed to achieve efficient model search and adaptive optimization for multiple image segmentation tasks.

Benefits of technology

It improves the flexibility and applicability of neural network architecture design, realizes the automatic search of the optimal network architecture in image segmentation tasks in different fields, significantly improving the accuracy and efficiency of segmentation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119204084B_ABST
    Figure CN119204084B_ABST
Patent Text Reader

Abstract

The present invention discloses a neural architecture search method for adaptive multi-domain image segmentation tasks. The method realizes efficient processing of tasks such as semantic segmentation, instance segmentation, panoptic segmentation, medical image segmentation, and edge detection by constructing a modular neural network architecture, combining adaptive neural architecture search and an improved particle swarm optimization algorithm. By optimizing the network structure and maintaining high spatial resolution, the present invention significantly reduces the computational complexity while improving segmentation accuracy, has broad application prospects, and is particularly suitable for complex image processing scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and deep learning, and in particular to a neural architecture search method for adaptive multi-domain image segmentation tasks. Background Art

[0002] With the rapid development of artificial intelligence and deep learning technology, neural networks have achieved remarkable results in many fields such as image processing, natural language processing, and speech recognition. Among them, image segmentation, as one of the key tasks of computer vision, has played an important role in applications such as autonomous driving, medical image analysis, satellite image processing, and augmented reality. The goal of the image segmentation task is to decompose the image into several meaningful regions or objects for further image analysis and understanding. However, designing an efficient neural network architecture suitable for a variety of image segmentation tasks is a complex and time-consuming challenge. Traditional neural network design methods rely on expert knowledge and a large number of experimental debugging processes, and often require repeated adjustments for different tasks and application scenarios. This not only consumes a lot of time and resources, but also makes it difficult to ensure the generalization ability of the designed network architecture in different tasks. With the increasing complexity of data sets and task requirements, the method of manually designing neural networks has gradually failed to meet actual needs. As a method for automatically designing neural network architectures, neural architecture search automatically searches for the optimal network architecture through algorithms, greatly reducing the burden of manual design. Traditional NAS methods, such as search strategies based on reinforcement learning and evolutionary algorithms, have shown great potential in tasks such as image classification. However, these methods usually have problems such as huge consumption of computing resources, long search process, lack of flexibility, etc., and are difficult to adapt to image segmentation tasks in different fields. More importantly, most of the existing NAS methods are optimized for specific tasks, lack universality, and cannot adaptively adjust the network architecture between a variety of different tasks. Summary of the invention

[0003] In view of the above-mentioned problems existing in the prior art, the present application is proposed. This method automatically designs and optimizes the neural network architecture suitable for a variety of image segmentation tasks by constructing a flexible network structure and a wide search space, thereby realizing efficient model search and adaptive optimization in a multi-task environment.

[0004] This paper proposes an adaptive neural architecture search method, which enables the network architecture to be adaptively adjusted according to the needs of different tasks by constructing a flexible and modular search space. Based on the design of a universal model, this method combines the particle swarm optimization algorithm (PSO), reward and punishment mechanism and multi-objective optimization strategy, which not only improves the search efficiency, but also enhances the flexibility and applicability of network architecture design. Finally, this method can automatically search for the optimal network architecture in image segmentation tasks in different fields, providing new ideas and tools for solving multi-task image segmentation problems.

[0005] The present invention relates to a neural architecture search method for adaptive multi-domain image segmentation tasks, which aims to efficiently process a variety of image segmentation tasks by combining modular design with an adaptive neural architecture search algorithm (NAS), including:

[0006] S1. Construct a unified description and generalized requirement model for multiple image segmentation tasks. Construct a general image segmentation task model to uniformly describe multiple image segmentation tasks such as semantic segmentation, instance segmentation, panoptic segmentation, medical image segmentation, and edge detection. By abstracting the requirements of these tasks, a unified input image format and output mask format are established, and the model is parameterized to adapt to different types of segmentation task requirements. This model can provide standardized task input, laying the foundation for subsequent network architecture design and optimization;

[0007] S2. Modular design in the adaptive neural architecture search algorithm. The entire neural network architecture is modularized through the adaptive neural architecture search algorithm. According to the unified demand model, the network is divided into multiple functional modules, and different functions are assigned to corresponding modules, so that each module can efficiently handle specific tasks. By analyzing the connection requirements between modules, the best connection method is determined, and the optimized network structure is constructed by adjusting the module connection to meet different task requirements. This modular design improves the flexibility and adaptability of the network and can better meet different image segmentation tasks;

[0008] S3. Design of cell-level architecture. A cell-level architecture is designed by representing neural network cells as directed acyclic graphs (DAGs) and building multiple blocks on this basis, each of which performs a specific neural network operation. The outputs of the blocks are combined through element addition operations to form complete cells, and the reuse and optimization of modules are achieved through a shared parameter mechanism. Through the combination and update of cells, the extraction of features at different levels is achieved, and efficient computing power is maintained in complex image segmentation tasks, achieving high spatial resolution maintenance and dynamic adaptation, as follows:

[0009] S31. Build a cell-level architecture to define the internal structure of cells, cell representation, cell update methods, and high information retention methods.

[0010] S32. Cell level architecture update,

[0011] S33. Design of high spatial resolution preservation method;

[0012] S4. Use the improved particle swarm optimization algorithm for neural architecture search. The improved particle swarm optimization algorithm (PSO) is used to perform global and local optimization of the network architecture. By initializing the position and velocity of the particles, the distribution of the particle swarm in the search space is established. The velocity and position of the particles are adjusted according to the inertia weight and acceleration coefficient, and the historical best position of each particle is optimized through fitness calculation. After the global search, the architecture parameters are further adjusted through the local optimization strategy, and finally the network architecture with the best performance on the validation set is selected, so as to achieve accurate processing of complex tasks. The specific steps are as follows,

[0013] S41. The initialization phase includes the initialization of particle positions and velocities.

[0014] S42. Global search and coarse adjustment stage,

[0015] S43. Local optimization and fine-tuning stage.

[0016] Furthermore, the present invention will be described in more detail, wherein the specific steps of constructing a unified description of multi-image segmentation tasks and a generalized demand model include:

[0017] S11. Provide a general description and definition for different image segmentation tasks, including the following steps:

[0018] S111. Define the semantic segmentation task. The semantic segmentation task classifies each pixel in an image into a predefined category. The input is one or more images to be processed, and the output is a mask image of the same size as the input image, where the value of each pixel represents the category it belongs to. Semantic segmentation is mainly used in fields such as autonomous driving and scene understanding;

[0019] S112. Define the instance segmentation task. The instance segmentation task distinguishes different instances in the same category based on pixel classification. The output is multiple instance masks, each mask represents a set of pixels of an independent object in the image. Instance segmentation is widely used in target detection and object recognition;

[0020] S113. Panoramic segmentation task. Panoramic segmentation combines semantic segmentation and instance segmentation, requiring each pixel in the image to be classified while distinguishing different instances of the same category. The output is a mask covering all objects in the image, including the background and foreground. Panoramic segmentation is often used to analyze complex scenes, such as urban street scenes;

[0021] S114. Define the task of medical image segmentation. Medical image segmentation extracts regions of interest, such as tumors or organ boundaries, from medical images. The output is a mask image that annotates specific medical structures. It is widely used in disease diagnosis and treatment planning, and has high accuracy requirements.

[0022] S115. Edge detection task. The edge detection task identifies and locates the edges of objects in an image. The output is a binary image, where "1" represents edge pixels and "0" represents non-edge pixels. Edge detection is often used as a pre-processing step for other segmentation tasks or applied independently to image analysis;

[0023] S116. Define other segmentation tasks. In addition to the above tasks, it also includes video frame segmentation, three-dimensional image segmentation, multispectral image segmentation, etc. The input may be a time series image, a stereo image, or a multi-channel image, and the output is a video frame sequence mask, a three-dimensional voxel mask, or a multispectral image segmentation result;

[0024] S12. Construction of generalized demand model,

[0025] S121. Abstract and unify the requirements of the above-mentioned image segmentation tasks to build a universal description model. The requirements of each task can be summarized as follows: the diversity of input images, the number of target object categories, the morphological characteristics of the target objects, and the output format of the task (such as binary mask, multi-channel mask, etc.).

[0026] S122. The input of all image segmentation tasks can be described as: a two-dimensional or three-dimensional image matrix with a specific resolution, where the matrix elements represent the grayscale or color values ​​of pixels or voxels. The input can be single-channel (grayscale image), three-channel (color image) or multi-channel (multispectral image). The unified description method is: , where Resolution represents the resolution of the image, Channels represents the number of channels, and Depth represents the bit depth of the image.

[0027] S123. The unified description of the output mask is:

[0028] , where Resolution is consistent with the input image, Classes indicates the number of categories, and Format defines the representation of the mask (such as a single-channel binary mask, a multi-channel classification mask, etc.). Different segmentation tasks are uniformly described by this model.

[0029] S124. According to the specific requirements of each segmentation task, the general requirement model is parameterized. The definable parameters include but are not limited to: TaskType (task type, such as semantic segmentation, instance segmentation, etc.), ObjectCategories (number of target object categories), InputChannels (number of input image channels), OutputFormat (output mask format), PrecisionRequirement (precision requirement), etc. By setting these parameters, the model can be flexibly configured to adapt to different segmentation tasks.

[0030] Furthermore, the specific steps of modular design in the adaptive neural architecture search algorithm include:

[0031] S21. Design a modular algorithm functional architecture, including the following sub-steps:

[0032] S211. According to the demand model of the image segmentation task, the overall algorithm is divided into several sub-modules while considering the independence and complementarity of each module. Each sub-module is responsible for processing a specific function or feature;

[0033] S212. According to the task requirements, different functions are assigned to corresponding submodules. Among them, the main function of the feature extraction module is to extract low-level and high-level features from the input image, the following information processing module is responsible for combining the global context information of the image, and the edge detection module mainly identifies and marks the edges of objects in the image;

[0034] S22. Combining and reusing sub-modules includes the following sub-steps:

[0035] S221. According to the overall needs of the network, by analyzing the input and output requirements of each sub-module, determine the connection method between the sub-modules so that information can flow efficiently between the modules;

[0036] S222. By changing the connection method of sub-modules or adjusting the information flow path, recombining sub-modules, and constructing a new network structure to meet the needs of different tasks, it can find the optimal network architecture suitable for specific tasks in a wide search space.

[0037] Furthermore, the specific steps of designing the cell-level architecture include:

[0038] S31. Construct a cell-level architecture to define the cell internal structure, cell representation, cell update method, and high information retention method, including the following sub-steps:

[0039] S311. Define the components in the cell-level architecture. Each cell is composed of multiple blocks, which are the basic units of the entire network. Each block is usually a binary structure that receives two input tensors and converts them into an output tensor. The basic operation of the block can be convolution, pooling, or other common neural network operations. The structure of each block can be regarded as a simple computing unit that is responsible for processing and transforming input features and passing the results to the next block. This binary structure enables each block to effectively combine and process feature information from different sources;

[0040] S312. The entire cell structure is represented as a directed acyclic graph (DAG), where each node represents an operation (such as convolution, pooling, etc.), each edge represents the flow of data between different operations, and allows free combination of multiple inputs and outputs. In the DAG, data flows from one node (operation) to the next node, gradually building and extracting feature information;

[0041] S313. The output of the block is combined by element addition operation. In the binary structure of each block, after the two input tensors are operated on respectively, the two output tensors generated are added by element addition to generate the final output tensor of the block;

[0042] S314. Combine multiple blocks in a certain order to form a complete cell, where each cell acts as an independent functional unit responsible for processing feature extraction at a specific level. In the cell structure, the input of each block can come from the output of the previous block, or the output of the previous two blocks, or even the input tensor of the entire cell;

[0043] S315. Reuse block or cell functions and reduce the complexity of the model by sharing parameters. Define different block structures, connection methods and combination strategies to support the search algorithm to find the optimal cell architecture in a wide search space;

[0044] S32. Design of a cell-level architecture update method, including the following sub-steps:

[0045] S321. A cell is defined as a small fully convolutional module that is repeated multiple times to form the entire neural network. Specifically,

[0046] Cells are made of The directed acyclic graph consists of blocks. Each block is a two-branch structure that maps from two input tensors to one output tensor. Blocks use five tuples To specify, where , is the choice of input tensor, , is the choice of layer type to be applied to the corresponding input tensor, the set of possible layer types There are 3×3 depthwise separable convolutions, 5×5 depthwise separable convolutions, 3×3 dilated convolutions with a dilation rate of 2, 5×5 dilated convolutions with a dilation rate of 2, 3×3 average pooling, 3×3 max pooling, skip connections, and no connections. The single output of the two branches is combined to form the output tensor of the block The output tensor of the cell It is simply the concatenation of the output tensors of the blocks Arrange in this order. The possible set of input tensors Include the output of the previous unit , the output of the previous unit And the output of each previous module in the current unit Therefore, as more modules are added to a unit, the choices for subsequent modules as input sources will also increase.

[0047] S322. At the Cell level modular design level, the output tensor of each block All connected to All hidden states in:

[0048] ,

[0049] in, Indicates Layer The output hidden state of each unit is Indicates connection to Layer The set of all input hidden states of units, Indicates that from Units to The operation function of each unit, in addition, Continuous relaxation To approximate, it is defined as:

[0050] ,

[0051] in:

[0052] ,

[0053] is with each operator The related normalized scalar is implemented by softmax, Indicates Operators act on The result, and Always included in In, and yes Combined with the above formula, The level updates are summarized as follows:

[0054] ,

[0055] S33. Design of a method for maintaining high spatial resolution, comprising the following steps:

[0056] S331. In the design of the network, a two-layer "stem" structure is first set up as the beginning of the network. Each layer of this structure reduces the spatial resolution of the input image by half to lay the foundation for subsequent deep processing. Next, the network will be divided into L layers, and the spatial resolution of each layer can be a multiple of 4, 8, 16 or 32. It is conducive to managing and optimizing the feature extraction process at different resolutions, so that the network can handle complex image segmentation tasks.

[0057] S332. In each layer of the network, a strategy of branching and skipping connections is adopted to improve the network's expressiveness. Specifically, branching allows the network to process information in parallel on different paths, so that it can focus on capturing different types of features. For example, some branches can focus on capturing local details of an image, while other branches process global contextual information. At the same time, skipping connections allow the network to pass the output of the previous layer directly to the subsequent layer to alleviate the gradient vanishing problem and promote the effective flow of information between different layers.

[0058] S333. A differential relaxation method is used to control the scalars of the connection strength between different hidden states. By incorporating these scalars into a differentiable computational graph and using the gradient descent algorithm, the values ​​of these scalars can be efficiently optimized, thereby converting the originally discrete architecture selection into a continuous optimization problem. At the same time, in order to adapt to the needs of different tasks, the network can choose different resolution paths at each layer, allowing flexible switching of different resolutions. Specifically, within a Cell, the spatial size of all tensors is the same, but at the network level, tensors may have different sizes. Therefore, in order to design continuous relaxation, each layer There are at most 4 hidden states ,The superscript in the upper left corner indicates the spatial resolution. A network-level continuous relaxation is designed to match the cell-level search space.

[0059]

[0060] ,

[0061] in, Indicates The spatial resolution of the layer is The hidden state output of

[0062] Indicates the resolution To resolution The weights define the relative importance of cross-resolution operations. represents the computational unit at the cell level, , Scalar Can be normalized and implemented through softmax, Controls the external network level, depending on the spatial size and layer index, each The scalar in controls a whole set of , The same architecture is specified, which does not depend on the spatial size or the index of the layer.

[0063] Furthermore, the specific steps of the improved particle swarm optimization algorithm for neural architecture search include:

[0064] S41. The initialization phase includes the initialization of particle positions and velocities, and includes the following steps:

[0065] S411. In particle position initialization, set the search space to a dimensional continuous space ,in Represents the number of architectural parameters that need to be optimized. In a neural network, the parameters that may be involved include the number of convolutional layers, the kernel size of each convolutional layer, the type of activation function, the connection method between layers, etc. These parameters can be represented by the various components of the particle. , where each All particles are The position of the particle corresponds to the initial architecture parameters of the model, the convolution kernel size, the number of layers, and the learning rate. For each particle , in is the total number of particles, and its position Randomly distributed in the search space, represented as

[0066] ,

[0067] in For the The particle in The initial position on the dimension is randomly generated and satisfies:

[0068] ,

[0069] in, and They are Minimum and maximum values ​​for dimension schema parameters,

[0070] S412. Particle velocity initialization, velocity vector It is also randomly initialized in each dimension, expressed as:

[0071] ,

[0072] in For the The particle in The initial position on the dimension satisfies:

[0073] ,

[0074] in, , Respectively The minimum and maximum values ​​of the dimension and the velocity bound ensure that the particle swarm has enough exploration ability in the initial stage while not moving too discretely or concentratedly in the search space.

[0075] S42. The global search and coarse adjustment phase includes the following steps:

[0076] S421. Inertia weight and acceleration coefficient adjustment. In the global search and rough adjustment stage, the improved PSO algorithm replaces some items in the traditional algorithm by introducing the average optimal solution, thereby gradually reducing the inertia weight and improving the accuracy of the algorithm. Used to control the influence of the particle's current velocity. The inertia weight can be set to a fixed value or adjusted dynamically:

[0077] ,

[0078] in, and are the maximum and minimum values ​​of the inertia weight, respectively. is the current iteration number, is the maximum number of iterations,

[0079] S422. Particle velocity update. In each iteration, the velocity of each particle is updated according to the inertia weight, acceleration coefficient and random factor. The formula is as follows:

[0080] ,

[0081] in, Represents particles exist The speed at the iteration, Represents the inertia weight, which is used to control the influence of the particle's current velocity. The value gradually decreases with iteration to balance exploration and development. It is a particle The current best position of , that is, the optimal neural network architecture parameters found by the current particle, is the global average optimal position, that is, the best architecture parameter among all particles in the particle swarm, and the acceleration coefficient and , i.e., the individual learning factor, determines the degree to which the particle follows its personal best position and the global average best position, and is set to a constant. The random factor and Used to introduce randomness so that particles have a certain exploration ability during the search process, randomly sampling from a uniform distribution, , Representative particles exist The position at the iteration, where the acceleration coefficient and Guide particles to move towards their own historical best architecture and the global best architecture of the entire group,

[0082] S423. Update particle position. Update the particle position according to the updated speed. The formula is as follows:

[0083] ,

[0084] After updating the position, check if the particle is out of the search space. If so, apply a speed penalty mechanism for each dimension of the particle. , if the position Exceeded the minimum value set or maximum value Then set the particle's velocity to zero and adjust its position in the opposite direction.

[0085]

[0086] Through this penalty mechanism, we can avoid crossing the boundary and ensure that particles always search within the valid search space.

[0087] S424. Fitness calculation and update, calculate the fitness of each particle, that is, evaluate the performance of the current architecture, assuming that the objective function is , then the fitness value is , each particle saves its historical best fitness value and its corresponding position,

[0088] renew and , if the current fitness Better than the best in history , then update the optimal position of the particle,

[0089] Update the global optimal position for all particles :

[0090] ,

[0091] When updating all particles Then, find the global best position ,

[0092] ,

[0093] in, It is a particle The historical best fitness value.

[0094] S43. Local optimization and fine-tuning phase, including the following steps:

[0095] S431. Select the local optimization area. After the global search, observe the distribution of the particle swarm in the search space, and select the area with higher density in the particle cluster as the target area for local optimization. By defining a bounding box that covers the part with the highest density in the particle cluster and using it as the search space for local optimization, the quality of the solution can be further improved.

[0096] S432. Gradient descent optimization, the global best position obtained in the PSO global search phase As the initial parameter of the gradient descent method . Set the initial learning rate for the gradient descent method , this learning rate determines the step size of parameter update at each iteration, and further optimizes the gradient descent method. Calculate the current architecture parameters right Gradient , using automatic differentiation tools to calculate the gradient;

[0097] ,

[0098] Each gradient It reflects the objective function of the first The architecture parameters are updated using gradient descent.

[0099] ,

[0100] in is the learning rate, is the gradient calculated under the current parameter configuration.

[0101] S433. Local optimal solution determination and final architecture selection. After the gradient descent method converges, observe the magnitude of parameter updates. If the parameter convergence has stabilized, the optimization can be stopped and the current optimal architecture parameters can be output. , and regard it as a local optimal solution. During the NAS search process, the final network architecture is selected through performance evaluation on the validation set. The performance indicators of all candidate architectures are recorded, and the architecture that suits the task requirements is finally selected. Comprehensive testing is performed on the validation set and test set to ensure that its performance on different data sets is stable and excellent.

[0102] The present invention provides a neural architecture search method for adaptive multi-domain image segmentation tasks. By constructing a modular neural network architecture, combined with an adaptive neural architecture search algorithm (NAS) and an improved particle swarm optimization algorithm (PSO), efficient processing of multiple image segmentation tasks is achieved. This method not only improves the adaptability and generalization ability of the model, but also effectively reduces the computational complexity while maintaining high spatial resolution, significantly improving the accuracy and efficiency of segmentation tasks, and is suitable for a variety of complex image processing scenarios, such as medical image analysis, autonomous driving and other fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0103] Figure 1 It is the structural diagram of the general image segmentation task model.

[0104] Figure 2 is a structural diagram of a modular neural network architecture.

[0105] Figure 3 It is a directed acyclic graph DAG of the cell-level architecture.

[0106] Figure 4 It is a flowchart of the improved particle swarm optimization algorithm PSO. DETAILED DESCRIPTION

[0107] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.

[0108] Example: This implementation proposes a neural architecture search method for adaptive multi-domain image segmentation tasks. This method optimizes the image segmentation model by constructing a modular neural network architecture and combining it with an adaptive neural architecture search (NAS) algorithm. The specific implementation steps are as follows:

[0109] S1. Construct a universal image segmentation task model to uniformly describe and model different types of image segmentation tasks (such as semantic segmentation, instance segmentation, panoptic segmentation, medical image segmentation, and edge detection), thereby providing standardized task input for subsequent algorithm design. Figure 1 This is a structural diagram of a general image segmentation task model, showing a unified description of the input image format (such as resolution, number of channels, bit depth) and output mask format (such as binary mask, multi-channel mask). The figure also shows how multiple image segmentation tasks (such as semantic segmentation, instance segmentation, panoptic segmentation, medical image segmentation, and edge detection) are modeled uniformly.

[0110] This step includes the following sub-steps:

[0111] S11. Provide a general description and definition for different image segmentation tasks, including the following steps:

[0112] S111. Define the semantic segmentation task. The semantic segmentation task classifies each pixel in an image into a predefined category. The input is one or more images to be processed, and the output is a mask image of the same size as the input image, where the value of each pixel represents the category it belongs to. Semantic segmentation is mainly used in fields such as autonomous driving and scene understanding.

[0113] S112. Define the instance segmentation task. The instance segmentation task distinguishes different instances in the same category based on pixel classification. The output is multiple instance masks, each mask represents a set of pixels of an independent object in the image. Instance segmentation is widely used in target detection and object recognition.

[0114] S113. Panoramic segmentation task. Panoramic segmentation combines semantic segmentation and instance segmentation, requiring each pixel in the image to be classified while distinguishing different instances of the same category. The output is a mask covering all objects in the image, including background and foreground. Panoramic segmentation is often used to analyze complex scenes, such as urban street scenes.

[0115] S114. Define the task of medical image segmentation. Medical image segmentation extracts regions of interest, such as tumors or organ boundaries, from medical images. The output is a mask image that annotates specific medical structures. It is widely used in disease diagnosis and treatment planning, and has high accuracy requirements.

[0116] S115. Edge detection task. The edge detection task identifies and locates the edges of objects in an image. The output is a binary image, where "1" represents edge pixels and "0" represents non-edge pixels. Edge detection is often used as a pre-processing step for other segmentation tasks or independently applied to image analysis.

[0117] S116. Define other segmentation tasks. In addition to the above tasks, other tasks include video frame segmentation, three-dimensional image segmentation, multispectral image segmentation, etc. The input may be a time series image, a stereo image, or a multi-channel image, and the output is a video frame sequence mask, a three-dimensional voxel mask, or a segmentation result of a multispectral image.

[0118] S12. Construction of generalized demand model

[0119] S121. Abstract and unify the requirements of the above-mentioned image segmentation tasks to build a universal description model. The requirements of each task can be summarized as follows: the diversity of input images, the number of target object categories, the morphological characteristics of the target objects, and the output format of the task (such as binary mask, multi-channel mask, etc.).

[0120] S122. The input of all image segmentation tasks can be described as: a two-dimensional or three-dimensional image matrix with a specific resolution, where the matrix elements represent the grayscale or color values ​​of pixels or voxels. The input can be single-channel (grayscale image), three-channel (color image) or multi-channel (multispectral image). The unified description method is:

[0121] , where Resolution represents the resolution of the image, Channels represents the number of channels, and Depth represents the bit depth of the image.

[0122] S123. The unified description of the output mask is:

[0123] , where Resolution is consistent with the input image, Classes indicates the number of categories, and Format defines the representation of the mask (such as a single-channel binary mask, a multi-channel classification mask, etc.). Different segmentation tasks are uniformly described by this model.

[0124] S124. According to the specific requirements of each segmentation task, the general requirement model is parameterized. The definable parameters include but are not limited to: TaskType (task type, such as semantic segmentation, instance segmentation, etc.), ObjectCategories (number of target object categories), InputChannels (number of input image channels), OutputFormat (output mask format), PrecisionRequirement (precision requirement), etc. By setting these parameters, the model can be flexibly configured to adapt to different segmentation tasks.

[0125] S2. Modular design in the adaptive neural architecture search algorithm, including the following sub-steps:

[0126] S21. Design modular algorithm functional architecture,

[0127] Contains the following sub-steps:

[0128] S211. According to the demand model of the image segmentation task, the overall algorithm is divided into several sub-modules while considering the independence and complementarity of each module. Each sub-module is responsible for processing a specific function or feature.

[0129] S212. According to the task requirements, different functions are assigned to corresponding submodules. Among them, the main function of the feature extraction module is to extract low-level and high-level features from the input image, the following information processing module is responsible for combining the global context information of the image, and the edge detection module mainly identifies and marks the edges of objects in the image. Figure 2 It is a structural diagram of a modular neural network architecture, showing the modular design in the adaptive neural architecture search algorithm. The diagram shows that the network architecture is divided into multiple functional modules, including the connection relationship and information flow between low-level feature extraction modules, high-level feature fusion modules, jump connection modules, and dynamic routing modules.

[0130] S22. Combining and reusing sub-modules includes the following sub-steps:

[0131] S221. According to the overall needs of the network, by analyzing the input and output requirements of each sub-module, determine the connection method between the sub-modules so that information can flow efficiently between the modules.

[0132] S222. By changing the connection method of sub-modules or adjusting the information flow path, recombining sub-modules, and constructing a new network structure to meet the needs of different tasks, it can find the optimal network architecture suitable for specific tasks in a wide search space.

[0133] S3. Design of cell-level architecture construction, Figure 3It is a directed acyclic graph (DAG) of the cell-level architecture, showing the basic structure of the cell-level architecture. The figure shows the node representation of multiple blocks in the DAG, and the flow of data between different operations (such as convolution and pooling);

[0134] Contains the following sub-steps:

[0135] S31. Construct a cell-level architecture to define the cell internal structure, cell representation, cell update method, and high information retention method, including the following sub-steps:

[0136] S311. Define the components in the cell-level architecture. Each cell is composed of multiple blocks, which are the basic units of the entire network. Each block is usually a binary structure that receives two input tensors and converts them into an output tensor. The basic operation of the block can be convolution, pooling, or other common neural network operations. The structure of each block can be regarded as a simple computing unit that is responsible for processing and transforming input features and passing the results to the next block. This binary structure enables each block to effectively combine and process feature information from different sources.

[0137] S312. The entire cell structure is represented as a directed acyclic graph (DAG), where each node represents an operation (such as convolution, pooling, etc.), each edge represents the flow of data between different operations, and allows free combination of multiple inputs and outputs. In the DAG, data flows from one node (operation) to the next node, gradually building and extracting feature information.

[0138] S313. The output of the block is combined through element-wise addition operations. In the binary structure of each block, after the two input tensors are subjected to their respective operations, the two output tensors generated are added together through element-wise addition to generate the final output tensor of the block.

[0139] S314. Combine multiple blocks in a certain order to form a complete cell, where each cell acts as an independent functional unit responsible for processing feature extraction at a specific level. In the cell structure, the input of each block can come from the output of the previous block, or the output of the previous two blocks, or even the input tensor of the entire cell.

[0140] S315. Reuse block or cell functions and reduce the complexity of the model by sharing parameters. Define different block structures, connection methods and combination strategies to support the search algorithm to find the optimal cell architecture in a wide search space.

[0141] S32. Design of a cell-level architecture update method, including the following sub-steps:

[0142] S321. A cell is defined as a small fully convolutional module that is repeated multiple times to form the entire neural network. Specifically, a cell is composed of The directed acyclic graph consists of blocks. Each block is a two-branch structure that maps from two input tensors to one output tensor. Blocks use five tuples To specify, where , is the choice of input tensor, , is the choice of layer type to be applied to the corresponding input tensor, the set of possible layer types There are 3×3 depthwise separable convolutions, 5×5 depthwise separable convolutions, 3×3 dilated convolutions with a dilation rate of 2, 5×5 dilated convolutions with a dilation rate of 2, 3×3 average pooling, 3×3 max pooling, skip connections, and no connections. The single output of the two branches is combined to form the output tensor of the block The output tensor of the cell It is simply the concatenation of the output tensors of the blocks Arrange in this order. The possible set of input tensors Include the output of the previous unit , the output of the previous unit And the output of each previous module in the current unit Therefore, as more modules are added to a unit, the choices for subsequent modules as input sources will also increase.

[0143] S322. At the Cell level modular design level, the output tensor of each block All connected to All hidden states in:

[0144] ,

[0145] in, Indicates Layer The output hidden state of each unit is Indicates connection to Layer The set of all input hidden states of units, Indicates that from Units to The operation function of each unit, in addition, Continuous relaxation To approximate, it is defined as:

[0146] ,

[0147] in:

[0148] ,

[0149] is with each operator The related normalized scalar is implemented by softmax, Indicates Operators act on The result, and Always included in In, and yes Combined with the above formula, The level updates are summarized as follows:

[0150] .

[0151] S33. Design of a method for maintaining high spatial resolution, comprising the following steps:

[0152] S331. In the design of the network, a two-layer "stem" structure is first set up as the beginning of the network. Each layer of this structure reduces the spatial resolution of the input image by half to lay the foundation for subsequent deep processing. Next, the network will be divided into L layers, and the spatial resolution of each layer can be a multiple of 4, 8, 16 or 32. It is conducive to managing and optimizing the feature extraction process at different resolutions, so that the network can handle complex image segmentation tasks.

[0153] S332. In each layer of the network, a strategy of branching and skipping connections is adopted to improve the network's expressiveness. Specifically, branching allows the network to process information in parallel on different paths, so that it can focus on capturing different types of features. For example, some branches can focus on capturing local details of an image, while other branches process global contextual information. At the same time, skipping connections allow the network to pass the output of the previous layer directly to the subsequent layer to alleviate the gradient vanishing problem and promote the effective flow of information between different layers.

[0154] S333. A differential relaxation method is used to control the scalars of the connection strength between different hidden states. By incorporating these scalars into a differentiable computational graph and using the gradient descent algorithm, the values ​​of these scalars can be efficiently optimized, thereby converting the originally discrete architecture selection into a continuous optimization problem. At the same time, in order to adapt to the needs of different tasks, the network can choose different resolution paths at each layer, allowing flexible switching of different resolutions. Specifically, within a Cell, the spatial size of all tensors is the same, but at the network level, tensors may have different sizes. Therefore, in order to design continuous relaxation, each layer There are at most 4 hidden states ,The superscript in the upper left corner indicates the spatial resolution. A network-level continuous relaxation is designed to match the cell-level search space.

[0155]

[0156] ,

[0157] in, Indicates The spatial resolution of the layer is The hidden state output of

[0158] Indicates the resolution To resolution The weights define the relative importance of cross-resolution operations. represents the computational unit at the cell level, , Scalar Can be normalized and implemented through softmax, Controls the external network level, depending on the spatial size and layer index, each The scalar in controls a whole set of , The same architecture is specified, which does not depend on the spatial size or the index of the layer.

[0159] S4. Using the improved particle swarm optimization algorithm (PSO),

[0160] Figure 4 This is a flowchart of the improved particle swarm optimization algorithm (PSO), which shows the various stages of the algorithm, including the initialization stage, global search and rough adjustment stage, fitness calculation and update stage, local optimization and fine adjustment stage, and optimal architecture selection stage. This figure details the specific process of particle position and velocity update, fitness calculation, and optimal architecture selection;

[0161] The algorithm is divided into the following stages:

[0162] S41. The initialization phase includes the initialization of particle positions and velocities, and includes the following steps:

[0163] S411. In particle position initialization, set the search space to a dimensional continuous space ,in Represents the number of architectural parameters that need to be optimized. In a neural network, the parameters that may be involved include the number of convolutional layers, the kernel size of each convolutional layer, the type of activation function, the connection method between layers, etc. These parameters can be represented by the various components of the particle. , where each All particles are The position of the particle corresponds to the initial architecture parameters of the model, the convolution kernel size, the number of layers, and the learning rate. For each particle , in is the total number of particles, and its position Randomly distributed in the search space, represented as

[0164] ,

[0165] in For the The particle in The initial position on the dimension is randomly generated and satisfies:

[0166] ,

[0167] in, and They are Minimum and maximum values ​​for dimension schema parameters,

[0168] S412. Particle velocity initialization, velocity vector It is also randomly initialized in each dimension, expressed as:

[0169] ,

[0170] in For the The particle in The initial position on the dimension satisfies:

[0171] ,

[0172] in, , Respectively The minimum and maximum values ​​of the dimension and the velocity bound ensure that the particle swarm has enough exploration ability in the initial stage while not moving too discretely or concentratedly in the search space.

[0173] S42. The global search and coarse adjustment phase includes the following steps:

[0174] S421. Inertia weight and acceleration coefficient adjustment. In the global search and rough adjustment stage, the improved PSO algorithm replaces some items in the traditional algorithm by introducing the average optimal solution, thereby gradually reducing the inertia weight and improving the accuracy of the algorithm. Used to control the influence of the particle's current velocity. The inertia weight can be set to a fixed value or adjusted dynamically:

[0175] ,

[0176] in, and are the maximum and minimum values ​​of the inertia weight, respectively. is the current iteration number, is the maximum number of iterations.

[0177] S422. Particle velocity update. In each iteration, the velocity of each particle is updated according to the inertia weight, acceleration coefficient and random factor. The formula is as follows:

[0178] ,

[0179] in, Represents particles exist The speed at the iteration, Represents the inertia weight, which is used to control the influence of the particle's current velocity. The value gradually decreases with iteration to balance exploration and development. It is a particle The current best position of , that is, the optimal neural network architecture parameters found by the current particle, is the global average optimal position, that is, the best architecture parameter among all particles in the particle swarm, and the acceleration coefficient and , i.e., the individual learning factor, determines the degree to which the particle follows its personal best position and the global average best position, and is set to a constant. The random factor and Used to introduce randomness so that particles have a certain exploration ability during the search process, randomly sampling from a uniform distribution, , Representative particles exist The position at the iteration, where the acceleration coefficient and Guide particles to move towards their own historical best architecture and the global best architecture of the entire group,

[0180] S423. Update particle position. Update the particle position according to the updated speed. The formula is as follows:

[0181] ,

[0182] After updating the position, check whether the particle is out of the search space. If so, apply a speed penalty mechanism for each dimension of the particle. , if the position Exceeded the minimum value set or maximum value Then set the particle's velocity to zero and adjust its position in the opposite direction.

[0183]

[0184] Through this penalty mechanism, we can avoid crossing the boundary and ensure that particles always search within the valid search space.

[0185] S424. Fitness calculation and update, calculate the fitness of each particle, that is, evaluate the performance of the current architecture, assuming that the objective function is , then the fitness value is , each particle saves its historical best fitness value and its corresponding position,

[0186] renew and , if the current fitness Better than the best in history , then update the optimal position of the particle,

[0187] Update the global optimal position for all particles :

[0188] ,

[0189] When updating all particles Then, find the global best position ,

[0190]

[0191] in, It is a particle The historical best fitness value.

[0192] S43. Local optimization and fine-tuning phase, including the following steps:

[0193] S431. Select local optimization area. After global search, observe the distribution of particle swarm in the search space and select the area with higher degree in particle cluster as the target area for local optimization. Define a bounding box that covers the part with the highest degree in particle cluster and use it as the search space for local optimization to further improve the quality of the solution.

[0194] S432. Gradient descent optimization. The global optimal position obtained in the PSO global search phase is As the initial parameter of the gradient descent method . Set the initial learning rate for the gradient descent method , this learning rate determines the step size of parameter update at each iteration, and further optimizes the gradient descent method. Calculate the current architecture parameters right Gradient , using automatic differentiation tools to compute the gradient.

[0195] ,

[0196] Each gradient It reflects the objective function of the first The architecture parameters are updated using gradient descent.

[0197] ,

[0198] in is the learning rate, is the gradient calculated under the current parameter configuration.

[0199] S433. Local optimal solution determination and final architecture selection. After the gradient descent method converges, observe the magnitude of parameter updates. If the parameter convergence has stabilized, the optimization can be stopped and the current optimal architecture parameters can be output. , and regard it as a local optimal solution. During the NAS search process, the final network architecture is selected through performance evaluation on the validation set. The performance indicators of all candidate architectures are recorded, and the architecture that suits the task requirements is finally selected. Comprehensive testing is performed on the validation set and test set to ensure that its performance on different data sets is stable and excellent.

[0200] The above description is only a specific implementation of the present application, so that those skilled in the art can understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest range consistent with the principles and novel features applied for herein.

Claims

1. A neural architecture search method for adaptive multi-domain image segmentation tasks, characterized in that: The method comprises the following steps: S1. Construct a unified description and requirement model for multi-image segmentation tasks, build a general image segmentation task model, abstract the task requirements, establish a unified input image format and output mask format, and parameterize the model to provide standardized task input. S2. Modular design in the adaptive neural architecture search algorithm. The adaptive neural architecture search algorithm modularizes the entire neural network architecture. According to the unified demand model, the network is divided into multiple functional modules, and different functions are assigned to corresponding modules, so that each module can efficiently handle the image segmentation task. By analyzing the connection requirements between modules, the best connection method is determined, and the module connection is adjusted to meet different task requirements to build an optimized network structure. S3. The design of the cell-level architecture is to represent the neural network cell as a directed acyclic graph (DAG) and build multiple blocks on this basis. Each block performs a neural network operation, and the output of the block is combined through element addition operations to form a complete cell. The module is reused and optimized through a shared parameter mechanism. Through the combination and update of cells, the extraction of features at different levels is achieved, and efficient computing power is maintained in complex image segmentation tasks, achieving high spatial resolution maintenance and dynamic adaptation. The details are as follows: S31. Constructing cell-level architecture, S32. Cell level architecture update, S33. Design of high spatial resolution preservation method; S4. Use the improved particle swarm optimization algorithm to search for neural architectures. Use the improved particle swarm optimization algorithm PSO to perform global and local optimization of the network architecture. Initialize the position and velocity of the particles, establish the distribution of the particle swarm in the search space, adjust the velocity and position of the particles according to the inertia weight and acceleration coefficient, and optimize the historical best position of each particle through fitness calculation. After the global search, adjust the architecture parameters through the local optimization strategy, and finally select the network architecture with the best performance on the verification set, so as to achieve accurate processing of complex tasks. The specific steps are as follows: S41. The initialization phase includes the initialization of particle positions and velocities. S42. Global search and coarse adjustment stage, S43. Local optimization and fine-tuning stage.

2. The neural architecture search method for adaptive multi-domain image segmentation tasks according to claim 1, characterized in that: Construct a unified description and requirement model for multi-image segmentation tasks. The specific steps include: S11. Provide a general description and definition for different image segmentation tasks, including the following steps: S111. Define a semantic segmentation task, which classifies each pixel in an image into a predefined category. The input is one or more images to be processed, and the output is a mask image with the same size as the input image, where the value of each pixel indicates the category to which it belongs. S112. Define an instance segmentation task. The instance segmentation task distinguishes different instances in the same category based on pixel classification and outputs multiple instance masks. Each mask represents a pixel set of an independent object in the image. S113. Panoramic segmentation task,Panoramic segmentation combines semantic segmentation and instance segmentation, requiring each pixel in the image to be classified, while distinguishing different instances of the same category, and outputting a mask covering all objects in the image, including background and foreground. S114. Define the task of medical image segmentation. Medical image segmentation extracts the region of interest from the medical image and outputs it as a mask image annotating the medical structure. S115. Edge detection task: The edge detection task identifies and locates the edges of objects in the image and outputs a binary image, where "1" represents edge pixels and "0" represents non-edge pixels. S116. Define other segmentation tasks, including video frame segmentation, three-dimensional image segmentation, and multispectral image segmentation, where the input is a time series image, a stereo image, or a multi-channel image, and the output is a video frame sequence mask, a three-dimensional voxel mask, or a segmentation result of a multispectral image; S12. The construction of the demand model is as follows: S121. Construct a general description model. The requirements of each task are summarized as follows: the diversity of input images, the number of target object categories, the morphological characteristics of target objects, and the output format of the task. S122. The input of all image segmentation tasks is described as: a two-dimensional or three-dimensional image matrix with resolution, the matrix elements represent the grayscale value or color value of the pixel or voxel, the input is a single-channel grayscale image, a three-channel color image, or a multi-channel multispectral image. The unified description method is: INputImage = {Resolution, Channels, Depth}, where Resolution represents the resolution of the image, Channels represents the number of channels, and Depth represents the bit depth of the image. S123. The unified description of the output mask is: OutputMask = {Resolution, Classes, Format}, where Resolution is consistent with the input image, Classes indicates the number of categories, and Format defines the representation of the mask. S124. According to the specific requirements of each segmentation task, the general requirement model is parameterized, and the defined parameters include task type, number of target object categories, number of input image channels, output mask format and accuracy requirements.

3. The neural architecture search method for adaptive multi-domain image segmentation tasks according to claim 1, characterized in that: The specific steps of modular design in the adaptive neural architecture search algorithm include: S21. Design a modular algorithm functional architecture, including the following sub-steps: S211. According to the demand model of the image segmentation task, the overall algorithm is divided into several sub-modules while considering the independence and complementarity of each module. S212. According to the task requirements, different functions are assigned to corresponding submodules, wherein the function of the feature extraction module is to extract low-level and high-level features from the input image, the following information processing module is responsible for combining the global context information of the image, and the edge detection module is responsible for identifying and marking the edges of objects in the image. S22. Combining and reusing sub-modules includes the following sub-steps: S221. According to the overall needs of the network, by analyzing the input and output requirements of each sub-module, determine the connection method between the sub-modules so that information can flow efficiently between modules. S222. By changing the connection method of sub-modules or adjusting the information flow path, the sub-modules are recombined to build a new network structure.

4. The neural architecture search method for adaptive multi-domain image segmentation tasks according to claim 1, characterized in that: The design of the cell-level architecture includes the following steps: S31. Construct a cell-level architecture to define the cell internal structure, cell representation, cell update method, and high information retention method, including the following sub-steps: S311. Define the components in the cell-level architecture. Each cell is composed of multiple blocks. Blocks are the basic units that constitute the entire network. Each block is a binary structure that receives two input tensors and converts them into an output tensor. The basic operations of the block are convolution and pooling. The structure of each block is regarded as a simple computing unit, which is responsible for processing and converting input features and passing the results to the next block. S312. The entire cell structure is represented as a directed acyclic graph (DAG), where each node represents an operation, each edge represents the flow of data between different operations, and allows free combination of multiple inputs and outputs. S313. The output of the block is combined by element addition operation. In the binary structure of each block, after the two input tensors are operated on respectively, the two output tensors generated are added by element addition to generate the final output tensor of the block. S314. Combine multiple blocks in order to form a complete cell, where each cell acts as an independent functional unit responsible for processing hierarchical feature extraction. In the cell structure, the input of each block comes from the output of the previous block, or the output of the previous two blocks, or the input tensor of the entire cell. S315. Reuse block or cell functions and reduce the complexity of the model by sharing parameters. Define different block structures, connection methods, and combination strategies to support the search algorithm to find the optimal cell architecture in a wide search space.

5. The neural architecture search method for adaptive multi-domain image segmentation tasks according to claim 1, characterized in that: S32. Cell-level architecture update, including the following sub-steps: S321. A cell is defined as a small fully convolutional module that is repeated multiple times to form the entire neural network. Specifically, a cell is a directed acyclic graph consisting of B blocks. Each block is a two-branch structure that maps from two input tensors to one output tensor. The i-th block in a cell is specified using a quintuple (I1, I2, O1, O2, C), where is a choice of input tensors, O1,O2∈O, is a choice of the type of layer to be applied to the corresponding input tensors, S322. At the Cell level modular design level, the output tensor of each block All connected to All hidden states in: in, represents the output hidden state of the ith unit in the lth layer, represents the set of all input hidden states connected to the ith unit in the lth layer, O j→i Represents the operation function from the jth unit to the ith unit. In addition, each O j→i Continuous relaxation To approximate, it is defined as: in: is with each operator O k ∈O related normalized scalar, implemented by softmax, Indicates that the kth operator acts on The result, H l-1 and H l-2 Always included in In, and H l yes Combined with the above formula, the update of Cell level is summarized as follows: H l =cell(H l-1 ,H l-2 ;α)。 6. The neural architecture search method for adaptive multi-domain image segmentation tasks according to claim 1, characterized in that: S33. Design of a method for maintaining high spatial resolution, comprising the following steps: S331. In the design of the network, a two-layer "stem" structure is first set up as the beginning of the network. The network will be divided into L layers, and the spatial resolution of each layer is a multiple of 4, 8, 16 or 32. S332. In each layer of the network, the strategy of branching and skipping connections is adopted to improve the network's expressive power. S333. The differential relaxation method is used to control the scalars of the connection strength between different hidden states. By incorporating these scalars into the differentiable computational graph, the gradient descent algorithm is used to efficiently optimize the values ​​of these scalars. Specifically, each layer l has a maximum of 4 hidden states {4H l , 8H l , 16H l , 32H l }, the superscript in the upper left corner indicates the spatial resolution, and the network-level continuous relaxation is designed to match the cell-level search space. in, s H l represents the hidden state output of the lth layer with a spatial resolution of S, Represents the weight from resolution X to resolution S, defining the relative importance of cross-resolution operations. Cell() represents the computational unit at the cell level. s=4, 8, 16, 32, 1=1, 2, ..., L, the scalar β is normalized and implemented by softmax. β controls the external network level, depending on the spatial size and the index of the layer. Each scalar in β controls a whole set of α, which specify the same architecture.

7. The neural architecture search method for adaptive multi-domain image segmentation tasks according to claim 1, characterized in that: The specific steps of the improved particle swarm optimization algorithm for neural architecture search include: S41. The initialization phase includes the initialization of particle positions and velocities, and includes the following steps: S411. In particle position initialization, the search space is set to a d-dimensional continuous space R d , where d represents the number of architectural parameters to be optimized, X i ={x i1 , x i2 , ..., x id }, where each X ij are the specific values ​​of the particle on the jth architecture parameter. The position of the particle corresponds to the initial architecture parameters of the model, the convolution kernel size, the number of layers, and the learning rate. For each particle i, i = 1, 2, ..., N, where N is the total number of particles, its position X i (0) Randomly distributed in the search space, denoted as X i (0)={x i1 (0), x i2 (0), ..., x id (0)} Among them, X ij (0) is the initial position of the i-th particle in the j-th dimension, which is randomly generated and satisfies: in, and are the minimum and maximum values ​​of the j-dimensional architecture parameters, respectively. S412. Particle velocity initialization, velocity vector V i (0) is also randomly initialized in each dimension, expressed as: V i (0)={v i1 (0),v i2 (0),...,v id (0)} where v ij (0) is the initial position of the i-th particle in the j-th dimension, satisfying: in, are the minimum and maximum values ​​of the j-th dimension respectively.

8. The neural architecture search method for adaptive multi-domain image segmentation tasks according to claim 1, characterized in that: S42. The global search and coarse adjustment phase includes the following steps: S421. Inertia weight and acceleration coefficient adjustment. In the global search and rough adjustment stage, the improved PSO algorithm introduces the average optimal solution. The inertia weight w(t) is used to control the influence of the current velocity of the particle. The inertia weight is set to a fixed value or adjusted dynamically: Among them, w max With w min are the maximum and minimum values ​​of the inertia weight, t is the current iteration number, and t max is the maximum number of iterations, S422. Particle velocity update. In each iteration, the velocity of each particle is updated according to the inertia weight, acceleration coefficient and random factor. The formula is as follows: Among them, V i (t+1) represents the speed of particle i at the t+1 iteration, w(t) represents the inertia weight, which is used to control the influence of the current speed of the particle. The value gradually decreases with the iteration to balance exploration and development. Pbest i is the current best position of particle i, that is, the optimal neural network architecture parameters found by the current particle, is the global average best position, i.e., the best architectural parameter among all particles in the swarm. The acceleration coefficients c1 and c2, i.e., the individual learning factors, determine the degree to which the particle follows its personal best position and the global average best position. They are set to constants. The random factors r1 and r2 are used to introduce randomness and are randomly sampled from a uniform distribution. r1, r2 ~ U(0, 1), X i (t) represents the position of particle i at iteration t, where the acceleration coefficients c1 and c2 guide the particle to move toward its own historical optimal architecture and the global optimal architecture of the entire group. S423. Update particle position. Update the particle position according to the updated speed. The formula is as follows: X i (t+1)=X i (t)+V i (t+1), After updating the position, check whether the particle is out of the search space. If so, apply the speed penalty mechanism. For each dimension D of the particle, if the position X i,D (t+1) exceeds the set minimum value X min or the maximum value X max Then set the particle's velocity to zero and adjust its position in the opposite direction. if then V i (t+1)=0, X i (t+1)=-X i (t+1), through this penalty mechanism, we can avoid crossing the boundary and ensure that the particle always searches within the valid search space. S424. Fitness calculation and update, calculate the fitness of each particle, that is, evaluate the performance of the current architecture. Let the objective function be J(X), then the fitness value is f i (t) = J(X i (t)), each particle saves its historical best fitness value Pbest i and its corresponding position, Update Pbest and Gbest, if the current fitness f i (t)Better than the best in history Then update the optimal position of the particle, Update the global best position Gbest for all particles: if then Gbest=Pbest i , After updating the Pbest of all particles, find the global best position Gbest. in, is the historical best fitness value of particle i.

9. The neural architecture search method for adaptive multi-domain image segmentation tasks according to claim 1, characterized in that: S43. Local optimization and fine-tuning phase, including the following steps: S431. Select a local optimization area. After global search, observe the distribution of the particle swarm in the search space, select the area with high particle cluster density as the target area for local optimization, and define a bounding box covering the part with the highest density in the particle cluster as the search space for local optimization. S432. Gradient descent optimization, using the global best position Gbest obtained in the PSO global search phase as the initial parameter θ0 of the gradient descent method, setting the initial learning rate η0 of the gradient descent method, and calculating the gradient of the current architecture parameter θ to J(θ) Use automatic differentiation tools to calculate the gradient, Each gradient It reflects the sensitivity of the objective function to the jth architecture parameter under the current parameter configuration, and uses the gradient descent method to update the architecture parameter where η t is the learning rate, is the gradient calculated under the current parameter configuration, S433. Local optimal solution determination and final architecture selection. After the gradient descent method converges, observe the amplitude of parameter update |θ t+1 -θ t |, if the parameter convergence has become stable, the optimization can be stopped, the current optimal architecture parameter θ is output, and it is regarded as the local optimal solution. In the NAS search process, the final network architecture is selected through performance evaluation on the validation set, the performance indicators of all candidate architectures are recorded, and finally the architecture suitable for the task requirements is selected.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements a neural architecture search method for adaptive multi-domain image segmentation tasks as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Cell image segmentation method based on particle swarm neural network

    CN106952275A

  • Defect classification method based on improved particle swarm wavelet neural network

    CN109934810A