Neural network architecture search method based on existing knowledge reasoning generation structure and coefficient

Through the neural network architecture search method based on existing knowledge, the hierarchical component structure and Bayesian theorem are used to reason, new structures and parameters are generated, and weights are generated through normal distribution functions, the problems of high computing resource consumption and low search efficiency in traditional methods are solved, and efficient neural network architecture search is achieved.

CN119940484APending Publication Date: 2025-05-06XIDIAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510008734.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When traditional neural network architecture search methods deal with large data sets and complex tasks, the computing resource consumption is high, the search efficiency is low, and the effective guidance mechanism is lacking, resulting in improper path selection and increasing computing cost and time consumption.

Method used

A neural network architecture search method based on the inference based on existing knowledge is adopted to generate structures and coefficients. By constructing a hierarchical component structure, simplifying it, inference is carried out, combining Bayesian theorem for inference, calculating the β value in the new structure, optimizing the neural network architecture search model, and generating weights for each layer of network through a normal distribution function to accelerate model convergence.

Benefits of technology

It significantly reduces resource consumption during the training process, improves search efficiency, reduces computational costs, ensures model performance and training efficiency, reduces dependence on random initialization, and improves the stability and global optimization of the architectural search process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940484A_ABST
    Figure CN119940484A_ABST
Patent Text Reader

Abstract

The invention discloses a neural network architecture search method based on an existing knowledge reasoning generation structure and coefficient, and the method comprises the following steps: 1, extracting part information according to the features of a data set and reference expert knowledge, and constructing a hierarchical part structure; step 2, simplifying the hierarchical part structure; step 3, selecting a part of components in the simplified and hierarchical component structure, training through a neural network architecture, and obtaining related structure and path information; 4, according to the path information, reasoning is conducted in combination with a hierarchical component structure and the Bayesian theorem, and a beta value in a new structure is obtained through calculation; and 5, through a normal distribution function, generating a corresponding weight for each beta value in each layer of network, and completing architecture search in the neural network structure search model. According to the method, the calculation consumption is reduced, and the improvement of the search efficiency is paid attention to.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of neural architecture search, and specifically relates to a neural network architecture search method for generating structures and coefficients based on existing knowledge reasoning. Background Art

[0002] When facing the problem of hyperparameter tuning, most studies of traditional machine learning models transform the structural design problem of the model into a hyperparameter optimization problem because the structure of the model is relatively simple and the number of hyperparameters is limited. These optimization problems are mainly solved by black-box optimization methods, such as evolutionary algorithms, Bayesian optimization, and reinforcement learning. The core idea of ​​the black-box optimization method is to regard the performance of the model as a function of the input hyperparameters, and to optimize the output performance of the objective function by continuously optimizing these hyperparameters. However, when the scale of the model is further expanded and the depth, width, and complexity of the network increase significantly, the number of hyperparameters also increases, making it difficult for traditional black-box optimization methods to meet actual needs in terms of efficiency and effect.

[0003] The introduction of NAS technology is precisely to solve this problem. NAS searches for the best performing model from a set of possible network architectures in an automated way. With the development of technology, NAS algorithms have made continuous progress in search efficiency and optimization accuracy. However, no matter what search optimization method is used, most NAS methods cannot avoid the process of model training and performance evaluation. This process usually involves training and verifying the model, and the computational cost is very high. In addition, different NAS methods differ mainly in module design and operation splicing methods. Some studies have attempted to further reduce the computational overhead of the model through methods such as pruning, distillation, or sharing parameters, but they still face great challenges in practical applications.

[0004] Traditional NAS methods usually require a lot of computing resources for architecture search. Especially when the search space is large, searching for network architectures often requires training tens of thousands of candidate models. Each search requires complete training and evaluation of the candidate architecture, which not only takes a lot of time, but also consumes a lot of computing resources. With the expansion of the scale of neural networks, especially the use of multi-layer network structures and complex models in deep learning, the required training time and computing resources are growing exponentially. Therefore, when dealing with large data sets and complex tasks, NAS algorithms usually need to rely on large-scale GPU clusters, which places very high demands on computing resources, making them difficult to apply in some environments with weaker computing power.

[0005] Traditional NAS methods usually use strategies such as random initialization or reinforcement learning for architecture search. These methods have relatively low search efficiency and large uncertainty. Random initialization makes it impossible for the search process to effectively utilize existing knowledge, resulting in a large number of invalid paths being repeatedly explored, further increasing the calculation time. In reinforcement learning, although the search efficiency can be improved by training the controller to predict the parameters of the network architecture, each architecture evaluation needs to be achieved through model training, and the computational overhead of this process is also huge. In addition, due to the lack of an effective guidance mechanism, the NAS algorithm may explore along paths with poor performance during the search process, wasting a lot of time and computing resources. In many cases, the search process of the network architecture lacks pertinence and guidance, resulting in poorly performing architectures being repeatedly trained and evaluated, further slowing down the search speed.

[0006] Many NAS methods start the search by randomly initializing the parameters at the beginning. Such randomness greatly increases the uncertainty of the training results. In the training process of neural networks, due to limited resources and limited training rounds, random initialization may lead to local optimal solutions in the search process and difficulty in converging quickly. In some cases, despite multiple rounds of training, the performance of the model is limited or even stagnant. Especially when the training time is long and the amount of data is large, the impact of random initialization may be more obvious, resulting in the training process failing to fully realize its potential. In addition, as the network structure and parameter space increase, the convergence of the optimization process becomes more complicated. The random initialization of NAS may make the optimization algorithm unable to find the best path in some complex architecture spaces, or it may take too long to converge, thereby slowing down the speed of the entire architecture search.

[0007] In the architecture search phase, NAS technology needs to explore multiple paths. Each path represents a different module or operation combination in the network architecture, which ultimately determines the structure of the model. Existing NAS methods often lack an effective guidance mechanism when selecting paths, resulting in the selected path being not the optimal path during training. In the retraining phase, further training is performed based on the path selected in the architecture search phase. Therefore, improper path selection will directly lead to a decrease in the performance of the network architecture and unsatisfactory model training results. More importantly, this inefficient path selection often leads to more training iterations and evaluations, further increasing computational costs and time consumption. This problem is particularly prominent when facing complex tasks and large-scale data. As the architecture space increases, the search space becomes larger, and path selection becomes more complex, while traditional search strategies are difficult to fully utilize existing knowledge and experience to improve search efficiency and accuracy. In the end, although path selection may increase model performance, the computational resource consumption and training time it brings are very high.

[0008] Specifically, as the model scale further expands, the number and types of modules that need to be searched in the NAS framework increase significantly. This not only significantly increases the computational cost and resource consumption, but also significantly prolongs the training time. Although model performance may be improved in such large-scale searches, the cost-effectiveness ratio is not high. Therefore, researchers began to focus on how to further improve search efficiency and reduce resource consumption while maintaining model performance. Summary of the invention

[0009] In order to overcome the deficiencies of the above-mentioned prior art, the purpose of the present invention is to provide a neural network architecture search method based on existing knowledge reasoning to generate structures and coefficients, while reducing computational consumption and focusing on improving search efficiency. By introducing the knowledge and experience gained from previous training, the search optimization strategy driven not only reduces excessive reliance on training data, but also further improves the prediction accuracy in the architecture evaluation stage. This method has universal applicability and is suitable for various deep learning tasks, including but not limited to image classification, target detection, natural language processing, etc.

[0010] In order to achieve the above object, the technical solution adopted by the present invention is:

[0011] A neural network architecture search method for generating structures and coefficients based on existing knowledge reasoning includes the following steps

[0012] Step 1: Based on the characteristics of the data set and referring to expert knowledge, component information is extracted and a hierarchical component structure is constructed;

[0013] Step 2: Simplify the hierarchical component structure;

[0014] Step 3: Select some components in the simplified hierarchical component structure and train them through the neural network architecture to obtain relevant structure and path information;

[0015] Step 4: Based on the path information, the hierarchical component structure and Bayesian theorem are combined for reasoning to calculate the β value in the new structure, and the β value is used to optimize the framework search model architecture in the neural network architecture search;

[0016] Step 5: Generate corresponding weights for each β value in each layer of the network through the normal distribution function. During the back propagation process, the β value will be updated along the gradient direction, thereby accelerating the convergence of the model and finally completing the architecture search in the neural network structure search model.

[0017] The step 1 decomposes the dataset image into important components and other subcomponents according to the root node, child nodes, and leaf nodes.

[0018] The specific steps of step 1 are:

[0019] (1) Analyze the characteristics of the data set: First, analyze the data set to identify the key characteristics that define the components;

[0020] (2) Refer to expert knowledge or domain-specific guidelines to understand which components are critical to the system or dataset;

[0021] (3) Based on expert knowledge, determine the main components that serve as the root nodes of the hierarchy;

[0022] (4) Identify child nodes and leaf nodes based on the data set and expert knowledge;

[0023] (5) Organize the components into a tree structure, with the root node at the top and child nodes and leaf nodes at the bottom. Each node represents a component and its relationship with other components, thus building a hierarchical component structure.

[0024] (6) Verify and optimize the structure: Verify the hierarchical component structure through domain knowledge and actual data to ensure its accuracy.

[0025] The step 2 is specifically as follows:

[0026] Simplify the resulting component tree structure. During the simplification process, only important structural information is retained, and identical or redundant parts are merged and omitted to output a streamlined hierarchical structure.

[0027] Merge: If two nodes are at the same level and their output features are very similar (such as the same spatial resolution and the same function), then merge the two nodes into one node;

[0028] Duplication of functions: If two components completely duplicate each other and they provide no additional information, merge them;

[0029] Depth similarity: If the depths of two nodes are very close and they belong to similar network modules, the two nodes are merged;

[0030] Deletion decision basis: In the hierarchy, if a node only makes invalid or duplicate transformations on the features of the previous layer, and deleting the node will not significantly affect the reasoning ability of the model, then the node is deleted.

[0031] In the initial stage of step 3, some components are trained through the neural network architecture, and each path from the root node to the leaf node in the hierarchical structure is selected for training. For each single component on the selected path, training is performed to obtain its β parameter information. In addition, all components on the path need to be jointly trained to extract the β parameter information of the entire path.

[0032] The step 3 is specifically as follows:

[0033] The starting point of the neural network architecture is a backbone structure consisting of two layers, and the spatial resolution of each layer decreases by a factor of 2;

[0034] The spatial resolution of each layer after the backbone structure will gradually decrease, and the spatial resolution of each layer can only be reduced by 2 times at most; each node represents a cell, and the spatial resolution of each layer will decrease by 2 times;

[0035] There are L layers in the neural network architecture, whose spatial resolution is unknown, and the maximum resolution is 4 times the downsampling, and the minimum resolution is 32 times the downsampling;

[0036] L layers means that each resolution layer contains L cells on the basis of reducing the resolution to 4 times;

[0037] The first layer after the backbone structure downsamples by 4 or 8 times, and so on, the second layer downsamples by 4, 8 or 16 times, with the ultimate goal of finding the best path in the network.

[0038] Each different Cell represents a different resolution. The current Cell receives input from three cells of different resolutions in the upper layer and obtains the output through weighted summation. The formula is expressed as:

[0039]

[0040] Where s represents the resolution multiple, s = 4, 8, 16, 32; a represents the number of network layers a = 1, 2, ..., L; H_l represents the feature map of the lth layer, H represents the spatial feature map, and l is the layer number; β_{l,s→s2}, β_{l,s→s}, β_{l,2s→s} are learned scalars (weights) that control the contribution of different resolution levels to the output of the current layer. These parameters are normalized to ensure that their sum is 1; Cell(·) represents the function or operation applied to the input, and the parameter α may control the parameters of the operation;

[0041] The parameter β is the weight of the resolution change path in the network space, which requires softmax, that is

[0042]

[0043] After multiple rounds of training, the model path of the current component is decoded.

[0044] The step 4 is specifically as follows:

[0045] The new β parameters are obtained by inferring the β parameters of each component according to the Bayesian formula;

[0046] Step (1): Bayes’ theorem describes how to update the probability through prior information and data. The formula is as follows:

[0047]

[0048] P(θ|D) is the posterior probability, which represents the probability distribution of parameter θ after observing data D;

[0049] P(D|θ) is the likelihood function, which indicates the probability of data D appearing under given parameters θ;

[0050] P(θ) is the prior probability, which represents the preliminary distribution of parameter θ in the absence of observed data;

[0051] P(D) is the marginal likelihood or evidence, which represents the overall probability of the data D.

[0052] Step (2): During the network architecture optimization and reasoning process, Bayesian theorem is used to infer the β parameters and optimize the path;

[0053] Step (3): Set the reasoning between nodes at the same level, obtain the relevant probabilities through training: P(V1), P(V1V2) and P(root|V1V2), and then use Bayes’ theorem to infer P(root|V2);

[0054] Calculate P(root|V2), that is, the conditional probability of the root node root given V2

[0055]

[0056] P(root|V1,V2) is the conditional probability of the root node given V1 and V2;

[0057] P(V1,V2) is the joint probability of V1 and V2;

[0058] P(V1) and P(V2) are the prior probabilities of V1 and V2 respectively;

[0059] Therefore, Bayes’ theorem updates the β parameters by calculating the conditional probabilities between each node;

[0060] Step (4): For node paths at different levels, infer the probability between nodes at each level in turn to obtain new β parameters;

[0061] Set the conditional probabilities such as P(root|V3) and P(V3|C1) obtained through training, calculate the probability of the path through Bayes' theorem, and obtain the new β parameter;

[0062]

[0063] P(root|V3) is the conditional probability of the root node root given V3;

[0064] P(V3) is the prior probability of V3;

[0065] P(V3|C1) is the conditional probability of V3 given C1;

[0066] P(C1) is the prior probability of C1;

[0067] Calculate P(root|C1) and then update and optimize the β parameter;

[0068] Step (5): Through the above reasoning process, the new β parameters are inferred from the prior information and training data using the Bayesian formula;

[0069] Step (6): By continuously reasoning and updating the β parameters, the optimal path can be selected among different paths. The updated β parameters eventually decode the new network space structure and the optimal path, thereby completing the optimization of the network architecture.

[0070] The step 5 is specifically as follows:

[0071] Through the normal distribution function, the corresponding weight w is generated for the β value in each layer of the network. The β value represents the transition probability of different resolutions. By adding the weight parameter w, the gradient descent process during training is ensured to be faster and more stable, so that the network can converge quickly and improve the performance of the overall model.

[0072] Specifically:

[0073] First, normalize the β value of a column and then plot its corresponding normal distribution;

[0074] Next, the corresponding coordinates of each β value on the normal distribution curve are obtained, and the corresponding normal distribution function value is extracted therefrom as w, that is, the weight value.

[0075] Beneficial effects of the present invention:

[0076] The present invention generates network structure and coefficients by introducing an inference mechanism, and generates new structures and coefficients only through partial training, rather than comprehensive training of all architectures. This optimization greatly reduces resource consumption during training, avoids the high computational overhead in traditional NAS methods, and effectively improves computational efficiency.

[0077] The present invention combines hierarchical information knowledge with the reasoning of architecture parameter β based on the parameters and path information obtained from existing training, analyzes the relationship between data sets using existing training results, selects paths through knowledge guidance, intervenes in the initial β parameter value, obtains new parameter values ​​as initial values ​​through calculation, avoids the exploration of invalid paths due to random initialization, and makes the search process more efficient. Randomness introduces great uncertainty and may cause trouble with local optimal solutions. This measure of the present invention effectively reduces the reliance on random initialization, thereby improving the stability of the architecture search process, reducing the reliance on local optimal solutions during training, and ensuring the global optimality of model search.

[0078] The present invention accelerates the process of path selection and model convergence by assigning weight values ​​to corresponding parameters, making the search process more efficient and being able to find the optimal architecture in a shorter time, thereby significantly shortening the training time.

[0079] In summary, the present invention significantly improves the computing resource consumption, search efficiency, random initialization uncertainty, and path selection accuracy in traditional NAS technology by inferring the generated structure and coefficients, guiding the reasoning path with knowledge, and introducing the weight parameter w, thereby reducing the computing cost while improving the performance and training efficiency of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 It is the overall system framework diagram of the present invention.

[0081] Figure 2 This is a component hierarchy diagram of the present invention.

[0082] Figure 3 This is a simplified hierarchical structure diagram of the present invention.

[0083] Figure 4 Schematic diagram of the network structure (search space) of the present invention.

[0084] Figure 5 It is a schematic diagram of the relationship between cells in different network layers in the network structure of the present invention.

[0085] Figure 6 This is a schematic diagram of the path searched after decoding by the present invention.

[0086] Figure 7 A normal distribution function diagram is generated for the present invention. DETAILED DESCRIPTION

[0087] The present invention will be further described in detail below in conjunction with the accompanying drawings.

[0088] like Figure 1As shown, the core structure of the neural network architecture search method for generating structures and coefficients based on existing knowledge reasoning is divided into six modules: a module for extracting a hierarchical structure of image components of a data set using expert knowledge, a module for simplifying hierarchical structures, a module for training node components, a module for obtaining paths, a module for generating new parameters and structures by reasoning, and a module for generating weights. Each module is interconnected to form the overall framework of the method for generating structures and parameters by reasoning proposed by the present invention.

[0089] Module for extracting component hierarchical structure of dataset images using expert knowledge: The task of this module is to extract component information and construct a hierarchical component structure based on the characteristics of the dataset and reference expert knowledge. Figure 2 In the figure, we can see that the hierarchical structure includes root nodes, child nodes, and leaf nodes. For example, important components such as decks and shells are used as root nodes, and other nodes are used as child nodes or leaf nodes to form a tree-like hierarchical structure.

[0090] The following is a more specific method and steps to extract component information:

[0091] (1) Analyze dataset features: First, the dataset is analyzed in detail to identify the key features that define the components. Depending on the type of dataset, these features can be geometric, visual, or semantic.

[0092] (2) Reference to expert knowledge: Reference expert knowledge or domain-specific guidelines to understand which components are critical to the system or dataset. For example, key structural components such as decks and hulls.

[0093] (3) Identify the root node: Based on expert knowledge, determine the main components that serve as the root nodes of the hierarchy. For example, components such as 'deck' and 'hull' can serve as root nodes.

[0094] (4) Define child nodes and leaf nodes: Based on data and expert knowledge, identify child nodes and leaf nodes. They are usually smaller and more complex components that depend on the root node or intermediate nodes.

[0095] (5) Build a hierarchy: Organize the components into a tree structure, with the root node at the top and child nodes and leaf nodes at the bottom. Each node represents a component and its relationship with other components.

[0096] (6) Verify and optimize the structure: Verify the constructed hierarchical structure through domain knowledge and actual data to ensure its accuracy. If necessary, adjust the structure to reflect the actual relationships and dependencies between components.

[0097] Hierarchical structure simplification module: For large data sets, the extraction of hierarchical structures often contains multiple smaller hierarchical structures. In order to facilitate subsequent reasoning calculations, these structures need to be simplified, retaining only the main structural information, and merging and omitting the same structures. Figure 3 Shows a simplified hierarchy.

[0098] Definition of smaller hierarchies: Smaller hierarchies mean that for each object, the component structure exposed by the same object at different viewing angles is always limited. For example, from one angle, a horse may only contain a horse head and a horse foot; but from another angle, a horse contains a horse head, a horse foot, and a horse body. The component information that an object can reveal from different viewing angles is called a "smaller hierarchy". A smaller hierarchy is formed by summarizing the component hierarchy of an object from all viewing angles and retaining key information.

[0099] The root nodes usually represent the most important and core parts, such as the main parts of an object (e.g., the horse's head, horse's feet, etc.) or the top level of the hierarchy. Under different viewing angles, these root nodes remain the same or change slightly, but in many cases, they are retained as the core information of an object.

[0100] Subnodes and leaf nodes are extensions of the root node, representing more detailed parts or subcomponents. For each object, the subnodes and leaf nodes may change under different viewing angles (for example, more parts of a horse's body can be seen from different angles, etc.). Therefore, subnodes and leaf nodes belong to smaller hierarchies, especially when they change under different viewing angles.

[0101] In a hierarchy, a root node may be preserved under all views, while child nodes and leaf nodes belong to smaller hierarchies because their definitions may change depending on the view of the object.

[0102] Node component training and path acquisition module: In the initial stage, some components need to be trained to obtain relevant knowledge about the structure and path. This process provides a basis for subsequent reasoning and path selection. By training single components or combined components, the corresponding network structure is decoded, and the corresponding path information is obtained through the β value.

[0103] Training of some components means selecting each path from the root node to the leaf node in the hierarchical structure for training. For each individual component on the selected path, training is performed to obtain its β parameter information. In addition, all components on the path need to be jointly trained to extract the β parameter information of the entire path. This dual training method ensures that both the individual information of each component and the combined information of the components on the entire path are taken into account, thereby optimizing the training effect.

[0104] In neural network architecture search, β is a key parameter in the framework search model. It can be obtained by directly saving β in the framework model training or by reading the β parameters in the trained framework model. β plays a vital role in the framework search process. It is used to select the optimal model path. When the framework model is trained for a single component or all components on a path in a hierarchical structure, the β parameters trained based on different data will contain the path information in the hierarchy. In this way, β can not only capture the relationship and dependency between components, but also reflect the organization and abstraction of each hierarchical structure in the model, thereby guiding the search process to find the optimal neural network architecture. This module provides the necessary β value, path data, etc. for the subsequent reasoning process.

[0105] Reasoning to generate new parameters and structure module: Based on the existing path information, the hierarchical structure and Bayesian theorem are combined for reasoning to calculate the β value in the new structure. These β values ​​will be used to obtain new network structures and paths, thereby optimizing the model architecture.

[0106] Weight generation module: Generate corresponding weights for each β value in each layer of the network through the normal distribution function. During the back propagation process, the β value will be updated along the gradient direction, thereby accelerating the convergence of the model.

[0107] The various modules work closely together to promote the generation of model architecture and parameters in order to improve search efficiency and model performance.

[0108] Principle

[0109] This invention analyzes the hierarchical relationship of image components in the data set, obtains parameter information and path structure of some components through training, and obtains new parameters and structures by inference, thereby reducing the resource consumption required for training, improving search efficiency, optimizing path selection, and enhancing model performance.

[0110] The node component training and path acquisition module is specifically:

[0111] In dense image prediction, the high resolution of the image needs to be maintained, so the following two principles need to be followed:

[0112] 1) The spatial resolution of the next layer may be expanded or reduced by a factor of two, or it may remain unchanged;

[0113] 2) The minimum resolution is 32 times downsampling.

[0114] Based on these two principles, the network structure is constructed. The starting point of the network structure is a "trunk" structure consisting of two layers. The two layers refer to the starting part of the network, that is, the backbone structure. The backbone structure consists of two layers, and the spatial resolution of each layer decreases by 2 times. Specifically, the role of these two layers is to define the preliminary feature extraction stage at the beginning of the network and set the foundation for subsequent network layers. These two layers form the foundation of the network and define the spatial resolution and downsampling rules of subsequent layers. The spatial resolution of each layer after the backbone structure will gradually decrease, and the spatial resolution of each layer can only be reduced by 2 times at most. For example, the first layer can only downsample 4 times or 8 times, and the second layer can choose 4 times, 8 times or 16 times.

[0115] Each node represents a cell, and the spatial resolution of each layer will decrease by 2 times. Next, there are L layers in the network. The two starting layers provide a starting point and preliminary setting of the spatial resolution for the subsequent layers of the network, and also provide a clear starting point for the optimal path search of the network. It should be noted that the "L layers" here does not mean that the resolution is directly reduced by L layers. In fact, the resolution is reduced by up to 4 times. The so-called "L layers" means that on the basis of the resolution being reduced to 4 times, each resolution layer contains L cells, which determines the structure of the network framework search process.

[0116] Its spatial resolution is unknown, and the maximum resolution is 4 times downsampling, and the minimum resolution is 32 times downsampling; since the spatial resolution of each layer differs by at most 2 times, the first layer after the backbone structure can only be downsampled by 4 times or 8 times, and so on, the second layer can only be downsampled by 4 times, 8 times, or 16 times. This can lead to possible network structures, such as Figure 4 The ultimate goal is to find the best path in this network.

[0117] Since the resolution changes exponentially, the tensor size between different layers may be different, as shown in 0. In this structure, each different Cell represents a different resolution. The current Cell (shown in red nodes) can receive input from three cells of different resolutions in the upper layer (shown in green nodes), and the output is obtained by weighted summation. The formula is expressed as:

[0118]

[0119] Among them, s represents the resolution multiple, s = 4, 8, 16, 32; a represents the number of network layers a = 1, 2, ..., L; H_l represents the feature map of the lth layer, H represents the spatial feature map, and l is the layer number; β_{l,s→s2}, β_{l,s→s}, β_{l,2s→s} are learned scalars (weights) that control the contribution of different resolution levels to the output of the current layer. These parameters are normalized to ensure that their sum is 1; Cell(·) represents the function or operation applied to the input, usually involving convolution or other feature extraction operations, and the parameter α may control the parameters of the operation.

[0120] The parameter β (the β in the above “obtaining the corresponding path information through the β value”) is the weight of the resolution change path in the network space, which requires softmax, that is

[0121]

[0122] After multiple rounds of training, the model path of the current component is decoded.

[0123] The decoding process generally involves the following steps:

[0124] (1) Extracting architectures from the search space: The search space is defined as a discrete set of different types of neural network layers, hyperparameter settings, etc. Each element in β corresponds to a choice in the search space.

[0125] (2) Mapping to specific operations: Once the value of β is obtained, it needs to be mapped to actual network layers, convolution types, activation functions, and other operations in some way.

[0126] (3) Constructing a specific network: Through this mapping relationship, the decoded β directly determines the specific structure of the network. For example, the decoded network may contain several convolutional layers, pooling layers, skip connections, etc. Their specific configuration (such as the number of layers, size, etc.) is determined by the parameter value in β.

[0127] The β parameter is a compressed representation of the neural network architecture, and the decoding process converts it into a specific network design and hyperparameter settings. In this way, deep learning models can be automatically searched and optimized in a huge search space, especially in image segmentation tasks, providing an efficient and automated architecture design process.

[0128] At this time, the β value is interpreted as the "transition probability" between different "states" (spatial resolution) and different "time steps" (number of network layers). In order to find the optimal path, the Viterbi algorithm is used to decode the path, so as to find the path with the highest probability in the search space, such as Figure 6 shown.

[0129] Key steps in the Viterbi decoding process

[0130] The search for the framework model is modeled through a hidden Markov model, where each "state" corresponds to a layer in the network architecture, or an important choice in the architecture (such as convolution type, pooling operation, stride, etc.). β is the encoding of these choices, and the goal of the Viterbi algorithm is to decode the most likely network architecture from the encoding β.

[0131] (1) Define hidden states and observation sequences

[0132] Hidden states: Hidden states usually represent a specific configuration of a layer in the architecture or a choice of a hyperparameter. For example, a hidden state might be the kernel size or number of channels of a specific convolutional layer.

[0133] Observations: Observations are data inputs to the model during training, or features related to architecture optimization. In some cases, observations may also represent specific parameters selected from the search space.

[0134] (2) Initialization probability and state transfer matrix

[0135] Initialization: The Viterbi algorithm starts decoding from a certain initial state. For each element in β, the Viterbi algorithm determines the starting point of the network architecture based on the initial probability distribution.

[0136] Transition Probabilities: The state transition matrix represents the probability of transitioning from one state to another. In neural architecture search, this usually involves the relationship between different layers or the dependencies between hyperparameters. For example, if the first convolutional layer chooses a kernel of a certain size, the choice of the second convolutional layer may be affected by this.

[0137] (3) Recursive calculation of the optimal path The Viterbi algorithm uses dynamic programming to gradually calculate the optimal path at each moment. For each moment t and each possible hidden state i, the Viterbi algorithm selects an optimal path and calculates the cumulative probability of the path. The specific process is as follows: For each possible β[i], calculate the probability of the optimal path from the previous state to the current state. Through recursive calculation, the optimal path is selected and the state is updated.

[0138] (4) Decoding the optimal architecture

[0139] Once the Viterbi algorithm has processed all states and moments, it will output the most likely hidden state sequence. This sequence is the decoded optimal network architecture.

[0140] Ultimately, the encoding in β will be mapped to the actual configuration of the network architecture. For example, the decoded result may tell you what convolution kernel to use, the number of layers, the number of nodes per layer, the activation function and other hyperparameters.

[0141] Through this process, the β parameters obtained from the training of a certain component and the decoding path are obtained, laying the foundation for the subsequent reasoning and coefficient generation.

[0142] The reasoning to generate new parameters and structure modules are specifically:

[0143] The new parameters generated by reasoning are the new β parameters obtained by reasoning the β parameters of each component in the hierarchical module according to the Bayesian formula;

[0144] Step (1): Basic formula of Yes's theorem

[0145] Bayes' theorem describes how to update probabilities using prior information and data. The formula is as follows:

[0146]

[0147] P(θ|D) is the posterior probability, which represents the probability distribution of parameter θ after observing data D;

[0148] P(D|θ) is the likelihood function, which indicates the probability of data D appearing under given parameters θ;

[0149] P(θ) is the prior probability, which represents the preliminary distribution of parameter θ in the absence of observed data;

[0150] P(D) is the marginal likelihood or evidence, which represents the overall probability of the data D.

[0151] Step (2): Application of Bayesian formula in reasoning

[0152] During network architecture optimization and inference, Bayesian theorem is used to infer the β parameters and optimize the paths. Here is how to calculate new β parameters between nodes at different levels using Bayesian theorem.

[0153] Step (3): Formula 1: Inference of new β parameters

[0154] It is assumed that reasoning is performed between nodes at the same level (for example, V1 and V2), and the relevant probabilities are obtained through training: P(V1), P(V1V2), and P(root|V1V2), and then P(root|V2) is inferred by applying Bayes’ theorem;

[0155] Formula derivation:

[0156] By using Bayes’ theorem, we can calculate P(root|V2), which is the conditional probability of the root node given V2.

[0157]

[0158] P(root|V1,V2) is the conditional probability of the root node given V1 and V2;

[0159] P(V1,V2) is the joint probability of V1 and V2;

[0160] P(V1) and P(V2) are the prior probabilities of V1 and V2 respectively.

[0161] Therefore, Bayes' theorem updates the β parameters by calculating the conditional probabilities between each node.

[0162] Step (4): Formula 2: Path reasoning between nodes at different levels

[0163] For node paths at different levels (e.g., root→V3→C1), the probabilities between nodes at each level are inferred in turn to obtain new β parameters.

[0164] Formula derivation: Set the conditional probabilities such as P(root|V3) and P(V3|C1) obtained through training, calculate the probability of the path through Bayes' theorem, and obtain the new β parameter.

[0165]

[0166] P(root|V3) is the conditional probability of the root node root given V3.

[0167] P(V3) is the prior probability of V3.

[0168] P(V3|C1) is the conditional probability of V3 given C1.

[0169] P(C1) is the prior probability of C1.

[0170] Calculate P(root|C1) and then update and optimize the β parameters.

[0171] Step (5): New β parameter generation

[0172] Through the above reasoning process, new β parameters are inferred from prior information and training data through the Bayesian formula. These new β parameters reflect the dependencies and conditional probabilities between various nodes and levels.

[0173] Step (6): Decoding new paths and network structures

[0174] By continuously reasoning and updating the β parameters, the network can select the optimal path between different paths. These updated β parameters eventually decode the new network space structure and the optimal path, thus completing the optimization of the network architecture.

[0175] According to 0, for nodes at the same level (such as V1 V2), reasoning is performed in combination with the root node root, and the following P(V1), P(V1V2), P(root|V1V2) are obtained through training:

[0176] ①

[0177] ②

[0178] ③

[0179] P(V1): Prior probability of node V1.

[0180] P(V1V2): The joint probability that nodes V1 and V2 exist at the same time.

[0181] P(root|V1V2): The conditional probability of the root node (root) given V1 and V2.

[0182] P(root|V2) is obtained by derivation of formula 1

[0183]

[0184] P(root|V1,V2): represents the conditional probability of the root node root given V1 and V2.

[0185] P(V1,V2): represents the joint probability of V1 and V2.

[0186] P(V1) and P(V2): the individual prior probabilities of V1 and V2 respectively

[0187] For nodes at different levels, that is, paths between hierarchical structures (such as root-V3-C1), the probability of a single component P(V1), the probability between two layers of nodes P(root|V1), and P(V1|C1) are obtained through training:

[0188] P(V1): Prior probability of a single component V1.

[0189] P(root|V1): Conditional probability of the root node (root) given V1.

[0190] P(V1|C1): Conditional probability of V1 given C1.

[0191] ①

[0192] ②

[0193] ③

[0194] Then use formula 2 to obtain P(root|C1).

[0195]

[0196] P(root|V3): The conditional probability of the root node (root) given V3.

[0197] P(V3): Prior probability of V3.

[0198] P(V3|C1): conditional probability of V3 given C1.

[0199] P(C1): prior probability of C1.

[0200] Through the above process, we can obtain the new β parameter, namely ( and. ) can also obtain new network space structures and decode new paths, thereby completing the search for structures.

[0201] The weight generation module is specifically:

[0202] Assuming that the network model includes 12 layers of network, the β parameter can be simplified to the following table:

[0203]

[0204] In this process, each column represents a layer of the network. A column is extracted from it, and the network model is expected to converge as quickly as possible. To this end, the larger β value (the larger β value selected by the decoding step) is further increased, and the smaller β value is further reduced to promote faster convergence. Therefore, a weight parameter w is introduced for each β value (the weight parameter w is introduced on the original β value so that the larger β value is further increased and the smaller β value is further reduced).

[0205] w is obtained indirectly through normal distribution. First, the β value of a column is normalized, and then its corresponding normal distribution is plotted (such as Figure 7 ). Next, the corresponding coordinates of each β value on the normal distribution curve are obtained, and the corresponding normal distribution function value is extracted from it as w, that is, the weight value. A larger β value will correspond to a larger w value, so that during back propagation, the gradient descent will also be larger, thereby achieving an accelerated descent effect, allowing the model to converge faster.

[0206] When generating the weight coefficient w, the core idea is to hope that larger β values ​​will get larger weights, while smaller β values ​​should correspond to smaller weights. The current solution uses the normal distribution function value to generate weights, but other generation methods can also be considered. For example, the slope of the normal distribution function value can be used as the weight value, thereby giving larger β values ​​a greater weight and further strengthening their influence in back propagation.

[0207] In addition, we can also consider using the reciprocal of the difference in β values ​​at the same position in the two layers of the network as the weight value. Since the range of β values ​​is (0,1), the difference between the two also falls within the range of (0,1), and taking its reciprocal will result in a value greater than 1. This method can effectively amplify the gap between β values, so that during back propagation, a larger difference can accelerate the propagation of the gradient, thereby helping the model converge faster.

[0208] The action relationships between modules are connected through the flow of knowledge and information, as described below:

[0209] The module extracts the component hierarchy of the dataset image based on the characteristics of the given dataset and combines the knowledge of domain experts to extract the information of each component in the image, and then builds a hierarchical component structure. This process is the starting point of the entire system because it provides basic component information and hierarchical relationships for subsequent reasoning and optimization. In actual operation, the dataset image will be decomposed into different components. For example, important components such as decks and shells will serve as root nodes, and other sub-components under them will serve as child nodes or leaf nodes, forming a tree structure (such as Figure 2 The construction of this structure is the basis for the subsequent modules to be able to effectively process and optimize.

[0210] The simplified hierarchy module simplifies the component tree structure formed by simplification. During the simplification process, only important structural information is retained;

[0211] 1. Define important structural information: Root nodes, key intermediate nodes, and leaf nodes are important structural information because they directly affect the final decision path. For some nodes, especially those located deep in the hierarchy, if they do not directly contribute to the final result, they may be regarded as redundant and can be deleted.

[0212] 2. The basis for determining deletion and merging;

[0213] Merge: Hierarchical consistency: If two nodes are at the same level and their output features are very similar (such as the same spatial resolution, the same function), then the two nodes can be merged into one node. For example, two different nodes both represent the same image region or part, and they provide similar feature information, and can be merged into a new node with weighted features.

[0214] Functionality duplication: If the functions of two components are completely repeated and they do not provide additional information, you can consider merging them. For example, if two components represent different perspectives of the same area, merging them can improve computational efficiency while retaining key feature information.

[0215] Depth similarity: If the depths of two nodes are very close and they belong to similar network modules (such as convolutional layers, pooling layers, etc.), the two nodes may be merged. For example, if the feature maps output by two adjacent convolutional layers are very similar and do not provide significantly different information, the two nodes can be merged.

[0216] Deletion decision basis: In a hierarchical structure, if a node only makes an invalid or repeated transformation on the features of the previous layer (such as an irrelevant pooling layer or downsampling layer), and deleting the node will not significantly affect the reasoning ability of the model, then the node can be deleted. For example, nodes with very small changes or irrelevant activations (such as repeated feature transformations) can be deleted. For some feature layers, especially those with simple feature maps or those with little impact on the final reasoning results, deletion can be considered. For example, in some neural networks, the feature maps of some intermediate layers have little effect on the final output, so they can be deleted. )Merge and omit identical or redundant parts to output a more streamlined hierarchy (such as Figure 3 This provides more concise and efficient structural information for the subsequent reasoning process.

[0217] Based on the simplified component hierarchy, the system needs to train some single-node components and combined node components in order to obtain relevant knowledge about the structure and path. The training process provides the necessary knowledge support for subsequent modules by training the principle of path acquisition, and lays the data foundation for the path acquisition and reasoning generation modules. Specifically, during the training process, the system provides important references for subsequent path acquisition and model optimization by learning the characteristics and patterns of node components. The training of node components helps the system capture key information that helps architecture search by analyzing the relationships and interactions between different components. This module outputs the model parameter values ​​obtained through training, as well as the decoded path information of components or combined components.

[0218] The module for inferring and generating new parameters and structures uses the model parameter values ​​and path information obtained in the node single component training module mentioned above, and generates new coefficients through inference. The path acquisition module obtains new network structures and parameters through calculation, and then further decodes new paths. The key task of this module is to obtain the path data of each layer in the network through the β value based on the parameter information obtained through training. By training single components or combined components, the system can select the appropriate path for each layer of the network, thereby inferring and generating new structures and parameters and decoding new paths.

[0219] After completing path reasoning and architecture generation, the last step is the weight generation module. In this module, the system generates corresponding weights w for the β values ​​in each layer of the network through the normal distribution function. These weights help the β values ​​to be updated along the gradient direction during the back propagation process, thereby accelerating the convergence of the model. Specifically, the β value represents the transition probability of different resolutions. By adding the weight parameter w, the gradient descent process during training is ensured to be faster and more stable, so that the network can converge quickly and improve the performance of the overall model. Since the weight of the β value directly affects the path selection and model optimization speed, this module plays a vital role in accelerating network training and improving search efficiency.

[0220] Combination Figure 1 , the relationship between each module in the whole system is closely linked and cooperates with each other. The present invention can significantly improve the efficiency of neural network architecture search, reduce the consumption of computing resources, while maintaining high model performance, and promote the automation process of neural network architecture design. The present invention proposes a method based on hierarchy and Bayesian theorem to generate new structures and parameters through trained parameters and paths. Specifically, the system first uses the hierarchy and Bayesian theorem to derive a new network architecture and its parameters, and generate a new path. When subsequent training is performed, the parameters used are no longer randomly initialized, but generated based on reasoning. This process is guided by the knowledge of the hierarchy, so that the relationship between the components can provide a clear guiding direction for the training process, thereby effectively improving the training efficiency and model performance.

[0221] The present invention normalizes the β value of each layer of the network, and then indirectly generates the weight parameter w through the normal distribution function. The specific operation is: first, the position of each β value is mapped through the normal distribution curve, and then the corresponding normal distribution function value is extracted as the weight w of the β value. These weight parameters will be used to adjust the amplitude of gradient update during the network training process, thereby accelerating the convergence of the model and optimizing the network architecture.

Claims

1. A neural network architecture search method for generating structures and coefficients based on existing knowledge reasoning, characterized in that: The following steps are included Step 1: Based on the characteristics of the data set and referring to expert knowledge, component information is extracted and a hierarchical component structure is constructed; Step 2: Simplify the hierarchical component structure; Step 3: Select some components in the simplified hierarchical component structure and train them through the neural network architecture to obtain relevant structure and path information; Step 4: Based on the path information, the hierarchical component structure and Bayesian theorem are combined for reasoning to calculate the β value in the new structure, and the β value is used to optimize the framework search model architecture in the neural network architecture search; Step 5: Generate corresponding weights for each β value in each layer of the network through the normal distribution function. During the back propagation process, the β value will be updated along the gradient direction, thereby accelerating the convergence of the model and finally completing the architecture search in the neural network structure search model.

2. A neural network architecture search method for generating structures and coefficients based on existing knowledge reasoning according to claim 1, characterized in that: The step 1 decomposes the dataset image into important components and other subcomponents according to the root node, child nodes, and leaf nodes.

3. A neural network architecture search method for generating structures and coefficients based on existing knowledge reasoning according to claim 2, characterized in that: The specific steps of step 1 are: (1) First, analyze the data set to identify the key features that define the components; (2) Refer to expert knowledge or domain-specific guidelines to understand which components are critical to the system or dataset; (3) Based on expert knowledge, determine the main components that serve as the root nodes of the hierarchy; (4) Identify child nodes and leaf nodes based on the data set and expert knowledge; (5) Organize the components into a tree structure, with the root node at the top and child nodes and leaf nodes at the bottom. Each node represents a component and its relationship with other components, thus building a hierarchical component structure. (6) Verify the hierarchical component structure through domain knowledge and actual data to ensure its accuracy.

4. A neural network architecture search method for generating structures and coefficients based on existing knowledge reasoning according to claim 3, characterized in that: The step 2 is specifically as follows: Simplify the resulting component tree structure, retain only important structural information, merge and omit identical or redundant parts, and output a streamlined hierarchical structure; Merge: If two nodes are at the same level and their output features are very similar, then merge the two nodes into one node; Duplication of functions: If two components completely duplicate each other and they provide no additional information, merge them; Depth similarity: If the depths of two nodes are very close and they belong to similar network modules, the two nodes are merged; Deletion decision basis: In the hierarchy, if a node only makes invalid or duplicate transformations on the features of the previous layer, and deleting the node will not significantly affect the reasoning ability of the model, then the node is deleted.

5. The neural network architecture search method for generating structures and coefficients based on existing knowledge reasoning according to claim 3, characterized in that: In the initial stage of step 3, some components are trained through the neural network architecture, and each path from the root node to the leaf node in the hierarchical structure is selected for training. For each single component on the selected path, training is performed to obtain its β parameter information. In addition, all components on the path need to be jointly trained to extract the β parameter information of the entire path.

6. A neural network architecture search method for generating structures and coefficients based on existing knowledge reasoning according to claim 5, characterized in that: The step 3 is specifically as follows: The starting point of the neural network architecture is a backbone structure consisting of two layers, and the spatial resolution of each layer decreases by a factor of 2; The spatial resolution of each layer after the backbone structure will gradually decrease, and the spatial resolution of each layer can only be reduced by 2 times at most; each node represents a cell, and the spatial resolution of each layer will decrease by 2 times; There are L layers in the neural network architecture, whose spatial resolution is unknown, and the maximum resolution is 4 times the downsampling, and the minimum resolution is 32 times the downsampling; L layers means that each resolution layer contains L cells on the basis of reducing the resolution to 4 times; The first layer after the backbone structure downsamples by 4 or 8 times, and so on, the second layer downsamples by 4, 8 or 16 times, with the ultimate goal of finding the best path in the network.

7. A neural network architecture search method for generating structures and coefficients based on existing knowledge reasoning according to claim 6, characterized in that: Each different Cell represents a different resolution. The current Cell receives input from three cells of different resolutions in the upper layer and obtains the output through weighted summation. The formula is expressed as: Where s represents the resolution multiple, s = 4, 8, 16, 32; a represents the number of network layers a = 1, 2, …, L; H_l represents the feature map of the lth layer, H represents the spatial feature map, and l is the layer number; β_{l,s→s2}, β_{l,s→s}, β_{l,2s→s} are learned scalars that control the contribution of different resolution levels to the output of the current layer. These parameters are normalized to ensure that their sum is 1; Cell(·) represents the function or operation applied to the input, and the parameter α may control the parameters of the operation; The parameter β is the weight of the resolution change path in the network space, which requires softmax, that is After multiple rounds of training, the model path of the current component is decoded.

8. A neural network architecture search method for generating structures and coefficients based on existing knowledge reasoning according to claim 7, characterized in that: The step 4 is specifically as follows: The new β parameters are obtained by inferring the β parameters of each component according to the Bayesian formula; Step (1): Bayes’ theorem describes how to update the probability through prior information and data. The formula is as follows: P(θ|D) is the posterior probability, which represents the probability distribution of parameter θ after observing data D; P(D|θ) is the likelihood function, which indicates the probability of data D appearing under given parameters θ; P(θ) is the prior probability, which represents the preliminary distribution of parameter θ in the absence of observed data; P(D) is the marginal likelihood or evidence, which represents the overall probability of data D; Step (2): During the network architecture optimization and reasoning process, Bayesian theorem is used to infer the β parameters and optimize the path; Step (3): Set the reasoning between nodes at the same level, obtain the relevant probabilities through training: P(V1), P(V1V2) and P(root|V1V2), and then use Bayes’ theorem to infer P(root|V2); Calculate P(root|V2), that is, the conditional probability of the root node root given V2 P(root|V1,V2) is the conditional probability of the root node given V1 and V2; P(V1,V2) is the joint probability of V1 and V2; P(V1) and P(V2) are the prior probabilities of V1 and V2 respectively; Bayesian theorem updates the β parameter by calculating the conditional probability between each node; Step (4): For node paths at different levels, infer the probability between nodes at each level in turn to obtain new β parameters; Set the conditional probabilities such as P(root|V3) and P(V3|C1) obtained through training, calculate the probability of the path through Bayes' theorem, and obtain the new β parameter; P(root|V3) is the conditional probability of the root node root given V3; P(V3) is the prior probability of V3; P(V3|C1) is the conditional probability of V3 given C1; P(C1) is the prior probability of C1; Calculate P(root|C1) and then update and optimize the β parameter; Step (5): Through the above reasoning process, the new β parameters are inferred from the prior information and training data using the Bayesian formula; Step (6): By continuously reasoning and updating the β parameters, the optimal path can be selected among different paths. The updated β parameters eventually decode the new network space structure and the optimal path, thereby completing the optimization of the network architecture.

9. A neural network architecture search method for generating structures and coefficients based on existing knowledge reasoning according to claim 8, characterized in that: The step 5 is specifically as follows: Through the normal distribution function, the corresponding weight w is generated for the β value in each layer of the network. The β value represents the transition probability of different resolutions. By adding the weight parameter w, the gradient descent process during training is ensured to be faster and more stable, so that the network can converge quickly and improve the performance of the overall model.

10. A neural network architecture search method for generating structures and coefficients based on existing knowledge reasoning according to claim 9, characterized in that: Specifically: First, normalize the β value of a column and then plot its corresponding normal distribution; Next, the corresponding coordinates of each β value on the normal distribution curve are obtained, and the corresponding normal distribution function value is extracted therefrom as w, that is, the weight value.

Citation Information

Cited By

  • Dynamic reasoning path optimization method based on neural architecture search

    CN120542571A