Method and system for constructing three-dimensional state space model search architecture with feedback

CN119991420AActive Publication Date: 2025-05-13INNER MONGOLIA UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510064511.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13
Estimated Expiration
2045-01-15

Smart Images

  • Figure CN119991420A_ABST
    Figure CN119991420A_ABST
Patent Text Reader

Abstract

The invention belongs to but not limited to the technical field of machine learning, and particularly relates to a three-dimensional state space model search architecture construction method and system with feedback, and the method comprises the steps: defining an operation space; constructing a common search unit and a down-sampling search unit in the operation space, and stacking the common search unit and the down-sampling search unit to form a super network; searching through a high-precision double-layer optimization method to obtain each edge in the directed acyclic graph of the super network and the architecture weight of each operation in the edge; the product of the architecture weight and the architecture weight of the corresponding operation serves as the final weight, the operation with the maximum final weight in each edge is obtained, the corresponding final weight serves as the final weight of the edge, the common search unit and the down-sampling search unit are stacked and updated according to the results of the edge and the operation, and the final model architecture is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to but is not limited to the field of machine learning technology, and in particular relates to a method and system for constructing a three-dimensional state space model search architecture with feedback. Background Art

[0002] Automated machine learning (AutoML) has lowered professional barriers and improved efficiency by automating processes such as data preprocessing, feature selection, model selection, and hyperparameter optimization. It is widely used in finance, healthcare, and other fields. Its core concept is to standardize and automate the machine learning development process, use search strategies to optimize model performance, and reduce human intervention. Neural architecture search (NAS) focuses on automatically designing the optimal neural network architecture through algorithms, and uses technologies such as reinforcement learning to efficiently find the best model in the network structure space to meet the needs of deep learning for efficient architecture. NAS is driving deep learning to develop in a more complex direction and is expected to become a dominant technology in the future. At present, NAS technology is still limited to the fields of convolutional neural networks and Transformers, and has not yet been combined with state-space deep learning models, forming a technical gap, which provides a new direction for the innovative application of AutoML and NAS.

[0003] In the industrial application of existing technologies, the three-dimensional state space model search architecture faces many technical problems, which limit its application effect and efficiency in complex tasks. Specifically, there are mainly the following technical problems:

[0004] 1. Limited receptive field and insufficient feature extraction:

[0005] Existing state-space models usually have limited receptive fields, which means that they cannot fully capture global and multi-scale feature information when processing high-dimensional image data. This limitation makes the model perform poorly in complex image segmentation and recognition tasks, and it is difficult to accurately capture the subtle structures and global relationships in the image.

[0006] 2. Low search efficiency and high consumption of computing resources:

[0007] Traditional search methods often rely on single-layer or shallow search units when constructing and optimizing state space models, resulting in a large search space and a complex optimization process. The high computational cost and long search process limit the real-time and scalability of the model in actual industrial applications, making it difficult to meet the needs of large-scale data processing.

[0008] 3. Insufficient fusion of multi-scale information and poor model generalization ability:

[0009] In existing architectures, the fusion of multi-scale information usually relies on simple splicing or weighted averaging methods, which cannot effectively integrate feature information at different scales. This deficiency results in poor generalization ability of the model when facing images with features of different scales, making it difficult to adapt to diverse application scenarios.

[0010] 4. Lack of feedback mechanism limits model performance improvement:

[0011] Traditional state-space models lack an effective feedback mechanism and cannot dynamically adjust and optimize the model structure based on the results of the model output. This lack of feedback limits the model's ability to adapt to complex tasks and makes it difficult to continuously optimize and improve performance during training.

[0012] 5. Difficulty in expanding the operating space and lack of adaptability:

[0013] Existing methods often require manual adjustment or redesign of search strategies when expanding the operation space, which lacks flexibility and automation. This inconvenience limits the rapid adaptation and optimization of the model when facing different task requirements, and reduces the overall efficiency and practicality of the system. Summary of the invention

[0014] In view of the problems existing in the prior art, the present invention provides a method and system for constructing a three-dimensional state space model search architecture with feedback.

[0015] The present invention is implemented as follows: a method for constructing a three-dimensional state space model search architecture with feedback, comprising:

[0016] Step 1: Define the operation space, add average pooling, maximum pooling and skip connection operations to the operation space; add a downsampling module before the state space to make the state space model obtain a double receptive field, add upsampling after the state space model to fuse the image information and the state space model output of the unsampled original receptive field, and upgrade the state space model to a three-dimensional state space model through anatomical scanning; introduce output feedback to improve the performance of the state space model.

[0017] Step 2: Construct a normal search unit and a downsampling search unit in the operation space, stack the normal search unit and the downsampling search unit, and add an image feature embedding module before the image is sent to the stacked network to form a supernet; the normal search unit and the downsampling search unit are both directed acyclic graphs with the same structure; the directed acyclic graph contains multiple nodes, each node represents a feature graph, and the edges between the nodes are mixed operations composed of all operations in the operation space. The direction of the edge represents the direction of the information flow, and each edge and each operation in the edge has corresponding architecture parameters;

[0018] Step 3: Search and obtain the architectural weight of each edge and each operation in the directed acyclic graph of the supernet through a high-precision two-layer optimization method;

[0019] Step 4: The final weight is obtained by multiplying the architecture weight and the corresponding operation architecture weight as the final weight, and the operation with the largest final weight in each edge is obtained. The corresponding final weight is used as the final weight of the edge. For units with input edges greater than two, only the first two edges with the largest final weight are retained. The results of the edge and operation are stacked and updated to the required number of layers for the normal search unit and the downsampling search unit to obtain the final model architecture. It should be noted that if the operation space needs to be expanded, new search results can be obtained by starting from step 1 again according to the receptive field tendency of the update unit, and the updated normal search unit and downsampling search unit are stacked to the required number of layers to obtain the final model architecture.

[0020] Further, step 3 specifically includes:

[0021] The architecture weights of each edge and each operation in the directed acyclic graph of the supernet are obtained through a high-precision two-layer optimization method search, including:

[0022] The training set images are sent to the supernet, the architecture parameters are fixed, the cross entropy loss is calculated according to the label after forward propagation, and the gradient is calculated by back propagation; the weight parameters of the supernet are optimized according to the gradient direction, and the weight parameters are repeated until convergence to obtain the final architecture weight of each edge and each operation in the directed acyclic graph of the supernet, that is, the architecture parameters and weight parameters are alternately optimized until convergence to obtain the final architecture weight of each edge and each operation in the directed acyclic graph of the supernet.

[0023] Furthermore, the training set images are fed into the supernet, including:

[0024] The training set images are randomly cropped, flipped, and normalized before being sent to the supernet.

[0025] Another object of the present invention is to provide a three-dimensional state space model search architecture construction system with feedback for implementing the three-dimensional state space model search architecture construction method with feedback, comprising:

[0026] Operation space definition module: used to define the operation space;

[0027] A supernet forming module, used for constructing a normal search unit and a downsampling search unit in an operation space, and stacking the normal search unit and the downsampling search unit to form a supernet;

[0028] An architecture weight search module is used to search and obtain the architecture weight of each edge and each operation in the directed acyclic graph of the supernet through a high-precision two-layer optimization method;

[0029] The final architecture acquisition module is used to obtain the operation with the largest final weight in each edge by multiplying the architecture weight and the corresponding operation architecture weight as the final weight, and use the corresponding final weight as the final weight of the edge. The normal search unit and the downsampling search unit are updated with the results of the edge and operation stacked to obtain the final model architecture.

[0030] Another object of the present invention is to provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for constructing a three-dimensional state space model search architecture with feedback.

[0031] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to execute the steps of the method for constructing a three-dimensional state space model search architecture with feedback.

[0032] Another object of the present invention is to provide an information data processing terminal, which includes the three-dimensional state space model search architecture construction system with feedback.

[0033] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:

[0034] First, the latest state-space deep learning model is the third basic model alongside the convolutional neural network model and the Transformer model. Currently, there is no technology that combines neural network architecture search with the state-space model, and there is still a blank in automated machine learning.

[0035] In common visual sequence classification models, Tranformer is used as the backbone network to represent complex patterns in visual data. Compared with CNN, ViT generally shows stronger modeling ability because it combines the self-attention mechanism to achieve global receptive field and dynamic prediction weight parameters. However, self-attention also brings a side effect, that is, the complexity of self-attention grows quadratically with the increase of input size. SSM has the advantage of being able to efficiently model visual data while retaining the self-attention mechanism relative to ViT. SSM is essentially a sequence model, and the first well-known work is the Kalman filter. In recent years, the series of papers by Gu et al. are impressive, in which the HiPPO operator compresses continuous signals and discrete time series online by projecting to a polynomial basis, and gives the S4 model the ability to establish long-range dependencies by initializing the matrix A. Vim is the earliest representative visual Mamba model, which uses position embedding to integrate spatial information. Later, VMamba combined convolution and spatial attention mechanisms and introduced the Cross-Scan Module (CSM) to achieve 1D selective scanning in a two-dimensional image space with a global receptive field. However, these Mamba visual applications are not designed for 3D images, and the large number of convolutions makes the classification ability of SSM questionable.

[0036] Inspired by VMamba, we designed a 3D State Space Supergraph (3DSSS), which searches for 3DSSS through a differentiable architecture to find a universal visual backbone consisting of blocks composed of concise SSMs for efficient classification of 3D lung nodules. The effectiveness of 3DSSS in reducing the complexity of attention computation is largely due to the selective scanning mechanism present in the Mamba model, also known as selective SSM. Unlike traditional attention computation methods, which allow dense information routing in context, Mamba requires that each element in a one-dimensional array (e.g., a text sequence) obtains contextual knowledge only through compressed hidden states, thereby reducing quadratic complexity to linear.

[0037] The present invention designs a supernet structure of a linear time series model, aiming to search for a visual backbone network of the linear time series model. By constructing a supernet, the present invention can quickly explore the combination of different state space expressions, thereby improving the search efficiency of the model. This method not only saves computing resources, but also ensures that the optimal state space model is found to better adapt to the classification task of three-dimensional lung nodules.

[0038] In order to improve the accuracy and model stability in three-dimensional image classification, the present invention designs an anatomical scanning and output feedback loop. This design aims to comprehensively analyze the same data from different angles through multiple scanning paths, enhance the model's sensitivity to different features, and effectively capture the key features of three-dimensional images. Through the feedback mechanism, the model can self-adjust and optimize the output of each scanning path, strengthen the recognition of important information, and improve classification accuracy and stability. In addition, the state-space module uses a pure state-space expression, which significantly compresses redundant manual design and improves the processing efficiency and generalization ability of the system. This multi-angle analysis and feedback optimization strategy significantly improves the classification effect of the model in practical applications and enhances the ability to accurately classify three-dimensional images.

[0039] The present invention uses upsampling and downsampling techniques to introduce multi-scale feature fusion capabilities to the state-space model. This innovation enables the model to extract useful features from different scales and combine them for classification, thereby more comprehensively understanding the structural information of the three-dimensional image. This multi-scale feature fusion strategy effectively improves the classification accuracy of the model, enabling effective recognition of objects of different sizes and shapes.

[0040] Design In the research of this invention, a supernet training method based on dynamic elimination mechanism is proposed to optimize the performance of the pulmonary nodule classification model by adjusting the Points value. This method allows the elimination rules to be dynamically adjusted according to the performance of the operation at each training stage, so that the model can fully exert its potential in different rounds.

[0041] One or more technical solutions provided in the present invention have at least the following technical effects or advantages:

[0042] 1. Define the operation space. In order to improve the stability of the operation space and prevent the model from being over-parameterized, the present invention adds average pooling, maximum pooling and skip connection operations in the operation space. Since the state space model does not have multi-scale characteristics, the present invention doubles the receptive field of the state space model by adding a downsampling module before the state space, and adds upsampling after the state space model to fuse the image information and the state space model output of the unsampled original receptive field. The existing state space model is a two-dimensional model. The present invention upgrades the state space model to a three-dimensional state space model through anatomical scanning. The existing state space model is unstable. The present invention introduces output feedback to improve the performance of the state space model; then construct a normal search unit and a downsampling search unit in the operation space. Since the state space model has only a single receptive field, that is, it cannot change the size and number of feature maps, in order for the supernet to downsample normally, the present invention designs two downsampling methods, such as Figure 4As shown in (b)(d) and (c)(e), (b)(d) can make the supernet more capable of downsampling and compressing information, and (c)(e) simplifies the supernet to significantly reduce the training video memory but sacrifices some of the downsampling and compressing information capabilities. The general search unit and the downsampling search unit are stacked and the image feature embedding module is added before the image is sent to the stacked network to form the final supernet; then the architecture weight of each edge and each operation in the directed acyclic graph of the supernet is searched through a high-precision two-layer optimization method; finally, the product of the architecture weight and the corresponding operation architecture weight is used as the final weight to obtain the operation with the largest final weight in each edge, and the corresponding final weight is used as the final weight of the edge. The normal search unit and the downsampling search unit are stacked and updated with the results of the edge and operation until the required number of layers to obtain the final model architecture.

[0043] 2. The operation weighting function in the present invention uses sigmoid to ensure fair competition among operations.

[0044] 3. The dynamic elimination mechanism in the present invention speeds up model search and reduces the truncation error of model selection to increase the accuracy of search.

[0045] Second, the present invention effectively solves many difficult problems in three-dimensional medical image data processing by innovatively constructing a three-dimensional state space classification model and introducing neural network architecture search (NAS) technology. Traditional two-dimensional image processing methods are difficult to fully exploit the spatial features in three-dimensional data, and the present invention, by combining sequence models, can not only capture the complex structure in three-dimensional data, but also model dynamic changes. Especially in the analysis of tumor growth, neurological diseases and cardiovascular lesions, the present invention can provide more accurate classification and diagnosis results. Through automated NAS technology, the model can quickly adapt to different tasks and data sets, significantly reducing the cost of manual parameter adjustment, while improving the flexibility and scalability of the model to meet the needs of medical image analysis for high precision and high efficiency.

[0046] The technical solution of the present invention is expected to bring significant economic and social benefits to the medical industry. By deeply mining the feature fusion of three-dimensional images and text data, this technology can improve classification accuracy and diagnostic accuracy in multimodal data analysis, provide patients with more accurate personalized treatment plans, and reduce the risk of misdiagnosis and overtreatment. In addition, the automated classification capability significantly reduces the labor costs of medical institutions and accelerates the efficiency of medical image analysis, allowing doctors to devote more energy to the treatment of complex cases. At the same time, the application of this technology in drug research and development and clinical trials will improve the analysis efficiency of experimental data and the reliability of results, accelerate the drug development cycle, and bring new value growth points to the pharmaceutical industry.

[0047] From the perspective of commercial value, the present invention fills the gap in the precise analysis technology of three-dimensional medical images and provides an efficient solution for the medical image analysis market. In the future, the technology based on the present invention can be developed into a SaaS platform to provide low-cost, high-efficiency artificial intelligence diagnostic tools to medical institutions through cloud services, further lowering the threshold for the application of industry technology. At the same time, the technology also has the potential for cross-industry applications, such as intelligent manufacturing, autonomous driving, and robotic vision, and can expand market space in multiple fields. The development trend of globalization also gives the technology the prospect of entering the international market, cooperating with global medical institutions to jointly promote the innovation and upgrading of medical image analysis.

[0048] Third, the present invention overcomes the application limitations of the prior art in three-dimensional image processing due to huge video memory overhead, model instability and low search efficiency by constructing a three-dimensional state space model search architecture with feedback. The supernet designed by the present invention adopts a combination of ordinary units and downsampling units, and through the residual connection structure and high-precision two-layer optimization method, it significantly reduces parameter redundancy and video memory overhead, while improving the training stability of the model. Compared with traditional convolutional models, the present invention effectively balances the structural design of depth and width, avoids the performance bottleneck caused by uneven depth, and provides an efficient and scalable solution for complex three-dimensional image data processing.

[0049] In addition, the present invention further optimizes the operation space of the supernet by introducing pooling operations and partial channel weighting mechanisms, thereby improving the stability and efficiency of the search process. Combined with the optimized architecture weight adjustment strategy, the present invention achieves accurate construction and efficient search of the three-dimensional state space model, significantly improving the performance and accuracy of three-dimensional image analysis. Its technological progress is not only reflected in reducing hardware resource requirements, but also in promoting the practical application of three-dimensional image processing algorithms in the fields of medical imaging, industrial detection, etc., accelerating the technological upgrading of related industries. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a flow chart of a method for constructing a three-dimensional state space model search architecture with feedback provided by an embodiment of the present invention;

[0051] Figure 2 It is a principle diagram of a three-dimensional anatomical scanning state space model with feedback provided by an embodiment of the present invention;

[0052] Figure 3 It is a schematic diagram of the state space model provided by an embodiment of the present invention to achieve a multi-scale effect;

[0053] Figure 4 It is a schematic diagram of a three-dimensional state space supernet search architecture with feedback provided by an embodiment of the present invention;

[0054] Figure 5is a schematic diagram of a sampling operation provided by an embodiment of the present invention;

[0055] Figure 6 is a schematic diagram of an operation elimination rule provided by an embodiment of the present invention;

[0056] Figure 7 is a search result graph in the LUNA16 three-dimensional data set provided by an embodiment of the present invention;

[0057] Figure 8 It is a structural diagram of a three-dimensional state space model search architecture construction system with feedback provided by an embodiment of the present invention.

[0058] Fig. 9 It is a performance comparison chart of the state space model designed manually on the fifth validation set in the cross-validation experiment of the LUNA16 three-dimensional data set provided by an embodiment of the present invention and an advanced state space model designed manually.

[0059] Fig.10 It is a performance comparison chart of the anatomical scanning in the present invention based on the 3D_VMamba model and 4-way scanning on the 5th validation set in the LUNA16 three-dimensional data set cross-validation experiment.

[0060] Fig.11 It is a performance comparison chart of the 3DSSS model with and without feedback on the 5th validation set in the LUNA16 three-dimensional data set cross-validation experiment provided by an embodiment of the present invention.

[0061] Fig.12 The lung nodule classification application example provided in the embodiment of the present invention does not include a class activation mapping visualization function diagram.

[0062] Fig.13 It is a developed class activation mapping visualization function diagram of the lung nodule classification application example of searching for convolutional neural networks provided by an embodiment of the present invention.

[0063] Fig.14 This is an image of a malignant lung nodule that is correctly diagnosed according to an embodiment of the present invention. In the figure, P is the confidence level, and D is the diameter of the lung nodule, with the unit of diameter being millimeters.

[0064] Fig.15 This is an image of a benign lung nodule that is correctly diagnosed according to an embodiment of the present invention. In the figure, P is the confidence level, and D is the diameter of the lung nodule, with the unit of diameter being millimeters.

[0065] Fig.16 This is the confidence distribution of correctly classifying lung nodules on the 5th validation set in the LUNA16 three-dimensional data set cross-validation experiment provided by an embodiment of the present invention.

[0066] Fig.17The embodiment of the present invention provides (a) the distribution of 1004 lung nodule images as feature dimensionality reduction, and (b) the distribution of 1004 lung nodules after feature dimensionality reduction after SSM learning.

[0067] All search results that appear in this document have normal units in the upper part and downsampled units in the lower part.

[0068] Fig.18 The search space of VSS provided by the embodiment of the present invention is PRIMITIVES=['Skip_Connect', 'VSS_12', 'VSS_Double_12', 'VSS_4', 'VSS_Double_4'], and the search result diagram is shown.

[0069] Fig.19 This is a network result diagram obtained by searching EMBED_DIM_32->96 according to an embodiment of the present invention.

[0070] Fig. 20 This is a network result graph searched by EMBED_DIM_96 provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0071] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0072] See also Figure 1 The method for constructing a neural network search architecture of a high-precision two-layer optimization method provided by an embodiment of the present invention includes:

[0073] S101: define the operation space;

[0074] In order to improve the stability of the operation space and prevent the model from being over-parameterized, the present invention adds average pooling, maximum pooling and skip connection operations in the operation space. Since the state space model does not have multi-scale characteristics, the present invention adds a downsampling module before the state space so that the state space model obtains a double receptive field, and adds upsampling after the state space model so that the image information and the state space model output of the unsampled original receptive field are fused. The existing state space model is a two-dimensional model. The present invention upgrades the state space model to a three-dimensional state space model through anatomical scanning. The existing state space model is unstable. The present invention introduces output feedback to improve the performance of the state space model.

[0075] The specific description of this step includes:

[0076] The state-space model can be viewed as a linear time-invariant system, which consists of hidden states Enter Mapping to output response They are usually expressed as linear ordinary differential equations, as follows:

[0077] h′(t)=Ah(t)+Bx(t)

[0078] y(t)=Ch(t)+Dx(t)

[0079] The domain of the coefficient matrix is

[0080] In order to integrate into the deep model, the continuous-time state space model needs to be discretized in advance, which can be achieved by solving ordinary differential equations and then performing a simple discretization process. The result is shown below:

[0081]

[0082] y k =Ch k +Dx k

[0083] The expression of the coefficient matrix is

[0084] In order to increase the stability of the state space model, the present invention introduces a feedback mechanism into the state space model and maintains the performance of the parallel scanning of cub::BlockScan in Mamba.

[0085] The continuous state space model after adjustment in actual use of the present invention is:

[0086] h′(t)=Ah(t)+Bx(t)

[0087] y(t)=C(h(t)+h′(t))+Dx(t)

[0088] Derivation of the continuous state space model with output feedback:

[0089]

[0090] Where K is the feedback factor.

[0091] In fact, since the feedback coefficient and the state space model coefficient of the present invention are learnable parameters or are calculated from learnable parameters, the present invention simplifies the feedback into the following formula and experimentally achieves performance improvement.

[0092] h′(t)=KCAh(t)+KC(B+KD)x(t)

[0093] While the sequential nature of Mamba is well aligned with natural language processing tasks involving temporal data, it poses significant challenges when applied to visual data, as visual data is inherently non-sequential and contains spatial information (e.g., local texture and global structure). To address this problem, the S4ND model reformulates the state-space model using convolution operations and directly extends the kernel from one dimension to two dimensions via outer products. However, this modification prevents the dynamics of the weights. Therefore, the present invention chooses a selective scanning method for data processing and proposes an anatomical scanning module.

[0094] like Figure 2 As shown, data transmission through a three-dimensional anatomical scanning state space model with output feedback includes three steps: three-dimensional anatomical scanning, SSM block selective scanning with output feedback, and regeneration and reconstruction. The present invention refers to three-dimensional anatomical scanning and regeneration and reconstruction as 3DAS-OF (3DAnatomical Scanning for SSM with OutputFeedback). Given the input data, 3DAS-OF first expands the image patches into sequences along twelve different traversal paths, processes each patch sequence in parallel using a separate state space model with feedback, and then regenerates the reconstruction result sequence to form an output image. 3DAS-OF adopts complementary traversal paths from three anatomical planes, so that each pixel in the image can effectively integrate the information of all other pixels from different directions, thereby promoting the establishment of a global receptive field.

[0095] like Figure 2 The sampling interval Δ shown is determined by the input to achieve dynamic weights, so the present invention deletes the following in the supernet design: Figure 3 The right spatial attention branch of Mamba is shown. The visual image has been projected to a high dimension through the embedding layer of the supernet. Figure 3 The projection layer in Mamba is also deleted from the supernet by the present invention. Since the state space equation does not have multi-scale characteristics, the present invention introduces a combination of up and down sampling in the supernet so that the state space equation can obtain a double receptive field before the down sampling unit, thereby realizing multi-scale feature fusion.

[0096] like Figure 3 As shown, the present invention retains the activation function in the Mamba structure to add nonlinearity to the sampled image. Overall, the present invention tries to remove artificial designs with unknown effects and difficult to explain, hoping to search for a space state equation network with unknown parameters in the supernet, and use ordinary units and downsampling units as a new backbone network. At the same time, some doubts about the performance of pure space state equation networks are explained. Figure 3 Medium SSMONLY Figure 2 Medium SSMONLY.

[0097] S102: constructing a normal search unit and a downsampling search unit in the operation space. Since the state space model has only a single receptive field, that is, it cannot change the size and number of feature maps, in order for the supernet to downsample normally, the present invention designs two downsampling methods, such as Figure 4 As shown in (b)(d) and (c)(e), (b)(d) can make the supernet more capable of downsampling and compressing information, and (c)(e) simplifies the supernet, which significantly reduces the training video memory but sacrifices some of the downsampling and compressing information capabilities. The normal search unit and the downsampling search unit are stacked and the image feature embedding module is added before the image is sent to the stacked network to form a supernet; the normal search unit and the downsampling search unit are both directed acyclic graphs with the same structure; the directed acyclic graph contains multiple nodes, each node represents a feature map, and the edges between the nodes are mixed operations composed of all operations in the operation space. The direction of the edge represents the direction of the information flow, and each edge and each operation in the edge has corresponding architecture parameters;

[0098] The specific description of this step includes:

[0099] like Figure 4 The supernet designed by the present invention shown in (a) is composed of two types of units, one is a common unit and the other is a downsampling unit. The units before and after the downsampling unit are marked with F and B respectively, indicating that special attention should be paid to the design of the supernet. From (a), it can be clearly noted that the three-dimensional state space supernet of the present invention has only one common unit before the downsampling unit, and the last four units of the total 8 units are divided into common units. In the VMamba experiment, the present invention found that the model underwent three downsamplings but the model depth was similar to [2,2,27,2]. Increasing the depth in the third stage improves the model performance, rather than uniformly increasing the model depth, which is very consistent with the supernet design. Inspired by VMamba, the search results of a common unit usually include multiple SSMs, so only one common unit is placed before the downsampling unit to imitate the successful experience of VMamba. The feature channel of the state space model is not the same as the feature channel model of the convolution. The feature channel of the state space model represents the encoding of the image. This leads to the idea of ​​randomly extracting part of the channel and sending it to the state space model, which does not fit the definition of the model and may cause risks such as model degradation, instability or non-convergence. However, lung nodule classification is a three-dimensional image, which inevitably leads to huge video memory overhead, which brings obstacles to increasing the depth of the supernet. Increasing the depth only after the second downsampling of the supernet can improve the model performance, which coincides with the design of the three-dimensional state space model supernet. The three-dimensional image uses Figure 5 In (a), the number of model parameters will be reduced by 4 times if the feature channels are doubled by downsampling once, and the unit memory overhead after the second downsampling will be 16 times less than that of the unsampled memory. Figure 5Because the lung nodule image is relatively small and in order to realize multi-scale fusion of spatial state equation, the supernet unit stacking of the present invention only performs two downsamplings.

[0100] like Figure 4 As shown in (b) and (c), two supernet sampling strategies are shown in the supernet. The present invention adopts the sampling strategy in (c), and the downsampling unit does not actually perform downsampling, but is an independent unit that needs a special unit to enhance the model performance after downsampling. The advantage of this is that it saves the sampling module before the state space model. Figure 1 Comparing (e) and (d), I used (d) to save the sampling operations of 0, 1 to 2-5 in (e). However, the sampling strategy of simplifying the network in this invention will also weaken the ability of the downsampling unit. From the results, it can be seen that the independent change of the downsampling unit can make up for this weakening.

[0101] The supernet designed by the present invention retains the residual structure of the units in the usual supernet. The formula is as follows. The jump connection neither introduces additional parameters nor increases the computational complexity. Where x is the input, y is the output, W is the learnable parameter of the i layer, and F is the operator of the layer. Experiments show that the residual network is easier to optimize and can obtain accuracy from a considerable depth.

[0102] y=F(x,{W i})+x Figure 4 (b) It should be noted that the node 6 of the Normal cell (F) to the node 0 of the Normal cell (B) needs to be downsampled. In fact, the number of feature channels of node 6 is 4 times the number of feature channels of the unit input. Special processing is required before inputting node 0 or 1. Figure 4 (e) The operation initiated by nodes 0 and 1 of the downsampling unit will perform downsampling, but the state space model has no step length concept, so the present invention adds a downsampling branch in the state space model operation, such as Figure 4 (e) as shown.

[0103] like Figure 4 As shown in (f), the present invention designs the operation space of the edge in the unit. In addition to the state space model and residual operations, due to the instability of the state space model training process, a large number of state space model operations are introduced, which makes the supernet difficult to train and causes a large amount of parameter redundancy. The pooling operation can have a certain invariance to some changes in the input image by statistically summarizing the features in the local area, increasing the stability of the supernet training while reducing the redundancy of the supernet parameters, so the present invention introduces average pooling and maximum pooling. The double weighting of p and q in (f) makes the supernet search process more stable. The use of Sigmoid weighted zero operation only requires implicit expression, so the present invention does not add zero operation in the operation space.

[0104] In the operation space, a normal search unit and a downsampling search unit are constructed. The operations on the edges and edges of the normal search unit and the downsampling search unit all use a partially decoupled operation weighting function as shown in the following formula:

[0105]

[0106] Among them, α is the architecture parameter and Sigmoid(α) is the architecture weight.

[0107] In the supernet, the output feature map of the edge of the search unit is shown in the following formula:

[0108]

[0109] in, is the output of the channel connection of edge (i,j).

[0110] In a supernet, the node feature graph is the weighted sum of the input edge results, as shown in the following formula:

[0111]

[0112] in, is a schema parameter.

[0113] S103: searching and obtaining the architecture weight of each edge and each operation in the directed acyclic graph of the supernet by a high-precision double-layer optimization method;

[0114] This step is specifically described. The architecture weights of each edge and each operation in the directed acyclic graph of the supernet are obtained by searching through a high-precision two-layer optimization method, including:

[0115] The training set images are sent to the supernet, the architecture parameters are fixed, the cross entropy loss is calculated according to the label after forward propagation, and the gradient is calculated by back propagation; the weight parameters of the supernet are optimized according to the gradient direction, and the weight parameters are repeated until convergence to obtain the final architecture weight of each edge and each operation in the directed acyclic graph of the supernet, that is, the architecture parameters and weight parameters are alternately optimized until convergence to obtain the final architecture weight of each edge and each operation in the directed acyclic graph of the supernet.

[0116] Specifically, the training set images are fed into the supernet, including:

[0117] The training set images are randomly cropped, flipped, and normalized before being sent to the supernet.

[0118] S104: The final weight is obtained by multiplying the architecture weight and the corresponding operation architecture weight as the final weight, and the operation with the largest final weight in each edge is obtained, and the corresponding final weight is used as the final weight of the edge. For the input edge in the unit greater than two, only the first two edges with the largest final weight are retained, and the results of the edge and operation are stacked and updated to the required number of layers for the normal search unit and the downsampling search unit to obtain the final model architecture. It should be noted that if the operation space needs to be expanded, new search results can be obtained by starting from step 1 again according to the receptive field tendency of the update unit, and the updated normal search unit and downsampling search unit are stacked to the required number of layers to obtain the final model architecture.

[0119] Specifically, the gradient of the architecture parameters of the high-precision two-layer optimization method is an approximate numerical method, as shown in the following formula:

[0120]

[0121] in, is the symbol for finding the gradient, L val (w * ,α) is the validation loss, L train (w,α) is the training loss, α is the architecture parameter, w * is the optimal network weight, w is the model parameter, ξ, ε are the hyperparameters for adjusting the approximate formula,

[0122] High-precision two-level optimization method uses w * (α 0 )initialization, Fit the height w * (α), thereby increasing the accuracy of subsequent numerical calculations. Note that all previous two-layer optimization initializations use Kaiming initialization, while this paper proposes a numerical method for two-layer optimization w * (α 0 )initialization.

[0123] w * (α 0 )The purpose of initialization is to better fit w * (α), and the purpose of Kaiming initialization is to stabilize the variance of the model output. Kaiming initialization satisfies the following formula:

[0124]

[0125] Where l is the layer index, y is the input response, n is the number of connections to the response, and w is the model parameters.

[0126] L train and L valIt is the cross entropy loss function, but regularization is added and the regularization maintains the physical properties of the loss, as shown in the following formula:

[0127]

[0128] L total =L val (w * (γ),γ)+w 0-1 L 0-1

[0129] Among them, L 0-1 is the regularization function, γ is all the architectural parameters including edges and operations, and w 0-1 is a hyperparameter for adjusting regularization.

[0130] like Figure 6 (a) is an example of a unit selecting edges. The present invention selects the edge with the largest p in each edge as a candidate operation. Then nodes 2 to 5 only select the edge with the largest product of p and q (orange). The elimination round is performed when selecting the edge with the largest p in each edge as a candidate operation. Figure 5 As shown in (b), during the supernet double-layer optimization process, when the p value is less than points, it will be replaced by the Zero operation, that is, it will be eliminated.

[0131] Point is between 0 and 1. The Points value of the elimination match is very important because the operations in the supernet show different capabilities at different training stages. The purpose of the present invention is to enable the target network to have the highest indicators of lung nodule classification. However, just like the entrance examination, not all the highest scorers will become scientists. Points are used as elimination rules, and the number of points eliminated determines the final performance tendency of the model. For example, if Points is initialized to 0.5, operations that do not perform well in the early stages of the competition will be eliminated. Because potential opponents are eliminated, the pooling operation has a huge advantage. Points values ​​close to 0 are theoretically the best, but this is just an overfitting phenomenon in the "exam". In other words, in order to search quickly and the data itself is limited, the "exam" cannot reflect the ability of the model in real data. For example, DARTS only uses half of the data to train the supernet. The present invention assumes that the operation ability will reach a peak in the same round, and the ideal P value can eliminate relatively poor operations at the peak of the ability of all operations without causing the supernet convergence to be extremely degraded. The present invention uses a dynamic Points value, which is initialized to 0.2, and the Points are increased to 0.25 and 0.3 in the last two eliminations. In order to make the network convergence more stable, the present invention performs elimination every 5 rounds.

[0132] The supernet designed by the present invention retains the residual structure of the units in the usual supernet, and the jump connection neither introduces additional parameters nor increases the computational complexity. Experiments show that the residual network is easier to optimize and can obtain accuracy from a considerable depth. The residual structure formula is as follows:

[0133] y=F(x,{W i})+x

[0134] like Figure 4 (b) It should be noted that the node 6 of the Normal cell (F) to the node 0 of the Normal cell (B) needs to be downsampled. In fact, the number of feature channels of node 6 is 4 times the number of feature channels of the unit input. Special processing is required before inputting node 0 or 1. Figure 4 (e) The operation initiated by nodes 0 and 1 of the downsampling unit will perform downsampling, but the state space model has no step length concept, so the present invention adds a downsampling branch in the state space model operation, such as Figure 4 (e) as shown.

[0135] In order to further reduce the memory overhead of the supernet, we also designed another simple sampling method such as Figure 4 (c) shows that we modify the downsampling operation of nodes 0 and 1 of the downsampling unit to perform the downsampling operation before inputting nodes 0 and 1, thus saving 56 downsampling operations to only two downsampling operations. If two downsampling units are used, 108 downsampling operations are saved, and so on. The strict supernet design is shown in Figure 4 (e) as shown.

[0136] Two different downsampling supernets Figure 4 (c) and Figure 4 The search space of (d) is as follows Figure 4 (f) The space shown.

[0137] The operation space of the edge in the unit is designed. In addition to the state space model and residual operations, the state space model training process is unstable. The introduction of a large number of state space model operations makes the supernet difficult to train and causes a large number of parameter redundancies. The pooling operation can statistically summarize the features in the local area, and has a certain invariance to some changes in the input image, which increases the stability of the supernet training and reduces the redundancy of the supernet parameters. Therefore, the present invention introduces average pooling and maximum pooling. Figure 4 The double weighting of p and q in (f) makes the supernet search process more stable. Using Sigmoid weighted zero operation only requires implicit expression, so the invention does not add zero operation in the operation space.

[0138] In the operation space, a normal search unit and a downsampling search unit are constructed. The operations on the edges and edges of the normal search unit and the downsampling search unit all use a partially decoupled operation weighting function as shown in the following formula:

[0139]

[0140] Among them, α is the architecture parameter and Sigmoid(α) is the architecture weight.

[0141] In the supernet, the output feature map of the edge of the search unit is shown in the following formula:

[0142]

[0143] in, is the output of the channel connection of edge (i,j).

[0144] In a supernet, the node feature graph is the weighted sum of the input edge results, as shown in the following formula:

[0145]

[0146] in, is a schema parameter.

[0147] The present invention searches for results in the LUNA16 three-dimensional data set, such as Figure 7 As shown, the upper picture is a normal unit searched in the LUNA16 3D data set, and the lower picture is a downsampled unit.

[0148] The final results of the model architecture in the LUNA16 3D dataset are compared with other advanced results as shown in Table 1. The best result and the second best result in each column are shown in bold and underline format, respectively. “-” means not mentioned.

[0149]

[0150] The searched model was compared with several existing state-of-the-art methods, and it was clear that the highest classification accuracy was achieved with the least model parameters. The high accuracy and F1 score indicate that 3DSSS has achieved a good balance between sensitivity and specificity.

[0151] The process of applying the neural network architecture in the example is:

[0152] Step 1: Extract lung nodules from lung window images;

[0153] This step is specifically described to extract lung nodules from the lung window image, including:

[0154] The center of the lung window image is set to -300, the width of the lung window image is adjusted to [-1200; 600], the contrast of the lung window image is linearly transformed to [0; 1], the lung image is obtained by segmentation annotation, and the lung nodules are extracted using the lung nodule annotation file, whose size is 36x36x36.

[0155] Step 2: Input the lung nodules into the built supernet model training model to obtain the model weights and architecture parameters;

[0156] This step is specifically described. The lung nodules are input into the built supernet model training model to obtain the model weights and architecture parameters, including:

[0157] The network search space, i.e., the edges in each unit in the supernet, is set. The present invention selects the state space model of the original receptive field anatomical scan and 4-way scan and the state space model of the double receptive field anatomical scan and 4-way scan; 3x3x3 maximum pooling and average pooling; identity mapping.

[0158] Build the normal search unit and downsampling search unit in the supernet, the topology is as follows Figure 7 As shown, the operations in both edges and in-edges of the present invention use sigmoid weighting.

[0159] Eight units are stacked as a supernet, where the second and fourth units are downsampling search units and the rest are normal search units.

[0160] Initialize the model parameters and architecture parameters, fix the architecture parameters, train the model parameters until convergence, and select the number of rounds with the highest accuracy in the validation set as a hyperparameter for the two-layer optimization.

[0161] Use the high-precision two-layer optimization method to train the supernet, initialize the model parameters and architecture parameters, fix the architecture parameters, and initialize the model parameters of the two-layer optimization: the number of rounds of training model parameters is the number of rounds in which the highest accuracy first appears in the validation set determined in the previous step, and start the two-layer optimization to alternately train the model parameters and architecture parameters.

[0162] Step 3: Select the model architecture through the model weights and architecture parameters to obtain the final normal unit and downsampling unit;

[0163] This step is explained in detail. The model architecture is selected through the model weight and architecture parameters to obtain the normal unit and the downsampling unit, including:

[0164] The architecture weight of the edge is multiplied by the architecture weight of the operation to get a metric parameter. The operation with the largest metric parameter of each edge is selected as the representative of the edge. When comparing the weights of the representatives of the node input edges, each node only retains the input edges with the top two metric parameters. Finally, the normal unit and downsampling unit after the search decision are obtained.

[0165] Step 4: Through the common units and downsampling units, stack them to the specified layer to input lung nodules, train the model weights, and obtain the final classification model.

[0166] This step is explained in detail. Through the common units and downsampling units, they are stacked to the specified layer to input lung nodules, train the model weights, and obtain the final classification model, including:

[0167] The common units and down-sampling units obtained in the previous step are stacked into 20 units, of which the second and fourth units are down-sampling units and the rest are common units. The model parameters are trained to obtain the final classification model for benign and malignant pulmonary nodules.

[0168] Step 5: Send the lung nodules obtained in the medical auxiliary diagnosis process to the benign and malignant lung nodule classification model obtained in step 4 to obtain the discrimination results and confidence levels as a criterion for assisting doctors in diagnosing lung cancer.

[0169] In summary, the embodiment of the present invention proposes a method for constructing a three-dimensional state space model search architecture with feedback, and the search results are used for 3D lung nodule classification. Compared with other advanced networks, 3DSSS achieves faster search speed and higher accuracy in this task. By designing 12 lung nodule scanning paths, the present invention realizes multi-angle analysis of different features and significantly enhances the diagnostic ability of the classification model. The stability of model training is increased by output feedback, which improves the final accuracy. In addition, combined with upsampling and downsampling techniques, the model introduces multi-scale feature fusion capabilities to fully capture the structural information of lung nodules, thereby improving classification accuracy. Finally, the supernet training method based on dynamic elimination mechanism proposed in the present invention further improves the model performance by optimizing the training process. These innovative methods jointly promote the development of three-dimensional lung nodule classification research and ensure the effectiveness and reliability of the model in practical applications. To the best of the knowledge of the present invention, this is the first 3D supernet of SSM. Experimental results on the LUNA16 benchmark dataset show that the model discovered by 3DSSS has excellent classification performance compared with existing methods.

[0170] Figure 4 (a) Uneven depth and Figure 4 The sampling strategy of (c) reduces a lot of video memory usage. The operation weighting function uses sigmoid to make the operations compete fairly. In addition, in the two-layer optimization, the model parameters are pre-trained to convergence, thereby improving the search accuracy of the two-layer optimization, which is reflected in the complete elimination of the problem of enrichment pooling operations in the search and refreshing the historical record of model accuracy.

[0171] An application embodiment of the present invention provides a computer device, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the method.

[0172] An application embodiment of the present invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the method.

[0173] An application embodiment of the present invention provides an information data processing terminal, which includes a system.

[0174] 1. Experiment

[0175] This study conducted a series of experiments to verify the proposed method and empirically analyzed the reasoning process. This paper will give the experimental settings and analyze the corresponding results below.

[0176] 1.1 Setup

[0177] Dataset and preprocessing

[0178] In the experiment, the present invention uses the settings of the LIDC-IDRI dataset and LUNA16. In particular, CT with slice thickness greater than 3 mm, inconsistent slice spacing or missing slices is removed from the LIDC-IDRI dataset. A total of 888 sequences with 1004 nodules are left, of which 450 nodules are positive. The LUNA16 dataset divides the 888 CT sequences into 10 subsets, numbered 0 to 9, which can be used for 10-fold cross validation. Accordingly, in each run, the present invention uses 9 subsets (containing approximately 900 samples) for training and the remaining 1 subset (containing approximately 100 samples) for testing. According to the settings in, the present invention evaluates the method of the present invention at 5 to 9 folds and reports the average performance.

[0179] In the data preprocessing stage, the present invention sets the center of the original data lung window to -600, adjusts the lung window width to [-1200; 600], and linearly transforms the image contrast to [0; 1]. The lung parenchyma is obtained using the segmentation annotation given by LUNA16, and finally the lung nodule annotation file is used to extract the lung nodules.

[0180] 1.2 Evaluation indicators

[0181] In order to evaluate the classification performance of lung nodules, the present invention uses four widely used indicators, accuracy, specificity, sensitivity and F1 score as shown in formulas (1)(2)(3)(5).

[0182] For the task of lung nodule classification, the sensitivity index is more important than other tasks. There are two types of errors in binary classification: false negatives and false positives. In extreme cases, assuming that the model has been able to completely miss positive nodules, doctors will pay less attention to non-positive nodules. Doctors will correct the false positives if they notice them, but this will lead to very fatal false negative errors, and positive nodules will be directly ignored, resulting in misdiagnosis. Therefore, the following index selects the F1 score instead of the AUC to select a model with as few false positives as possible.

[0183]

[0184] Sensitivity is also called the true positive rate, which indicates the ability of the classifier to judge positive examples as positive examples. Specificity is also called the true negative rate, which indicates the ability of the classifier to judge negative examples as negative examples. In the formula, TP, FN, FP, and TN are true positive, false negative, false positive, and true negative, respectively. The larger the values ​​of these criteria, the better the performance. The F1 score evaluates the trade-off between sensitivity and specificity. The F1 score hopes that all positive nodules can be detected. The harmonic mean is sensitive to small numbers. The sensitivity is between 0 and 1 as shown in formula (3) and the precision is between 0 and 1 as shown in formula (4), so this indicator uses the harmonic mean. The higher the F1 score, the better the performance.

[0185]

[0186] Implementation details

[0187] All experiments were performed on a GeForce RTX 4090 with 24GB of video memory. The hyperparameters of the final model were set to batch_size=4, epochs=600, WARMUP_EPOCHS=20, layers=20, BASE_LR=2e-4, MIN_LR=8e-5.

[0188] Searching for architectures on the Luna16 dataset

[0189] The search results of the present invention in the three-dimensional state space supernet are as follows: Figure 7 As shown. The feature compression between the common unit C_{k} and the unit realizes multi-scale fusion. In addition, the present invention searches for modules directly on Luna16, and there is no problem of applicable proxy data sets.

[0190] The search results are as follows Figure 7 As shown, k is the unit output. It can be seen that the unit has a residual design, that is, k-2 is part of the input. The large number of lines in the ordinary unit is not the result of degradation. The ordinary unit tends to choose a single-layer SSM_12 operation. The information is simultaneously fused through nodes 1 and 2. The 3-node realizes a jump connection structure. The fusion of 0 and 1, 2 information realizes a multi-scale effect.

[0191] Table 1 compares with existing methods on LUNA16 dataset for folds 5-9. Accuracy, Sensitivity, Specificity, F1 Score, and Parameters are denoted as Accu., Sens., Spec., F1 S., and Para.

[0192]

[0193] Note: The best and second best results in each column are highlighted in bold and underlined, respectively. A dash ("-") indicates that information is not mentioned.

[0194] First, the model of this study is compared with several existing advanced methods, including Multi-crop CNN, Nodule-level 2D CNN, Vanilla 3D CNN, ADNN, DeepLung, AE-DPN, NASLung and NAS-qa. The performance of the SSM network of this study is reported in Table 1.

[0195] Obviously, the 3DSSS network of the present invention achieves the highest accuracy and the minimum model parameters. The high accuracy and F1 score show that the 3DSSS network has achieved a good balance between sensitivity and specificity. In other words, the 3DSSS network can correctly classify most nodules, which will greatly reduce the burden on doctors. Although this study did not weigh the model parameters, the natural model parameters of the 3DSSS network are relatively small, which is also a major advantage of the model of the present invention.

[0196] exist Fig.14 and Fig.15 The present invention shows the validation results of malignant and benign pulmonary nodules in Luna16 set 5 in a cross-validation experiment. There are 92 pulmonary nodules in set 5, which can be found from Fig.14 and Fig.15 From the distribution of confidence, we can see that most of the lung nodules have a classification confidence of more than 0.9, which shows that the model has a strong classification performance. In order to fully demonstrate the performance of the model, the present invention extracts 12 lung nodules with different confidences. Among the selected lung nodules, 6 are malignant and 6 are benign. The confidence in the figure is the result of the two values ​​output by the model through the Softmax function and the maximum value. Confidence is a partial explanation of the black box model and a reference indicator for doctors. Fig.16The confidence distribution of correctly classified lung nodules on the 5th validation set in the LUNA16 three-dimensional data set cross-validation experiment provided by the embodiment of the present invention is shown. Through the display of confidence, the model's ability to classify benign and malignant lung nodules can be proved, and lung nodules with lower confidence can also attract the attention of doctors, further narrowing the range of lung nodules that doctors need to pay attention to, which is in line with the expectations of the present invention.

[0197] 1.3 Comparative Experiment and Analysis

[0198] Before designing the three-dimensional state space supernet, the present invention modifies the artificially designed excellent state space model into three dimensions to be suitable for the task of lung nodule classification. The present invention found in the experiment that the image scanning path of the input state space will greatly affect the final classification accuracy of the model. As shown in Table 2.

[0199] Table 2 Performance of the results of different search methods and the performance of the results after the search space is expanded.

[0200]

[0201] Table 3 Model hyperparameter settings

[0202]

[0203] ViM is one of the earliest Mamba visual models. In this paper, ViM is modified into 3D_ViM to classify lung nodules. After a large number of hyperparameters are debugged, DEPTHS=24, EMBED_DIM=192, STATE=16, and PATCH_SIZE=4 are the best results in the experiment, as shown in Table 2. The results show that the early two-way Mamba classification model has weak ability in lung nodule classification.

[0204] VMamba adds a scan path before the state space module. After the present invention changes VMamba to 3D_VMamba_4, the accuracy performance in the lung nodule classification task can be ranked among the excellent models. After a large number of hyperparameter debugging, DEPTHS = [2,2,20,2], EMBED_DIM = 96, STATE = 16, PATCH_SIZE = 2 is the best combination in the experiment, and the experimental results are shown in Table 2. Here, STATE = 16 performs better than 32. When STATE = 32, the model vibrates at a large learning rate and the model does not converge after the learning rate is reduced.

[0205] After controlling the variables, 3D_VMamba_4 was changed to 3D_VMamba_12 as shown in the table, and the anatomical scan was Figure 2As shown in Table 2, the overall performance of the model is improved. The hyperparameters of 3D_VMamba_12 are consistent with those of 3D_VMamba_4, with the only difference being STATE=32. This is not because the present invention has no control variables, but because the accuracy of the 3D_VMamba_4 model with STATE=32 fluctuates roughly between 71.73% and 80.43%. On the contrary, 3D_VMamba_12 can achieve 99% training accuracy with fast convergence at a large learning rate at STATE=32, indicating that 3D_VMamba_12 has relatively flat features near the extreme value and has strong stability. 3D_VMamba_4 and 3D_VMamba_12 demonstrate the superiority of the present invention in anatomical scanning.

[0206] Regarding the setting of hyperparameters, the present invention finds that the reason why the model will experience oscillation gradient explosion when the number of EMBED_DIM and STATE combinations is too large remains to be explained. In other aspects, a key parameter of 3D_VMamba_12 is PATCH_SIZE=2, which makes [1,32,32,32] encoded into [96,16,16,16]. This is relative to PATCH_SIZE=4, that is, [96,8,8,8], and the sequence length is doubled, which is more in line with the advantages of Mamba long sequence modeling. The hyperparameters in the supernet design of the present invention are mostly borrowed from the settings of 3D_VMamba_12. For example, in the supernet, two fewer units are placed in the first two stages, more units are placed in the third stage, PATCH_SIZE=2 and EMBED_DIM=96, etc.

[0207] Table 4. Comparison between VSS operation and SSMOnly in supernet search

[0208]

[0209] As shown in Table 3, the number of parameters of 3D_VMamba_12 reaches 64.74M. The present invention believes that the Δ in the discretized state space model given by mamba already has the effect of the attention mechanism. Compared with VSS in 3D_VMamba_12, the present invention designs a supernet multi-scale SSM structure to delete the 3D_VMamba_12 spatial attention mechanism, convolution lifting dimension, separable convolution, FFN, that is, the present invention only retains the SSM part, and introduces up and down sampling to make up for the problem of fixed SSM field of view. As shown in Table 4, only the SSM operation still performs well in the end. The search space of VSS in the table is PRIMITIVES = ['Skip_Connect', 'VSS_12', 'VSS_Double_12', 'VSS_4', 'VSS_Double_4'], and the search results are as follows Fig.18 ; Due to the limitation of RTX4090 video memory, the number of unit stacking here is 10 layers.

[0210] Table 5. Comparison of two sampling strategies in SuperNet.

[0211]

[0212] Table 6 Comparison of knockout strategies used in supernet search

[0213]

[0214] Sample1 imitates DARTS's solution of halving the size of feature maps and doubling their number, but SSM does not have the ability to change the size and number of feature maps. Therefore, when designing Sample1, the present invention simply loads the convolution before the SSM module to obtain sampling capability. In this way, the supernet needs to add 56 convolution operations, which is contrary to the idea of ​​optimizing VSS of the present invention. Therefore, the present invention designs Sample2 to perform sampling between cells so that 56 convolutions are compressed into 2 convolution operations. The relative downsampling learning ability may be weakened, but experiments show that it has no effect on accuracy. As shown in Table 5, Sample2 performs better.

[0215] Based on Sample2, we conducted an elimination round, which made the supernet closer to the final selected model to improve the search accuracy. After several rounds of elimination rounds, the network was simplified to reduce the search time. As shown in Table 6, the model searched after the introduction of the elimination round has better accuracy and faster speed.

[0216] Table 7 Effect of EMBED_DIM hyperparameters on network search results

[0217]

[0218] EMBED_DIM_32->96 means that the present invention sets EMBED_DIM in the supernet to 32, and uses EMBED_DIM=96 for the model training searched out. This continues the DARTS convolutional supernet search result training by increasing the dimension to reduce the supernet search time and improve the performance of the searched network. However, this idea does not perform well in the three-dimensional state space supernet. Based on a large number of preliminary experiments in Table 2, the present invention finds that there is a certain constraint relationship between EMBED_DIM and STATE, which cannot be explained by simply increasing the feature dimension by convolution. As shown in the results, if EMBED_DIM=96 is used directly to search the supernet, the supernet will select lines and pooling to simplify the network, which is why pooling must be added to the supernet design of the present invention. When EMBED_DIM is set to 32, the searched network tends to select a more complex model. When the searched network EMBED_DIM is set to 96, the convergence of the searched network training under a generally large learning rate is extremely unstable, and it is only possible to converge under a smaller learning rate. Unfortunately, in order to search for a better model, the present invention sets EMBED_DIM to 96. Although the search time is prolonged, the overall search accuracy is higher.

[0219] The network results of the search for EMBED_DIM_32->96 are as follows Fig.19 ; EMBED_DIM_96 search results are as follows Fig. 20 。

[0220] 1.4 Visualization of Classification Feature Distribution

[0221] Visualize the output of the penultimate layer in the SSM network. We use t-SNE to reduce each feature vector to 2 dimensions and Fig.17 Visualization. Fig.17 The orange (50% transparency) points are malignant nodules, and the blue (50% transparency) points are benign nodules. The present invention uses the scikit-learn toolbox in its implementation. Since the current research on class activation mapping has not been extended to the SSM model, the present invention uses t-SNE to visualize the feature distribution learned by the SSM network, and makes a visual explanation of the capabilities of the SSM network model based on the visualization of feature distribution. Fig.17 As shown in (a), the three-dimensional lung nodule image is directly projected into two dimensions. The benign and malignant lung nodules are not obvious, and it is even impossible to distinguish between benign and malignant lung nodules through projection. Fig.17As shown in (b), the features learned by the SSM network can be projected into low dimensions to effectively distinguish benign and malignant lung nodules, proving the effectiveness of SSM network learning. At the same time, the model was generated by the low 5-fold cross-validation of ten folds, and 1004 lung nodules are all samples. (b) also reflects the consistency of training accuracy and validation accuracy.

[0222] Of course, there are features that are learned that are not necessarily the parts that need to be focused on, and they can also be classified. For example, there is a model that learns to classify airplanes based on the blue sky, so further research is needed to explain the relationship between the learned features and the input images.

[0223] Existing automatic machine learning is already relatively complete based on traditional algorithms. For example, Auto-sklearn poses a huge challenge to manually designed algorithms on relatively simple data. Auto-sklearn has already entered the field of industrial algorithm design.

[0224] Neural network architecture search is a subtask of AutoML. The present invention designs a supernet for searching SSM models on a three-dimensional dataset of lung nodules to fill the gap of AutoML and lay the foundation for the automatic integration of deep learning models after the further development of computing devices in the future. The cross-validation method used in the present invention also follows the evaluation method of the existing AutoML training model. Deep learning shows obvious performance bottlenecks in the classification of three-dimensional lung nodules. After a lot of research, the present invention believes that on the one hand, there is uncertainty in the benign and malignant annotation of lung nodules, and on the other hand, the amount of data is too small to meet the needs of the model. Specifically, malignant lung nodules are in various forms, such as solid nodules, intermediate solid nodules, subsolid nodules, etc. The present invention finds that the probability of misclassification of subsolid nodules is relatively high. The overall ratio of negative and positive lung nodules selected by the present invention is 1 to 0.81. The data is relatively uniform from the perspective of binary classification, but subsolid nodules in positive samples are relatively rare and difficult to classify. From the data distribution of positive samples, the samples are extremely uneven. Are negative lung nodule samples far more than positive samples in nature? Obviously, the present invention balances the data set. At present, deep learning performs poorly on unbalanced data sets, and traditional machine learning methods cannot automatically learn features. If positive lung nodule features are not found, the advantages of traditional machine learning on unbalanced data sets cannot be brought into play. The current bottleneck in the classification of benign and malignant lung nodules does not mean the temporary end of the research, but rather heralds the real beginning of the research. In the future, this invention will find the cause of the bottleneck in the performance of benign and malignant lung nodules, and improve the classification performance from multiple angles such as data sets, preprocessing, models, and medical analysis.

[0225] 2. Conclusion

[0226] A supernet for searching SSM networks, called 3D Anatomical Multi-Scale SSM Supergraph (3DSSS), is designed, and the search results are used for 3D lung nodule classification. Compared with other advanced networks, 3DSSS achieves faster search speed and higher accuracy in this task. By designing 12 lung nodule scanning paths, the present invention realizes multi-angle analysis of different features, significantly enhancing the diagnostic ability of the classification model. In addition, combined with upsampling and downsampling techniques, the model introduces multi-scale feature fusion capabilities to comprehensively capture the structural information of lung nodules, thereby improving classification accuracy. Finally, the supernet training method based on dynamic elimination mechanism proposed in the present invention further improves the model performance by optimizing the training process. These innovative methods jointly promote the development of 3D lung nodule classification research and ensure the effectiveness and reliability of the model in practical applications. To the best of the knowledge of the present invention, this is the first 3D supernet of SSM. Experimental results on the LUNA16 benchmark dataset show that the model discovered by 3DSSS has excellent classification performance compared with existing methods.

[0227] 3DSSS can be used to automatically analyze and classify three-dimensional objects, scenes or environments. This has important applications in autonomous driving, smart security, robot navigation and other fields. For example, in autonomous driving, 3DSSS can help identify and classify objects in complex road scenes, such as pedestrians, vehicles, obstacles, etc., thereby improving the vehicle's decision-making ability.

[0228] This technology can be used for automatic classification and diagnosis of 3D medical images. For example, 3DSSS can accurately classify organs, tumors or lesions in 3D medical images such as CT and MRI, helping doctors to make more efficient diagnoses and develop treatment plans.

[0229] In virtual reality or augmented reality applications, 3DSSS can help quickly identify and classify different elements in the virtual environment and enhance the user interaction experience. For example, in education, entertainment or industrial training, the system can adjust and present the corresponding 3D scene in real time according to user needs.

[0230] In industrial automation and robotic control, 3DSSS can be used to classify and analyze the robot's interaction with its surroundings, especially for precise operations in complex three-dimensional spaces. For example, 3DSSS can help industrial robots automatically sort, assemble, and transport goods in warehouses.

[0231] 3DSSS can be applied in geographic information systems to automatically analyze and classify three-dimensional terrain data or urban building models. In this way, efficient data analysis and decision support can be provided for urban planning, disaster prediction and environmental protection.

[0232] In the field of game development, 3DSSS can be used to intelligently generate and optimize the three-dimensional virtual world in the game, improve the details and interactivity of the game environment, and make the game more realistic and immersive.

[0233] Just as automatic machine learning has transformed traditional machine learning by automating model design, neural network architecture search technology (NAS) will revolutionize deep learning in a similar way. Although AutoML has been widely adopted by the industry to automatically search and integrate traditional machine learning models, it is a counterpart to deep learning that can automatically discover the best neural network architecture. As deep learning continues to develop, NAS is expected to play a leading role in model design. With the continuous advancement of computing hardware, NAS can eventually automatically search and integrate all existing deep neural network architectures, reflecting the impact of AutoML on traditional machine learning.

[0234] The model trained on LUNA16 can be integrated into the hospital's computer tomography-assisted diagnosis system, for example Fig.12 The search method of the present invention is applied to network application examples found by convolutional hypernet search. Fig.13 shown.

[0235] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with the technical field within the technical scope disclosed by the present invention and within the spirit and principle of the present invention should be covered by the protection scope of the present invention.

Claims

1. A method for constructing a three-dimensional state space model search architecture with feedback, characterized in that: include: Step 1: Define the operation space; Average pooling, maximum pooling and skip connection operations are added to the operation space; By adding a downsampling module before the state space, the state space model obtains a double receptive field, and adding upsampling after the state space model to fuse the image information and the state space model output of the unsampled original receptive field, the state space model is upgraded to a three-dimensional state space model through anatomical scanning; output feedback is introduced to improve the performance of the state space model; Step 2: Construct a normal search unit and a downsampling search unit in the operation space, stack the normal search unit and the downsampling search unit, and add an image feature embedding module before the image is sent to the stacked network to form a supernet; The normal search unit and the downsampling search unit are both directed acyclic graphs with the same structure; the directed acyclic graph contains multiple nodes, each node represents a feature graph, the edges between the nodes are mixed operations composed of all operations in the operation space, the direction of the edge represents the direction of the information flow, and each edge and each operation in the edge has corresponding architecture parameters; Step 3: Search and obtain the architectural weight of each edge and each operation in the directed acyclic graph of the supernet through a high-precision two-layer optimization method; Step 4: The final weight is obtained by multiplying the architecture weight and the corresponding operation architecture weight, and the operation with the largest final weight in each edge is obtained. The corresponding final weight is used as the final weight of the edge. For units with input edges greater than two, only the first two edges with the largest final weight are retained. The results of the edge sum operation are stacked to update the normal search unit and the downsampling search unit until the required number of layers is obtained to obtain the final model architecture. If the operation space needs to be expanded, new search results can be obtained by starting from step 1 again according to the receptive field tendency of the update unit, and the updated normal search units and downsampling search units can be stacked to the required number of layers to obtain the final model architecture.

2. The method for constructing a three-dimensional state space model search architecture with feedback according to claim 1, characterized in that: Step 1 specifically includes: The state space model can be viewed as a linear time-invariant system, which consists of hidden states Enter Mapping to output response They are usually expressed as linear ordinary differential equations, as follows: h′(t)=Ah(t)+Bx(t) y(t)=Ch(t)+Dx(t) The domain of the coefficient matrix is In order to integrate into the deep model, the continuous-time state-space model needs to be discretized beforehand, which can be achieved by solving ordinary differential equations and then performing a simple discretization process; the result is shown below: y k =Ch k +Dx k The expression of the coefficient matrix is In order to increase the stability of the state space model, the present invention introduces a feedback mechanism into the state space model and maintains the performance of the parallel scanning of cub::BlockScan in Mamba; The continuous state space model adjusted for practical use: h′(t)=Ah(t)+Bx(t) y (t)=C(h(t)+h′(t))+Dx(t) Derivation of the continuous state space model with output feedback: Where K is the feedback coefficient; The feedback is kept in a certain calculation structure and simplified into the following formula and the performance is improved experimentally; h'(t)=KCAh(t)+KC(B+KD)x(t).

3. The method according to claim 1, characterized in that There are two sampling schemes for the downsampling unit in the supernet. One is to add downsampling before each operation that does not have downsampling capability in the edges starting from nodes 0 and 1 in the sampling unit; the other is to add downsampling after all edges of the input sampling unit, and the sampling unit is only used as an independent change unit to adapt to sampling; the downsampling operation is achieved by adjusting the resolution and number of feature channels of the input feature map, and the resolution of the feature map after sampling is halved and the number of feature channels is doubled.

4. The method according to claim 1, characterized in that: In the high-precision two-layer optimization method, the architecture weight of the supernet is determined by the following steps: Fix the architecture parameters, input the training set into the hypernet, calculate the cross entropy loss of the training set, and train until the model parameters converge; Fix the model parameters, input the training set into the hypernet, and calculate the cross entropy loss of the training set; Compute the gradients of the architecture weights using backpropagation; Adjust a round of architecture weights according to the gradient direction. Fix the architecture parameters, input the training set into the hypernet, and calculate the cross entropy loss of the training set; Compute the gradients of the architecture weights using backpropagation; Adjust the model weights for one round according to the gradient direction; Alternately optimize the architecture weights and model weights until convergence.

5. The method according to claim 1, characterized in that The search unit in the supernet includes a directed acyclic graph structure composed of multiple edges, and the feature output of each edge is obtained by weighted summation of features.

6. The method according to claim 1, characterized in that In the supernet design, the common unit and the downsampling unit respectively include a residual connection structure, and the residual connection is realized by a jump connection between nodes, which does not introduce additional parameters and maintains the stability of the network.

7. A system for constructing a three-dimensional state space model search architecture with feedback for implementing the method for constructing a three-dimensional state space model search architecture with feedback as claimed in any one of claims 1 to 6, characterized in that: include: Operation space definition module: used to define the operation space; A supernet forming module, used for constructing a normal search unit and a downsampling search unit in an operation space, and stacking the normal search unit and the downsampling search unit to form a supernet; An architecture weight search module is used to search and obtain the architecture weight of each edge and each operation in the directed acyclic graph of the supernet through a high-precision two-layer optimization method; The final architecture acquisition module is used to obtain the operation with the largest final weight in each edge by multiplying the architecture weight and the corresponding operation architecture weight as the final weight, and use the corresponding final weight as the final weight of the edge. The normal search unit and the downsampling search unit are updated with the results of the edge and operation stacked to obtain the final model architecture.

8. A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for constructing a three-dimensional state space model search architecture with feedback as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to execute the steps of the method for constructing a three-dimensional state space model search architecture with feedback as described in any one of claims 1 to 6.

10. An information data processing terminal, comprising the three-dimensional state space model search architecture construction system with feedback according to claim 7.

Citation Information

Patent Citations

  • Search method for three-dimensional model of mixing characteristic based on feedback

    CN101359342A

  • Causal analysis and graph neural network data generation method based on causal decision tree

    CN116882558A

  • Interactive method for resolving exterior orientation elements of camera based on three-dimensional model

    CN118331474A