Method and System for Constructing a Three-Dimensional State Space Model Search Architecture with Feedback
By constructing a three-dimensional state space model search architecture with feedback, combined with upsampling and downsampling technologies, the existing state space model has solved the problems of limited receptive field, low search efficiency, and insufficient multi-scale fusion in three-dimensional image processing, and efficient and accurate three-dimensional image classification and diagnosis are achieved.
Patent Information
- Application Number
- CN202510064511.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-01-15
AI Technical Summary
The existing state space model feels limited receptive field when processing high-dimensional image data, lacks feature extraction, low search efficiency, insufficient multi-scale information fusion, lacks feedback mechanisms, and difficult to expand the operation space, resulting in poor performance in complex tasks and difficult to adapt to diverse application scenarios.
By building a three-dimensional state space model search architecture with feedback, introducing an output feedback mechanism, combining upsampling and downsampling technologies, designing a hypernet structure, using high-precision two-layer optimization method to search the architecture weights, using pooling operations and jump connections, building ordinary search units and downsampling search units, forming directed acyclic graphs, realizing multi-scale feature fusion and dynamic adjustment.
It improves the classification accuracy and stability of the model to three-dimensional images, enhances the sensitivity to different features, improves the search efficiency and generalization ability of the model, adapts to different task requirements, and meets the high precision and high efficiency needs of three-dimensional medical image data processing.
Smart Images

Figure CN119991420B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to, but is not limited to, the technical field of machine learning, and particularly relates to a method and system for constructing a three-dimensional state space model search architecture with feedback. Background Art
[0002] Automated Machine Learning (AutoML) standardizes and automates the machine learning development process, including data preprocessing, feature selection, model selection, and hyperparameter optimization, reducing the professional threshold and improving efficiency. It is widely used in fields such as finance and healthcare. Its core concept is to optimize model performance using search strategies while reducing human intervention. Neural Architecture Search (NAS) focuses on automatically designing optimal neural network architectures through algorithms, using techniques such as reinforcement learning to efficiently search for the best models in the network structure space to meet the requirements of deep learning for efficient architectures. NAS is driving the development of deep learning towards more complex directions and is expected to become a dominant technology in the future. Currently, NAS technology is still limited to the fields of convolutional neural networks and Transformers and has not been combined with state space deep learning models, creating a technical gap that provides a new direction for the innovative application of AutoML and NAS.
[0003] In the industrial applications of the existing technology, the three-dimensional state space model search architecture faces many technical problems, which limit its application effect and efficiency in complex tasks. Specifically, the following technical problems mainly exist:
[0004] 1. Limited receptive field and insufficient feature extraction:
[0005] Existing state space models usually have a limited receptive field, resulting in the inability of the model to fully capture global and multi-scale feature information when processing high-dimensional image data. This limitation makes the model perform poorly in complex image segmentation, recognition, and other tasks and difficult to accurately capture the fine structures and global relationships in images.
[0006] 2. Low search efficiency and high computational resource consumption:
[0007] Traditional search methods often rely on single-layer or shallow search units when constructing and optimizing state space models, resulting in a large search space and a complex optimization process. The high computational cost and long search process limit the real-time performance and scalability of the model in actual industrial applications and are difficult to meet the requirements of large-scale data processing.
[0008] 3. Insufficient multi-scale information fusion and poor model generalization ability:
[0009] In the existing architecture, the fusion of multi-scale information usually relies on simple concatenation or weighted average methods, which cannot effectively integrate the feature information at different scales. This deficiency leads to poor generalization ability of the model when facing images with different scale features and makes it difficult to adapt to diverse application scenarios.
[0010] 4. Lack of feedback mechanism, limited improvement of model performance:
[0011] Traditional state space models lack an effective feedback mechanism and cannot dynamically adjust and optimize the model structure according to the results output by the model. This feedback-lacking design limits the adaptive ability of the model in complex tasks and makes it difficult to continuously optimize and improve performance during the training process.
[0012] 5. Difficulty in expanding the operation space, insufficient adaptability:
[0013] When existing methods expand the operation space, they often need to manually adjust or re-design the search strategy, lacking flexibility and automation. This inconvenience limits the rapid adaptation and optimization of the model when facing different task requirements, reducing the overall efficiency and practicality of the system. Summary of the Invention
[0014] Aiming at the problems existing in the prior art, the present invention provides a method and system for constructing a three-dimensional state space model search architecture with feedback.
[0015] The present invention is implemented as follows. A method for constructing a three-dimensional state space model search architecture with feedback includes:
[0016] Step 1: Define the operation space, and add average pooling, max pooling, and skip connection operations to the operation space; make the state space model obtain a double receptive field by adding a downsampling module in front of the state space, and add upsampling after the state space model to fuse the image information with the output of the state space model with the original unsampled receptive field, and elevate the state space model to a three-dimensional state space model through anatomical scanning; introduce output feedback to improve the performance of the state space model.
[0017] Step 2: Construct a normal search unit and a downsampling search unit in the operation space, stack the normal search unit and the downsampling search unit, and add an image feature embedding module before the image is fed into the stacked network to form a super network; both the normal search unit and the downsampling search unit are directed acyclic graphs with the same structure; the directed acyclic graph contains multiple nodes, each node represents a feature map, the edges between the nodes are composed of all operations in the operation space to form a hybrid operation, the direction of the edge represents the direction of information flow, and each edge and each operation in the edge have corresponding architecture parameters;
[0018] Step 3: Search for the architectural weights of each edge and each operation in the edge of the directed acyclic graph of the supernet through a high-precision double-layer optimization method;
[0019] Step 4: Use the product of the architectural weight and the corresponding operation architectural weight as the final weight, obtain the operation with the largest final weight in each edge, and use the corresponding final weight as the final weight of the edge. For units with more than two input edges, only keep the top two edges with the largest final weight. Stack the results of the edges and operations to update the ordinary search unit and the downsampling search unit until the required number of layers to obtain the final model architecture. It should be noted that if it is necessary to expand the operation space, the process can start from Step 1 again according to the receptive field tendency of the updated unit to obtain a new search result, and stack the updated ordinary search unit and downsampling search unit to the required number of layers to obtain the final model architecture.
[0020] Further, Step 3 specifically includes:
[0021] Search for the architectural weights of each edge and each operation in the edge of the directed acyclic graph of the supernet through a high-precision double-layer optimization method, including:
[0022] Send the training set images into the supernet, fix the architectural parameters, calculate the cross-entropy loss according to the labels after forward propagation, and perform backpropagation to calculate the gradients; optimize the weight parameters of the supernet according to the gradient direction, and repeat the weight parameters until convergence to obtain the final architectural weights of each edge and each operation in the edge of the directed acyclic graph of the supernet, that is, alternately optimize the architectural parameters and weight parameters until convergence to obtain the final architectural weights of each edge and each operation in the edge of the directed acyclic graph of the supernet.
[0023] Further, sending the training set images into the supernet includes:
[0024] Send the training set images into the supernet after random cropping, flipping, and normalization.
[0025] Another object of the present invention is to provide a three-dimensional state space model search architecture construction system with feedback for implementing the above-mentioned three-dimensional state space model search architecture construction method with feedback, including:
[0026] Operation space definition module: used to define the operation space;
[0027] Supernet formation module, used to construct ordinary search units and downsampling search units in the operation space, stack the ordinary search units and downsampling search units to form a supernet;
[0028] Architectural weight search module, used to search for the architectural weights of each edge and each operation in the edge of the directed acyclic graph of the supernet through a high-precision double-layer optimization method;
[0029] Final architecture acquisition module, which is used to obtain the operation with the largest final weight in each edge by taking the product of the architecture weight and the corresponding operation architecture weight as the final weight, and use the corresponding final weight as the final weight of the edge, and update the ordinary search unit and the downsampling search unit with the stack of the edge and the operation result to obtain the final model architecture.
[0030] Another object of the present invention is to provide a computer device, which includes a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor is caused to execute the steps of the method for constructing a three-dimensional state space model search architecture with feedback.
[0031] Another object of the present invention is to provide a computer-readable storage medium, which stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the method for constructing a three-dimensional state space model search architecture with feedback.
[0032] Another object of the present invention is to provide an information data processing terminal, which includes the three-dimensional state space model search architecture construction system with feedback.
[0033] Combined with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:
[0034] First, the latest state space deep learning model is the third basic model on a par with the convolutional neural network model and the Transformer model. At present, there is no technology that combines neural network architecture search and state space model, and it is still blank in automated machine learning.
[0035] In common visual sequence classification models, Tranformer is used as the backbone network to represent complex patterns in visual data. Compared with CNN, ViT generally shows stronger modeling ability because it combines the self-attention mechanism to achieve global receptive field and dynamic prediction weight parameters. However, self-attention also brings a side effect, that is, the complexity of self-attention grows quadratically with the increase of input size. SSM has the advantage of being able to efficiently model visual data while retaining the self-attention mechanism relative to ViT. SSM is essentially a sequence model, and the first well-known work is the Kalman filter. In recent years, the series of papers by Gu et al. are impressive, in which the HiPPO operator compresses continuous signals and discrete time series online by projecting to a polynomial basis, and gives the S4 model the ability to establish long-range dependencies by initializing the matrix A. Vim is the earliest representative visual Mamba model, which uses position embedding to integrate spatial information. Later, VMamba combined convolution and spatial attention mechanisms and introduced the Cross-Scan Module (CSM) to achieve 1D selective scanning in a two-dimensional image space with a global receptive field. However, these Mamba visual applications are not designed for 3D images, and the large number of convolutions makes the classification ability of SSM questionable.
[0036] Inspired by VMamba, we designed a 3D State Space Supergraph (3DSSS), which searches for 3DSSS through a differentiable architecture to find a universal visual backbone consisting of blocks composed of concise SSMs for efficient classification of 3D lung nodules. The effectiveness of 3DSSS in reducing the complexity of attention computation is largely due to the selective scanning mechanism present in the Mamba model, also known as selective SSM. Unlike traditional attention computation methods, which allow dense information routing in context, Mamba requires that each element in a one-dimensional array (e.g., a text sequence) obtains contextual knowledge only through compressed hidden states, thereby reducing quadratic complexity to linear.
[0037] The present invention designs a supernet structure of a linear time series model, aiming to search for a visual backbone network of the linear time series model. By constructing a supernet, the present invention can quickly explore the combination of different state space expressions, thereby improving the search efficiency of the model. This method not only saves computing resources, but also ensures that the optimal state space model is found to better adapt to the classification task of three-dimensional lung nodules.
[0038] To improve the accuracy and model stability in 3D image classification, the present invention designs an anatomical scan and an output feedback loop. This design aims to comprehensively analyze the same data from different angles through multiple scan paths, enhancing the model's sensitivity to different features, thereby effectively capturing the key features of 3D images. Through the feedback mechanism, the model can self-adjust and optimize the output of each scan path, strengthening the recognition of important information and improving the classification accuracy and stability. In addition, the state space module adopts a pure state space expression, significantly compressing redundant manual designs and improving the system's processing efficiency and generalization ability. This strategy of multi-angle analysis and feedback optimization significantly enhances the classification effect of the model in practical applications and strengthens the ability to accurately classify 3D images.
[0039] The present invention adopts upsampling and downsampling techniques, introducing multi-scale feature fusion capabilities into the state space model. This innovation enables the model to extract useful features from different scales and combine them for classification, thus more comprehensively understanding the structural information of 3D images. This multi-scale feature fusion strategy effectively improves the classification accuracy of the model, enabling effective recognition of objects of different sizes and shapes.
[0040] In the research of the present invention, a supernet training method based on a dynamic elimination mechanism is proposed, optimizing the performance of the pulmonary nodule classification model by adjusting the Points value. This method allows the elimination rules to be dynamically adjusted according to the performance of the operations at each training stage, enabling the model to fully exert its potential in different rounds.
[0041] One or more technical solutions provided in the present invention have at least the following technical effects or advantages:
[0042] 1. Define the operation space. To improve the stability of the operation space and prevent the model from being over-parameterized, average pooling, max pooling, and skip connection operations are added to the operation space in the present invention. Since the state space model does not have multi-scale characteristics, the present invention enables the state space model to obtain a doubled receptive field by adding a downsampling module before the state space, and fuses the image information and the output of the state space model with the original receptive field without sampling by adding upsampling after the state space model. The existing state space model is a two-dimensional model, and the present invention elevates the state space model to a three-dimensional state space model through anatomical scanning. The existing state space model has instability, and the present invention introduces output feedback to improve the performance of the state space model; then ordinary search units and downsampling search units are constructed in the operation space. Since the state space model has only a single receptive field, that is, it cannot change the feature map size and the number of feature maps, in order for the supernet to normally downsample samples, the present invention designs two downsampling methods, such as Figure 4As shown in (b)(d) and (c)(e), (b)(d) enables the supernetwork to have a stronger ability to downsample and compress information, while (c)(e) simplifies the supernetwork, significantly reducing the training video memory but sacrificing part of the downsampling and information compression ability. Stack ordinary search units and downsampling search units, and add an image feature embedding module before the image is fed into the stacked network to form the final supernetwork. Then, use the high-precision double-layer optimization method to search for the architecture weights of each edge and each operation in the directed acyclic graph of the supernetwork. Finally, use the product of the architecture weights and the corresponding operation architecture weights as the final weights, obtain the operation with the largest final weight in each edge, and use the corresponding final weight as the final weight of the edge. Stack and update the ordinary search units and downsampling search units with the results of the edges and operations until the required number of layers to obtain the final model architecture.
[0043] 2. In the present invention, the operation weighting function uses sigmoid to enable fair competition among operations.
[0044] 3. The dynamic elimination mechanism in the present invention accelerates model search and reduces the truncation error in model selection, increasing the search accuracy.
[0045] Second, the present invention effectively solves many problems in the processing of three-dimensional medical image data by innovatively constructing a three-dimensional state space classification model and introducing neural network architecture search (NAS) technology. Traditional two-dimensional image processing methods are difficult to fully exploit the spatial features in three-dimensional data, while the present invention can not only capture the complex structures in three-dimensional data but also model dynamic changes by combining sequence models. Especially in the analysis of tumor growth, neurological diseases, and cardiovascular lesions, the present invention can provide more accurate classification and diagnostic results. Through the automated NAS technology, the model can quickly adapt to different tasks and datasets, significantly reducing the cost of manual hyperparameter tuning, while enhancing the flexibility and scalability of the model to meet the requirements of high precision and high efficiency in medical image analysis.
[0046] The technical solution of the present invention is expected to bring significant economic and social benefits to the medical industry. By deeply mining the feature fusion of three-dimensional images and text data, this technology can improve the classification accuracy and diagnostic accuracy in multimodal data analysis, provide more accurate personalized treatment plans for patients, and reduce the risks of misdiagnosis and over-treatment. In addition, the automated classification ability significantly reduces the labor costs of medical institutions and accelerates the efficiency of medical image analysis, enabling doctors to devote more energy to the treatment of complex cases. Meanwhile, the application of this technology in drug research and development and clinical trials will improve the analysis efficiency and result reliability of experimental data, accelerate the drug development cycle, and bring new value growth points to the pharmaceutical industry.
[0047] From a commercial value perspective, the present invention fills the gap in the precise analysis technology of three-dimensional medical images and provides an efficient solution for the medical imaging analysis market. In the future, the technology based on the present invention can be developed into a SaaS platform to provide low-cost and high-efficiency artificial intelligence diagnostic tools for medical institutions through cloud services, further reducing the threshold of industry technology application. At the same time, the technology also has the potential for cross-industry applications, such as intelligent manufacturing, autonomous driving, and robot vision, and can expand the market space in multiple fields. The trend of globalization also enables the technology to have the prospect of entering the international market, cooperating with global medical institutions to jointly promote the innovation and upgrading of medical imaging analysis.
[0048] Thirdly, by constructing a three-dimensional state space model search architecture with feedback, the present invention overcomes the application limitations in three-dimensional image processing caused by huge video memory overhead, unstable models, and low search efficiency in the prior art. The designed supernetwork of the present invention combines ordinary units and downsampling units, and at the same time, through the residual connection structure and high-precision double-layer optimization method, significantly reduces parameter redundancy and video memory overhead, while improving the training stability of the model. Compared with traditional convolutional models, the present invention effectively balances the structural design of depth and width, avoids performance bottlenecks caused by uneven depth, and provides an efficient and scalable solution for complex three-dimensional image data processing.
[0049] In addition, by introducing pooling operations and a partial channel weighting mechanism, the present invention further optimizes the operation space of the supernetwork, improving the stability and efficiency of the search process. Combined with the optimized architecture weight adjustment strategy, the present invention realizes the precise construction and efficient search of the three-dimensional state space model, significantly improving the performance and accuracy of three-dimensional image analysis. Its technological progress is not only reflected in reducing the demand for hardware resources but also in promoting the practical application of three-dimensional image processing algorithms in fields such as medical imaging and industrial inspection, accelerating the technological upgrading of related industries. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flowchart of a method for constructing a three-dimensional state space model search architecture with feedback provided by an embodiment of the present invention;
[0051] Figure 2 is a schematic diagram of a three-dimensional anatomical scan state space model with feedback provided by an embodiment of the present invention;
[0052] Figure 3 is a schematic diagram of the principle of the state space model to achieve multi-scale effects provided by an embodiment of the present invention;
[0053] Figure 4 is a schematic diagram of a three-dimensional state space supernetwork search architecture with feedback provided by an embodiment of the present invention;
[0054] Figure 5It is the schematic diagram of the sampling operation provided by the embodiment of the present invention;
[0055] Figure 6 It is the schematic diagram of the operation elimination rule provided by the embodiment of the present invention;
[0056] Figure 7 It is the search result diagram in the LUNA16 three-dimensional dataset provided by the embodiment of the present invention;
[0057] Figure 8 It is the structural diagram of the three-dimensional state space model search architecture construction system with feedback provided by the embodiment of the present invention.
[0058] Figure 9 It is the performance comparison diagram between the fifth validation set and the state-of-the-art manually designed state space model in the cross-validation experiment of the LUNA16 three-dimensional dataset provided by the embodiment of the present invention.
[0059] Figure 10 It is the performance comparison diagram between the fifth validation set and the 4-way scan based on the 3D_VMamba model in the cross-validation experiment of the anatomical scan in the LUNA16 three-dimensional dataset provided by the present invention.
[0060] Figure 11 It is the performance comparison diagram between the fifth validation set and the 3DSSS model without feedback in the cross-validation experiment of the LUNA16 three-dimensional dataset provided by the embodiment of the present invention.
[0061] Figure 12 It is the function diagram without class activation mapping visualization in the application example of lung nodule classification provided by the embodiment of the present invention.
[0062] Figure 13 It is the function diagram with developed class activation mapping visualization in the application example of lung nodule classification by searching the convolutional neural network provided by the embodiment of the present invention.
[0063] Figure 14 It is the image of a malignant lung nodule correctly diagnosed in the embodiment of the present invention. In the figure, P is the confidence level and D is the diameter of the lung nodule, and the diameter unit is millimeter.
[0064] Figure 15 It is the image of a benign lung nodule correctly diagnosed in the embodiment of the present invention. In the figure, P is the confidence level and D is the diameter of the lung nodule, and the diameter unit is millimeter.
[0065] Figure 16 It is the confidence distribution of correctly classified lung nodules in the fifth validation set in the cross-validation experiment of the LUNA16 three-dimensional dataset provided by the embodiment of the present invention.
[0066] Figure 17(a) is the distribution of 1,004 lung nodule images for feature dimensionality reduction, and (b) is the distribution of feature dimensionality reduction of 1,004 lung nodules after SSM learning provided by an embodiment of the present invention.
[0067] In all search results appearing in this document, the upper part is the ordinary unit and the lower part is the downsampling unit.
[0068] Figure 18 The search space of VSS provided by an embodiment of the present invention is PRIMITIVES = ['Skip_Connect', 'VSS_12', 'VSS_Double_12', 'VSS_4', 'VSS_Double_4'], and the search result graph.
[0069] Figure 19 The network result graph searched by EMBED_DIM_32 -> 96 provided by an embodiment of the present invention.
[0070] Figure 20 The network result graph searched by EMBED_DIM_96 provided by an embodiment of the present invention. Detailed implementation manners
[0071] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0072] See Figure 1 , the neural network search architecture construction method of the high-precision double-layer optimization method provided by an embodiment of the present invention includes:
[0073] S101: Define the operation space;
[0074] In order to improve the stability of the operation space and prevent the model from being over-parameterized, average pooling, max pooling and skip connection operations are added to the operation space in the present invention. Since the state space model does not have multi-scale characteristics, the present invention makes the state space model obtain a double receptive field by adding a downsampling module before the state space, and adds an upsampling after the state space model to fuse the image information with the output of the state space model with the original receptive field without sampling. The existing state space model is a two-dimensional model, and the present invention upgrades the state space model to a three-dimensional state space model through anatomical scans. The existing state space model has instability, and the present invention introduces output feedback to improve the performance of the state space model.
[0075] The specific description of this step includes:
[0076] The state space model can be regarded as a linear time-invariant system, which passes through the hidden state Map the input to the output response They are usually represented as linear ordinary differential equations as follows:
[0077] h′(t) = Ah(t) + Bx(t)
[0078] y(t) = Ch(t) + Dx(t)
[0079] where the domain of the coefficient matrix is
[0080] To integrate into the deep model, the continuous-time state space model needs to be discretized in advance, which can be achieved by solving the ordinary differential equation and then performing a simple discretization process. The result is as follows:
[0081]
[0082] y k = Ch k + Dx k
[0083] where the expression of the coefficient matrix is
[0084] To increase the stability of the state space model, the present invention introduces a feedback mechanism into the state space model and maintains the performance of cub::BlockScan parallel scan in Mamba.
[0085] The adjusted continuous state space model in the actual use of the present invention:
[0086] h′(t) = Ah(t) + Bx(t)
[0087] y(t) = C(h(t) + h′(t)) + Dx(t)
[0088] Derivation of adding output feedback to the continuous state space model:
[0089]
[0090] where K is the feedback coefficient.
[0091] Actually, since the feedback coefficient and the state space model coefficient of the present invention are learnable parameters or calculated from learnable parameters, the present invention simplifies the feedback into the following formula and experimentally obtains performance improvement.
[0092] h′(t) = KCAh(t) + KC(B + KD)x(t)
[0093] Although the sequence properties of Mamba are well-aligned with natural language processing tasks involving temporal data, when applied to visual data, it poses significant challenges because visual data itself is non-sequential and contains spatial information (such as local texture and global structure). To address this issue, the S4ND model reformulates the state space model with convolutional operations and directly extends the kernel from one-dimensional to two-dimensional via the outer product. However, this modification prevents the dynamics of the weights. Therefore, the present invention selects the selective scanning method for data processing and proposes the anatomical scanning module.
[0094] As Figure 2 shown, passing data through a three-dimensional anatomical scanning state space model with output feedback includes three steps: three-dimensional anatomical scanning, selective scanning of the SSM block with output feedback, and regenerative reconstruction. The present invention collectively abbreviates three-dimensional anatomical scanning and regenerative reconstruction as 3DAS-OF (3DAnatomical Scanning for SSM with Output Feedback). Given the input data, 3DAS-OF first unfolds the image patches into sequences along twelve different traversal paths, processes each patch sequence in parallel using a state space model with individual feedback, and subsequently regeneratively reconstructs the result sequence to form the output map. 3DAS-OF enables each pixel in the image to effectively integrate information from all other pixels in different directions by adopting complementary traversal paths from three anatomical planes, thereby facilitating the establishment of a global receptive field.
[0095] As Figure 2 shown, the sampling interval Δ is determined by the input to achieve dynamic weights, so in the hypernetwork design, the present invention deletes the right spatial attention branch of Mamba as Figure 3 shown. The visual image has been projected to a high-dimensional space through the embedding layer of the hypernetwork, Figure 3 and the projection layer in Mamba is also deleted from the hypernetwork by the present invention. Since the state space equation does not have the property of multi-scale, the present invention introduces an upsampling and downsampling combination in the hypernetwork so that the state space equation can obtain a double receptive field before the downsampling unit, thereby achieving multi-scale feature fusion.
[0096] As Figure 3 shown, the present invention retains the activation function in the Mamba structure to add non-linearity to the sampled image. Overall, the present invention tries to remove the artificially designed elements with unknown effects and difficult to interpret, hoping that the hypernetwork can search for a parameter-efficient unknown spatial state equation network, which consists of ordinary units and downsampling units as the new backbone network. It also makes explanations for some doubts about the performance of the pure spatial state equation network. Figure 3 SSMONLY in Figure 2 corresponds to
[0097] S102: Construct ordinary search units and downsampling search units in the operation space. Since the state space model has only a single receptive field, that is, it cannot change the feature map size and the number of feature maps. In order for the supernet to normally downsample samples, two downsampling methods are invented and designed. As shown in Figure 4 (b)(d) and (c)(e), (b)(d) can make the supernet have a stronger ability to downsample and compress information, and (c)(e) simplifies the supernet, resulting in a significant reduction in training video memory, but sacrificing part of the downsampling and compressing information ability. Stack ordinary search units and downsampling search units and add an image feature embedding module before the image is fed into the stacked network to form a supernet; both the ordinary search unit and the downsampling search unit are directed acyclic graphs of the same structure; the directed acyclic graph contains multiple nodes, each node represents a feature map, and the edges between the nodes are composed of all operations in the operation space to form a mixed operation. The direction of the edge represents the direction of the information flow, and each edge and each operation in the edge have corresponding architecture parameters;
[0098] Specific description of this step includes:
[0099] As Figure 4 shown in (a) of, the supernet designed by the present invention is composed of two types of units, one is an ordinary unit, and the other is a downsampling unit. The units before and after the downsampling unit are marked with F and B respectively after the unit to indicate that special attention is needed during the design of the supernet. It can be clearly noted from (a) that different from the convolutional supernet architecture, the three-dimensional state space supernet of the present invention has only one ordinary unit before the downsampling unit. Among the total 8 units, the last four units are all divided into ordinary units. The present invention found in the VMamba experiment that the model undergoes three downsamplings, but the model depth is similar to [2, 2, 27, 2]. The depth is increased in the third stage to improve the model performance, rather than increasing the model depth uniformly, which is very consistent with the supernet design. Inspired by VMamba, the search result of an ordinary unit usually includes multiple SSMs. Therefore, only one ordinary unit is placed before the downsampling unit to imitate the successful experience of VMamba. The feature channels of the state space model do not have the same meaning as those of the convolutional feature channel model. The feature channels of the state space model represent the encoding of images. This leads to the idea of randomly extracting some channels and feeding them into the state space model not being consistent with the model definition, which may cause risks such as model degradation, instability, or non-convergence. However, the classification of pulmonary nodules is a three-dimensional image, which will inevitably lead to a huge video memory overhead, posing an obstacle to increasing the depth of the supernet. Increasing the depth only after the second downsampling of the supernet can improve the model performance, and at the same time, it exactly fits the design of the three-dimensional state space model supernet. For three-dimensional images, as shown in Figure 5 (a) of, when the feature channels are doubled by downsampling by a factor of two, the model parameter quantity will be reduced by 4 times. Then, the video memory overhead of the unit after the second downsampling will be saved by 16 times compared with the video memory without sampling. Weighing the computational overhead, the present invention uses Figure 5The downsampling method in (b). Since the pulmonary nodule images are small and for the purpose of implementing multi-scale fusion of the spatial state equation, the supernet units of the present invention are stacked only for two downsamplings.
[0100] As Figure 4 Shown in (b) and (c) of the figure, two supernet sampling strategies are shown in the supernet. The present invention adopts the sampling strategy in (c). In the downsampling unit, no actual downsampling is performed, but a special unit is independently varied for the need after downsampling to enhance the model performance. The advantage of this is that it saves the sampling module before the state space model, such as Figure 1 The comparison between (e) and (d) in the figure. By using the method in (d), I saved the sampling operations from 0, 1 to 2-5 in (e). However, this simplified network sampling strategy of the present invention will also cause the weakening of the downsampling unit's ability. Judging from the results, the independent variation of the downsampling unit can make up for this weakening.
[0101] The supernet designed by the present invention retains the residual structure of the units in the ordinary supernet. The formula is shown as follows. The skip connection neither introduces additional parameters nor increases the computational complexity, where x is the input, y is the output, W is the learnable parameter of the i-th layer, and F is the operator of this layer. Experiments show that the residual network is easier to optimize and can obtain accuracy from a relatively large depth.
[0102] y = F(x, {W i}) + x, as Figure 4 (b) It should be noted that downsampling is required from node 6 of Normal cell (F) to node 0 of Normal cell (B). Actually, the number of feature channels of node 6 is 4 times that of the input feature channels of the unit, and special processing is required before inputting to node 0 or 1, such as Figure 4 (e) shown. The operations starting from nodes 0 and 1 of the downsampling unit will perform downsampling, but the state space model has no concept of stride, so the present invention adds a downsampling branch in the state space model operation, such as Figure 4 (e) shown.
[0103] As Figure 4 Shown in (f) of the figure, the present invention designs the operation space of the edges in the unit. In addition to the operations of the state space model and the residual, due to the instability in the training process of the state space model, introducing a large number of state space model operations makes the supernet difficult to train and causes a large amount of parameter redundancy. The pooling operation can summarize the statistics of features in the local area and has a certain invariance to some changes in the input image, increasing the stability of the supernet training while reducing the parameter redundancy of the supernet. Therefore, the present invention introduces average pooling and max pooling. The double weighting of p and q in (f) makes the supernet search process more stable. Using the Sigmoid weighted zero operation only requires implicit expression, so the present invention does not add zero operation in the operation space.
[0104] Construct ordinary search units and downsampled search units in the operation space. For the edges and operations within the edges of the ordinary search units and downsampled search units, partial decoupled operation weighting functions are used as shown in the following formula:
[0105]
[0106] where α is an architecture parameter and Sigmoid(α) is the architecture weight.
[0107] In the supernet, the output feature map of the edge of the search unit is as shown in the following formula:
[0108]
[0109] where is the output of the partial channel connection of edge (i, j).
[0110] In the supernet, the node feature map is the weighted sum of the input edge results, as shown in the following formula:
[0111]
[0112] where is an architecture parameter.
[0113] S103: Search for the architecture weights of each edge and each operation in the directed acyclic graph of the supernet through a high-precision two-layer optimization method;
[0114] Specifically explain this step. Searching for the architecture weights of each edge and each operation in the directed acyclic graph of the supernet through a high-precision two-layer optimization method includes:
[0115] Send the training set images into the supernet, fix the architecture parameters, calculate the cross-entropy loss according to the labels after forward propagation, and perform backpropagation to calculate the gradients; optimize the weight parameters of the supernet according to the gradient direction, and repeat the weight parameters until convergence to obtain the final architecture weights of each edge and each operation in the directed acyclic graph of the supernet, that is, alternately optimize the architecture parameters and weight parameters until convergence to obtain the final architecture weights of each edge and each operation in the directed acyclic graph of the supernet.
[0116] Specifically, sending the training set images into the supernet includes:
[0117] Send the training set images into the supernet after random cropping, flipping, and normalization.
[0118] S104: Obtain the operation with the largest final weight in each edge by using the product of the architecture weight and the corresponding operation architecture weight as the final weight, and use the corresponding final weight as the final weight of the edge. For units with more than two input edges, only keep the top two edges with the largest final weight of the edge. Update the normal search unit and the downsampling search unit with the results of the edges and operations stacked until the required number of layers to obtain the final model architecture. It should be noted that if it is necessary to expand the operation space, the process can start from step 1 again according to the receptive field tendency of the updated unit to obtain new search results, and stack the updated normal search unit and downsampling search unit to the required number of layers to obtain the final model architecture.
[0119] Specifically, the gradient of the architecture parameters of the high-precision double-layer optimization method is an approximate numerical method, as shown in the following formula:
[0120]
[0121] Among them, is the symbol for finding the gradient, L val (w * ,α) is the validation loss, L train (w,α) is the training loss, α is the architecture parameter, w * is the optimal network weight, w is the model parameter, ξ, ε are hyperparameters for adjusting the approximate formula,
[0122] The high-precision double-layer optimization method is initialized with w * (α0), to highly fit w * (α), thereby increasing the precision of subsequent numerical calculations. Note that all previous double-layer optimization initializations used Kaiming initialization, while the present invention proposes w * (α0) initialization for the numerical method of double-layer optimization.
[0123] The purpose of initializing w * (α0) is to better fit w * (α), while the purpose of Kaiming initialization is to stabilize the variance of the model output. Kaiming initialization satisfies the following formula:
[0124]
[0125] where l is the layer index, y is the input response, n is the number of connections of the response, and w is the model parameter.
[0126] L train and L val are cross-entropy loss functions, but regularization is added and the regularization maintains the physical properties of the loss, as shown in the following formula:
[0127]
[0128] L total = L val (w * (γ), γ) + w 0-1 L 0-1
[0129] Among them, L 0-1 is a regularization function, γ is all architecture parameters including edges and operations, and w 0-1 is a hyperparameter for adjusting regularization.
[0130] For example Figure 6 in (a) is an example of unit selection edges. The present invention selects the one with the largest p in each edge as the candidate operation. Then, for nodes 2 to 5, only the two edges (orange) of the input nodes with the largest product of p and q are selected. The knockout tournament is carried out when selecting the one with the largest p in each edge as the candidate operation. As shown in Figure 5 (b), during the hypernetwork double-layer optimization process, when the p value is less than points, it will be replaced with the Zero operation, that is, eliminated.
[0131] Point is between 0 and 1. The Points value of the knockout tournament is very important because the operations in the hypernetwork show different capabilities at different training stages. The purpose of the present invention is to make the indicators of pulmonary nodule classification of the target network the highest. However, just like the entrance examination, the one with the highest score does not necessarily become a scientist. As the elimination rule, the eliminated points determine the final performance tendency of the model. For example, initializing Points to 0.5 will eliminate the operations with poor performance in the early stage of the competition. Because all potential opponents are eliminated, the pooling operation gains a huge advantage. Theoretically, the Points value approaching 0 is the best, but this is just an overfitting phenomenon that occurs in the "examination". In other words, in order to have a fast search speed and limited data itself, the "examination" cannot reflect the ability of the model in real data. For example, DARTS only uses half of the data to train the hypernetwork. The present invention assumes that the operation ability will reach its peak within the same round. The ideal P value can eliminate the relatively poor operations at the peak of the ability of all operations without causing the convergence of the hypernetwork to deteriorate extremely. The present invention uses a dynamic Points value, which is initialized to 0.2, and is increased to 0.25 and 0.3 respectively in the last two eliminations. In order to make the network convergence more stable, the present invention conducts an elimination every 5 rounds.
[0132] The hypernetwork designed by the present invention retains the residual structure of the units in the usual hypernetwork. The skip connection neither introduces additional parameters nor increases the computational complexity. Experiments show that the residual network is easier to optimize and can obtain accuracy from a relatively large depth. The residual structure formula is as follows:
[0133] y = F(x, {W i ) + x
[0134] As Figure 4 (b) It should be noted that the nodes from 6 of Normal cell(F) to 0 of Normal cell(B) need to be downsampled. Actually, the number of feature channels of node 6 is 4 times the number of input feature channels of the unit. Special processing is required before inputting to node 0 or 1, such as Figure 4 (e) shown. The operations starting from nodes 0 and 1 of the downsampling unit will perform downsampling. However, the state space model has no concept of stride, so the present invention adds a downsampling branch in the state space model operation, as shown in Figure 4 (e).
[0135] To further reduce the video memory overhead of the supernet, we also designed another simple sampling method, as shown in Figure 4 (c). We modified the operation of performing downsampling starting from nodes 0 and 1 of the downsampling unit to performing downsampling operations before inputting to nodes 0 and 1. In this way, the number of downsampling operations is reduced from 56 to only two. If two downsampling units are used, 108 downsampling operations are saved, and so on. The strict supernet design is as shown in Figure 4 (e).
[0136] Two different downsampled supernets Figure 4 (c) and Figure 4 (d) have a search space as shown in Figure 4 (f).
[0137] The operation space of the edges in the unit is designed. In addition to the operations of the state space model and the residual, due to the instability during the training process of the state space model, introducing a large number of state space model operations makes the supernet difficult to train and causes a large amount of parameter redundancy. Pooling operations can summarize the statistics of features in a local area and have a certain invariance to some changes in the input image, increasing the stability of supernet training while reducing supernet parameter redundancy. Therefore, the present invention introduces average pooling and max pooling. Figure 4 (f) The double weighting of p and q makes the supernet search process more stable. Using the Sigmoid weighted zero operation only requires an implicit expression, so the present invention does not add a zero operation in the operation space.
[0138] In the operation space, ordinary search units and downsampling search units are constructed. The operations in the edges and within the edges of the ordinary search units and downsampling search units all use a partially decoupled operation weighting function as shown in the following formula:
[0139]
[0140] Among them, α is the architecture parameter, and Sigmoid(α) is the architecture weight.
[0141] In the supernet, the output feature map of the edge of the search unit is shown in the following formula:
[0142]
[0143] Among them, is the output of the partial channel connection of edge (i,j).
[0144] In the supernet, the node feature map is the weighted sum of the input edge results, as shown in the following formula:
[0145]
[0146] Among them, is the architecture parameter.
[0147] The search results of the present invention in the LUNA16 3D dataset are as Figure 7 shown. The upper figure is the ordinary unit searched in the LUNA16 3D dataset, and the lower figure is the downsampling unit.
[0148] The final results of the model architecture in the LUNA16 3D dataset compared with other advanced results are shown in Table 1. The best result and the second-best result in each column are shown in bold and underlined formats respectively. "-" indicates not mentioned.
[0149]
[0150] The present invention compares the searched model with several existing state-of-the-art methods. Obviously, it achieves the highest classification accuracy with the fewest model parameters. The high accuracy and F1 score indicate that 3DSSS achieves a good balance between sensitivity and specificity.
[0151] The process of applying the neural network architecture in the example is as follows:
[0152] Step 1: Extract pulmonary nodules from the lung window image;
[0153] Specifically explain this step. Extracting pulmonary nodules from the lung window image includes:
[0154] Set the center of the lung window image to -300, adjust the width of the lung window image to [-1200; 600], linearly transform the contrast of the lung window image to [0; 1], obtain the lung image using segmentation annotation, and extract pulmonary nodules with a size of 36x36x36 using the pulmonary nodule annotation file.
[0155] Step 2: Input the pulmonary nodules into the built supernet model to train the model, and obtain the weights and architecture parameters of the model;
[0156] Specifically describe this step. Input the pulmonary nodule into the constructed supernet model training model to obtain the weight and architecture parameters of the model, including:
[0157] Set the network search space, that is, the edges in each unit of the supernet. The state space models selected in the present invention are the original receptive field anatomical scan and the 4-way scan state space model, and the double receptive field anatomical scan and the 4-way scan state space model; 3x3x3 max pooling and average pooling; identity mapping.
[0158] Build the ordinary search unit and the downsampling search unit in the supernet. The topological structure is as Figure 7 shown. In the present invention, sigmoid weighting is used for the operations in both the edges and the edges.
[0159] Stack 8 units as the supernet, where the second unit and the fourth unit are downsampling search units, and the rest are ordinary search units.
[0160] Initialize the model parameters and architecture parameters, fix the architecture parameters, train the model parameters until convergence, and select the number of rounds when the accuracy is the highest for the first time in the validation set as a hyperparameter for double-layer optimization.
[0161] Use the high-precision double-layer optimization method to train the supernet. Initialize the model parameters and architecture parameters, fix the architecture parameters, and initialize the model parameters for double-layer optimization: the number of rounds of training the model parameters is the number of rounds when the accuracy is the highest for the first time in the validation set determined in the previous step, and start the double-layer optimization to alternately train the model parameters and architecture parameters.
[0162] Step 3: Select the model architecture through the weight and architecture parameters of the model to obtain the final ordinary unit and downsampling unit;
[0163] Specifically describe this step. Select the model architecture through the weight and architecture parameters of the model to obtain the ordinary unit and downsampling unit, including:
[0164] Multiply the architecture weight of the edge by the architecture weight of the operation to obtain a metric parameter. Select the operation with the largest metric parameter for each edge as the representative of the edge, and compare the weights of the representatives of the input edges of the node. Only the top two input edges with the metric parameter are retained for each node. Finally, obtain the ordinary unit and downsampling unit after search decision.
[0165] Step 4: Stack the ordinary unit and the downsampling unit to the specified layer, input the pulmonary nodule, and train the model weight to obtain the final classification model.
[0166] Specifically describe this step. Stack the ordinary unit and the downsampling unit to the specified layer, input the pulmonary nodule, and train the model weight to obtain the final classification model, including:
[0167] Stack twenty units of the ordinary units and downsampling units obtained in the previous step, where the second unit and the fourth unit are downsampling units, and the rest are ordinary units. Train the model parameters to obtain the final benign and malignant classification model for pulmonary nodules.
[0168] Step 5: Send the pulmonary nodules obtained during the medical auxiliary diagnosis process into the benign and malignant classification model for pulmonary nodules obtained in Step 4 to obtain the discrimination result and confidence level as the criterion for assisting doctors in diagnosing lung cancer.
[0169] In summary, the embodiment of the present invention proposes a method for constructing a search architecture of a three-dimensional state space model with feedback, and the search results are used for 3D pulmonary nodule classification. Compared with other advanced networks, 3DSSS achieves a faster search speed and higher accuracy in this task. By designing 12 scanning paths for pulmonary nodules, the present invention realizes multi-angle analysis of different features, significantly enhancing the diagnostic ability of the classification model. The stability of model training is increased by output feedback, improving the final accuracy. In addition, by combining upsampling and downsampling techniques, the model introduces the ability of multi-scale feature fusion to comprehensively capture the structural information of pulmonary nodules, thereby improving the classification accuracy. Finally, the supernet training method based on the dynamic elimination mechanism proposed by the present invention further improves the model performance by optimizing the training process. These innovative methods jointly promote the development of three-dimensional pulmonary nodule classification research, ensuring the effectiveness and reliability of the model in practical applications. As far as the present invention is concerned, this is the first 3D supernet of SSM. Experimental results on the LUNA16 benchmark dataset show that the model discovered by 3DSSS has excellent classification performance compared with existing methods.
[0170] Figure 4 (a) Uneven depth and Figure 4 (c)'s sampling strategy reduces a large amount of video memory occupancy. The operation weighting function uses sigmoid to make operations compete fairly. In addition, in the double-layer optimization, the model parameters are pre-trained to convergence, thus improving the search accuracy of the double-layer optimization, manifested in completely eliminating the problem of enrichment pooling operations that occur in the search and refreshing the historical record of the model accuracy.
[0171] The application embodiment of the present invention provides a computer device, which includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the method.
[0172] The application embodiment of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the method.
[0173] The application embodiment of the present invention provides an information data processing terminal, and the information data processing terminal includes a system.
[0174] 1. Experiment
[0175] This study conducted a series of experiments to verify the proposed method and empirically analyzed the reasoning process. This paper will give the experimental settings and analyze the corresponding results below.
[0176] 1.1 Setup
[0177] Dataset and preprocessing
[0178] In the experiment, the present invention uses the settings of the LIDC-IDRI dataset and LUNA16. In particular, CT with slice thickness greater than 3 mm, inconsistent slice spacing or missing slices is removed from the LIDC-IDRI dataset. A total of 888 sequences with 1004 nodules are left, of which 450 nodules are positive. The LUNA16 dataset divides the 888 CT sequences into 10 subsets, numbered 0 to 9, which can be used for 10-fold cross validation. Accordingly, in each run, the present invention uses 9 subsets (containing approximately 900 samples) for training and the remaining 1 subset (containing approximately 100 samples) for testing. According to the settings in, the present invention evaluates the method of the present invention at 5 to 9 folds and reports the average performance.
[0179] In the data preprocessing stage, the present invention sets the center of the original data lung window to -600, adjusts the lung window width to [-1200; 600], and linearly transforms the image contrast to [0; 1]. The lung parenchyma is obtained using the segmentation annotation given by LUNA16, and finally the lung nodule annotation file is used to extract the lung nodules.
[0180] 1.2 Evaluation indicators
[0181] In order to evaluate the classification performance of lung nodules, the present invention uses four widely used indicators, accuracy, specificity, sensitivity and F1 score as shown in formulas (1)(2)(3)(5).
[0182] For the task of lung nodule classification, the sensitivity index is more important than other tasks. There are two types of errors in binary classification: false negatives and false positives. In extreme cases, assuming that the model has been able to completely miss positive nodules, doctors will pay less attention to non-positive nodules. Doctors will correct the false positives if they notice them, but this will lead to very fatal false negative errors, and positive nodules will be directly ignored, resulting in misdiagnosis. Therefore, the following index selects the F1 score instead of the AUC to select a model with as few false positives as possible.
[0183]
[0184] Sensitivity, also known as the true positive rate, represents the ability of a classifier to classify positive examples as positive. Specificity, also known as the true negative rate, represents the ability of a classifier to classify negative examples as negative. In the formula, TP, FN, FP, and TN are true positive, false negative, false positive, and true negative in sequence. The larger the values of these criteria, the better the performance. The F1 score evaluates the trade-off between sensitivity and specificity. The F1 score hopes to detect all positive nodules. The harmonic mean is sensitive to small numbers. Since sensitivity, as shown in formula (3), and precision, as shown in formula (4), are between 0 and 1, this metric uses the harmonic mean. The higher the F1 score, the better the performance.
[0185]
[0186] Implementation details
[0187] All experiments were completed on a GeForce RTX 4090 with 24GB of video memory. The hyperparameters of the final model were set as batch_size = 4, epochs = 600, WARMUP_EPOCHS = 20, layers = 20, BASE_LR = 2e-4, MIN_LR = 8e-5.
[0188] Searching for architectures on the Luna16 dataset
[0189] The search results of the present invention in the three-dimensional state space supernet are as Figure 7 shown. Feature compression between ordinary units C_{k} and units realizes multi-scale fusion. And the present invention directly searches for modules on Luna16, without the problem of applying proxy datasets.
[0190] The search results are as Figure 7 shown, where k is the unit output. It can be seen that the unit has a design similar to a residual, that is, k - 2 is used as part of the input. The relatively large number of lines in the ordinary unit is not the result of deterioration. The ordinary unit prefers to select the single-layer SSM_12 operation. The information passes through nodes 1 and 2 simultaneously and then fuses. Node 3 realizes a skip connection structure, and the fusion of information from nodes 0 and 1, 2 realizes a multi-scale effect.
[0191] Table 1 compares with existing methods for folds 5 - 9 on the LUNA16 dataset. Accuracy, Sensitivity, Specificity, F1 Score, and Parameters are represented as Accu., Sens., Spec., F1 S., and Para.
[0192]
[0193] Note: The best and second-best results in each column are highlighted in bold and underlined respectively. A dash (“-”) indicates that the information is not mentioned.
[0194] First, the model of this study was compared with several existing advanced methods, including Multi-crop CNN, Nodule-level 2D CNN, Vanilla 3D CNN, ADNN, DeepLung, AE-DPN, NASLung, and NAS-qa. The performance of the SSM network of this study was reported in Table I.
[0195] Obviously, the 3DSSS network of the present invention achieves the highest accuracy and has the smallest number of model parameters. The high accuracy and F1 score indicate that the 3DSSS network achieves a good balance between sensitivity and specificity. In other words, the 3DSSS network can correctly classify most nodules, which will greatly reduce the burden on doctors. Although this study did not make a trade-off on model parameters, the 3DSSS network naturally has a relatively small number of model parameters, which is also a major advantage of the model of the present invention.
[0196] In Figure 14 and Figure 15 this invention demonstrates the verification results of malignant and benign pulmonary nodules in Set 5 of the Luna16 dataset in the cross-validation experiment. There are 92 pulmonary nodules in Set 5. As can be seen from the Figure 14 and Figure 15 confidence distributions, the classification confidence of most pulmonary nodules is above 0.9, indicating that the model has strong classification performance. To comprehensively demonstrate the model performance, this invention extracts 12 pulmonary nodules with different confidences. Among the selected pulmonary nodules, 6 are malignant and 6 are benign. The confidence in the figure is the result obtained by taking the maximum value after passing the two values output by the model through the Softmax function. Confidence is a partial explanation of the black-box model and also a reference index provided to doctors. Figure 16 shows the confidence distribution of correctly classified pulmonary nodules on the 5th validation set in the cross-validation experiment of the LUNA16 three-dimensional dataset provided by the embodiment of this invention. Through the display of confidence, it can be proved that the model has the classification ability for the benign and malignant of various pulmonary nodules, and the pulmonary nodules with lower confidence can also attract the attention of doctors, further narrowing the range of pulmonary nodules that doctors need to pay attention to, which meets the expectations of this invention.
[0197] 1.3 Comparative Experiments and Analysis
[0198] Before designing the three-dimensional state space super network, this invention modified an excellent manually designed state space model to three dimensions to apply to the pulmonary nodule classification task. This invention found in the experiment that the image scanning path of the input state space will greatly affect the final classification accuracy of the model. As shown in Table II.
[0199] Table 2 Performance of Results of Different Search Methods and Results after Search Space Expansion
[0200]
[0201] Table 3 Model Hyperparameter Settings
[0202]
[0203] As one of the earliest Mamba vision models, ViM is modified by the present invention into 3D_ViM for classifying pulmonary nodules. After a large number of hyperparameter adjustments, DEPTHS = 24, EMBED_DIM = 192, STATE = 16, and PATCH_SIZE = 4 are the best results in the experiment, which are shown in Table 2. The results show that the early dual-path Mamba classification model has weak ability in pulmonary nodule classification.
[0204] VMamba adds a scanning path in front of the state space module. After the present invention modifies VMamba into 3D_VMamba_4, its accuracy performance in the pulmonary nodule classification task can rank among excellent models. After a large number of hyperparameter adjustments, DEPTHS = [2, 2, 20, 2], EMBED_DIM = 96, STATE = 16, and PATCH_SIZE = 2 are the best combinations in the experiment, and the experimental results are shown in Table 2. Here, the performance of STATE = 16 is better than 32. When STATE = 32, the model vibrates at a large learning rate and does not converge after the learning rate decreases.
[0205] As shown in the table after controlling variables, changing 3D_VMamba_4 to 3D_VMamba_12, the anatomical scan is as Figure 2 shown, which improves the comprehensive performance of the model as shown in Table 2. The hyperparameters of 3D_VMamba_12 are the same as those of 3D_VMamba_4. The only difference is that STATE = 32. It is not that the present invention does not control variables, but that the accuracy of the STATE = 32 model of 3D_VMamba_4 oscillates roughly in the range of 71.73% to 80.43%. Instead, 3D_VMamba_12 with STATE = 32 converges quickly at a large learning rate, and the training accuracy can reach 99%. This shows that the feature of 3D_VMamba_12 is relatively flat near the extreme value and has strong stability. 3D_VMamba_4 and 3D_VMamba_12 prove the superiority of the anatomical scan of the present invention.
[0206] Regarding the hyperparameter settings, the present invention finds that when the combined quantity of EMBED_DIM and STATE is too large, the model will experience oscillations and gradient explosion, and the reason remains to be explained. In other aspects, a key parameter of 3D_VMamba_12 is PATCH_SIZE = 2, which encodes [1, 32, 32, 32] into [96, 16, 16, 16]. Compared with PATCH_SIZE = 4, that is, [96, 8, 8, 8], the sequence length increases by two times, which better conforms to the advantages of Mamba for long sequence modeling. Most of the hyperparameters in the hypernetwork design of the present invention draw on the settings of 3D_VMamba_12. For example, fewer units are placed in the first two stages of the hypernetwork, more units are placed in the third stage, PATCH_SIZE = 2, and EMBED_DIM = 96, etc.
[0207] Table IV Comparison between the hypernetwork search VSS operation and SSMOnly
[0208]
[0209] As shown in Table III, the parameter quantity of 3D_VMamba_12 reaches 64.74M. The present invention believes that Δ in the discretized state space model given by mamba already has the effect of the attention mechanism. Compared with VSS in 3D_VMamba_12, the hypernetwork multi-scale SSM structure designed by the present invention deletes the spatial attention mechanism in 3D_VMamba_12, including convolutional dimension elevation and reduction, separable convolution, and FFN. That is, the present invention only retains the SSM part and introduces upsampling and downsampling to make up for the problem of fixed SSM vision. As shown in Table IV, only the SSM operation still performs excellently in the end. The search space of VSS in the table is PRIMITIVES = ['Skip_Connect', 'VSS_12', 'VSS_Double_12', 'VSS_4', 'VSS_Double_4'], and the search results are as Figure 18 ; Limited by the video memory of RTX4090, the number of unit stacks here is 10 layers.
[0210] Table V Comparison of two sampling strategies of the hypernetwork
[0211]
[0212] Table VI Comparison of the hypernetwork search using the knockout strategy
[0213]
[0214] Sample1 mimics the scheme of halving the feature map size and doubling the number in DARTS. However, SSM does not have the ability to change the feature map size and number. Therefore, when designing Sample1, the present invention simply loads the convolution before the SSM module to obtain the sampling ability. In this way, 56 additional convolution operations are required for the supernet, which is contrary to the idea of optimizing VSS in the present invention. Therefore, the present invention designs Sample2 to perform sampling between cells, so that 56 convolutions are compressed into 2 convolution operations. Relatively speaking, the downsampling learning ability may be weakened, but experiments show that it has no impact on the accuracy. As shown in Table 5, Sample2 performs better.
[0215] Based on Sample2, a knockout tournament is carried out, which makes the supernet closer to the finally selected model to improve the search accuracy. And after several rounds of knockout tournaments, the network is simplified, reducing the search time. As shown in Table 6, the model searched after introducing the knockout tournament has better accuracy and faster speed.
[0216] Table 7 Influence of the EMBED_DIM hyperparameter of the supernet on the network search results
[0217]
[0218] For EMBED_DIM_32->96, the present invention sets EMBED_DIM in the supernet to 32, and uses EMBED_DIM = 96 for training the searched model. Here, it continues the practice in DARTS of increasing the dimension during the training of the convolution supernet search results to reduce the supernet search time and improve the performance of the searched network. However, this does not perform well in the three-dimensional state space supernet. Based on a large number of preliminary experiments in Table 2, the present invention finds that there is a certain constraint relationship between EMBED_DIM and STATE, which cannot be explained simply by the way of increasing the feature dimension with convolution. As shown in the results, directly using EMBED_DIM = 96 to search the supernet will cause the supernet to choose lines and pooling to simplify the network, which is also the reason why pooling must be added in the design of the present invention's supernet. The network searched with EMBED_DIM set to 32 tends to choose more complex models. When the EMBED_DIM of the searched network is simply set to 96, the training of the searched network converges extremely unstably under a relatively large learning rate and can only converge under a relatively small learning rate. Unfortunately, here the present invention sets EMBED_DIM to 96 for searching a better model. Although it prolongs the search time, the overall search accuracy is higher.
[0219] The network results searched with EMBED_DIM_32->96 are as Figure 19 ; The network results searched with EMBED_DIM_96 are as Figure 20 。
[0220] 1.4 Visualization of Classification Feature Distribution
[0221] Visualize the output of the penultimate layer in the SSM network. The present invention uses t-SNE to reduce each feature vector to 2D and visualize it in Figure 17 . Figure 17 In, the orange (50% transparency) dots are malignant nodules, and the blue (50% transparency) dots are benign nodules. The present invention uses the scikit-learn toolbox in the implementation. Since the research on class activation mapping has not extended to the SSM model yet, the present invention visualizes the feature distribution learned by the SSM network through t-SNE, and makes a visual explanation of the capabilities of the SSM network model based on the visualization of the feature distribution. As Figure 17 shown in (a), directly projecting the three-dimensional lung nodule image into two dimensions, the malignancy and benignancy of the lung nodules are not significant, and it is even impossible to distinguish the malignancy and benignancy through the projection. As Figure 17 shown in (b), projecting the features learned by the SSM network into a low dimension can more effectively distinguish the malignancy and benignancy of lung nodules, proving the effectiveness of the SSM network learning. At the same time, the model is generated by the lower 5 folds in the ten-fold cross-validation, and 1004 lung nodules are all the samples. (b) also reflects the consistency of the training accuracy and the validation accuracy.
[0222] Of course, there are cases where the learned features that do not need to be concerned can also be classified. For example, there is a model that learns the blue sky to classify airplanes. Therefore, further research is needed to achieve an explanation of the relationship between the learned features and the input pictures.
[0223] Existing automated machine learning has been relatively complete in traditional algorithms. For example, Auto-sklearn poses a huge challenge to manually designed algorithms on relatively simple data, and Auto-sklearn has entered the industrial algorithm design and use.
[0224] Neural network architecture search is a sub-task of AutoML. In this invention, a supernet for searching the SSM model is designed on a three-dimensional lung nodule dataset, filling the gap in AutoML and laying a foundation for automatically integrating deep learning models after the further development of computing devices in the future. The way of using cross-validation in this invention also follows the evaluation method of existing AutoML training models. Deep learning shows obvious performance bottlenecks in three-dimensional lung nodule classification. Through a large amount of research, this invention believes that on the one hand, there is uncertainty in the benign and malignant annotation of lung nodules, and on the other hand, the quantity of data is too small to meet the requirements of the model. Specifically, malignant lung nodules have various morphologies, such as solid nodules, semi-solid nodules, and sub-solid nodules, etc. This invention finds that the probability of misclassification of sub-solid nodules is relatively high. Overall, the ratio of negative to positive lung nodules selected in this invention is 1 to 0.81. From the perspective of binary classification, the data is relatively uniform, but the number of sub-solid nodules in the positive samples is small, making them difficult classification samples. From the data distribution of the positive samples, the samples are extremely uneven. Naturally, the number of negative lung nodule samples far exceeds that of positive samples. Obviously, this invention has balanced the dataset. Currently, deep learning performs poorly on imbalanced datasets, while traditional machine learning methods cannot automatically learn features. Without finding the features of positive lung nodules, the advantages of traditional machine learning on imbalanced datasets cannot be exerted. The current bottleneck in the benign and malignant classification of lung nodules does not mean the temporary end of this research, but indicates the real start of the research. In the future, this invention will find the reasons for the performance bottleneck of the benign and malignant of lung nodules and improve the classification performance from multiple perspectives such as datasets, preprocessing, models, and medical analysis.
[0225] 2. Conclusion
[0226] A supernet for searching the SSM network, called the 3D Anatomical multi-scale SSM supergraph (3DSSS), is designed, and the search results are used for 3D lung nodule classification. Compared with other advanced networks, 3DSSS achieves faster search speed and higher accuracy in this task. By designing 12 lung nodule scanning paths, this invention realizes multi-angle analysis of different features, significantly enhancing the diagnostic ability of the classification model. In addition, combined with upsampling and downsampling techniques, the model introduces multi-scale feature fusion ability to comprehensively capture the structural information of lung nodules, thus improving the classification accuracy. Finally, the supernet training method based on the dynamic elimination mechanism proposed in this invention further improves the model performance by optimizing the training process. These innovative methods jointly promote the development of three-dimensional lung nodule classification research and ensure the effectiveness and reliability of the model in practical applications. As far as this invention knows, this is the first 3D supernet of SSM. Experimental results on the LUNA16 benchmark dataset show that the model discovered by 3DSSS has excellent classification performance compared with existing methods.
[0227] 3DSSS can be used for the automated analysis and classification of three-dimensional objects, scenes, or environments. This has important applications in fields such as autonomous driving, intelligent security, and robot navigation. For example, in autonomous driving, 3DSSS can help identify and classify objects in complex road scenes, such as pedestrians, vehicles, obstacles, etc., thereby improving the vehicle's decision-making ability.
[0228] This technology can be used for the automatic classification and diagnosis of three-dimensional medical images. For example, 3DSSS can accurately classify organs, tumors, or lesion areas in three-dimensional medical images such as CT and MRI, helping doctors to make more efficient diagnoses and treatment plans.
[0229] In the applications of virtual reality or augmented reality, 3DSSS can help quickly identify and classify different elements in the virtual environment, enhancing the user interaction experience. For example, in education, entertainment, or industrial training, the system can adjust and present corresponding three-dimensional scenes in real time according to user needs.
[0230] In industrial automation and robot control, 3DSSS can be used to classify and analyze the interaction between robots and the surrounding environment, especially for precise operations in complex three-dimensional spaces. For example, 3DSSS can help industrial robots automatically sort, assemble, and transport goods in a warehouse.
[0231] 3DSSS can be applied in geographic information systems for the automatic analysis and classification of three-dimensional terrain data or urban building models. Through this method, it can provide efficient data analysis and decision-making support for urban planning, disaster prediction, and environmental protection.
[0232] In the field of game development, 3DSSS can be used to intelligently generate and optimize the three-dimensional virtual world in games, enhancing the details and interactivity of the game environment, making the game more realistic and immersive.
[0233] Just as automated machine learning has transformed traditional machine learning through automated model design, neural network architecture search technology (NAS) will also revolutionize deep learning in a similar way. Although AutoML has been widely adopted in the industry for automatically searching and integrating traditional machine learning models, as its counterpart in deep learning, it can automatically discover the best neural network architecture. With the continuous development of deep learning, NAS is expected to play a leading role in model design. With the continuous progress of computing hardware, NAS can ultimately automatically search and integrate all existing deep neural network architectures, reflecting the impact of AutoML on traditional machine learning.
[0234] The model trained by the present invention on LUNA16 can be integrated into the computer tomography-assisted diagnosis system of a hospital, and examples are Figure 12As shown. An example of the network application obtained by the search method of the present invention through convolutional supernet search is as Figure 13 shown.
[0235] As mentioned above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for constructing a three-dimensional state space model search architecture with feedback, characterized in that Including: Step 1: Define the operation space; Add average pooling, max pooling, and skip connection operations to the operation space; By adding a downsampling module before the state space, the state space model obtains a doubled receptive field. After the state space model, add upsampling to fuse the image information with the output of the state space model with the original unsampled receptive field. Upgrade the state space model to a three-dimensional state space model through anatomical scanning; introduce output feedback to improve the performance of the state space model; Step 2: Construct ordinary search units and downsampling search units in the operation space, stack the ordinary search units and downsampling search units, and add an image feature embedding module before the image is fed into the stacked network to form a supernet; Both the ordinary search unit and the downsampling search unit are directed acyclic graphs of the same structure; the directed acyclic graph contains multiple nodes, each node represents a feature map, the edges between the nodes are composed of all operations in the operation space to form a mixed operation, the direction of the edge represents the direction of the information flow, and each edge and each operation in the edge have corresponding architecture parameters; Step 3: Search for the architecture weights of each edge and each operation in the directed acyclic graph of the supernet through a high-precision two-layer optimization method; Step 4: Use the product of the architecture weight and the corresponding operation architecture weight as the final weight, obtain the operation with the largest final weight in each edge, and use the corresponding final weight as the final weight of the edge. For units with more than two input edges, only keep the top two edges with the largest final weight of the edge. Stack and update the ordinary search unit and the downsampling search unit according to the results of the edge and the operation until the required number of layers to obtain the final model architecture; If it is necessary to expand the operation space, start from Step 1 again according to the receptive field tendency of the updated unit to obtain new search results, stack and update the ordinary search unit and the downsampling search unit to the required number of layers to obtain the final model architecture; Step 1 specifically includes: The state space model is regarded as a linear time-invariant system, which maps the input x(t) ∈ N to the output response y(t) ∈ L through the hidden state h(t) ∈ L ; they are usually expressed as linear ordinary differential equations as follows: h′(t) = Ah(t) + Bx(t) y(t) = Ch(t) + Dx(t) where the domain of the coefficient matrix is \(A\in N×N \), \(B, C\in N \), \(D\in 1 \).
2. The method for constructing a three-dimensional state space model search architecture with feedback according to claim 1, characterized in that, To integrate it into a deep model, it is necessary to discretize the continuous-time state space model in advance, which is achieved by solving ordinary differential equations and then performing a simple discretization process; the results are as follows: y k = Ch k + Dx k Among them, the expression of the coefficient matrix is To increase the stability of the state space model, introduce a feedback mechanism in the state space model and maintain the performance of cub::BlockScan parallel scanning in Mamba; In actual use, the adjusted continuous state space model: h′(t) = Ah(t) + Bx(t) y(t) = C(h(t) + h′(t)) + Dx(t) Derivation of adding output feedback to the continuous state space model: Where K is the feedback coefficient; Simplify the feedback to a certain calculation structure into the following formula and experimentally obtain a performance improvement; h′(t) = KCAh(t) + KC(B + KD)x(t).
3. The method according to claim 1, wherein There are two alternative sampling schemes for the downsampling unit in the supernet; one is to add downsampling before each operation without downsampling ability among the edges starting from input nodes 0 and 1 in the sampling unit; the other is to add downsampling after all the edges input to the sampling unit, and the sampling unit only serves as an independent variable unit adapting to sampling; the downsampling operation is achieved by adjusting the resolution and the number of feature channels of the input feature map, and the resolution of the sampled feature map is halved while the number of feature channels is doubled.
4. The method according to claim 1, wherein In the high-precision double-layer optimization method, the architecture weights of the supernet are determined through the following steps: Fix the architecture parameters, input the training set into the supernet, calculate the cross-entropy loss of the training set, and train until the model parameters converge; Fix the model parameters, input the training set into the supernet, and calculate the cross-entropy loss of the training set; Use backpropagation to calculate the gradient of the architecture weights; Adjust the architecture weights by one round according to the gradient direction; Fix the architecture parameters, input the training set into the supernet, and calculate the cross-entropy loss of the training set; Use backpropagation to calculate the gradient of the architecture weights; Adjust the model weights by one round according to the gradient direction; Alternately optimize the architecture weights and the model weights until convergence.
5. The method according to claim 1, characterized in that, The search unit in the supernet includes a directed acyclic graph structure composed of multiple edges, and the feature output of each edge is obtained through weighted sum of features.
6. The method according to claim 1, characterized in that, In the supernet design, the ordinary unit and the downsampling unit respectively include a residual connection structure, and the residual connection is realized through skip connections between nodes, without introducing additional parameters and maintaining the stability of the network.
7. A three-dimensional state space model search architecture construction system with feedback for implementing the method for constructing a three-dimensional state space model search architecture with feedback as described in any one of claims 1 to 6, characterized in that It includes: An operation space definition module: used to define the operation space; A supernet formation module, used to construct ordinary search units and downsampling search units in the operation space, stack the ordinary search units and the downsampling search units to form a supernet; An architecture weight search module, used to search for the architecture weights of each edge and each operation in the directed acyclic graph of the supernet through the high-precision double-layer optimization method; A final architecture obtaining module, used to use the product of the architecture weights and the corresponding operation architecture weights as the final weights, obtain the operation with the largest final weight in each edge, use the corresponding final weight as the final weight of the edge, and stack and update the ordinary search units and the downsampling search units with the results of the edge and the operation to obtain the final model architecture.
8. A computer device, the computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the method for constructing a three-dimensional state space model search architecture with feedback as described in any one of claims 1 to 6.
9. A computer-readable storage medium, storing a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of the method for constructing a three-dimensional state space model search architecture with feedback as described in any one of claims 1 to 6.
10. An information data processing terminal, the information data processing terminal includes the system for constructing a three-dimensional state space model search architecture with feedback as described in claim 7.
Citation Information
Patent Citations
Search method for three-dimensional model of mixing characteristic based on feedback
CN101359342A
Causal analysis and graph neural network data generation method based on causal decision tree
CN116882558A