Lymphoma focus positioning method based on nnFormer energization
Through the nnFormer model, the feedforward neural network combining factorization and channel attention mechanisms, the problem of low recognition accuracy of lymphoma lesions is solved, and the efficient, precise positioning and detailed information of lymphoma lesions are achieved, which improves the accuracy and efficiency of diagnosis.
Patent Information
- Application Number
- CN202510157658.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-19
- Filing Date
- 2025-02-13
- Publication Date
- 2025-07-11
AI Technical Summary
The existing deep learning models have problems with insufficient recognition accuracy and insufficient processing capabilities for complex lesions in lymphoma lesions. Traditional imaging diagnosis methods rely on doctor experience and are susceptible to human factors, making it difficult to achieve accurate positioning and quantitative analysis.
The nnFormer model is used to combine factorization and channel attention mechanism feedforward neural network, through data cleaning, filtering technology and regional suggestion network, combined with self-attention and channel attention mechanism, to realize automated processing and in-depth analysis of tumor lesions, and use the regional suggestion network to generate candidate regions and accurately locate lesions through regression algorithms.
It realizes efficient and precise positioning of lymphoma lesions, improves identification accuracy and processing capabilities for complex lesions, provides detailed key lesions information, reduces artificial errors, and improves diagnosis accuracy and efficiency.
Smart Images

Figure CN120298488A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for lesion localization, and more particularly to a lymphoma lesion localization system empowered by nnFormer. Background Art
[0002] Lymphoma is a malignant tumor originating from the lymphohematopoietic system, and its early detection and accurate diagnosis are crucial for the treatment and prognosis of patients. Traditional lymphoma diagnosis methods mainly rely on medical imaging techniques such as ultrasound, computed tomography (CT), magnetic resonance imaging (MRI), and positron emission tomography (PET). Although these techniques can show the approximate location and morphology of lymphoma lesions, they have limitations in the precise localization and quantitative analysis of lesions.
[0003] Traditional imaging diagnosis methods rely on doctors' experience and expertise, and the recognition accuracy and diagnosis efficiency of lesions are affected by human factors. In addition, the morphology and location of lymphoma lesions are diverse and may be confused with other tissues or lesions, increasing the difficulty of diagnosis. Therefore, it is of great significance to develop a system that can automatically, accurately, and efficiently identify and localize lymphoma lesions.
[0004] With the development of artificial intelligence and deep learning technologies, image recognition technologies based on neural network models provide new solutions for the precise localization of lymphoma lesions. Deep learning algorithms can automatically learn the features of image data and, through the training of a large amount of data, achieve high-precision recognition and localization of lymphoma lesions. However, there are still some problems in the existing deep learning models for lymphoma lesion recognition, such as insufficient recognition accuracy and insufficient processing ability for complex lesions. Summary of the Invention
[0005] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide an innovative lymphoma lesion ultra-precise localization system empowered by nnFormer, aiming to use advanced deep learning algorithms and neural network models to achieve automatic, accurate, and efficient recognition and localization of lymphoma lesions, providing strong support for the precise treatment of lymphoma.
[0006] To achieve the above purpose, the present invention provides the following technical solution: A lymphoma lesion localization system empowered by nnFormer, characterized by comprising:
[0007] A user layer that interacts with external users, used to collect image data from external users or feedback the localization result to external users;
[0008] A data management layer that communicates with the user layer, used to receive the collected image data and perform storage management on the image data;
[0009] The AI intelligent processing module is communicatively connected to the data management layer to call the image data stored in the data management layer, and processes the image data to output a lesion localization image;
[0010] The front-end medical view module communicates with the user layer. The front-end medical view module includes view basic tools, a window level control module, a multi-section control module, and a color control module, and is used to be displayed on the user layer display screen for the user to select corresponding modules and options;
[0011] Among them, the specific steps for the AI intelligent processing module to process the image data are as follows: First, perform data cleaning on the image data, convert the image data to data combined with the XY coordinate axes, and then input it into the tumor molecular imaging segmentation model. Through the tumor molecular imaging segmentation model, automated processing and in-depth analysis are performed to identify and locate the tumor lesion position, and then an image with the tumor lesion position displayed is output.
[0012] As a further improvement of the present invention, the tumor molecular imaging segmentation model adopts the nnFormer model, and factorization is introduced into the feed-forward neural network structure of the nnFormer model to perform feed-forward network calculations. At the same time, a channel attention mechanism module is embedded before the feed-forward neural network to integrate the channel attention mechanism after the local window attention mechanism processes the feature map x.
[0013] As a further improvement of the present invention, the factorization method in the feed-forward neural network structure is to decompose the weight matrix of the traditional fully connected layer into multiple low-rank matrices, and use two smaller fully connected layers to implement this decomposition process, specifically as follows: In a standard two-layer MLP, the forward propagation of the feed-forward network can be expressed as:
[0014] h = α(W1x + b1)
[0015] y = W2h + b2
[0016] In the formula, x is the input vector; W1 and W2 are weight matrices; b1 and b2 are bias vectors; α is the activation function; h is the output of the hidden layer; y is the network output;
[0017] In the factored FFN, the weight matrices W1 and W2 are decomposed into multiple lower-rank matrices, that is, each weight matrix is decomposed into the product of two smaller matrices.
[0018] As a further improvement of the present invention, the specific method of decomposing the weight matrices W1 and W2 into multiple lower-rank matrices is: decompose W1 into U1V1, and decompose W2 into U2V2. In this way, the forward propagation process can be re-expressed as:
[0019] Z1 = U1X
[0020] h = σ(V1Z1 + b1)
[0021] Z2 = U2h
[0022] y = V2Z2 + b2
[0023] Wherein, U1 and V1 are two smaller matrices used to decompose W1, U2 and V2 are two smaller matrices used to decompose W2, and Z1 and Z2 are intermediate variables.
[0024] As a further improvement of the present invention, the specific manner in which the local window attention mechanism processes the feature map X is as follows: The local volume self-attention mechanism performs self-attention calculation within a three-dimensional local volume and combines with the multi-head self-attention mechanism to form a skip attention structure. Assuming the input feature map is x ∈ R H×W×C , where H and W are the height and width of the feature map respectively, and C is the number of channels, the local window attention mechanism can be expressed as:
[0025]
[0026] Wherein, Q, K, and V are the query, key, and value matrices respectively, and B is the relative position encoding, which are usually obtained from the input feature map x through linear transformation.
[0027] As a further improvement of the present invention, the skip attention structure is based on the three-dimensional local volume self-attention mechanism. During the process of upsampling the low-resolution feature map to a high-resolution feature map, the feature map is combined with the encoder features through the skip attention structure as follows:
[0028] Assume that the output of the l-th layer Transformer block of the encoder is After linear projection, the key matrix is obtained and the value matrix
[0029]
[0030] Wherein, LP represents the linear projection operation. Assume that the feature map output after upsampling in the l*-th layer of the decoder is which is regarded as the query matrix The skip connection attention structure can be summarized as:
[0031]
[0032] Wherein, l* represents the upsampling layer number of the decoder.
[0033] As a further improvement of the present invention, the method for integrating the channel attention mechanism is as follows: for the upsampled feature maps at each level, calculate their local multi-head self-attention and channel attention respectively, and fuse the spatial context information and global feature information through residual connection, specifically as follows:
[0034] Apply the local window attention mechanism to the input feature map x to obtain the output feature map x attn .
[0035] For x attn Apply global average pooling to obtain the channel descriptor z ∈ R C :
[0036]
[0037] Learn the correlation between channels through a fully connected layer and a ReLU activation function to obtain the intermediate feature s ∈ R C :
[0038] s = σ(W2δ(W1z))
[0039] where and are the weight matrices of the fully connected layer, r is the scaling ratio, δ is the ReLU activation function, and σ is the sigmoid activation function;
[0040] Obtain the weight of each channel through the sigmoid activation function and use it to recalibrate the channels of x attn :
[0041]
[0042] Apply a feed-forward neural network to the feature map x recalibrated by the channel attention mechanism to obtain the final output; and the specific way to perform residual connection on the spatial context information captured by the self-attention mechanism and the global feature information captured by the channel attention is as follows: attn
[0043]
[0044] x output = Concat(x channel_atten , x self_attention ).
[0045] As a further improvement of the present invention, the way the AI intelligent processing module performs data cleaning is to use filtering technology to remove noise in medical images, improve image quality, and then standardize images from different imaging devices or different patients to eliminate image deviations caused by differences in imaging conditions, making subsequent analysis more consistent. After that, techniques such as histogram equalization are applied to enhance the contrast of the images, making tumor lesions more obvious in the images. Finally, the input images are uniformly adjusted to the size required by the model to ensure that the deep learning model can process them correctly.
[0046] As a further improvement of the present invention, the specific way to implement automated processing and in-depth analysis through the tumor molecular imaging segmentation model to identify and locate the position of tumor lesions is as follows: First, classify the extracted features to identify whether there are tumor lesions in the image, and use the region proposal network to generate candidate regions that may contain tumors, providing a basis for subsequent positioning. Set a threshold to screen the candidate regions to ensure that only high-confidence tumor recognition results are retained.
[0047] As a further improvement of the present invention, the data management layer adopts a hierarchical storage architecture, classifies and stores data according to access frequency and importance, including hot data and cold data, encrypts all sensitive data before storage, records audit logs of all data access and operations, and also has automated regular backups, providing a secure remote access channel for the user layer. Through VPN or SSL encryption technology, it ensures the security of data during transmission.
[0048] The beneficial effect of the present invention is the improvement of the FFN structure. To further enhance the non-linear expression ability of the model, on the basis of the original standard two-layer multi-layer perceptron (MLP), a factorized feed-forward network (FFN) is introduced to perform feed-forward network calculations. By decomposing the weight matrix of the traditional fully connected layer into multiple low-rank matrices and using two smaller fully connected layers to implement this decomposition process, and after processing the feature map x by the local window attention mechanism, a channel attention mechanism is integrated. To further improve the model's understanding ability of the complex relationships between different channels within each window, a channel attention mechanism module is embedded before the subsequent feed-forward neural network (FFN). The tumor lesion precise positioning and analysis module in the AI intelligent processing module relies on the innovative nnFormer architecture after the above optimization and upgrade, that is, the tumor molecular imaging segmentation model, to perform automated processing and in-depth analysis on medical images obtained by ultra-low radioactivity imaging technology. It can efficiently and accurately identify and locate the position of tumor lesions and at the same time provide detailed key information of the lesions. Brief Description of the Drawings
[0049] Figure 1 It is the overall structural module block diagram of this system.
[0050] Figure 2 It is a schematic diagram of an image processing architecture that includes ID input and a complex segmentation stage.
[0051] Figure 3 It is a schematic diagram of the tumor lesion segmentation result of the system.
[0052] Figure 4 It is a schematic diagram of the right sidebar of the system.
[0053] Figure 5 It is a schematic diagram of the main interface of this system.
[0054] Figure 6 It is a schematic diagram of the comparative experiment results of the innovative nnFormer. Detailed implementation manners
[0055] Next, the present invention will be further described in detail with reference to the embodiments given in the accompanying drawings.
[0056] Refer to Figure 1 as shown Figure 1 It shows a complex flowchart for web - end users, mainly serving groups such as intermediaries, clinicians, clinical technicians, scientific researchers, and students. The whole process from data acquisition to model application is described in detail. In the data acquisition stage, users can collect, clean, and extract data features through various methods. Then, in the data processing stage, the data will be transformed and enhanced to improve its diversity and robustness. Next, in the model construction stage, users can use the training dataset to train the model and verify it through the validation dataset to evaluate the performance and accuracy of the model. In the model evaluation stage, users will comprehensively evaluate the model and optimize the model according to the evaluation results. Finally, in the model application stage, users can deploy the trained and optimized model to the actual application scenario to solve practical problems.
[0057] In addition, the flowchart also includes various elements, terms, and functional modules related to the system. For example, icons such as system architecture, user interface, and algorithm framework show the overall structure and components of the system. At the same time, the flowchart also involves specific terms or API interface names such as DICOM, backend / uploac, length, uploadTumorCT, as well as data or status information such as image data, user accounts, lab segmentation status, login status, and permission status that are stored or managed. In addition, the system also provides rich image processing or operation functions, such as highlighting lesions, oil forwarding slices, rotation, etc., and components or functional modules of the front - end interface, such as medical views, window level control modules, etc.
[0058] Therefore Figure 1 the following system construction is specifically disclosed in
[0059] The user layer interacts with external users, used to collect image data from external users or feedback the positioning result to external users;
[0060] The data management layer communicates with the user layer, used to receive the collected image data and perform storage management on the image data;
[0061] The AI intelligent processing module is communicatively connected to the data management layer to call the image data stored in the data management layer and process the image data to output a lesion location image;
[0062] The front-end medical view module communicates with the user layer. The front-end medical view module includes view basic tools, window level control module, multi-section control module and color control module, and is used to be displayed on the user layer display screen for the user to select the corresponding modules and options;
[0063] Among them, the specific steps for the AI intelligent processing module to process the image data are as follows: First, perform data cleaning on the image data, convert the image data to data combined with the XY coordinate axes, and then input it into the tumor molecular imaging segmentation model. Through the tumor molecular imaging segmentation model, automated processing and in-depth analysis are performed to identify and locate the tumor lesion position. After that, an image with the tumor lesion position marked is output. The tumor molecular imaging segmentation model used in this embodiment is an improved version of the nnFormer model. The structure of the nnFormer model is based on the U-shaped structure of U-Net, mainly composed of three parts: an encoder, a bottleneck, and a decoder. At the same time, it also has skip connections and multi-head self-attention (MSA), etc. The design of the nnFormer model structure enables it to effectively process 3D medical image segmentation tasks. By combining the advantages of convolution and Transformer and introducing local and global self-attention mechanisms, the performance and efficiency of the model are improved. Therefore, the existing structure and functions of the nnFormer model are prior art, so they will not be elaborated in this embodiment. In this embodiment, the improvements to the model in this embodiment are mainly described, and there are the following points:
[0064] Such as Figure 1 and Figure 2As shown, in this embodiment, the structure of the feed-forward network (FFN) is first optimized. Traditional FFNs usually use a standard two-layer multi-layer perceptron (MLP) to perform feed-forward network calculations. However, to further enhance the model's non-linear representation ability, the present invention introduces a factorized feed-forward network. The core of this improvement lies in decomposing the weight matrix of the traditional fully-connected layer into multiple lower-rank matrices, and implementing this decomposition process through two smaller fully-connected layers. Such a design not only reduces the number of model parameters but also improves the model's computational efficiency while maintaining a strong non-linear mapping ability.
[0065] At the same time, in addition to the improvement of the FFN structure, this embodiment also integrates a channel attention mechanism to further enhance the model's performance. After the local window attention mechanism processes the feature map, we embed a channel attention mechanism module. This module can capture the complex relationships between different channels within each window, thereby enhancing the model's understanding ability of the features of each channel. Combined with the factorized FFN, the model can more effectively process the features of the input image and extract more representative information. Such a design not only improves the model's accuracy but also enhances the model's generalization ability for complex image data. Through these two major improvements, the model performs excellently in four metrics: Dice, Jaccard, Precision, and Recall. In particular, it reaches the highest score in the Precision metric, indicating its very high accuracy in predicting positive samples.
[0066] This embodiment preferably adopts the following solutions for the above improvements: optimizing the structure of the feed-forward network (FFN) to enhance the model's non-linear representation ability by decomposing the weight matrix of the fully-connected layer, while reducing the number of parameters and computational complexity. Decompose the weight matrix into multiple lower-rank matrices and implement this decomposition process through two or more smaller fully-connected layers.
[0067] In a standard two-layer MLP, the forward propagation of the feed-forward network can be expressed as:
[0068] h = α(W1x + b1)
[0069] y = W2h + b2
[0070] Where x is the input vector; W1 and W2 are weight matrices; b1 and b2 are bias vectors; α is the activation function; h is the output of the hidden layer; y is the output of the network.
[0071] In the factorized FFN, the weight matrices W1 and W2 are decomposed into multiple lower-rank matrices. For simplicity, each weight matrix is decomposed into the product of two smaller matrices.
[0072] Specifically, decompose W1 into U1V1 and W2 into U2V2. In this way, the forward propagation process can be re-expressed as:
[0073] Z1 = U1X
[0074] h = σ(V1Z1 + b1)
[0075] Z2 = U2h
[0076] y = V2Z2 + b2
[0077] Where U1 and V1 are two smaller matrices used to decompose W1, U2 and V2 are two smaller matrices used to decompose W2, and Z1 and Z2 are intermediate variables.
[0078] It can be seen that through factorization, the number of parameters and computational complexity can be significantly reduced. Assume that the size of the original weight matrix W1 is d h ×d x , and the size of W2 is d y ×d h . If each weight matrix is decomposed into two matrices of size d h ×r and r×d x , then the number of parameters will be reduced from d h d x +d y d h to 2d h r + 2rd x +d y r. When r < min(d x , d h , d y ), the number of parameters will be significantly reduced.
[0079] Furthermore, the factored feedforward network enhances the model's non-linear representation ability while reducing the number of parameters and computational complexity by decomposing the weight matrix of the fully connected layer into multiple lower-rank matrices and implementing this decomposition process through two smaller fully connected layers. This method helps reduce the overfitting risk and computational cost of the model while maintaining its performance.
[0080] Integrating the channel attention mechanism to further improve the model's performance is to use the enhanced model that combines the three-dimensional local volume self-attention mechanism and the channel attention mechanism to improve the model's understanding ability of the input feature map;
[0081] Specifically, the three-dimensional local volume self-attention mechanism effectively captures the long-term dependencies and hierarchical features in high-dimensional data by performing self-attention calculations within a three-dimensional local volume and combining the multi-head self-attention mechanism. Assume the input feature map is x ∈ R H×W×C, where H and W are the height and width of the feature map respectively, and C is the number of channels. The local window attention mechanism can be expressed as:
[0082]
[0083] where Q, K, and V are the Query, Key, and Value matrices respectively, and B is the relative position encoding, which are usually obtained from the input feature map x through linear transformation. In 3D local window attention, these matrices are divided into multiple small windows for processing.
[0084] Furthermore, the skip attention structure is based on the 3D local volume self-attention mechanism. During the process of upsampling the low-resolution feature map to a high-resolution feature map, the feature map is combined with the encoder features through the skip attention structure, so as to better capture semantic and fine-grained information. Suppose the output of the l-th Transformer block in the encoder is After linear projection, the key matrix is obtained and the value matrix
[0085]
[0086] where LP represents the linear projection operation. Suppose the feature map output after upsampling in the l-th * layer of the decoder is which is regarded as the query matrix The skip connection attention structure can be summarized as:
[0087]
[0088] where l * represents the upsampling layer number of the decoder.
[0089] After applying the skip attention structure to the feature map, the channel attention mechanism is further integrated to enhance the model's ability to understand the complex relationships between channels. Specifically, for each level of upsampled feature map, its local multi-head self-attention and channel attention are calculated respectively, and the spatial context information and global feature information are fused through residual connection. Then, the local window attention mechanism is applied to the input feature map x to obtain the output feature map x attn .
[0090] Applying global average pooling to x attn obtains the channel descriptor z ∈ R C :
[0091]
[0092] Learn the correlation between channels through a fully connected layer (or two fully connected layers) and the ReLU activation function to obtain the intermediate feature s ∈ R C :
[0093] s = σ(W2δ(W1z))
[0094] where and are the weight matrices of the fully connected layers, r is the scaling ratio, δ is the ReLU activation function, and σ is the sigmoid activation function.
[0095] Obtain the weight of each channel through the sigmoid activation function and use it to recalibrate the channels of x attn :
[0096]
[0097] Apply a feed-forward neural network to the feature map x recalibrated by the channel attention mechanism to obtain the final output. Finally, perform a residual connection on the spatial context information captured by the self-attention mechanism and the global feature information captured by the channel attention. The whole process can be summarized as: attn Apply a feed-forward neural network to the feature map x recalibrated by the channel attention mechanism to obtain the final output. Finally, perform a residual connection on the spatial context information captured by the self-attention mechanism and the global feature information captured by the channel attention. The whole process can be summarized as:
[0098]
[0099] x output = Concat(x channel_atten , x self_attention )
[0100] In this way, the model can better capture the complex relationships between different channels in the feature map, thereby improving the overall performance.
[0101] As Figure 3 shown, relying on the optimized and upgraded innovative nnFormer architecture, the present invention implements automated processing and in-depth analysis of medical images obtained by ultra-low radiation imaging technology. In the center of the interface, a medical image clearly marked with the location of the tumor lesion and its key information is prominently presented. Thanks to the high efficiency and accuracy of the nnFormer architecture, the lesion is accurately located, and key features such as size, shape, and location are presented in detail, providing intuitive and detailed diagnostic basis for doctors and greatly improving the accuracy and personalization level of medical services.
[0102] The process of the above-mentioned automated processing and in-depth analysis is as follows: First, in this embodiment, filtering techniques (such as Gaussian filtering and median filtering) are used to remove noise in medical images and improve image quality. Then, the images from different imaging devices or different patients are standardized to eliminate image deviations caused by differences in imaging conditions, making subsequent analysis more consistent. After that, techniques such as histogram equalization are applied to enhance the contrast of the images, making tumor lesions more obvious in the images. Finally, the input images are uniformly adjusted to the size required by the model to ensure that the deep learning model can process them correctly.
[0103] Specifically, the extracted features are classified by a trained deep learning model to identify whether there are tumor lesions in the image. The Region Proposal Network (RPN) is used to generate candidate regions that may contain tumors, providing a basis for subsequent localization. A threshold is set to screen the candidate regions to ensure that only high-confidence tumor recognition results are retained. Based on the candidate regions, the boundary of the tumor is accurately determined through a regression algorithm, and the coordinates of the optimal bounding box are output. Combining semantic segmentation technology, the shape and position of the tumor are further refined to provide a clearer lesion contour map for doctors. Finally, the recognition and localization results are graphically displayed on the user interface for doctors to facilitate subsequent diagnosis and analysis.
[0104] As Figure 4 shown, the front end of this system is a highly integrated and intelligent human-computer interaction interface, aiming to optimize various business processes and improve work efficiency. It integrates advanced hardware and software technologies and achieves high flexibility and scalability through modular design. The system uses cloud computing and big data analysis technologies to be able to process massive data in real time, providing accurate business insights and predictive analysis. The user interface is friendly and intuitive, supporting multi-terminal access to ensure that users can easily operate and manage regardless of time and location.
[0105] Specifically, the UI adopts a graphical design and displays the localization and analysis results of tumor lesions through intuitive charts, images, and icons. For example, three-dimensional images or two-dimensional slices can be used to show the location, size, and shape of the lesions. Through color coding, arrow indication, heat maps, etc., complex medical data is converted into visual elements that are easy to understand. This helps users quickly identify the key features of the lesions. Interactive tools such as zooming, panning, and rotating are provided, enabling users to freely explore and analyze the detailed information of the lesions. In this embodiment, to avoid excessive information interference, key functions and buttons should be placed in prominent and easily accessible positions. The present invention provides clear navigation menus and tab pages, enabling users to easily find the required functions and settings. When the user is operating, the UI should provide intelligent prompts and feedback to help the user understand the current state, avoid incorrect operations, and provide necessary help information.
[0106] As Figure 5As shown in the left sidebar, it is specifically responsible for comprehensively managing and securely storing all medical imaging data and its associated information, ensuring that the data is properly safeguarded. To achieve this goal, the module integrates an efficient data backup mechanism and a refined management strategy, thus effectively guaranteeing the integrity and recoverability of the data. It is specifically responsible for comprehensively managing and securely storing all medical imaging data and its associated information, ensuring that these valuable data are properly safeguarded. To achieve this goal, the module integrates an efficient data backup mechanism and a refined management strategy, thus effectively guaranteeing the integrity and recoverability of the data.
[0107] This system uses MySQL as the database management system to provide stable and reliable data storage services. At the same time, the backend part is developed using the Python Flask framework to ensure the flexibility and scalability of the system. Such a technology selection enables the core module of data storage and management to give full play to its advantages and provide users with an efficient, secure, and convenient medical imaging data management solution.
[0108] The management method of the above database management system adopts a hierarchical storage architecture, classifying and storing data according to access frequency and importance, including hot data (frequently accessed) and cold data (seldom accessed), to improve storage efficiency. All sensitive data is encrypted before storage to ensure the security of data in both static and dynamic states. Audit logs of all data access and operations are recorded for easy tracking and review.
[0109] Furthermore, this embodiment also sets an automated regular backup strategy to ensure that data is backed up regularly to a secure storage location. A secure remote access channel is provided for users, and through VPN or SSL encryption technology, the security of data during transmission is ensured.
[0110] As Figure 5 and Figure 6 shown, by cleverly integrating a feedback mechanism into the user interface design, doctors and nurses can instantly provide valuable opinions on the system's diagnostic results, suggestions, and overall usage experience. In addition, with the help of reinforcement learning algorithms, the system can continuously adjust its decision-making logic based on the feedback received in actual applications. For example, when processing specific types of medical images, the system can self-optimize the diagnostic process, thereby further improving the accuracy of diagnosis. The present invention sets a clear reward mechanism for the model, giving corresponding rewards according to its performance in actual scenarios (such as the improvement in the accuracy of tumor localization, the significant decrease in the misjudgment rate, etc.), so as to motivate the model to continuously pursue excellence and achieve continuous improvement.
[0111] Specifically, a user feedback function is embedded through the interface design, allowing doctors and nurses to provide real-time feedback on the system's diagnosis results, suggestions, and usage experience. The feedback information will be organized and stored for subsequent analysis. The latest medical imaging research results are regularly updated and integrated to ensure that the system is always based on the latest scientific evidence. Furthermore, through the reinforcement learning algorithm, the system can adjust its strategies according to the feedback during actual use. For example, for specific types of images, the system can self-optimize the diagnostic logic to improve the accuracy rate. A reward mechanism is set for the model, and rewards are given according to its performance in real scenarios (such as accurate tumor localization, reduced false positive rate, etc.) to promote the continuous improvement of the model.
[0112] In this embodiment, in view of the high sensitivity and privacy protection requirements of medical imaging data, a privacy protection and data security enhancement module is specifically added to the system to ensure the strict protection of data. This module uses cutting-edge encryption technology to encrypt all medical imaging data in all links of storage and transmission, thereby effectively blocking unauthorized access behaviors and preventing the risk of data leakage.
[0113] Specifically, in this embodiment, advanced encryption technology is used to encrypt all medical imaging data during storage and transmission to prevent unauthorized access and data leakage. The system requires users to perform multi-factor authentication when logging in, including passwords, fingerprint recognition, or one-time verification codes, to enhance account security. The present invention records all user access behaviors and operation records, including login time, data types accessed, modification records, etc., to ensure that the data access history can be traced. Real-time monitoring is implemented to detect abnormal activities, such as multiple failed login attempts or abnormal access to a large amount of data, and the system administrator is notified in a timely manner.
[0114] In summary, this embodiment provides an innovative lymphoma lesion ultra-precise positioning system empowered by nnFormer, with the following characteristics: Characteristic 1: Improvement of the FFN structure. Originally, a standard two-layer multi-layer perceptron (MLP) was used to perform feed-forward network calculations. To further enhance the model's non-linear expression ability, a factored feed-forward network (FFN) was introduced. The weight matrix of the traditional fully connected layer was decomposed into multiple lower-rank matrices, and this decomposition process was achieved through two smaller fully connected layers; Characteristic 2: After processing the feature map x with the local window attention mechanism, a channel attention mechanism was integrated. To further improve the model's understanding ability of the complex relationships between different channels within each window, a channel attention mechanism module was embedded before the subsequent feed-forward neural network (FFN); Characteristic 3: Tumor lesion precise positioning and analysis module. This module relies on the innovative nnFormer architecture optimized and upgraded through Characteristic 1 and Characteristic 2, and performs automated processing and in-depth analysis on medical images obtained by ultra-low radiation imaging technology. It can efficiently and accurately identify and lock the location of tumor lesions, and at the same time, provide detailed key information about the lesions; Characteristic 4: User interface design, equipped with an intuitive graphical operation platform, compatible with diverse interaction modes. In addition, this interface also supports a wide range of image formats and input and output functions; Characteristic 5: Data storage and management core module. As the center of the system, it is specifically responsible for comprehensively managing and securely storing all medical image data and their associated information. This module integrates an efficient data backup mechanism and a refined management strategy to ensure data integrity and recoverability. It is equipped with an advanced retrieval and query engine that can quickly respond to complex data requests. Furthermore, this module supports remote access and seamless sharing functions, enabling medical image data to cross geographical limitations and achieve efficient circulation and collaboration; Characteristic 6: System continuous learning and optimization module. As the intelligent upgrade engine of the entire system, this module is responsible for continuously collecting user feedback, clinical verification results, and the latest medical image research data. Through automatic or semi-automatic means, it iteratively updates the nnFormer model, tumor lesion positioning algorithm, and AI-assisted diagnosis logic. In addition, this module also uses advanced technologies such as reinforcement learning to enable the system to continuously optimize its performance, improve accuracy, and generalization ability in actual use.
[0115] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An nnFormer-empowered lymphoma lesion localization system, characterized in that: Including: A user layer that interacts with external users, used to collect image data from external users or feedback positioning results to external users; A data management layer that communicates with the user layer, used to receive the collected image data and perform storage management on the image data; An AI intelligent processing module that is communicatively connected to the data management layer to call the image data stored in the data management layer and process the image data to output a lesion localization image; A front-end medical view module that communicates with the user layer. The front-end medical view module includes view basic tools, window level control module, multi-section control module and color control module, and is used to be displayed on the user layer display screen for users to select corresponding modules and options; Among them, the specific steps for the AI intelligent processing module to process image data are as follows: first, perform data cleaning on the image data, convert the image data to data combined with the XY coordinate axes, and then input it into the tumor molecular imaging segmentation model. Through the tumor molecular imaging segmentation model, automated processing and in-depth analysis are performed to identify and locate the tumor lesion position, and then an image with the tumor lesion position displayed is output.
2. The nnFormer-empowered lymphoma lesion localization system according to claim 1, characterized in that: The tumor molecular imaging segmentation model adopts the nnFormer model, and factorization is introduced into the feed-forward neural network structure of the nnFormer model to perform feed-forward network calculations. At the same time, a channel attention mechanism module is embedded before the feed-forward neural network to integrate the channel attention mechanism after the local window attention mechanism processes the feature map x.
3. The nnFormer-empowered lymphoma lesion localization system according to claim 2, characterized in that: The factorization method in the feed-forward neural network structure is to decompose the weight matrix of the traditional fully connected layer into multiple low-rank matrices and use two smaller fully connected layers to implement this decomposition process. Specifically as follows: in a standard two-layer MLP, the forward propagation of the feed-forward network can be expressed as: h = α(W1x + b1) y = W2h + b2 In the formula, x is the input vector; W1 and W2 are weight matrices; b1 and b2 are bias vectors; α is the activation function; h is the output of the hidden layer; y is the network output; In the factored FFN, the weight matrices W1 and W2 are decomposed into multiple lower-rank matrices, that is, each weight matrix is decomposed into the product of two smaller matrices.
4. The nnFormer-empowered lymphoma lesion localization system according to claim 3, characterized in that: The specific method of decomposing the weight matrices W1 and W2 into multiple lower-rank matrices is: decompose W1 into U1V1 and decompose W2 into U2V2. In this way, the forward propagation process can be re-expressed as: Z1 = U1X h = σ(V1Z1 + b1) Z2 = U2h y = V2Z2 + b2 In the formula, U1 and V1 are two smaller matrices used to decompose W1, U2 and V2 are two smaller matrices used to decompose W2, and Z1 and Z2 are intermediate variables.
5. The nnFormer-empowered lymphoma lesion localization system according to any one of claims 2 to 4, characterized in that: The specific way in which the local window attention mechanism processes the feature map X is as follows: The local volumetric self-attention mechanism performs self-attention calculation within a three-dimensional local volume and combines with the multi-head self-attention mechanism to form a skip attention structure. Assuming the input feature map is x ∈ R H×W×C , where H and W are the height and width of the feature map respectively, and C is the number of channels. The local window attention mechanism can be expressed as: Among them, Q, K, and V are the query, key, and value matrices respectively, and B is the relative position encoding, which are usually obtained from the input feature map x through linear transformation.
6. The nnFormer-empowered lymphoma lesion localization system according to claim 5, wherein: The skip attention structure is based on a three-dimensional local volume self-attention mechanism. During the process of upsampling the low-resolution feature map to a high-resolution feature map, the feature map is combined with the encoder features through the skip attention structure. Specifically as follows: Suppose the output of the l-th layer Transformer block of the encoder is After linear projection, the key matrix is obtained and the value matrix Among them, LP represents the linear projection operation. Assume that the feature map output after upsampling on the l-th layer of the decoder is * which is regarded as the query matrix The skip connection attention structure can be summarized as: Among them, l * represents the upsampling layer number of the decoder.
7. The nnFormer-empowered lymphoma lesion localization system according to claim 6, wherein: The method of integrating the channel attention mechanism is as follows: for the upsampled feature maps at all levels, calculate their local multi-head self-attention and channel attention respectively, and fuse the spatial context information and global feature information through residual connections, specifically as follows: Apply the local window attention mechanism to the input feature map x to obtain the output feature map x attn . For x attn Apply global average pooling to obtain a channel descriptor z ∈ R C : Learn the correlation between channels through the fully connected layer and the ReLU activation function to obtain the intermediate feature s ∈ R C : s = σ(W2δ(W1z)) wherein, and are the weight matrices of the fully connected layers, r is the scaling ratio, δ is the ReLU activation function, and σ is the sigmoid activation function; The weights for each channel are obtained through the sigmoid activation function and are used to recalibrate the channels of x attn : For the feature map x recalibrated by the channel attention mechanism attn Apply a feed-forward neural network to obtain the final output; the specific way of performing a residual connection between the spatial context information captured by the self-attention mechanism and the global feature information captured by the channel attention mechanism is as follows: x output = Concat(x channel_atten , x self_attention ).
8. The nnFormer-empowered lymphoma lesion localization system according to claim 7, characterized in that: The way the AI intelligent processing module performs data cleaning is to use filtering technology to remove noise in medical images and improve image quality. Then, standardize the images from different imaging devices or different patients to eliminate image deviations caused by differences in imaging conditions, making subsequent analysis more consistent. After that, apply techniques such as histogram equalization to enhance the contrast of the images, making tumor lesions more obvious in the images. Finally, uniformly adjust the input images to the size required by the model to ensure that the deep learning model can process them correctly.
9. The nnFormer-empowered lymphoma lesion localization system according to claim 8, wherein: The specific way to implement automatic processing and in-depth analysis through the tumor molecular imaging segmentation model to identify and locate the position of tumor lesions is as follows: first, classify the extracted features to identify whether there are tumor lesions in the image, and use the region proposal network to generate candidate regions that may contain tumors, providing a basis for subsequent localization. Set a threshold to screen the candidate regions to ensure that only the tumor recognition results with high confidence are retained.
10. The nnFormer-empowered lymphoma lesion localization system according to claim 9, characterized in that: The data management layer adopts a hierarchical storage architecture, classifies and stores data according to access frequency and importance, including hot data and cold data, encrypts all sensitive data before storage, records the audit logs of all data access and operations, and also has automatic regular backups, providing a secure remote access channel for the user layer. Through VPN or SSL encryption technology, ensure the security of data during transmission.