Systems and methods for processing brain images
By working together with the signal acquisition unit, encoder, language big data model and image decoder, the problem of inaccurate brain images in the prior art is solved, and high-quality brain image generation is achieved, which is suitable for brain state detection and judgment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIONGAN ANYING TECHNOLOGY CO LTD
- Filing Date
- 2025-10-10
- Publication Date
- 2026-06-02
Smart Images

Figure CN122134941A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this application relate to the field of computer model-based image processing, and particularly to a system and method for processing brain images. Background Technology
[0002] The statements herein are provided merely as background information in connection with this application and do not necessarily constitute prior art.
[0003] Brain detection usually requires the use of brain images. However, due to the complexity of the brain structure, current technologies cannot directly obtain brain images, or the brain images obtained by existing technologies deviate significantly from the actual brain images. Summary of the Invention
[0004] A brief overview of this application is provided below to offer a basic understanding of certain aspects thereof. It should be understood that this overview is not an exhaustive summary of the application. It is not intended to identify key or essential parts of the application, nor is it intended to limit its scope. Its purpose is merely to present certain concepts in a simplified form as a prelude to the more detailed description that follows.
[0005] This application provides a system for processing brain images, comprising: a signal acquisition unit configured to acquire electromagnetic signals from the brain; an encoder configured to receive electromagnetic signals and encode them into words; a language big model configured to receive words and perform deep semantic processing on them to obtain processed words; and an image decoder configured to receive the processed words to obtain a brain image.
[0006] This application also provides a method for processing brain images, which uses the aforementioned system and includes the following steps: S1: acquiring electromagnetic signals from the brain; S2: compressing the electromagnetic signals and extracting effective information to obtain word units; S3: inputting the word units into a language model to obtain processed word units; S4: inputting the processed word units into an image decoder to obtain a brain image.
[0007] The system for processing brain images provided in the embodiments of this application utilizes an encoder and a large language model to encode the electromagnetic signals acquired by the signal acquisition unit into words and perform deep semantic processing on the words so that the image decoder can process the processed words and obtain an image. The encoder, the large language model, and the image decoder work together to convert the electromagnetic signals into accurate brain images.
[0008] The method for processing brain images provided in the embodiments of this application compresses electromagnetic signals and extracts effective information therein to obtain lexical units, which can ensure the authenticity of the signals and reduce the amount of data; by inputting the lexical units into a large language model for processing, the brain images obtained after decoding the processed lexical units are more accurate. Attached Figure Description
[0009] To further illustrate the above and other advantages and features of this application, the specific embodiments of this application will be described in more detail below with reference to the accompanying drawings. The drawings, together with the following detailed description, are included in and form a part of this specification. Elements having the same function and structure are indicated by the same reference numerals. It should be understood that these drawings only depict typical examples of this application and should not be considered as limiting the scope of this application.
[0010] Figure 1 This is a schematic diagram of the structure of a system for processing brain images provided in an embodiment of this application; Figure 2 This is a schematic flowchart of a method for processing brain images provided in an embodiment of this application; Figure 3 It is a schematic diagram of an actual brain image; Figure 4 This is a schematic diagram of a brain image obtained using an existing method; Figure 5 This is a schematic diagram of a brain image obtained using another existing method; Figure 6 This is a schematic diagram of a brain image obtained by a system or method for processing brain images provided in an embodiment of this application.
[0011] Explanation of reference numerals in the attached figures: 1. Signal acquisition unit; 2. Encoder; 3. Language big model; 4. Image decoder; 100. System for processing brain images. Detailed Implementation
[0012] Exemplary embodiments of this application will be described below with reference to the accompanying drawings. For clarity and brevity, not all features of actual implementations are described in the specification. However, it should be understood that many implementation-specific decisions must be made in the development of any such actual embodiment to achieve the developer's specific goals, such as complying with constraints related to the system and business, and these constraints may vary depending on the implementation. Furthermore, it should be understood that while development work can be very complex and time-consuming, such development work is merely a routine task for those skilled in the art who benefit from the content of this application.
[0013] It should also be noted that, in order to avoid obscuring this application with unnecessary details, only the equipment structure and / or processing steps closely related to the solution according to this application are shown in the accompanying drawings, while other details that are not closely related to this application are omitted.
[0014] The following disclosure provides several different implementations or examples for carrying out this application. To simplify the disclosure of this application, specific examples of components and methods are described below. Of course, these are merely examples and are not intended to limit this application. In the description of the embodiments of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0015] Currently, the brain images obtained by existing methods differ significantly from actual brain images, posing an obstacle to the detection and assessment of brain states. Therefore, a system or method is needed to obtain accurate brain images to assist in the detection and assessment of brain states.
[0016] One aspect of the embodiments of this application provides a system for processing brain images. Figure 1 A schematic diagram of the structure of a system for processing brain images provided in an embodiment of this application is shown, as follows: Figure 1 As shown, it includes: a signal acquisition unit 1, configured to acquire electromagnetic signals from the brain; an encoder 2, configured to receive electromagnetic signals and encode them into words; a language big model 3, configured to receive words and perform deep semantic processing on them to obtain processed words; and an image decoder 4, configured to receive processed words to obtain brain images.
[0017] The system 100 for processing brain images provided in the embodiments of this application uses an encoder 2 and a language big model 3 to encode the electromagnetic signals acquired by the signal acquisition unit 1 into words and perform deep semantic processing on the words so that the image decoder 4 can process the processed words and obtain an image. The encoder 2, the language big model 3, and the image decoder 4 work together to convert the electromagnetic signals into accurate brain images.
[0018] In some embodiments, by acquiring brain electromagnetic signals and encoding them into words, and then performing deep semantic processing on the words, the acquired brain images can be made closer to actual brain images. Figure 3 A schematic diagram showing an actual brain image is provided. Figure 4 This diagram illustrates a brain image obtained using an existing method. Figure 5 This illustrates a brain image obtained using another existing method, such as... Figures 3-5As shown, the brain images obtained by existing methods differ significantly from the actual brain images, with problems such as unclear brain tissue texture and low contrast of brain tissue, making it impossible to identify the true state of the areas in the brain images that need to be judged. Figure 6 The illustration shows a schematic diagram of a brain image obtained by a system or method for processing brain images provided in an embodiment of this application, such as... Figure 6 As shown, the brain image obtained using the image processing method provided in the embodiments of this application is compared with... Figure 1 The images shown have a high degree of similarity to the actual brain images. Figure 4 and Figure 5 Compared to existing methods, the brain images shown have clearer brain tissue texture and higher contrast at the edges, which can better assist in the detection and judgment of brain state.
[0019] In some embodiments, the encoder 2 includes a variational neural network and a U-shaped neural network, which are connected in series. The electromagnetic signal is first input to the variational neural network, processed by it, and then the processing result is input to the U-shaped neural network. After processing by the U-shaped neural network, word units are obtained. The variational neural network can extract semantic information from the electromagnetic signal, while the U-shaped neural network can further extract higher-level semantic information. Using both variational and U-shaped neural networks to extract semantic information from the electromagnetic signal at multiple scales optimizes the structural consistency of the electromagnetic signal.
[0020] By processing electromagnetic signals using a cascaded variational neural network and a U-shaped neural network, the obtained lexical units can include multimodal data. Through hierarchical interaction, cross-modal feature alignment can be achieved, thereby overcoming the limitations of a single modality and assisting in the accurate delineation of complex abnormal regions in images (e.g., glioma infiltration boundaries).
[0021] The processing method of converting electromagnetic signals into tokens can dynamically crop redundant areas in images, thereby focusing the processing on key areas containing effective information in the image. Compared with traditional neural network model processing methods, it can reduce memory usage and improve processing efficiency, making it suitable for edge device deployment and facilitating real-time image processing.
[0022] A multi-layer random gating structure is set between encoder 2 and image decoder 4. In some embodiments, the multi-layer random gating structure includes multiple learnable gating units, each of which includes two parallel paths. The first path is set as a feature preservation path, used to directly transmit key feature information between encoder 2 and image decoder 4; the second path is set as a feature transformation path, which realizes feature transformation through one-dimensional convolution and nonlinear activation.
[0023] In some embodiments, the switching state of the gating unit is randomly selected through differentiable Gumbel-Softmax sampling, thereby enabling dynamic adjustment of the information flow path based on the input samples, which enhances the transmission of effective features while suppressing noise.
[0024] Furthermore, this invention improves upon the fixed prior distribution in traditional variational networks by replacing it with a conditional prior generation mechanism. Through a lightweight auxiliary network, the parameters (mean and variance) of the prior distribution are predicted in real time based on the input data. This enables dynamic adjustment of the latent spatial distribution. This mechanism significantly alleviates the "posterior collapse" problem common in variational neural networks, improving not only the clarity and diversity of generated images in image generation tasks but also drastically reducing the convergence time.
[0025] In some embodiments, a bidirectional feature interaction module is set in the bottleneck layer of the U-shaped neural network. This module establishes a bidirectional association from shallow features of the decoder to deep features of the encoder through a cross-attention mechanism, realizing long-range dependency modeling between multi-scale features.
[0026] Specifically, this module uses the initial features of the decoder as the query and the deep features of encoder 2 as the key and value to calculate attention weights and reconstruct features, thereby utilizing both high-level semantic information and low-level detail information during information transmission. Using the above setup, the peak signal-to-noise ratio (PSNR) can be improved by approximately 2.1 dB in image reconstruction tasks, which is significantly better than the traditional unidirectional skip connection structure.
[0027] In some embodiments, encoder 2 is configured to train a concatenated variational neural network and a U-shaped neural network using a 1-structured similarity index. Compared to traditional neural network models trained using mean squared error, training the concatenated variational neural network and U-shaped neural network using a 1-structured similarity index can avoid structural differences caused by the different absolute sizes of different images.
[0028] In some embodiments, the large language model 3 includes a word segmenter, multiple intermediate layers, and a language output head; the word segmenter, multiple intermediate layers, and language output head are connected in series; wherein word units are input to the word segmenter, and processed word units are output from the language output head. By setting up the above-mentioned large language model 3, feature extraction and semantic understanding of word units can be achieved, which is beneficial to subsequent image recognition; and by using a series approach, information loss in feature transmission in traditional pipeline models can be avoided, thereby making the obtained brain images more conducive to the judgment of the results of subsequent steps.
[0029] In some embodiments, when the processed image is a brain image, the word segmenter, multi-layer intermediate layer, and language output head are configured to be connected in series, which can improve the recognition accuracy of small abnormal areas (e.g., early microbleeds) in brain images.
[0030] In some embodiments, the word segmenter is configured to process words into a structured sequence, enabling multi-scale feature encoding of images, including local texture and global anatomical structure.
[0031] In some embodiments, the language output head is configured to map the features of lexical units to interpretable output, thereby enabling the automated generation of structured reports to facilitate subsequent image decoding by the image decoder 4.
[0032] In some embodiments, interpretable output may include data such as segmentation masks and pathological classification probabilities to enhance the understandability of the output results.
[0033] In some embodiments, the multi-layer intermediate layer includes an attention mechanism network and a gated delta neural network, which are connected in parallel. The output of the word segmenter is input to both the attention mechanism network and the gated delta neural network, and the outputs of both are input to the language output head. Through the attention mechanism and convolution operations, spatial and semantic relationships between word units can be established, thereby significantly improving the context awareness capability of a predetermined region of the image.
[0034] In some embodiments, in order to overcome the information transmission barrier caused by the mismatch between the input and output dimensions of existing modules, a trainable dimension projection layer is set in the first intermediate layer of the word segmenter and the multi-layer intermediate layer. The dimension projection layer is configured to include a fully connected linear transformation unit and a layer normalization operation unit, which are connected sequentially. It can automatically project the word embedding vector output by the word segmenter from the original dimension to the input dimension required by the subsequent intermediate layer, and ensure numerical stability through normalization processing.
[0035] In some embodiments, the multi-layer intermediate layer is further provided with a dynamic pooling layer based on attention weights. This layer performs global modeling of the variable-length sequence through a trainable query vector, calculates the importance weight of each position, and aggregates them accordingly to generate a fixed-length semantic representation, thereby retaining the most critical information while compressing the sequence dimension.
[0036] Another aspect of the embodiments of this application provides a method for processing brain images, which employs the aforementioned system. Figure 2A schematic flowchart illustrating a method for processing brain images provided in an embodiment of this application is shown, such as... Figure 2 As shown, it includes the following steps: S1: Acquire electromagnetic signals from the brain; S2: Compress electromagnetic signals and extract effective information to obtain word units; S3: Input word units into a language model to obtain processed word units; S4: Input processed word units into an image decoder to obtain an image.
[0037] The method for processing brain images provided in the embodiments of this application compresses electromagnetic signals and extracts effective information therein to obtain lexical units, which can ensure the authenticity of the signals and reduce the amount of data; by inputting the lexical units into a large language model for processing, the brain images obtained after decoding the processed lexical units are more accurate.
[0038] In some embodiments, step S2 further includes the following steps: S21: performing a short-time Fourier transform on the electromagnetic signal to decompose the electromagnetic signal into time-frequency domain features; S22: compressing the time-frequency domain features, extracting the compressed features and concatenating them to obtain lexical units. The above processing method can perform multi-level information extraction and achieve structural consistency optimization between multi-scale features and electromagnetic signals.
[0039] In some embodiments, step S3 further includes the following steps: S31: using a sub-word-based segmentation algorithm to segment the word units; S32: using a differentiable intermediate layer to map the segmentation results obtained in step S31 into a high-dimensional vector; S33: mapping the high-dimensional vector into an interpretable output. By processing the word units using the above method, the obtained word units can include multimodal data, overcoming the limitations of a single modality and assisting in the accurate delineation of complex abnormal regions in images (e.g., glioma infiltration boundaries).
[0040] In some embodiments, in step S31, the lexical units can be processed into a structured sequence, which enables the encoding of multi-scale features of the image, including local texture and global anatomical structure.
[0041] In some embodiments, in step S32, positional encodings that can be differentiated are superimposed on the intermediate layer, and the temporal information of the lexical units is preserved by addition or concatenation so that the gradient of the lexical units can be backpropagated.
[0042] In some embodiments, differential position coding can accurately determine the spatial coordinates of voxels in an image, thereby enabling the reconstruction of the three-dimensional anatomical structure of the brain. The differential characteristics of position coding can automatically optimize the spatial correspondence between modalities, thereby improving the signal-to-noise ratio of the fused image and enhancing image quality.
[0043] In some embodiments, step S4 further includes the following steps: S41: converting lexical units into high-dimensional vectors; S42: converting high-dimensional vectors into tensors of image matrix dimensions; S43: determining brain images based on tensors of image matrix dimensions.
[0044] In some embodiments, in step S4, a precise, high-fidelity cross-modal mapping from the text semantic space to the medical image pixel space is achieved by transforming text lexical units into brain images.
[0045] Regarding the embodiments of this application, it should also be noted that, without conflict, the embodiments of this application and the features in the embodiments can be combined with each other to obtain new embodiments.
[0046] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. The scope of protection of this application shall be determined by the scope of the claims.
Claims
1. A system for processing brain images, characterized in that, It includes: The signal acquisition unit is configured to acquire electromagnetic signals from the brain; An encoder configured to receive the electromagnetic signal and encode the electromagnetic signal into words; A large language model is configured to receive the word units and perform deep semantic processing on them to obtain processed word units; An image decoder is configured to receive processed words to obtain the brain image.
2. The system according to claim 1, characterized in that, The encoder includes: a variational neural network and a U-shaped neural network. The variational neural network and the U-shaped neural network are configured in series. The electromagnetic signal is first input to the variational neural network, and after processing by the variational neural network, the processing result is input to the U-shaped neural network. After processing by the U-shaped neural network, the word element is obtained.
3. The system according to claim 2, characterized in that, The encoder is configured to train the concatenated variational neural network and the U-shaped neural network using a 1-structured similarity index.
4. The system according to claim 1, characterized in that, The large language model includes a word segmenter, multiple intermediate layers, and a language output head; The word segmenter, the multi-layer intermediate layer, and the language output head are connected in series; The word units are input to the word segmenter, and the processed word units are output from the language output head.
5. The system according to claim 4, characterized in that, The word segmenter is configured to process the tokens into a structured sequence.
6. The system according to claim 4, characterized in that, The language output head is configured to map the features of the lexical units into interpretable output.
7. The system according to claim 4, characterized in that, The multi-layer intermediate layer includes an attention mechanism network and a gated delta neural network, wherein the attention mechanism network and the gated delta neural network are connected in parallel; The output of the word segmenter is input to the attention mechanism network and the gated delta neural network, respectively. The outputs of the attention mechanism network and the gated delta neural network are respectively input to the language output head.
8. The system according to claim 1, characterized in that, The image decoder comprises a linear connection layer, an element rearrangement layer, and a lightweight U-shaped neural network, which are connected in series. The word is input to the linear connection layer, and after processing by the linear connection layer, a high-dimensional vector is output. The high-dimensional vector is input to the element rearrangement layer, and after processing by the element rearrangement layer, a tensor of the image matrix dimension is output. The tensor of the image matrix dimension is input into the U-shaped neural network to obtain the processed brain image.
9. A method for processing brain images, characterized in that, The method employs the system described in any one of claims 1-8, and includes the following steps: S1: Acquire electromagnetic signals from the brain; S2: Compress the electromagnetic signal and extract effective information to obtain word elements; S3: Input the lexical units into the language model to obtain the processed lexical units; S4: Input the processed lexical units into the image decoder to obtain the brain image.
10. The method according to claim 9, characterized in that, Step S2 also includes the following steps: S21: Perform a short-time Fourier transform on the electromagnetic signal to decompose the electromagnetic signal into time-frequency domain features; S22: Compress the time-frequency domain features, extract the compressed features, and concatenate them to obtain the word units.
11. The method according to claim 9, characterized in that, In step S3, also The steps include: S31: Using a sub-word-based word segmentation algorithm, the word units are segmented; S32: Employ a differentiable intermediate layer to map the word segmentation results obtained in step S31 into a high-dimensional vector; S33: Map the high-dimensional vector into an interpretable output.
12. The method according to claim 11, characterized in that, In step S31, the lexical units are processed into a structured sequence.
13. The method according to claim 12, characterized in that, In step S32, differentially identifiable positional codes are superimposed on the intermediate layer, and the temporal information of the lexical units is preserved by addition or concatenation so that the gradient of the lexical units can be backpropagated.
14. The method according to any one of claims 9-13, characterized in that, Step S4 also includes the following steps: S41: Convert the lexical units into high-dimensional vectors; S42: Convert the high-dimensional vector into a tensor of image matrix dimensions; S43: Determine the brain image based on the tensor of the image matrix dimension.