Data cleaning method suitable for massive neuron three-dimensional imaging data

By combining the segmentation model and classification model of attention mechanism and gating mechanism, massive neuronal data are automatically cleaned, solving the problems of low data cleaning efficiency and poor stability in the existing technology, and efficient and accurate neuronal data cleaning is achieved.

CN120147529AActive Publication Date: 2025-06-13HUST SUZHOU INST FOR BRAINMATICS
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510218440.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-24
Filing Date
2025-02-26
Publication Date
2025-06-13
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

When processing massive neuronal data cleaning methods, the existing neuronal data cleaning methods need to cut the entire brain in advance, consuming a lot of calculation and spatial resources; classification methods are difficult to obtain stable results on different individuals; manual observation and re-examination are difficult and inefficient.

Method used

A segmentation model combining attention mechanism and gating mechanism is adopted to enhance the neuronal signal of the projection block in the maximum projection map and weaken the non-neuron signal. Through the classification model, the projection block containing neuronal signals is classified with the projection block without neuronal signals, and the label information of the projection block containing neurons is automatically cleaned and derived. Finally, the three-dimensional data blocks of the target shape are obtained through multi-threading processing of slicing and re-splicing.

Benefits of technology

Effectively suppress the impact of noise, improve the efficiency and accuracy of data cleaning, reduce manual participation time, improve the degree of automation, and ensure the stability and high recall of massive neuronal data cleaning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147529A_ABST
    Figure CN120147529A_ABST
Patent Text Reader

Abstract

The invention discloses a data cleaning method suitable for massive neuron three-dimensional imaging data. The method comprises the following steps: acquiring two-dimensional brain image sequence data; obtaining a maximum value projection drawing, removing invalid data outside the brain contour, dividing the data in the brain contour into projection blocks, and naming the projection blocks according to the positions of the projection blocks in the current maximum value projection drawing; the neuron signals are highlighted, and the projection blocks are classified; exporting label information of all projection blocks containing neurons, and generating a cleaned label information set; and segmenting and re-splicing all the projection blocks containing the neurons to obtain a three-dimensional data block with a target shape. According to the method, the stability and accuracy of a classification result can be ensured while neuron data cleaning of the whole brain of the macaque can be completed in an extremely short time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing, and particularly to a data cleaning method applicable to massive three-dimensional imaging data of neurons. Background Art

[0002] Neurons, namely nerve cells, are the most basic structural and functional units of the nervous system and play an extremely important role in the basic function research of the brain. Neurons support numerous high-level functions such as cognition, emotion, and behavior through complex network connections. The diameter of the neuron cell body is between 10μm and 30μm, and the diameter of the protrusions initiated from the cell body reaches the sub-micron level at the minimum. In order to obtain clearer and more complete neuron morphological information, sub-micron imaging of the whole brain is required; and high-resolution imaging is also accompanied by a large amount of imaging data, among which the imaging data of the whole macaque brain is as high as hundreds of TB. For the sparsely labeled whole macaque brain, the proportion of neuron signals is less than 5%, and a large amount of redundant data brings difficulties in data storage, transmission, and calculation. In addition, traditional biomedical image processing software such as Amira and Fiji cannot efficiently complete the manual recheck of neuron data cleaning tasks and requires a large amount of human resources. Therefore, cleaning the neuron data of the sparsely labeled whole macaque brain and extracting the data blocks containing neurons will be crucial for the subsequent rapid analysis of neuron signals.

[0003] Existing neuron data cleaning methods, such as the invention patent with the publication number of CN117218144A, disclose a data cleaning method based on neuron signal recognition. The data cleaning method based on neuron signal recognition can reduce manual time consumption, achieve a good cleaning effect with only a small number of intermediate files generated, greatly reduce the data volume, and effectively avoid waste of unnecessary computing and storage resources. This data cleaning method first obtains the sparse labeling data and cytoarchitecture data of the brain, then divides the sparse labeling data into several data blocks of the same size, performs maximum intensity projection on the data blocks within the brain contour, outputs the maximum intensity projection map, then uses methods such as adaptive threshold segmentation to perform binary classification on all the maximum projections, and finally checks and makes up for deficiencies in the binary classification results through manual observation.

[0004] This data cleaning method is applicable to the processing of relatively small storage data, such as mouse neuron data processing. However, when dealing with massive neuron data (such as macaque neuron data processing), it still faces the following problems: First, it is necessary to pre-cut the whole brain in advance, which consumes a large amount of computing and space resources; Second, when applied to the processing of large and complex data volumes, the above classification method is difficult to obtain stable results on different individuals; Third, the imaging data blocks of macaques are as high as millions, and it is difficult to check and make up for deficiencies in the classification results through manual observation. Summary of the Invention

[0005] Therefore, to solve the above problems, the present invention provides a data cleaning method applicable to massive three-dimensional imaging data of neurons.

[0006] The present invention is implemented through the following technical solutions:

[0007] A data cleaning method applicable to massive three-dimensional imaging data of neurons, comprising the following steps:

[0008] Obtain two-dimensional brain image sequence data;

[0009] Obtain the maximum intensity projection map of the two-dimensional brain image sequence data, remove the invalid data outside the brain contour in each maximum intensity projection map, and divide the data inside the brain contour to obtain projection blocks, and name them according to their positions in the current maximum intensity projection map;

[0010] Combining the attention mechanism and the gating mechanism, enhance the neuron signals in the projection blocks in the maximum intensity projection map and weaken the non-neuron signals to highlight the distribution positions of the neuron signals, classify the projection blocks containing neuron signals and the projection blocks without neuron signals into different projection block types, and obtain the classification result;

[0011] According to the classification result, export the label information of all projection blocks containing neurons, and the label information includes at least the name of the projection block, and generate a cleaned label information set;

[0012] Extract each projection block containing neurons from the label information set, and cut and re-stitch all the projection blocks containing neurons to obtain a three-dimensional data block with the target shape.

[0013] Preferably, after obtaining the classification result, recheck the classification result, reclassify the misclassified projection blocks, and obtain the final classification result.

[0014] Preferably, the step of "obtaining the maximum intensity projection map of the two-dimensional brain image sequence data, removing the invalid data outside the brain contour in the maximum intensity projection map, and dividing the data inside the brain contour" includes the following steps:

[0015] Through viral sparse labeling, make the brightness of neurons in the imaging data greater than the background information;

[0016] Taking the preset number of image layers as a unit, sequentially perform maximum intensity projection on the two-dimensional brain image sequence data of the whole brain to obtain a number of maximum intensity projection maps;

[0017] Sequentially perform downsampling on each maximum intensity projection map, and then upsample to the original resolution after segmenting the brain contour to obtain the original resolution segmentation map of each maximum intensity projection map;

[0018] The brain contour region of each of the original resolution segmentation maps is divided into a plurality of projection blocks.

[0019] Preferably, before dividing the projection block types, a segmentation model and a classification model are obtained; wherein, both the segmentation model and the classification model are trained using a plurality of projection block data containing neurons and a plurality of projection block data without neurons. Among them, the above segmentation model is used to enhance the neuron signals in the projection blocks within the maximum intensity projection map and weaken the non-neuron signals to highlight the distribution positions of the neuron signals, and the classification model is used to classify the projection blocks containing neuron signals and the projection blocks without extracted neuron signals.

[0020] Preferably, the segmentation model is an Attention Unet model. The Attention Unet model is trained using the projection block dataset, and taking neuron signals as the target feature, each projection block in the maximum intensity projection map to be detected is input into the trained Attention Unet model, and the attention of the neuron signals is increased through the attention mechanism in the model, thereby realizing the extraction of the neuron signals in the current maximum intensity projection map.

[0021] Preferably, the classification model is a ResNet-18 convolutional neural network. The ResNet-18 convolutional neural network is trained using the projection block dataset training set. Each projection block in the maximum intensity projection map to be detected is input into the trained ResNet-18 convolutional neural network, and each projection block is classified using neuron signals as the classification feature, and the classification result is output.

[0022] Preferably, the "cutting and re-stitching all projection blocks containing neurons to obtain a three-dimensional data block of the target shape" means using multi-threaded processing to divide each projection block containing neurons into several small data blocks, creating a thread for each small data block respectively, each thread reads its corresponding small data block respectively, and writes the small data block to the correct position in the merged file, so that several small data blocks are re-stitched to obtain a three-dimensional data block of the target shape.

[0023] Preferably, the small data blocks are obtained by dividing the thickness of the projection blocks, and their lengths and widths are equal.

[0024] Preferably, the "re-checking the classification result" is realized by a cleaning tool, and the cleaning tool includes:

[0025] An input module, the input module includes a projection block data input port, a label information input port, and a downsampled image input port; the projection block data input port is used to input the projection block data within the brain contour in the maximum intensity projection map, including projection blocks containing neurons and projection blocks without neurons, the label information input port is used to input the cleaned label information set, and the downsampled image input port is used to input the 50-μm downsampled image of the maximum intensity projection map;

[0026] A display module, which is used to display the currently operating maximum intensity projection map and the selected area in the current maximum intensity projection map, and the visualization range of the selected area can be adjusted through buttons;

[0027] An adjustment module, which is used to display an enlarged view of the selected area, and display all projection blocks within the enlarged view, and the misclassified projection blocks can be reclassified by the mouse to obtain the final classification result after reexamination;

[0028] An export module, which is used to export the final classification result as a txt file, and the txt file contains the names of all neuron projection blocks.

[0029] The beneficial effects of the technical solution of the present invention are mainly reflected in:

[0030] 1. First, the maximum intensity projection map of the two-dimensional brain image sequence data is obtained, and invalid data is removed through data preprocessing. After automatically cleaning the data, the projection blocks containing neurons are distinguished, and only the projection block data containing neurons needs to be three-dimensionally sliced and spliced, saving a large amount of data processing time. When cleaning neuron data, a combination of segmentation and classification is adopted to effectively suppress the influence of noise, facilitating the completion of neuron data cleaning of the whole macaque brain in a very short time while ensuring the stability and accuracy of the classification result.

[0031] 2. Before cleaning neuron data, a large amount of data is first used to train the segmentation model and the classification model, thereby improving the robustness of the model. Among them, the segmentation model combines the attention mechanism and the gating mechanism to enhance the neuron signals of the projection blocks within the maximum intensity projection map and weaken the non-neuron signals, so as to highlight the distribution position of the neuron signals, facilitating the classification model to quickly classify and improving work efficiency.

[0032] 3. The present method provides a data cleaning tool. The data cleaning tool can adjust the range of the visualization area through buttons, and by clicking the mouse, the projection blocks judged by the user to contain neurons can be added or the projection blocks that do not contain neurons can be removed, so that the misclassified projection blocks can be reclassified, facilitating rapid manual reexamination, ensuring a high recall rate of the data cleaning result, and finally exporting the final classification result in txt format. Description of the Drawings

[0033] Figure 1 It is a flowchart of a data cleaning method applicable to massive three-dimensional imaging data of neurons;

[0034] Figure 2 It is a working schematic diagram (including the re-inspection step) of a data cleaning method applicable to massive three-dimensional imaging data of neurons;

[0035] Figure 3 It is a schematic diagram of the data preprocessing step (including maximum projection, removing data outside the brain contour, and dividing data inside the brain contour) in a data cleaning method applicable to massive three-dimensional imaging data of neurons;

[0036] Figure 4 It is a schematic diagram of the segmentation and classification step (automatic cleaning) in a data cleaning method applicable to massive three-dimensional imaging data of neurons;

[0037] Figure 5 It is a schematic diagram of the cleaning tool interface;

[0038] Figure 6 It is a schematic diagram of the cutting and re - stitching step in a data cleaning method applicable to massive three-dimensional imaging data of neurons. Detailed implementation manners

[0039] To clearly and detailedly show the purpose, advantages and features of the present invention, it will be illustrated and explained through the non - restrictive description of the following preferred embodiments. This embodiment is only a typical example of applying the technical solution of the present invention. Any technical solution formed by equivalent replacement or equivalent transformation falls within the scope of protection required by the present invention. At the same time, it is declared that in the description of the solution, it should be noted that in the present invention, the terms "a plurality of", "several" mean two or more, unless otherwise clearly and specifically defined, so it cannot be construed as a limitation to the present invention.

[0040] The present invention discloses a data cleaning method applicable to massive three - dimensional imaging data of neurons. As Figure 1 , Figure 2 shown, starting from the original two - dimensional sequence data and ending with obtaining all three - dimensional data blocks containing neurons, it specifically includes the following steps:

[0041] Obtain two - dimensional brain image sequence data;

[0042] Obtain the maximum projection map of the two - dimensional brain image sequence data, remove the invalid data outside the brain contour in each maximum projection map, and divide the data inside the brain contour to obtain projection blocks, and name them according to their positions in the current maximum projection map.

[0043] In some embodiments, this step specifically includes:

[0044] Through viral sparse labeling, the brightness of neurons in the imaging data is greater than the background information. Among them, assuming that the size of the two-dimensional image sequence data of the monkey brain obtained is 100000 * 80000 * 20000, that is, 20000 two-dimensional image data with 100000 * 80000 pixels, and the imaging resolution is 0.65 * 0.65 * 3 μm. Through viral sparse labeling, the brightness of neurons in the imaging data is usually greater than the background information. Therefore, the morphological information of neurons can be significantly observed through maximum projection.

[0045] Taking the preset number of image layers as a unit, perform maximum projection on the two-dimensional brain image sequence data of the whole brain in sequence to obtain a number of maximum projection maps. In one embodiment, taking 360 layers as a unit, perform maximum projection on the whole brain of the macaque, and finally obtain 56 maximum projection maps.

[0046] Perform downsampling on each maximum projection map in sequence, and after segmenting the brain contour, upsample it to the original resolution to obtain the original resolution segmentation map of each maximum projection map. Specifically, in the maximum projection map, there will be some tissue debris with autofluorescence characteristics remaining outside the brain contour. These debris will affect the accuracy of neuron recognition. Therefore, it is necessary to remove the invalid data outside the brain contour before dividing the maximum projection map. Both the brain contour segmentation and the data clearing outside the brain contour adopt existing segmentation and clearing methods, which will not be elaborated here. Among them, first downsample the projection map with the original resolution to 50 μm, and then upsample the 50 μm image after segmentation to the original resolution. Downsampling reduces the complexity of the image by reducing the resolution or size of the image. Its purpose is to reduce the number of pixels in the image, thereby reducing the impact of noise on the result, reducing the computational amount, and increasing the receptive field size of the convolutional kernel.

[0047] Divide the brain contour area of each of the original resolution segmentation maps into multiple projection blocks. In a preferred embodiment, the projection blocks are square small blocks with uniform thickness and side length. Combining the above embodiments, the thickness and side length of the projection blocks are both 360 pixels.

[0048] Specifically, the "naming the projection block according to its position in the current maximum projection map" means that after dividing the projection blocks, name each projection block respectively. Locate the position of the current projection block in the current maximum projection map through the name of the projection block, and at the same time generate the label information of each projection block, and store the name of the projection block into the corresponding label information.

[0049] In the above embodiment, the thickness and side length of the projection block are both 360 pixels, and its name can display its coordinate information. For example, for the data block named "59_60_27", its upper left corner coordinates in the original two-dimensional sequence data are [59 * 360, 60 * 360, 27 * 360], and the coordinates of the lower right corner are

[0050] [60 * 360, 61 * 360, 28 * 360]. The position of the data block can be accurately located according to the upper left corner and lower right corner coordinates of the data block.

[0051] After obtaining the label information of each projection block, combining the attention mechanism and the gating mechanism, enhance the neuron signals of the projection blocks in the maximum intensity projection map and weaken the non-neuron signals to highlight the distribution position of the neuron signals, classify the projection blocks containing neuron signals and those without neuron signals into different projection block types, and obtain the classification result;

[0052] In some embodiments, before classifying the projection block types, obtain the segmentation model and the classification model; wherein, both the segmentation model and the classification model are trained with multiple projection block data containing neurons and multiple projection block data without neurons. In a preferred embodiment, before training the model, first construct a projection block data set composed of 8000 projection blocks containing neurons and 100000 projection blocks without neurons, and use this projection block data set to train the segmentation model and the classification model, thereby effectively improving the robustness of the model.

[0053] Among them, use the above segmentation model to enhance the neuron signals of the projection blocks in the maximum intensity projection map and weaken the non-neuron signals to highlight the distribution position of the neuron signals. In a preferred embodiment, the segmentation model is the Attention Unet model. The Attention Unet model introduces the attention mechanism on the basis of the traditional Unet architecture, which can improve the performance of the image segmentation task. By dynamically focusing on and selecting the image regions of interest, it can improve the accuracy and fineness of the segmentation and effectively suppress the irrelevant regions; train the Attention Unet model with the projection block data set, and use the neuron signals as the target features. Input each projection block in the maximum intensity projection map to be detected into the trained Attention Unet model. Through the attention mechanism in the model, generate and adjust the attention coefficients with different convolution operations, so as to focus on the neuron signals and realize the extraction of the neuron signals in the current maximum intensity projection map.

[0054] In addition, a classification model is used to classify the projection blocks containing neuron signals and the projection blocks from which neuron signals have not been extracted; in a preferred embodiment, the classification model is a ResNet-18 convolutional neural network. The ResNet-18 convolutional neural network is trained using the projection block dataset training set. Each projection block in the maximum intensity projection map to be detected is input into the trained ResNet-18 convolutional neural network, and each projection block is classified using the neuron signal as the classification feature, and the classification result is output.

[0055] In some embodiments, after obtaining the classification result, the classification result is rechecked, the misclassified projection blocks are reclassified, and the final classification result is obtained; and after obtaining the classification result, the label information of all projection blocks containing neurons is derived according to the classification result, and a cleaned label information set is generated.

[0056] In a preferred embodiment of the present invention, the "rechecking the classification result" is implemented by a cleaning tool, and the cleaning tool includes:

[0057] An input module, the input module includes a projection block data input port, a label information input port, and a downsampled image input port; the projection block data input port is used to input the projection block data within the brain contour in the maximum intensity projection map, including projection blocks containing neurons and projection blocks without neurons, the label information input port is used to input the cleaned label information set, and the downsampled image input port is used to input a 50-μm downsampled image of the maximum intensity projection map;

[0058] A display module, which is used to display the maximum intensity projection map being operated and the selected area in the current maximum intensity projection map, and the visualization range of the selected area can be adjusted by pressing a key;

[0059] An adjustment module, which is used to display an enlarged view of the selected area, and display all projection blocks within the enlarged view, and the misclassified projection blocks can be reclassified by using a mouse, so as to obtain the final classification result after rechecking;

[0060] An export module, which is used to export the final classification result as a txt file, and the txt file contains the names of all neuron projection blocks; using this cleaning tool can achieve fast semi-artificial rechecking, thereby improving the efficiency of artificial rechecking.

[0061] After exporting the final classification result, each projection block containing neurons is extracted from the label information set, and all projection blocks containing neurons are cut and re-stitched to obtain a three-dimensional data block with the target shape.

[0062] Specifically, the label information of each projection block containing neurons is extracted from the label information set, and these projection blocks are located by name. Subsequently, these projection blocks containing neurons are sliced and re - stitched. In a preferred embodiment, the statement "all projection blocks containing neurons are sliced and re - stitched to obtain a three - dimensional data block of the target shape" means that multi - threading processing is adopted. Each projection block containing neurons is evenly divided into several small data blocks. A thread is created for each small data block. Each thread reads its corresponding small data block and writes the small data block to the correct position in the merged file, so that several small data blocks are re - stitched to obtain a three - dimensional data block of the target shape. Among them, the small data block is obtained by dividing the thickness of the projection block, and their lengths and widths are equal; Combining the above - mentioned embodiments, when the thickness and side length of each projection block are both 360 pixels, each projection block is first cut into small data blocks with a thickness of about 10 pixels and other side lengths equal to the size of the projection block (that is, the length and width both remain 360 pixels); Subsequently, multiple small data blocks are stitched into a large data block of the target shape through multi - threading. The target shape of the large data block can be set according to requirements and will not be elaborated here. This process only cuts the data blocks containing neurons in the original data, greatly saving computing and storage resources compared with the existing technical process.

[0063] To verify the effect of the data cleaning method applicable to three - dimensional imaging data of a large number of neurons disclosed in the present invention, whole - brain images of 5 macaques are provided, and the following experiments are respectively carried out:

[0064] Experiment 1: From the whole brains of 5 macaques (Brain1 - Brain5), some data are selected to prepare projection blocks, and the data cleaning method applicable to three - dimensional imaging data of a large number of neurons is adopted. The projection blocks are classified through a segmentation model and a classification model for data cleaning tests. The test results are shown in Table 1.

[0065] Table 1: Quantitative evaluation of the cleaning results of samples Brain1 - Brain5;

[0066]

[0067]

[0068] The test results in Table 1 show that the automatic cleaning model of the present invention has a high recall rate and precision rate on different individuals, indicating that the segmentation model and classification model used in the present invention have high robustness.

[0069] Experiment 2: The whole - brain two - dimensional sequence data of 5 macaques are classified by the data cleaning method applicable to three - dimensional imaging data of a large number of neurons, and a cleaning tool is used for re - inspection. The cleaning results are shown in Table 2:

[0070] Table 2: Results of whole-brain data cleaning after manual recheck;

[0071]

[0072] As shown in Table 2 above, the data within the brain contour on average occupies about 30% of the whole brain. Therefore, removing the brain contour can greatly improve the efficiency of data cleaning; while the proportion of neuron data blocks in the whole brain is usually less than 1%, which indicates that neuron data cleaning can greatly reduce the content unrelated to neurons in the original data and provide a more efficient data basis for subsequent rapid neuron analysis.

[0073] Experiment 3: The data cleaning method for massive three-dimensional neuron imaging data disclosed in the present invention and the existing data cleaning method were respectively used to clean the whole-brain neuron data of a rhesus monkey. In this experiment, as Figure 2 shown, in the rapid data cleaning method of the present invention, data preprocessing includes: maximum projection, removing data outside the brain contour, and dividing data within the brain contour; the automatic cleaning refers to classifying the projection blocks using a segmentation model and a classification model; at the same time, rechecking with a data cleaning tool; the existing data cleaning method directly cuts the whole-brain data into blocks, performs maximum projection on the data within the brain contour, and uses the threshold segmentation method to perform binary classification on all maximum projections, and checks and makes up for deficiencies in the binary classification results through manual observation; the cleaning results of each method are as follows:

[0074] Table 3: Time consumption of the present invention for cleaning the whole-brain neuron data of a single rhesus monkey;

[0075]

[0076] Table 4: Time consumption of the traditional method for cleaning the whole-brain neuron data of a single rhesus monkey;

[0077]

[0078] According to the above cleaning results, it can be seen that in the data cleaning method for massive three-dimensional neuron imaging data disclosed in the present invention, the time of manual participation is only about 20 hours, and the total time required to clean the whole-brain neuron data of a rhesus monkey is only about 5 days; while the time of manual participation (i.e., the time consumed for rechecking) and the total cleaning time of the existing data cleaning method are both multiples of the method of the present invention; this shows that the present invention can greatly improve the efficiency of cleaning the whole-brain neuron data of rhesus monkeys, effectively reduce the degree of manual participation, and improve the degree of automation.

[0079] There are still various implementation manners of the present invention. All technical solutions formed by equivalent transformation or equivalent substitution fall within the protection scope of the present invention.

Claims

1. A data cleaning method applicable to massive neuron three-dimensional imaging data, characterized in that: The following steps are involved: Acquire two-dimensional brain image sequence data; Obtaining the maximum projection map of the two-dimensional brain image sequence data, removing the invalid data outside the brain contour in each maximum projection map, and dividing the data within the brain contour to obtain projection blocks, and naming the projection blocks according to their positions in the current maximum projection map; Combining the attention mechanism and the gating mechanism, the neuron signal of the projection block in the maximum projection map is enhanced and the non-neuron signal is weakened to highlight the distribution position of the neuron signal, and the projection blocks containing the neuron signal and the projection blocks without the neuron signal are classified into different projection block types, and the classification results are obtained; According to the classification results, the label information of all projection blocks containing neurons is derived, wherein the label information at least includes the name of the projection block, and a cleaned label information set is generated; Each projection block containing neurons is extracted from the label information set, and all projection blocks containing neurons are segmented and reconnected to obtain a three-dimensional data block of the target shape.

2. The data cleaning method applicable to massive neuron three-dimensional imaging data according to claim 1, characterized in that: After obtaining the classification results, the classification results are reviewed, the incorrectly classified projection blocks are reclassified, and the final classification results are obtained.

3. The data cleaning method applicable to massive neuron three-dimensional imaging data according to claim 1, characterized in that: The method of "obtaining a maximum projection map of two-dimensional brain image sequence data, removing invalid data outside the brain contour in the maximum projection map, and dividing the data within the brain contour into blocks" comprises the following steps: Through viral sparse labeling, the brightness of neurons in imaging data is greater than background information; Taking the preset number of image layers as the unit, the maximum projection is performed on the two-dimensional brain image sequence data of the whole brain in sequence to obtain a number of maximum projection images; Each maximum projection image is downsampled in turn, and upsampled to the original resolution after segmenting the brain contour to obtain the original resolution segmentation image of each maximum projection image; The brain contour area of ​​each of the original resolution segmentation images is segmented into a plurality of projection blocks.

4. The data cleaning method applicable to massive neuron three-dimensional imaging data according to claim 1, characterized in that: Before dividing the projection block types, a segmentation model and a classification model are obtained; wherein the segmentation model and the classification model are both trained using a plurality of projection block data containing neurons and a plurality of projection block data not containing neurons, wherein the segmentation model is used to enhance the neuron signal of the projection block in the maximum projection image and weaken the non-neuron signal to highlight the distribution position of the neuron signal, and the classification model is used to classify the projection blocks containing neuron signals and the projection blocks without neuron signals.

5. The data cleaning method applicable to massive neuron three-dimensional imaging data according to claim 4, characterized in that: The segmentation model is an Attention Unet model. The projection block data set is used to train the Attention Unet model, and the neuron signal is used as the target feature. Each projection block in the maximum projection map to be detected is input into the trained Attention Unet model, and the attention of the neuron signal is increased through the attention mechanism in the model, thereby realizing the extraction of the neuron signal in the current maximum projection map.

6. The data cleaning method applicable to massive neuron three-dimensional imaging data according to claim 4, characterized in that: The classification model is a ResNet-18 convolutional neural network. The projection block data set is used to train the ResNet-18 convolutional neural network. Each projection block in the maximum value projection map to be detected is input into the trained ResNet-18 convolutional neural network. Each projection block is classified using the neuron signal as the classification feature, and the classification result is output.

7. The data cleaning method applicable to massive neuron three-dimensional imaging data according to claim 1, characterized in that: The "segmenting and rejoining all projection blocks containing neurons to obtain a three-dimensional data block of a target shape" means using multi-threaded processing to divide each projection block containing neurons into several small data blocks, creating a thread for each small data block, each thread reads the small data block corresponding to it, and writes the small data block to the correct position of the merged file, so that several small data blocks are rejoined to obtain a three-dimensional data block of the target shape.

8. The data cleaning method applicable to massive neuron three-dimensional imaging data according to claim 7, characterized in that: The small data blocks are obtained by dividing the thickness of the projection block, and the length and width of the small data blocks are equal.

9. The data cleaning method applicable to massive neuron three-dimensional imaging data according to claim 2, characterized in that: The "rechecking the classification results" is achieved by a cleaning tool, which includes: An input module, the input module comprising a projection block data input port, a label information input port and a down-sampled image input port; the projection block data input port is used to input projection block data located within the brain contour in the maximum projection map, including projection blocks containing neurons and projection blocks without neurons; the label information input port is used to input a cleaned label information set; the down-sampled image input port is used to input a 50 μm down-sampled image of the maximum projection map; A display module is used to display the maximum value projection map being operated and the selected area in the current maximum value projection map. The visualization range of the selected area can be adjusted by pressing buttons; An adjustment module is used to display an enlarged view of the selected area and all projection blocks in the enlarged view, and to reclassify the misclassified projection blocks by using the mouse, so as to obtain the final classification result after re-inspection; The export module is used to export the final classification result as a txt file, wherein the txt file contains the names of all neuron projection blocks.

Citation Information

Patent Citations

  • Data cleaning method based on neuron signal identification

    CN117218144A

  • Three-dimensional neuron image segmentation method based on segmentation and super-resolution joint model

    CN118134949A

  • Method and device for detecting pulmonary nodule in computed tomography image, and computer-readable storage medium

    US20200005460A1

  • Object Location Method, Device and Storage Medium Based on Image Segmentation

    US20200057917A1

  • Image classification using machine learning including weakly-labeled data

    WO2024059291A1