A neuroimaging data preprocessing system and a neuroimaging data preprocessing method

By designing a neuroimage data preprocessing system, the problem of diversity of data formats and naming methods is solved, and the standardized processing and automated preprocessing of neuroimage data are realized, which significantly improves the processing efficiency.

CN119621316BActive Publication Date: 2025-07-01CHANGPING NAT LAB
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411687062.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-07-01
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

In the prior art, the data formats and naming methods of neural image data are diverse, which leads to difficulty in data sharing, integration and big data analysis, and the preprocessing speed is slow, which cannot meet the actual application needs.

Method used

A neural image data preprocessing system is designed, including file input module, workflow module, workflow scheduling module and file output module. Through steps such as format recognition and conversion, task generation, calculation node allocation and result storage, standardized processing of neural image data is realized.

Benefits of technology

By introducing format standardization technology, the preprocessing process of neural image data is accelerated, the entire process of data is automated, and the processing efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119621316B_ABST
    Figure CN119621316B_ABST
Patent Text Reader

Abstract

The present invention provides a neuroimaging data preprocessing system and a neuroimaging data preprocessing method. The neuroimaging data preprocessing system and the neuroimaging data preprocessing system include: a file input module, a workflow module, a workflow scheduling module, and a file output module. The file input module is used to perform format recognition and conversion on the original neuroimaging data to obtain neuroimaging data in a standard format; the workflow module is used to generate corresponding tasks to be processed according to the data information of the neuroimaging data in the standard format; the workflow scheduling module is used to allocate computing nodes for the tasks to be processed to complete the preprocessing of the neuroimaging data in the standard format; the file output module is used to perform standardized naming and storage on the preprocessing results corresponding to the original neuroimaging data. The present invention provides a neuroimaging data preprocessing system and a neuroimaging data preprocessing method, which improve the processing efficiency of neuroimaging data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a system and method for preprocessing neuroimaging data. Background Art

[0002] Neuroimaging data can be used to reflect the information set of brain structure, function, and related physiological and pathological states, and has been widely used in the fields of clinical medicine and psychology.

[0003] At present, there are already some preprocessing software for neuroimaging data, which have the preprocessing functions for functional neuroimaging data and structural neuroimaging data. In the prior art, different research institutions and researchers use different devices and software to collect and process data, resulting in various data formats, naming methods, etc. of neuroimaging data, which brings great difficulties to the sharing, integration, and subsequent big data analysis of neuroimaging data, resulting in a slow preprocessing speed of neuroimaging data and unable to meet the requirements of practical applications. Summary of the Invention

[0004] In view of the problems in the prior art, embodiments of the present invention provide a system and method for preprocessing neuroimaging data, which can at least partially solve the problems existing in the prior art.

[0005] In a first aspect, the present invention proposes a system for preprocessing neuroimaging data, including a file input module, a workflow module, a workflow scheduling module, and a file output module, wherein:

[0006] The file input module is used to perform format recognition and conversion on the original neuroimaging data to obtain neuroimaging data in a standard format;

[0007] The workflow module is used to generate corresponding tasks to be processed according to the data information of the neuroimaging data in the standard format;

[0008] The workflow scheduling module is used to allocate computing nodes to the tasks to be processed corresponding to the neuroimaging data in the standard format so that the computing nodes execute the tasks to be processed corresponding to the neuroimaging data in the standard format, and complete the preprocessing of the neuroimaging data in the standard format;

[0009] The file output module is used to perform standardized naming and storage on the preprocessing results corresponding to the original neuroimaging data.

[0010] In a second aspect, the present invention provides a method for preprocessing neuroimaging data, which is applied to the system for preprocessing neuroimaging data according to any one of the above embodiments, and includes:

[0011] Identify and convert the original neuroimaging data to obtain neuroimaging data in a standard format;

[0012] Generate corresponding tasks to be processed according to the data information of the neuroimaging data in the standard format;

[0013] Allocate a computing node for the task to be processed so that the computing node executes the task to be processed corresponding to the neuroimaging data in the standard format to complete the preprocessing of the neuroimaging data in the standard format;

[0014] Standardize the naming and storage of the preprocessing results corresponding to the original neuroimaging data.

[0015] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory. The processor executes the program to implement the neuroimaging data preprocessing method according to any one of the above embodiments.

[0016] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program / instructions, and when the computer program / instructions are executed by a processor, the neuroimaging data preprocessing method according to any one of the above embodiments is implemented.

[0017] In a fifth aspect, the present invention provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the neuroimaging data preprocessing method according to any one of the above embodiments is implemented.

[0018] The neuroimaging data preprocessing system and method provided by the embodiments of the present invention include a file input module, a workflow module, a workflow scheduling module, and a file output module. The file input module is used to identify and convert the original neuroimaging data to obtain neuroimaging data in a standard format; the workflow module is used to generate corresponding tasks to be processed according to the data information of the neuroimaging data in the standard format; the workflow scheduling module is used to allocate a computing node for the task to be processed corresponding to the neuroimaging data in the standard format so that the computing node executes the task to be processed corresponding to the neuroimaging data in the standard format to complete the preprocessing of the neuroimaging data in the standard format; the file output module is used to standardize the naming and storage of the preprocessing results corresponding to the original neuroimaging data, introduce the format standardization technology of neuroimaging data, accelerate the overall preprocessing process, and realize the full-process automation of neuroimaging data processing, improving the processing efficiency of neuroimaging data. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:

[0020] Figure 1 is a schematic structural diagram of a neuroimaging data preprocessing system provided by the first embodiment of the present invention.

[0021] Figure 2 is a schematic structural diagram of a neuroimaging data preprocessing system provided by the second embodiment of the present invention.

[0022] Figure 3 is a schematic flowchart of a neuroimaging data preprocessing method provided by the third embodiment of the present invention.

[0023] Figure 4 is a schematic diagram of each functional unit of the structural image T1w provided by the fourth embodiment of the present invention.

[0024] Figure 5 is a comparison diagram of the data preprocessing effects between the present application and the prior art provided by the fifth embodiment of the present invention.

[0025] Figure 6 is a schematic physical structure diagram of a computer device provided by the sixth embodiment of the present invention. Detailed Embodiments

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the following will further elaborate on the embodiments of the present invention in conjunction with the drawings. Herein, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention. It should be noted that, without conflict, the embodiments and features in the embodiments of the present application can be combined arbitrarily. All the data acquisition, storage, use, processing, etc. in the technical solutions of the present application comply with the relevant regulations of laws and regulations. The user information in the embodiments of the present application is obtained through legal and compliant channels, and the acquisition, storage, use, processing, etc. of the user information have obtained the authorization and consent of the customers.

[0027] To facilitate the understanding of the technical solutions provided by the present application, the following will first explain the relevant content of the technical solutions of the present application.

[0028] In the prior art, the preprocessing speed of neuroimaging software is slow, making it difficult to meet the requirements of the big data era. Moreover, when preprocessing a large volume of neuroimaging data, manual intervention is required to set up the software, and it is impossible to automatically preprocess a large volume of neuroimaging data, reducing the processing efficiency of a large volume of neuroimaging data.

[0029] The present invention proposes a neuroimaging data preprocessing system. Through module splitting technology, deep learning technology, Graphics Processing Unit (GPU) resource scheduling technology, and neuroimaging data standardization technology, the above problems are effectively solved. It can process neuroimaging data quickly, automatically, and standardly, significantly improving the speed of neuroimaging data preprocessing and the preprocessing efficiency. Thus, researchers can focus on data analysis and scientific research without spending a lot of time on neuroimaging data preprocessing.

[0030] For the preprocessing of neuroimaging data, this application designs a standardized storage standard for neuroimaging data, and the entire preprocessing pipeline depends on this standard for file reading and writing. Since the amount of neuroimaging data read and written in the pipeline is frequent and huge, and the standardized neuroimaging data is stored uniquely, there will be no problems of storage redundancy and repeated copying during the execution of different modules, significantly reducing the number of pipeline file reads and writes and improving the retrieval efficiency. This application not only solves the cumbersome and lengthy file retrieval steps caused by non-standard file formats but also avoids the problems of low speed and efficiency caused by redundant file reads and writes.

[0031] Figure 1 It is a schematic structural diagram of the neuroimaging data preprocessing system provided by the first embodiment of the present invention. As Figure 1 shown, the neuroimaging data preprocessing system provided by the embodiment of the present invention includes a file input module 101, a workflow module 102, a workflow scheduling module 103, and a file output module 104, where:

[0032] The file input module 101 is used to identify and convert the format of the original neuroimaging data to obtain neuroimaging data in a standard format;

[0033] The workflow module 102 is used to generate corresponding tasks to be processed according to the data information of the neuroimaging data in the standard format;

[0034] The workflow scheduling module 103 is used to allocate computing nodes for the tasks to be processed corresponding to the neuroimaging data in the standard format so that the computing nodes execute the tasks to be processed corresponding to the neuroimaging data in the standard format and complete the preprocessing of the neuroimaging data in the standard format;

[0035] The file output module 104 is used to standardize the naming and storage of the preprocessing results corresponding to the original neuroimaging data.

[0036] Specifically, the file input module 101 can identify the data format of the original neuroimaging data and determine whether the data format of the original neuroimaging data is the standard format. If it is not the standard format, the original neuroimaging data will be format-converted into neuroimaging data in the standard format. The neuroimaging data in the standard format is beneficial for data sharing and data analysis, and the neuroimaging data in the standard format will be transmitted to the workflow module 102 for subsequent processing. If it is not the standard format, the file input module 101 can output a prompt message to prompt the user to standardize the neuroimaging data. Among them, the standard format is preset, for example, the Brain Imaging DataStructure (BIDS) standard format is adopted.

[0037] For example, if the data format of the original neuroimaging data is the DICOM format, it can be converted into the NIfTI format through the format conversion tool dcm2niix, and information such as the subject information, modality type, acquisition time, equipment, subject number, scan type, and scan sequence of the original data can be parsed. Then, according to the requirements of the standard format, it is named and stored in a standardized manner to obtain neuroimaging data in the standard format. It is also possible to convert the data format of the original neuroimaging data into the standard format through the format conversion tools dcm2bids or BIDScoin.

[0038] For example, if the standard format of the neuroimaging data adopts the BIDS standard format, the BIDS Validator software tool can be used to identify whether the data format of the original neuroimaging data is the BIDS standard format. If the original neuroimaging data is already in the BIDS standard format, then no format conversion is required and it can be directly provided to the workflow module 102 for processing.

[0039] The workflow module 102 generates a to-be-processed task corresponding to the neuroimaging data in the standard format according to the data information of the neuroimaging data in the standard format. The to-be-processed task corresponding to the neuroimaging data in the standard format may include multiple to-be-processed tasks and the dependency relationships between the to-be-processed tasks. Each to-be-processed task is used to complete one or more processes in the neuroimaging data. The to-be-processed tasks with dependency relationships need to be executed in sequence, and the to-be-processed tasks without dependency relationships can be executed in parallel. For different modalities of neuroimaging data, the corresponding to-be-processed tasks are different. The task information of the to-be-processed task may include required computing resources, model identifiers, etc. For example, the to-be-processed task corresponding to the neuroimaging data in the standard format may include multiple to-be-processed tasks, each to-be-processed task corresponds to a functional unit, and each functional unit is obtained by splitting based on the preprocessing process of the neuroimaging data. When generating the to-be-processed task, a functional unit is assigned to each to-be-processed task.

[0040] The workflow scheduling module 103 can obtain the computing resources of each computing node. According to the computing resources of each computing node and the computing resources required by the to-be-processed task, the to-be-processed task is assigned to the computing node for data preprocessing. The computing node assigned with the to-be-processed task will execute the assigned to-be-processed task, complete the corresponding data processing, and then return the data processing result to the file output module 104. Among them, the computing resources include but are not limited to CPU, memory, video memory, etc., which are set according to actual needs, and are not limited in the embodiments of the present invention. The computing node can be a local node or a computing node on different distributed platforms.

[0041] After the to-be-processed task corresponding to the neuroimaging data in the standard format is executed, the preprocessing result corresponding to the original neuroimaging data can be obtained. The file output module 104 performs standardized naming and storage on the preprocessing result corresponding to the original neuroimaging data. The standardized naming method and storage directory are set according to actual needs, and are not limited in the embodiments of the present invention.

[0042] For example, the preprocessing result corresponding to the original neuroimaging data is stored in the BIDS standard format and stored with a predefined directory structure, predefined directory naming, and predefined file naming.

[0043] The neuroimaging data preprocessing system provided by the embodiments of the present invention includes a file input module, a workflow module, a workflow scheduling module, and a file output module. The file input module is used to perform format recognition and conversion on the original neuroimaging data to obtain neuroimaging data in a standard format. The workflow module is used to generate corresponding tasks to be processed according to the data information of the neuroimaging data in the standard format. The workflow scheduling module is used to allocate computing nodes for the tasks to be processed corresponding to the neuroimaging data in the standard format so that the computing nodes execute the tasks to be processed corresponding to the neuroimaging data in the standard format, and complete the preprocessing of the neuroimaging data in the standard format. The file output module is used to perform standardized naming and storage on the preprocessing results corresponding to the original neuroimaging data, introducing the format standardization technology of neuroimaging data, accelerating the overall preprocessing process, and realizing the full-process automation of neuroimaging data processing, improving the processing efficiency of neuroimaging data.

[0044] On the basis of the above embodiments, further, the tasks to be processed corresponding to the neuroimaging data in the standard format include multiple tasks to be processed and the dependency relationships between the respective tasks to be processed. Each task to be processed corresponds to a functional unit, and the respective functional units are obtained by splitting based on the preprocessing process of the neuroimaging data.

[0045] Specifically, the tasks to be processed include multiple tasks to be processed and the dependency relationships between the respective tasks to be processed. A task to be processed is the smallest task unit that can be allocated to a computing node for data processing. The dependency relationships between the respective tasks to be processed refer to whether there is a sequential relationship in the execution of the respective tasks to be processed. If there is no sequential relationship, then the tasks to be processed without a sequential relationship can be executed in parallel to improve the data processing efficiency. If there is a sequential relationship between multiple tasks to be processed, then the above multiple tasks to be processed need to be executed in the sequential relationship. The dependency relationships between the respective tasks to be processed are determined based on the dependency relationships between the functional units corresponding to the respective tasks to be processed, and the dependency relationships between the respective functional units are predetermined based on the preprocessing process of the neuroimaging data.

[0046] Each task to be processed corresponds to a functional unit, and the respective functional units are obtained by splitting based on the preprocessing process of the neuroimaging data. Based on the preprocessing process of the neuroimaging data, the preprocessing process is split into multiple sub-processes. Each sub-process corresponds to a functional unit. Each functional unit corresponds to at least one sub-process, and the functional unit will be configured with a corresponding data processing program to implement the data processing of the corresponding sub-process. Each functional unit obtained after splitting will obtain functional unit attributes, and the functional unit attributes may include input files, output files, maximum memory consumption, average CPU consumption, running time, and GPU video memory consumption.

[0047] The preprocessing process based on neuroimaging data can be split into individual functional units according to the following principles:

[0048] (1) Each preprocessing step in a functional unit has a file dependency. If there is a file dependency between two adjacent steps in the preprocessing process, then these two steps correspond to the same functional unit; if there is no file dependency between two adjacent steps in the preprocessing process, then these two steps can be split into different functional units. For example, if the preprocessing process includes three steps: step A, step B, and step C in sequence, and the execution of step B does not require the output result of step A, but the execution of step C requires the output result of step B, then the above preprocessing process can be split into [step A] and [step B, step C]. [step A] corresponds to one functional unit, and [step B, step C] corresponds to another functional unit.

[0049] (2) When each step in a functional unit runs, whether the resource consumption of every two adjacent steps is roughly the same. For example, if a functional unit includes four steps A, B, C, and D that are executed in sequence, A uses 100% of the applied resources when running, B uses 80% of the applied resources when running, C uses 60% of the applied resources when running, and D uses 20% of the applied resources when running. The functional unit can be further split into three functional units. Functional unit 1 corresponds to steps [A, B], functional unit 2 corresponds to step [C], and functional unit 3 corresponds to step [D]. After the functional unit is split, functional unit 1 can apply for 100% of the resources of the original functional unit, functional unit 2 can apply for 60% of the resources of the original functional unit, and functional unit 3 can apply for 20% of the resources of the original functional unit. Compared with the original functional unit, the resources required to be applied for by functional unit 2 and functional unit 3 are reduced after splitting, avoiding resource waste and improving resource utilization.

[0050] In the process of splitting the preprocessing process of neuroimaging data to obtain each functional unit, the splitting results can be continuously adjusted according to the above two principles, so that the resource consumption of adjacent steps within the functional unit is approximately the same, and there is a dependency relationship between adjacent steps. After each split, it is necessary to update the attribute information of the functional unit, including but not limited to input files, output files, maximum memory consumption, average CPU consumption, and running time, etc. Based on the input files and output files, it can be determined whether there is a dependency relationship between each functional unit. If the output file or part of the output file of a functional unit is used as the input file of another functional unit, then there is a dependency relationship between these two functional modules; based on the maximum memory consumption and average CPU consumption, the resource consumption of each step in the functional unit can be determined. After determining the resource consumption of each step in the functional unit, the computing resources required by the functional unit can be determined. When generating the tasks to be processed, the computing resources required by the functional unit will be used as the computing resources required by the corresponding tasks to be processed. The computing resources can include CPU requirements, memory requirements, video memory requirements, etc., which are set according to actual needs and are not limited in the embodiments of the present invention.

[0051] The data processing flow corresponding to each functional unit is obtained by splitting the preprocessing process of the structural image T1w magnetic resonance image (hereinafter referred to as T1w MRI) in this application as follows: After entering the process, the head movement during the scanning process will be corrected first (corresponding to the head movement correction unit), then the bias field correction based on N4 will be performed (corresponding to the bias field correction unit), then the brain tissue will be segmented into 95 cortical and subcortical regions (corresponding to the brain tissue segmentation unit), then the cortical reconstruction of the cerebral white matter and gray matter will be performed (corresponding to the cortical reconstruction unit), then the registration of the cortical surface will be performed (corresponding to the cortical surface registration unit), and finally the result of the cortical structure partition will be output (cortical structure partition unit); in addition, the registration and projection of the volume (corresponding to the volume data registration unit) have no dependency relationship with the above-mentioned each process.

[0052] The specific splitting process is as follows. For the preprocessing process of T1w MRI, to obtain the above results, since there is no dependency between cortical data processing and volume data processing, it is split into two functional units, the cortical data processing unit and the volume data registration unit. Taking the cortical data processing unit as an example, it is found that the resource consumption varies greatly from the start of execution to the brain tissue segmentation step, from the end of the brain tissue segmentation step to the cortical reconstruction step, and from the end of the cortical reconstruction step to the cortical structure partitioning step. Moreover, after the cortical reconstruction step, the subsequent processing of the left and right cerebral hemispheres has no dependency and can be parallelized. Therefore, the cortical processing unit is divided into three functional units, the brain tissue segmentation unit, the cortical reconstruction unit, and the cortical structure partitioning unit. Subsequently, in the brain tissue segmentation unit, it is found that the resource consumption of the head motion correction step, the bias field correction step, and the brain tissue segmentation step varies greatly, so the brain tissue segmentation unit is split into the head motion correction unit, the bias field correction unit, and the brain tissue segmentation unit. Subsequently, it is found that the computational resources of the cortical surface registration step and the subsequent cortical structure partitioning step in the cortical functional structure partitioning unit vary greatly, so it is divided into the cortical surface registration unit and the cortical structure partitioning unit. Finally, the above-mentioned preprocessing process of T1w MRI corresponding to each functional unit is formed.

[0053] By analyzing the splitting results of the neuroimaging preprocessing process, it is found that the running time of some functional units is long. The corresponding functions of the functional units are implemented through deep learning technology, maintaining consistent file input and output, and these modules with long running time can be accelerated without affecting the neuroimaging preprocessing process. For example, a deep learning-based cerebral cortex registration model and a deep learning-based brain volume registration model are used. The file interfaces of these two models are the same as those of the previous functional units, and the preprocessing process can be accelerated without affecting the overall process. In addition, the neuroimaging preprocessing process can be improved. By introducing deep learning technology, on the basis of accelerating the corresponding functional units, due to the reduction of dependencies, the steps of the neuroimaging preprocessing process can be reduced, and the neuroimaging preprocessing process can be accelerated. For example, using a deep learning-based brain structure segmentation model and a cerebral cortex reconstruction model not only accelerates the original functional units but also reduces the time-consuming units.

[0054] On the basis of the above embodiments, further, the data processing programs corresponding to the functional units are encapsulated into containers to support cross-platform use.

[0055] Specifically, for each functional unit, a corresponding data processing program is configured. The data processing programs corresponding to the functional units can be encapsulated into containers to facilitate cross-platform use. Container technologies such as Docker and Singularity can be used to implement each functional unit.

[0056] Based on the above embodiments, further, among the multiple functional units, there is a functional unit that applies a deep learning model.

[0057] Specifically, for the preprocessing process of neuroimaging data, some of the split processes can be implemented using a deep learning model, and the corresponding functional unit is the functional unit that applies the deep learning model. If the functional unit is a functional unit that applies a deep learning model, a model identifier can be set for the functional unit, and the model identifier corresponding to the functional unit will be used as the model identifier for the corresponding task to be processed.

[0058] For example, brain tissue segmentation can be implemented using deep learning models such as FastSurferCNN, FastSurferVINN, and SynthSeg to segment brain tissue in multi-modal neuroimaging data, improving the efficiency of brain tissue segmentation; cerebral cortex reconstruction can be implemented using deep learning models such as FastCSR and CorticalFlow++, enabling rapid cerebral cortex reconstruction while ensuring topological correctness. Compared with a neuroimaging preprocessing system that does not use a deep learning model, the process of repairing the topology of the cerebral cortex surface is reduced, saving a great deal of time. Cerebral cortex surface registration can be implemented using deep learning models such as S3Reg and SUGAR, reducing the time for cerebral cortex registration and improving the efficiency of cerebral cortex registration. Spatial normalization can be implemented using deep learning models such as VoxelMorph and SynthMorph, reducing time consumption and enabling rigid and non-rigid spatial transformation between different modal neuroimaging data.

[0059] Based on the above embodiments, further, the data information includes the modal type of the original neuroimaging data, and the workflow module 102 is specifically configured to obtain the respective functional units corresponding to the modal type and the dependency relationships between the respective functional units according to the modal type.

[0060] Specifically, neuroimaging data includes neuroimaging data of different modal types, and different modal types of neuroimaging data have different preprocessing processes. The modal types of neuroimaging data and the respective functional units required for preprocessing the neuroimaging data and the dependency relationships between the respective functional units are preset. The workflow module 102 can obtain the modal type of the neuroimaging data in standard format from the data information of the neuroimaging data in standard format, and then query and obtain the respective functional units corresponding to the modal type and the dependency relationships between the respective functional units according to the modal type of the neuroimaging data in standard format. When the workflow module 102 generates the tasks to be processed corresponding to the neuroimaging data in standard format, the tasks to be processed correspond one-to-one with the functional units, and the dependency relationships between the respective functional units are used as the dependency relationships between the respective tasks to be processed.

[0061] Figure 2 is a schematic structural diagram of the neuroimage data preprocessing system provided by the second embodiment of the present invention. As Figure 2 shown, on the basis of the above embodiments, further, the workflow scheduling module 103 includes a queue unit 1031 and a resource allocation unit 1032, where:

[0062] The queue unit 1031 is used to store multiple tasks to be processed;

[0063] The resource allocation unit 1032 is used to allocate a computing node for each task to be processed, and after determining that the multiple tasks to be processed include GPU tasks, according to the video memory size of the GPU of the computing node, the free video memory size of the GPU of the computing node, and the required video memory size of the GPU task, allocate a computing node for the GPU task.

[0064] Specifically, the tasks to be processed corresponding to the neuroimage data in the standard format include multiple tasks to be processed. The queue unit 1031 can store multiple tasks to be processed, and the resource allocation unit 1032 will allocate a computing node for each task to be processed. Each task to be processed has a requirement for computing resources, and the resource allocation unit 1032 will allocate a computing node with sufficient computing resources to process the task to be processed for each task to be processed. The resource allocation unit 1032 can distribute each task to be processed to a computing node that meets the computing resource requirements of the task to be processed in a distributed manner for processing.

[0065] For example, the computing resources required by the task to be processed include CPU requirements and memory requirements. The resource allocation unit 1032 will screen a computing node that meets the CPU requirements and memory requirements of the task to be processed according to the CPU and memory usage conditions of each computing node, and allocate the task to be processed. After receiving the task to be processed, the computing node will execute the task to be processed, complete the data processing, and send the processing result to the file output module 104. Among them, the CPU requirement can be represented by the number of CPUs. For tasks to be processed that can be multi-threaded processed, multiple CPUs can be applied for to improve the processing efficiency. The memory requirement can be represented by the required memory size.

[0066] The resource allocation unit 1032 can determine whether the task to be processed is a GPU task according to the task information of the task to be processed. If the task information of the task to be processed includes a model identifier, it indicates that the data processing process of the task to be processed needs to apply a deep learning model, and this task to be processed is a GPU task, and it is more efficient to use a GPU for processing. If the task information of the task to be processed does not include a model identifier, then this task to be processed is not a GPU task and can be regarded as a regular task to be processed. The resource allocation unit 1032 will obtain the video memory size and the free video memory size of the GPUs of each computing node, and compare the free video memory size of the computing node with the required video memory size of the GPU task. If the free video memory size of the computing node is greater than or equal to the required video memory size of the GPU task, then the GPU task is allocated to this computing node; if the free video memory size of the computing node is less than the required video memory size of the GPU task, and the video memory size of the GPU of the computing node is greater than the required video memory size of the GPU task, the GPU task can be allocated to this computing node, but it is necessary to wait for more free video memory to appear on this computing node. When the free video memory of this computing node is greater than the required video memory size of the GPU task, the GPU task will be processed. If the video memory size of the computing node is less than the required video memory size of the GPU task, then the GPU task will not be allocated to this computing node. It can be seen that the resource allocation unit 1032 will traverse each computing node, screen out the computing nodes with free video memory size greater than the required video memory size of the GPU task, and allocate the GPU task; if there is no computing node with free video memory size greater than the required video memory size of the GPU task, then screen out the computing nodes with video memory size greater than the required video memory size of the GPU task and allocate the GPU task; if there is no computing node with free video memory size greater than the required video memory size of the GPU task, then the GPU task is allocated to the computing node as a regular task to be processed, that is, the video memory size is no longer considered.

[0067] Since GPU models are different, there are significant differences in video memory size. There are also large differences in the video memory occupied by deep learning models. If a GPU is exclusively occupied during the use of a deep learning model, it will reduce the utilization rate of the GPU and increase the preprocessing time of neuroimaging data. Therefore, in the process of applying a deep learning model to process data in this application, the resource allocation unit 1032 schedules GPU tasks based on the video memory usage of computing nodes, and preferentially processes GPU tasks through computing nodes with sufficient video memory resources, which is beneficial to improving the processing efficiency of neuroimaging data.

[0068] Based on the above embodiments, further, the neuroimaging data in the standard format is stored in a preset directory and named in a standard manner.

[0069] Specifically, the present application proposes a standardized storage standard for neuroimaging data. For the configuration items of the configuration file of the neuroimaging data in the standard format, the file storage directory, file naming method, auxiliary file format, storage format, etc. are specified in the file configuration items. During the preprocessing process of neuroimaging data, according to the file configuration items of the neuroimaging data, the storage locations of the relevant files required for the neuroimaging data can be quickly and accurately retrieved, and then read and write operations can be performed. The file storage directory included in the file configuration items is the preset directory for storing the neuroimaging data in the standard format, and the neuroimaging data in the standard format is standardized named according to the file naming method included in the file configuration items.

[0070] For example, the neuroimaging data can be named using the subject number, experimental session, data run, and modality type.

[0071] Based on the above embodiments, further, the neuroimaging data preprocessing system provided by the embodiment of the present invention further includes a visualization module, and the visualization module is used to summarize the preprocessing results of each neuroimaging data and output a visualization report.

[0072] For example, picture outputs can be designed for the processing results of each main functional unit, such as the results of brain tissue segmentation of T1w MRI, the results of cerebral cortex reconstruction of T1w MRI, the head motion correction results of BOLD MRI, and the tSNR (temporal signal-to-noise ratio) of BOLD MRI. Finally, the collected pictures are summarized into a visualization report to help users quickly check the correctness of the preprocessing results. The main functional units are set according to actual needs, and the embodiments of the present invention do not make limitations.

[0073] The present application designs a file-based trigger interface, and realizes a fully automated preprocessing process without manual intervention by managing the input and output files required by each module. A single module will first quickly locate to the specific file directory according to the incoming file configuration items, and retrieve all the required files in this directory. The file configuration items support fuzzy search for a large number of related files and also support precise search for a single target file. After the module retrieves the required target file, it can directly read the content of the target file and the content of the auxiliary file according to the specified format. After processing the data, the module will name the result data file according to the specified format and file configuration in the file configuration items, store it in the specified directory file, and finally transmit the corresponding configuration information of the output file to the dependent module for further processing. Finally, a fully automated preprocessing pipeline is realized.

[0074] Figure 3 is a schematic flowchart of the neuroimaging data preprocessing method provided by the third embodiment of the present invention, asFigure 3 As shown in Figure 3 , the neuroimaging data preprocessing method provided by the embodiments of the present invention is applied to the neuroimaging data preprocessing system described in any of the above embodiments, and includes:

[0075] S301. Identify and convert the original neuroimaging data to obtain neuroimaging data in a standard format;

[0076] Specifically, the file input module can identify the data format of the original neuroimaging data and determine whether the data format of the original neuroimaging data is the standard format. If it is not the standard format, the original neuroimaging data will be format-converted into neuroimaging data in the standard format.

[0077] S302. Generate corresponding tasks to be processed according to the data information of the neuroimaging data in the standard format;

[0078] Specifically, the workflow module will generate tasks to be processed corresponding to the neuroimaging data in the standard format according to the data information of the neuroimaging data in the standard format. The tasks to be processed corresponding to the neuroimaging data in the standard format may include multiple tasks to be processed and the dependency relationships between the respective tasks to be processed. Each task to be processed is used to complete one or more processes in the neuroimaging data.

[0079] S303. Allocate computing nodes for the tasks to be processed so that the computing nodes execute the tasks to be processed corresponding to the neuroimaging data in the standard format to complete the preprocessing of the neuroimaging data in the standard format;

[0080] Specifically, the workflow scheduling module can obtain the computing resources of each computing node, and allocate the tasks to be processed to the computing nodes for data preprocessing according to the computing resources of each computing node and the computing resources required by the tasks to be processed.

[0081] S304. Standardize the naming and store the preprocessing results corresponding to the original neuroimaging data.

[0082] Specifically, the file output module can receive the preprocessing results corresponding to the original neuroimaging data, and then standardize the naming and store the preprocessing results corresponding to the original neuroimaging data.

[0083] The neuroimaging data preprocessing method provided by the embodiments of the present invention can perform format recognition and conversion on the original neuroimaging data to obtain neuroimaging data in a standard format; generate corresponding tasks to be processed according to the data information of the neuroimaging data in the standard format; allocate computing nodes to the tasks to be processed so that the computing nodes execute the tasks to be processed corresponding to the neuroimaging data in the standard format to complete the preprocessing of the neuroimaging data in the standard format; standardize the naming and storage of the preprocessing results corresponding to the original neuroimaging data, realizing the full-process automated processing of neuroimaging data and improving the processing efficiency of neuroimaging data.

[0084] For the embodiments of the neuroimaging data preprocessing method provided by the embodiments of the present invention, reference may specifically be made to the detailed description of the embodiments of the above system, which will not be elaborated herein.

[0085] Next, taking the neuroimaging data preprocessing system provided by the embodiments of the present invention as an example, the specific implementation of the technical solution of the present invention will be described by taking the preprocessing of the structural image T1w as an example.

[0086] Based on the preprocessing process of T1w MRI, the functional units obtained for the structural image T1w include: head motion correction unit, bias field correction unit, brain tissue segmentation unit, cortical reconstruction unit, volume data registration unit, cortical surface registration unit, and cortical structure partitioning unit, as well as the dependencies between the head motion correction unit, bias field correction unit, brain tissue segmentation unit, cortical reconstruction unit, volume data registration unit, cortical surface registration unit, and cortical structure partitioning unit, as Figure 4 shown. Among them, the brain tissue segmentation unit, cortical reconstruction unit, cortical surface registration unit, cortical structure partitioning unit, and volume data registration unit each apply a deep learning model and will set a model identifier. Corresponding the above functional units and the dependencies of each functional unit to the modal type of T1w MRI. And obtaining the computing resource requirements of each of the above functional units.

[0087] The collected original T1w MRI is input into the file input module 101. The file input module 101 identifies whether the data format of the original T1w MRI is the BIDS standard format. If it is not the BIDS standard format, the original T1w MRI will be format-converted to obtain the T1w MRI in the standard format. If it is the BIDS standard format, subsequent processing will be directly performed. The T1w MRI in the standard format will be stored in a preset directory and named in the standard way of subject number (subject), experimental session (session), data run (run), and modality type (type). The file input module 101 can parse and obtain information such as subject information, experimental session, data run, modality type, acquisition time, equipment, subject number, scan type, and scan sequence from the original data header information of the original T1w MRI.

[0088] The workflow module 102 generates corresponding tasks to be processed for the T1w MRI in the standard format. The workflow module 102 obtains the modality type from the data information of the T1w MRI in the standard format. According to the modality type of the T1w MRI, the head motion correction unit, bias field correction unit, brain tissue segmentation unit, cortical reconstruction unit, cortical surface registration unit, cortical structure partition unit, and volume data registration unit can be queried. And the dependency relationships of the head motion correction unit, bias field correction unit, brain tissue segmentation unit, cortical reconstruction unit, cortical registration unit, and cortical structure partition unit are sequential. The volume data registration unit can be performed independently. Corresponding tasks to be processed are generated for the head motion correction unit, bias field correction unit, brain tissue segmentation unit, cortical reconstruction unit, cortical surface registration unit, cortical structure partition unit, and volume data registration unit, obtaining the task to be processed 1 corresponding to the head motion correction unit, the task to be processed 2 corresponding to the bias field correction unit, the task to be processed 3 corresponding to the brain tissue segmentation unit, the task to be processed 4 corresponding to the cortical reconstruction unit, the task to be processed 5 corresponding to the cortical surface registration unit, the task to be processed 6 corresponding to the cortical structure partition unit, and the task to be processed 7 corresponding to the volume data registration unit. And based on the sequential relationship of the head motion correction unit, bias field correction unit, brain tissue segmentation unit, cortical reconstruction unit, cortical surface registration unit, and cortical structure partition unit, it is obtained that the task to be processed 1, the task to be processed 2, the task to be processed 3, the task to be processed 4, the task to be processed 5, and the task to be processed 6 are executed sequentially, and the task to be processed 7 can be executed independently. The task information of the task to be processed 1 and the task to be processed 2 does not include the model identifier. The task information of the task to be processed 3, the task to be processed 4, the task to be processed 5, the task to be processed 6, and the task to be processed 7 includes the model identifier.

[0089] The queue unit 1031 of the workflow scheduling module 103 stores the execution order of the to-be-processed tasks 1, 2, 3, 4, 5, and 6 into the to-be-processed task queue in sequence, and stores the to-be-processed task 7 into the to-be-processed task queue.

[0090] The resource allocation unit 1032 of the workflow scheduling module 103 allocates computing nodes for each to-be-processed task. The resource allocation unit 1032 determines that the to-be-processed tasks 3, 4, and 5 are GPU tasks, and the to-be-processed tasks 1 and 2 are regular to-be-processed tasks according to the model identifiers included in the task information of the above five to-be-processed tasks. The resource allocation unit 1032 allocates computing nodes for the to-be-processed tasks 1 and 2 in sequence according to the computing resource requirements of the to-be-processed tasks 1 and 2. The resource allocation unit 1032 allocates computing nodes for the to-be-processed tasks 3, 4, 5, 6, and 7 in sequence according to the required video memory sizes of the to-be-processed tasks 3, 4, 5, 6, and 7, and allocates a computing node for the to-be-processed task 7.

[0091] The computing node that allocates the to-be-processed task 1 calls the container of the data processing program corresponding to the deployed head motion correction unit to execute the allocation of the to-be-processed task 1, and obtains the processing result of the allocation of the to-be-processed task 1. The computing node that allocates the to-be-processed task 2 calls the container of the data processing program corresponding to the deployed bias field correction unit to execute the allocation of the to-be-processed task 2, and obtains the processing result of the allocation of the to-be-processed task 2. The computing node that allocates the to-be-processed task 3 calls the container of the data processing program corresponding to the deployed brain tissue segmentation unit to execute the allocation of the to-be-processed task 3, and obtains the processing result of the allocation of the to-be-processed task 3. The computing node that allocates the to-be-processed task 4 calls the container of the data processing program corresponding to the deployed cortex reconstruction unit to execute the allocation of the to-be-processed task 4, and obtains the processing result of the allocation of the to-be-processed task 4. The computing node that allocates the to-be-processed task 5 calls the container of the data processing program corresponding to the deployed cortex surface registration unit to execute the allocation of the to-be-processed task 5, and obtains the processing result of the allocation of the to-be-processed task 5. The computing node that allocates the to-be-processed task 6 calls the container of the data processing program corresponding to the deployed cortex structure partitioning unit to execute the allocation of the to-be-processed task 6, and obtains the processing result of the allocation of the to-be-processed task 6. The computing node that allocates the to-be-processed task 7 calls the container of the data processing program corresponding to the deployed volume data registration unit to execute the allocation of the to-be-processed task 7, and obtains the processing result of the allocation of the to-be-processed task 7.

[0092] The file output module 104 can perform standardized naming and storage on the processing results of the task to be processed 1, the processing results of the task to be processed 2, the processing results of the task to be processed 3, the processing results of the task to be processed 4, the processing results of the task to be processed 5, the processing results of the task to be processed 6, and the processing results of the task to be processed 7. It can store the processing results of the task to be processed 1, the processing results of the task to be processed 2, the processing results of the task to be processed 3, the processing results of the task to be processed 4, the processing results of the task to be processed 5, the processing results of the task to be processed 6, and the processing results of the task to be processed 7 in the same root directory, store them in different folders, store them in the BIDS standard format, and name them with the agreed file names. The processing results can be standardized named using the subject number, experimental session, data run, modality type, and processing status.

[0093] The open-source software toolkit for magnetic resonance imaging (MRI) data preprocessing in the prior art and the neuroimaging data preprocessing system of this application are respectively used to preprocess the same neuroimaging data. As Figure 5 shown, in the case of serial processing, the processing speed of the technical solution of this application for neuroimaging data is more than 10 times that of the prior art; in the case of batch processing, the cumulative number of measured objects processed by this application is more than 10 times that of the prior art; in the case of cluster processing, the resource consumption of this application is at most 15% of the prior art; in the case of clinical data processing, the completion rate and accuracy rate of this application are significantly higher than those of the prior art. Figure 5 In, the blue corresponds to the technical solution of this application, and the gray corresponds to the prior art.

[0094] Figure 6 is a schematic diagram of the physical structure of the computer device provided by the sixth embodiment of the present invention. As Figure 6As shown in the figure, the computer device may include: a processor 601, a communications interface 602, a memory 603, and a communication bus 604. Among them, the processor 601, the communications interface 602, and the memory 603 communicate with each other through the communication bus 604. The processor 601 may call the logical instructions in the memory 603 to execute the methods provided in the foregoing method embodiments. For example, it includes: performing format recognition and conversion on the original neuroimaging data to obtain neuroimaging data in a standard format; generating corresponding processing tasks to be processed according to the data information of the neuroimaging data in the standard format; allocating computing nodes for the processing tasks to be processed so that the computing nodes execute the processing tasks corresponding to the neuroimaging data in the standard format to complete the preprocessing of the neuroimaging data in the standard format; performing standardized naming and storage on the preprocessing results corresponding to the original neuroimaging data.

[0095] In addition, when the logical instructions in the foregoing memory 603 are implemented in the form of software functional units and sold or used as independent products, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs, Read-Only Memories), random access memories (RAMs, Random Access Memories), magnetic disks, or optical discs.

[0096] This embodiment discloses a computer program product. The computer program product includes computer programs / instructions stored on a non-transitory computer-readable storage medium. When the computer programs / instructions are executed by a computer, the computer can execute the methods provided in the foregoing method embodiments. For example, it includes: performing format recognition and conversion on the original neuroimaging data to obtain neuroimaging data in a standard format; generating corresponding processing tasks to be processed according to the data information of the neuroimaging data in the standard format; allocating computing nodes for the processing tasks to be processed so that the computing nodes execute the processing tasks corresponding to the neuroimaging data in the standard format to complete the preprocessing of the neuroimaging data in the standard format; performing standardized naming and storage on the preprocessing results corresponding to the original neuroimaging data.

[0097] This embodiment provides a computer-readable storage medium that stores computer programs / instructions, which cause the computer to execute the methods provided in the above method embodiments. For example, it includes: performing format recognition and conversion on the original neuroimaging data to obtain neuroimaging data in a standard format; generating corresponding tasks to be processed according to the data information of the neuroimaging data in the standard format; allocating computing nodes to the tasks to be processed so that the computing nodes execute the tasks to be processed corresponding to the neuroimaging data in the standard format to complete the preprocessing of the neuroimaging data in the standard format; and performing standardized naming and storage on the preprocessing results corresponding to the original neuroimaging data.

[0098] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0099] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0100] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0101] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process Figure 1 one process or a plurality of processes and / or blocks Figure 1 or steps for realizing the functions specified in a block or a plurality of blocks.

[0102] In the description of the present specification, descriptions with reference to the terms "one embodiment", "a specific embodiment", "some embodiments", "for example", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In the present specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0103] The above specific embodiments have further elaborated on the object, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A neural imaging data preprocessing system, characterized in that: It includes a file input module, a workflow module, a workflow scheduling module and a file output module, wherein: The file input module is used to identify and convert the format of the original neural imaging data to obtain the neural imaging data in a standard format; The workflow module is used to generate corresponding tasks to be processed according to the data information of the neuroimaging data in the standard format; The workflow scheduling module is used to allocate computing nodes to the pending tasks corresponding to the neural imaging data in the standard format so that the computing nodes execute the pending tasks corresponding to the neural imaging data in the standard format and complete the preprocessing of the neural imaging data in the standard format; The file output module is used to standardize the naming and storage of the preprocessing results corresponding to the original neuroimaging data; Among them, the data information includes the modal type of the original neural imaging data, and the workflow module is specifically used to obtain the functional units corresponding to the modal type and the dependency relationship between the functional units according to the modality type, and each task to be processed corresponds to a functional unit; each functional unit is obtained by splitting based on the preprocessing process of the neural imaging data, and the dependency relationship between the functional units is predetermined based on the preprocessing process of the neural imaging data.

2. The neural imaging data preprocessing system according to claim 1, characterized in that: The tasks to be processed corresponding to the neuroimaging data in the standard format include multiple tasks to be processed and dependencies between the tasks to be processed.

3. The neural image data preprocessing system according to claim 2, characterized in that: The data processing program corresponding to the functional unit is encapsulated into a container to support cross-platform use.

4. The neural image data preprocessing system according to claim 2, characterized in that: The functional units include a functional unit for applying a deep learning model.

5. The neural image data preprocessing system according to claim 1, characterized in that: The workflow scheduling module includes a queue unit and a resource allocation unit, wherein: The queue unit is used to store multiple tasks to be processed; The resource allocation unit is used to allocate a computing node to each task to be processed, and after determining that the multiple tasks to be processed include a GPU task, allocate a computing node to the GPU task according to the video memory size of the GPU of the computing node, the free video memory size of the GPU of the computing node, and the required video memory size of the GPU task.

6. The neural imaging data preprocessing system according to any one of claims 1 to 5, characterized in that: The neuroimaging data in the standard format are stored in a preset directory and named in a standard manner.

7. A method for preprocessing neural image data based on the neural image data preprocessing system according to any one of claims 1 to 6, characterized in that: include: Perform format recognition and conversion on raw neuroimaging data to obtain neuroimaging data in a standard format; Generate corresponding tasks to be processed according to the data information of the neuroimaging data in the standard format; Allocating computing nodes to the tasks to be processed so that the computing nodes execute the tasks to be processed corresponding to the neural imaging data in the standard format to complete preprocessing of the neural imaging data in the standard format; The preprocessing results corresponding to the original neuroimaging data are named and stored in a standardized manner.

8. A computer device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the method of claim 7.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program / instruction, and when the computer program / instruction is executed by a processor, the method of claim 7 is implemented.

10. A computer program product comprising a computer program / instructions, characterized in that The computer program / instructions implement the method of claim 7 when executed by a processor.

Citation Information

Patent Citations

  • Distributed CT imaging and intelligent diagnosis and treatment system and method

    CN116646061A