A Distributed Reconstruction Resource Management Method Based on Cloud Devices
通过在云端的分布式重建系统中动态分配资源,解决了PET重建工作站硬件限制的问题,实现了高效的重建任务处理和负载均衡,缩短了重建时间。
Patent Information
- Application Number
- CN202210379557.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-04-12
AI Technical Summary
The existing PET reconstruction workstation hardware resources are limited and multiple reconstruction tasks cannot be processed at the same time, resulting in too long waiting time, and the computing complexity of advanced algorithms increases and the reconstruction time is extended under the same hardware configuration.
The distributed reconstruction resource management method based on cloud devices is adopted, and the available resource information is periodically broadcasted in the broadcast domain, and the computing tasks are dynamically allocated, and multiple nodes are used to jointly handle the reconstruction tasks to achieve load balancing.
It improves the processing efficiency of reconstruction tasks, reduces waiting time, rationally utilizes the idle resources of the distributed system, and optimizes the allocation of computing resources.
Smart Images

Figure CN114721828B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed reconstruction, and particularly to a distributed reconstruction resource management method based on cloud devices. Background Art
[0002] PET-CT is a nuclear medicine imaging device that combines PET and CT. Among them, PET (Positron Emission Tomography) is responsible for collecting PET sequences with functional imaging functions; CT (X-ray Computed Tomography) is responsible for collecting CT sequences with structural imaging functions. After the PET data acquisition is completed, the algorithm module uses the PET raw data for reconstruction to generate PET images. During the reconstruction process, the CT image is also used for attenuation correction of the PET image reconstruction.
[0003] The raw data obtained during PET scanning is transmitted to the PET reconstruction workstation in real time. This workstation is responsible for storing the PET raw data and various system files required for reconstruction, as well as reconstructing the PET raw data into images. Taking a relatively common hardware configuration as an example, such as an Intel Xeon W series 12-core CPU and 32GB of memory, with this hardware configuration, for the data volume obtained from a 2-minute trunk scan, using the OSEM (Ordered Subsets Expectation Maximization method) + TOF (Time of Flight) reconstruction method, with 2 iterations and 10 subsets, it takes about 2 minutes to complete the reconstruction. Moreover, due to the limited hardware resources of the PET reconstruction workstation, only one reconstruction can be processed at the same time. Once there are multiple reconstructions in the scan, or the user needs to perform multiple post-reconstructions on the existing data during the scan, the reconstructions that cannot start can only enter the queue and wait, and this kind of waiting often lasts for a long time. In recent years, with the continuous progress of reconstruction algorithms, various advanced algorithms have gradually emerged, such as image optimization algorithms using deep learning technology. Due to the complexity of internal calculations, these advanced algorithms consume more time under the same hardware configuration.
[0004] In order to shorten the reconstruction time and improve the reconstruction efficiency, how to use cloud computing to achieve remote reconstruction has become a technical problem that needs to be solved urgently at present. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a distributed reconstruction resource management method based on cloud devices, which is used to shorten the reconstruction time and improve the reconstruction efficiency in cloud reconstruction, and solve the technical problem of unbalanced computing resources.
[0007] (2) Technical Solutions
[0008] To achieve the above object, the main technical solutions adopted by the present invention include:
[0009] In a first aspect, an embodiment of the present invention provides a distributed reconstruction resource management method based on cloud devices. The distributed reconstruction resource management method is applied to a distributed reconstruction system for medical image reconstruction in a medical imaging device. A plurality of nodes implementing a data reconstruction program are configured in the distributed reconstruction system. All nodes are located in a broadcast domain of a distributed reconstruction system and periodically broadcast available resource information in a broadcast manner. The method includes:
[0010] S10. The first node periodically obtains available resource information in the first node and broadcasts it to other nodes in the broadcast domain, and receives available resource information broadcast by each node in the broadcast domain;
[0011] S20. The first node determines whether a local reconstruction task sent by an imaging device connected to the first node is received;
[0012] S30. If the first node receives a local reconstruction task, the local reconstruction task is segmented and threads for the segmented data are created according to a resource allocation strategy, the available resource information of the current first node, and the available resource information of other nodes in the broadcast domain, so as to distributively complete the local reconstruction task;
[0013] The resource allocation strategy is used to determine one or more nodes capable of processing the local reconstruction task according to the processing information of the local reconstruction task, and the local reconstruction task is segmented according to the selected nodes, the processing efficiency of the nodes, and the available resource information of the nodes.
[0014] Optionally, the method further includes:
[0015] After all nodes processing the local reconstruction task complete the processing of the data assigned to them, the available resource information inside the nodes is updated and broadcast;
[0016] Among them, when the first node determines one or more nodes capable of processing the local reconstruction task, the priority of the first node is the highest priority.
[0017] Optionally, a data structure of available resource information of each node in the broadcast domain and a node efficiency information table of each node are maintained in each node;
[0018] Among them, the information in the node efficiency information table of each node stored in the first node is the processing efficiency of each node calculated by the first node based on the data packets processed by each node within a specified duration.
[0019] Optionally, SOCKET communication mode is adopted for communication between any two nodes;
[0020] The available resource information includes one or more of the following:
[0021] Total number of CPU cores, currently available cores; memory information of the node, free memory information, response speed, supported algorithms.
[0022] In a second aspect, an embodiment of the present invention further provides a server, which belongs to a node of a distributed reconstruction system for medical image reconstruction in a medical imaging device. The server is located in a broadcast domain and broadcasts available resource information to other nodes in the distributed reconstruction system in a broadcast manner. The server includes:
[0023] A distributed computing and processing module, configured to communicate with other nodes in the distributed reconstruction system, receive a local reconstruction task transmitted by an imaging device connected to the server, and perform data segmentation on the local reconstruction task and create threads and / or processes for the segmented data based on a resource allocation policy, available resource information of the current node, and available resource information of other nodes in the broadcast domain;
[0024] A reconstruction module, configured to process the segmented local reconstruction task according to threads and / or processes of at least one node determined by the distributed computing and processing module by invoking a distributed service interface in the distributed reconstruction system to complete the reconstruction;
[0025] The resource allocation policy is used to determine more than one node capable of processing the local reconstruction task according to the processing information of the local reconstruction task, and segment the local reconstruction task according to the selected node, the processing efficiency of the node, and the available resource information of the node.
[0026] Optionally, the distributed computing and processing module includes:
[0027] A network communication unit, configured to communicate with other nodes in the distributed reconstruction system;
[0028] A computing resource management unit, configured to periodically obtain available resource information of other nodes in the distributed reconstruction system, and maintain a data structure of available resource information of each node in the broadcast domain and a node efficiency information table of each node stored inside the server;
[0029] A computing resource dynamic utilization unit, configured to determine more than one node capable of processing the local reconstruction task based on a resource allocation policy, available resource information of the current node, and available resource information of other nodes in the broadcast domain;
[0030] A data splitting and merging unit, configured to split the local reconstruction task according to all selected nodes, the processing efficiency of each node among all the nodes, and the available resource information, and distribute it to the selected nodes for reconstruction processing.
[0031] Optionally, the network communication unit is configured to communicate using the SOCKET communication method; or communicate using the HTTPSOCKET communication method;
[0032] A computing resource management unit, configured to send the available resource information and reconstruction task information of the local node in the form of a broadcast to the broadcast domain; and receive broadcast packets sent by other nodes in the broadcast domain;
[0033] The reconstruction task information includes information of the local reconstruction task or information of the reconstruction tasks distributed by other nodes;
[0034] A computing resource dynamic utilization unit, configured to, after receiving the local reconstruction task and determining the reconstruction algorithm, obtain the available resource information that can be used by the local node; and the available computing resource information of other nodes in the broadcast domain, and determine the available resource information of the local server and the available resource information of other nodes in the broadcast domain according to the local reconstruction concurrency limit and the reconstruction concurrency limit within the broadcast domain, select all nodes for processing the local reconstruction task, and create multiple threads and / or processes for processing the local reconstruction task;
[0035] A data splitting and merging unit, configured to receive the threads and / or processes sent by the computing resource dynamic utilization unit, the available resource information of all selected nodes, and the processing efficiency of each node among all selected nodes, split the local reconstruction task, and determine the data size allocated to each selected node for distribution;
[0036] The threads and / or processes for processing the split data run partially on the local node and partially on other nodes in the broadcast domain.
[0037] Optionally, it further includes:
[0038] The computing resource dynamic utilization unit creates the maximum number of threads for processing the local reconstruction task based on the available resource information of all selected nodes.
[0039] In a third aspect, an embodiment of the present invention further provides a distributed reconstruction system based on cloud devices, including multiple servers according to any one of the second aspects. Each server forms a decentralized broadcast domain as a node, and each server corresponds to one or more imaging devices located in a hospital.
[0040] (III) Advantageous Effects
[0041] The method of the embodiment of the present invention broadcasts the available resource information of each node in the broadcast domain in real time or periodically, so that when a node has a local reconstruction task, it can be jointly processed by more than one node, thereby effectively improving the processing efficiency of the local reconstruction task and reducing the waiting time. At the same time, the node that receives the local reconstruction task is preferably processed, resources are reasonably distributed, and idle resources in the distributed reconstruction system are reasonably utilized.
[0042] In addition, in the embodiment of the present invention, the service of each node in the broadcast domain has mastered the dynamic computing resource information of all nodes in its broadcast domain. At this time, the domain composed of all nodes is decentralized, and the nodes in the domain can be flexibly increased or decreased dynamically without affecting the overall distributed computing function. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A schematic diagram of a process flow of a distributed reconstruction resource management method based on cloud devices provided by an embodiment of the present invention;
[0044] Figure 2 A schematic diagram of a flow chart of a distributed reconstruction resource management method based on cloud devices provided in another embodiment of the present invention;
[0045] Figure 3 A schematic diagram of the structure of a distributed reconstruction system based on cloud devices provided by an embodiment of the present invention;
[0046] Figure 4 A schematic diagram of a time slice used by a node in a distributed reconstruction system provided by an embodiment of the present invention;
[0047] Figure 5 A schematic diagram of dividing a task to be reconstructed into data packets provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0048] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation modes in conjunction with the accompanying drawings.
[0049] With the development of cloud computing over the years and the gradual popularization of new-generation communication technologies represented by 5G, using powerful cloud computing resources and high-speed networks to achieve remote reconstruction has become a feasible technical solution.
[0050] In other words, many medical imaging devices such as PET-CT and PET will no longer be configured with local workstations, but will use dedicated workstations in the cloud. All imaging data is completed using cloud computing. At this time, the problem of uneven utilization of computing resources is brought about.
[0051] For example, assume there is Company A that sells and installs 100 devices, corresponding to 100 cloud workstations. Since these 100 devices are distributed in different hospitals in different regions, the number of patients and busy times vary. During a certain period, among the 100 cloud workstations, some nodes (i.e., cloud workstations) will be busy and some will be idle. Even some nodes will be idle most of the time due to the small number of patients. In this case, it is entirely possible to utilize the computing resources of the idle nodes to share part of the computing tasks for the busy nodes, improve the data processing speed, return the results to the imaging device faster, and improve the data processing performance.
[0052] The name descriptions involved in the embodiments of the present invention are as follows:
[0053] Node: A service in the cloud used for data reconstruction of imaging devices, that is, a cloud computing workstation for image reconstruction calculation. All such workstations in the embodiments of the present invention are referred to as nodes (i.e., cloud devices).
[0054] Broadcast number: A unique identifier unique to each node. The service running on each node for processing image reconstruction tasks uses this number to identify itself and is used to be recognized by the services of other nodes.
[0055] Broadcast domain: A broadcast domain is a set or category of broadcast numbers. The service of each node can use the broadcast method to send information to one or more broadcast domains.
[0056] Local reconstruction concurrency limit: A set of values used to specify the maximum number of CPU cores that the service can utilize locally on the node and the corresponding maximum concurrency.
[0057] Intra-domain reconstruction concurrency limit: A set of values used to specify the maximum number of CPU cores that the service can utilize in its broadcast domain except locally on the node and the corresponding maximum concurrency.
[0058] Service: A program running mode. The functions described in this embodiment run on the node in the form of a service. The services mentioned in the following embodiments can all refer to the functional scope running on the node in the form of a service.
[0059] Time slice: When performing the reconstruction task planning, it is the time interval for adjusting the task allocation to the nodes, such as Figure 5As shown in the figure. Within a time slice, the usage mode of a specific node (the number of CPU cores occupied and the size of data sent) remains unchanged. Since the start and end times of the reconstruction tasks for each node are not synchronized, the time slice in the embodiments of the present invention is planned independently for each node. If the time slice is too short, it will cause the thread task not to have enough time to complete, and the algorithm will consume a large amount of time on the dynamic programming task. If the time slice is too long, it will cause the node resources to be occupied by non-local reconstruction tasks for a long time, resulting in the local task being unable to effectively use the local node computing resources. In the embodiments of the present invention, the time slice information of each node is determined based on the node processing efficiency of each node.
[0060] Embodiment 1
[0061] As Figure 1 and Figure 2 As shown in the figure, the embodiments of the present invention provide a distributed reconstruction resource management method based on cloud devices. In this embodiment, the distributed reconstruction resource management method is applied to a distributed reconstruction system for medical image reconstruction in medical imaging devices (such as PET-CT, PET, CT). A plurality of nodes (i.e., cloud devices) that implement data reconstruction programs are configured in the distributed reconstruction system. All nodes are located in the broadcast domain of a distributed reconstruction system and periodically broadcast available resource information in a broadcast manner. The method includes:
[0062] S10. The first node periodically obtains the available resource information in the first node and broadcasts it to other nodes in the broadcast domain, and receives the available resource information broadcast by each node in the broadcast domain.
[0063] Specifically, the services on each node can broadcast the available resource information to the broadcast domain at a predetermined time interval (such as 1s / 2s / 3s, etc.). The available resource information may include: the computing resource information of the node, such as the number of cores and load of the CPU, and the memory occupancy.
[0064] It should be noted that the first node in the above steps will also receive the available resource information broadcast by other nodes in the broadcast domain. By broadcasting the available resource information, the services of each node in the broadcast domain can obtain the dynamic computing resource information of all nodes in the broadcast domain in real time. Thus, all nodes in the broadcast domain are decentralized, and the nodes in the broadcast domain can be flexibly added or reduced dynamically without affecting the overall distributed computing function.
[0065] S20. The first node periodically checks whether it receives a local reconstruction task.
[0066] Generally, each node within a broadcast domain corresponds to more than one imaging device installed locally in a hospital, which is used to receive data uploaded by the imaging device for reconstruction. The service within each node is used to reconstruct the received data. In this embodiment, the service within the node reconstructs a to-be-reconstructed task received based on the resources within the broadcast domain.
[0067] S30. If the first node receives a local reconstruction task, it performs data segmentation on the local reconstruction task and creates threads / or processes for the segmented data according to the resource allocation policy, the available resource information of the current first node, and the available resource information of other nodes within the broadcast domain, so as to complete the local reconstruction task in a distributed manner;
[0068] The resource allocation policy is used to determine more than one node capable of processing the local reconstruction task according to the processing information of the local reconstruction task, and segment the local reconstruction task according to the selected node, the processing efficiency of the node, and the available resource information of the node.
[0069] When a set of data, its reconstruction method, and reconstruction parameters of the data waiting for reconstruction are sent from an imaging device locally in the hospital to the corresponding node, the service obtains the available resource information for implementing the reconstruction task from the local of the node and other nodes within the broadcast domain according to the local reconstruction concurrency limit and the intra-domain reconstruction concurrency limit. If all the available resource information within the broadcast domain cannot reach the upper limit set by the intra-domain reconstruction concurrency limit, then obtain the maximum computing resources (i.e., available resource information) that can be provided within the broadcast domain; then, provide all the available resource information to the service within the first node for implementing the reconstruction task (which can be called the original data reconstruction service).
[0070] For example, if the first node is providing computing resources to other nodes within the broadcast domain and thus cannot reach the upper limit of the local reconstruction concurrency limit, then obtain the maximum computing resources within the first node. That is, the computing resources within the first node are preferentially provided to the local reconstruction task received by the first node. If all the computing resources within the broadcast domain cannot reach the upper limit set by the intra-domain reconstruction concurrency limit, then obtain the maximum available resource information that can be provided within the broadcast domain, and provide the available resource information to the original data reconstruction service within the first node.
[0071] At this time, the original data reconstruction service within the first node creates threads / or processes according to the obtained available resource information to process the original data to be reconstructed in parallel.
[0072] In this embodiment, the maximum number of threads that the original data reconstruction service can create is to match the local reconstruction concurrency limit plus the intra-domain reconstruction concurrency limit. The maximum number of threads can be determined by the available resource information within the broadcast domain that the original data reconstruction service within the first node can obtain.
[0073] Based on the above description, it can be understood that the method of this embodiment may further include the following step S40 not shown in the figure:
[0074] S40: After the selected node finishes processing the local reconstruction task, update the available resource information inside the node and broadcast it;
[0075] Among them, when the first node determines more than one node capable of processing the local reconstruction task, the priority of the first node is the highest priority.
[0076] In this embodiment, each node comprehensively stores the available resource information and node processing efficiency of each node in the broadcast domain; the node processing efficiency is the amount of data processed by a thread in a time slice. That is, for the raw data reconstruction service, according to the obtained available resource information, threads are created to process the data to be reconstructed in parallel. The maximum number of threads that the raw data reconstruction service can create.
[0077] In this embodiment, SOCKET communication mode is adopted for communication between any two nodes in the broadcast domain; or, HTTPSOCKET communication mode is adopted for communication between any two nodes;
[0078] The available resource information (which can be called computing resource / computing resource information) includes one or more of the following: total number of CPU cores, current available cores; memory information of the node, free memory information, response speed, supported algorithms.
[0079] It can be understood that during the process of the service reconstructing data in each node, the service also needs to determine the segmentation granularity of the data to be reconstructed (that is, the raw data of the local reconstruction task received by the first node). Here, the granularity refers to the length of each data segment after the raw data is segmented.
[0080] In this embodiment, the segmentation algorithm of the first node can determine the segmentation granularity of the local reconstruction task based on the segmentation and merging efficiency and the single-task processing time. If the granularity is too small, the service in each node will consume a large amount of resources on the segmentation of the raw data and the integration of the calculation results. If the granularity is too large, it will lead to too long single-data processing time, causing the computing resources of the nodes in the broadcast domain to be occupied by other nodes for a long time, resulting in the inability to obtain local computing resources in time when there is a local reconstruction task in the node.
[0081] The reconstruction resource management service of this embodiment can exist as a middleware between the post-reconstruction program and the operating system. For the post-reconstruction program, it is not clear whether the threads executed in parallel are specifically executed on the CPU of the local node or the CPUs of other nodes in other cloud devices. The service integrates cloud nodes through a unique distributed computing algorithm, makes full use of the computing resources of idle nodes, and improves the post-reconstruction speed.
[0082] Example 2
[0083] In this embodiment, when using a cloud computing workstation to execute the PET reconstruction task / PET-CT reconstruction task locally in a hospital, it makes full use of the computing resources of the cloud workstation to achieve load balancing and improve the reconstruction speed.
[0084] The server in this embodiment, i.e., the cloud device, belongs to a node in a distributed reconstruction system for medical image reconstruction in medical imaging devices. The service (i.e., program) within the node is used to implement the reconstruction of PET-CT data. The server is located in a broadcast domain and broadcasts available resource information to other nodes in the distributed reconstruction system in a broadcast manner. All the nodes constitute the distributed reconstruction system in this embodiment.
[0085] Combined with Figure 3 As shown, the server in this embodiment may include: a distributed computing and processing module A10 and a reconstruction module A20.
[0086] Specifically, the distributed computing and processing module A10 is used to communicate with other nodes in the distributed reconstruction system, periodically check the local reconstruction task, and based on the resource allocation strategy, the available resource information of the current node, and the available resource information of other nodes in the broadcast domain, perform data segmentation on the local reconstruction task and create threads, processes, etc. for the segmented data;
[0087] The reconstruction module A20 is used to process the segmented local reconstruction task according to the threads, processes, etc. of at least one node determined by the distributed computing and processing module through the distributed service interface in the distributed reconstruction system to complete the reconstruction.
[0088] It should be noted that the reconstruction module A20 can draw on the functions of the reconstruction module in the existing cloud workstation, and realize the reconstruction of the local reconstruction task by calling the interfaces of the distributed services corresponding to the computing resources of the processes, threads, etc. created by the distributed computing and processing module A10.
[0089] The resource allocation strategy in this embodiment is used to determine more than one node capable of processing the local reconstruction task according to the processing information of the local reconstruction task, and segment the local reconstruction task according to the selected node, the processing efficiency of the node, and the available resource information of the node.
[0090] In a possible implementation manner, the above-mentioned distributed computing and processing module A10 may include: a network communication unit A11, a computing resource management unit A12, a computing resource dynamic utilization unit A13, and a data segmentation and merging unit A14.
[0091] Among them, the network communication unit A11 is used to communicate with other nodes in the distributed reconstruction system;
[0092] The computing resource management unit A12 is used to periodically obtain the available resource information of other nodes in the distributed reconstruction system; in this embodiment, the computing resource management unit A12 maintains multiple types of tables. For example, a hash table of the available resource information of each node updated in real time, a record table of the node processing efficiency of each node updated in real time, a packet information summary table of the number of data packets currently processed by each node, etc. In other embodiments, the tables maintained by the computing resource management unit A12 may be a data packet information table corresponding to the task to be reconstructed and a node efficiency information table, etc. This embodiment is not limited and can be adjusted according to actual needs.
[0093] For example, the data packet information table contains each piece of information included in the data packet header and the number of data packets recorded in this table. There are multiple data packet information tables, and each data processing node corresponds to a data packet information table, which is used to analyze the processing efficiency of each node. The data packet information table in this embodiment may be a linked list. The characteristic of a linked list is that it is extremely efficient in deleting and adding nodes at the head and tail. It is suitable for the situation where the table data changes in real time.
[0094] The node efficiency information table includes but is not limited to the following information: node name, broadcast number of the node, data processing time of the node in the past period of time, size of the data segment processed by the node in the past period of time, highest data processing speed of the node, lowest data processing speed of the node. The "past period of time" here refers to a period of time calculated forward from the current time, which can be set according to the actual situation and is generally 60 seconds.
[0095] The computing resource management unit A12 can calculate the processing efficiency related information of the corresponding node in the corresponding time period based on the information in the internally maintained table and update this information to the node efficiency information table in real time.
[0096] In a possible implementation manner, the above node processing efficiency can also be calculated and updated in real time by the computing resource management unit according to the following formulas (1) and (2).
[0097] In this embodiment, the node processing efficiency of each node includes: the average processing time and average processing speed of the node for a data packet;
[0098] Formula (1) is the average processing time,
[0099]
[0100] Formula (2) is the average processing speed;
[0101]
[0102] A is a node representation, n is the total number of data packets to be processed, and T k is the processing time of the k-th data packet; D k is the number of bytes of the k-th data packet.
[0103] The computing resource dynamic utilization unit A13 is configured to determine, based on a resource allocation policy, available resource information of the current node, and available resource information of other nodes within the broadcast domain, one or more nodes capable of processing the local reconstruction task;
[0104] The data splitting and merging unit A14 is configured to split the local reconstruction task according to the selected node, the processing efficiency of the node, and the available resource information of the node, and distribute it to the selected nodes for processing.
[0105] Embodiment III
[0106] For a better understanding of the functions of the units in the distributed computing processing module A10 in Embodiment II, in combination with Figure 3 and Figure 4 each unit will be described in detail.
[0107] The network communication unit A11 is responsible for communication between nodes. Since all nodes are in the cloud, a virtual subnet can be formed between nodes. When the virtual subnet is working properly, SOCKET communication can be used. When the virtual subnet fails, it can be automatically switched to HTTPSOCKET, that is, the network communication unit A11 supports two communication methods based on HTTPSOCKET and SOCKET. The communication method in this embodiment can ensure the high availability and scalability of the network communication unit.
[0108] The computing resource management unit A12 is configured to manage the available computing resource information (i.e., available resource information) of all computing nodes in the distributed reconstruction system.
[0109] In practical applications, the available computing resource information includes one or more of the following: node identifier, node broadcast number, node IP, node service port number, CPU information of the node, response speed of the node, supported algorithm version, etc.
[0110] The CPU information of the node includes one or more of the following: total number of CPU cores, currently available cores; memory information of the node, etc.;
[0111] The memory information of the node includes one or more of the following: total memory size, currently free memory size, etc.
[0112] The computing resource management unit A12 maintains a table / data structure / hash table containing the above-mentioned available computing resource information for each node for use by the distributed computing processing module A10.
[0113] Generally, the computing resource management unit A12 dynamically maintains the above table, for example, refreshing it at fixed time intervals (usually less than 1 second) to ensure that the real-time status of each node is maintained. In specific processing, the above table can be implemented using a hash table structure to improve the speed of element access and modification.
[0114] The data splitting and merging unit A14 is responsible for splitting the data belonging to the local reconstruction task and then distributing it to the node that calls and processes the split data for that node to process. After splitting the data, the data splitting and merging unit A14 forms a data packet from the split data. The structure of the data packet is divided into two parts: the data packet header (PackageHead) and the data (PackageData). The data packet header contains, but is not limited to, the following information: data size (in bytes, how many bytes is the data length), data index (the position of the data in the entire post-reconstructed original data), sending node, receiving node, data sending time (sending node), data recycling time (sending node), data receiving time (receiving node), data sending time (receiving node), data processing time consumption (empty when sending, filled by the node that receives and processes the data). After the data packet header, there is the data, as Figure 4 shown.
[0115] In this embodiment, the distributed computing processing module A10 will appropriately fine-tune the amount of data and the data packet size allocated to the node according to the processing efficiency of different nodes to improve the overall processing efficiency.
[0116] Embodiment 4
[0117] For the distributed reconstruction system described in the above embodiments, the following describes the usage process of the distributed reconstruction system in detail in conjunction with steps 01 to 07.
[0118] Step 01: Each node in the distributed reconstruction system sends the computing resource information and / or the data to be reconstructed of this node to the broadcast domain in the form of a broadcast at a specified frequency.
[0119] At the same time, each node is also used to receive broadcast packets with the same information sent by other nodes in the distributed reconstruction system.
[0120] Each node reads the information in the broadcast packet and updates the corresponding information table (such as the above hash table) maintained locally.
[0121] Step 02: If there is information indicating that the distributed computing processing module A10 within a node is called (i.e., when a node in the distributed reconstruction system receives its local reconstruction task), obtain the computing resources that the distributed computing processing module A10 can utilize (i.e., available resource information). Determine the computing resources that can be utilized on the local node currently, as well as the computing resources that can be utilized on other nodes within the broadcast domain. For example, the distributed computing processing module A10 of the first node that receives the local reconstruction task determines the computing resources that can be utilized on the local node, as well as the computing resources that can be utilized on other nodes within the broadcast domain, according to the pre-maintained hash table, and
[0122] the distributed computing processing module A10 of this first node determines the actually used local computing resources and the computing resources within the broadcast domain based on the local reconstruction concurrency limit and the reconstruction concurrency limit within the broadcast domain. The broadcast domain here is the distributed reconstruction system.
[0123] Step 03: The distributed computing processing module A10 of this first node creates a reconstruction processing process that contains multiple threads.
[0124] In this embodiment, when creating the process, the local reconstruction task and the information of the available computing resources are given to the data splitting and merging unit A14, and the data size allocated to each node is determined by the data splitting and merging unit A14;
[0125] Then, the split data is given to each thread for processing. Some of these threads run on the local node, and some run on other nodes within the broadcast domain. In this embodiment, based on the broadcast method, the distributed computing processing module A10 of the first node can obtain the reconstruction progress information of each node in real time.
[0126] For example, at the moment when the local node starts a new local reconstruction task and is providing computing resource services to other nodes and this computing is not yet completed, only the idle computing resources of the local node are temporarily utilized.
[0127] During the reconstruction process, if other nodes occupy the computing resources of the local node, after the remote thread ends, the computing resources occupied by this thread are recycled for local reconstruction use. Other nodes learn through the broadcast packet that the computing resources that the local node can provide gradually decrease.
[0128] During the reconstruction process, if it is learned through the broadcast packet of other nodes that other nodes have started local reconstruction tasks, gradually reduce the number of threads allocated to that node.
[0129] After other nodes start their local reconstruction tasks, they will gradually recycle the computing resources for their local use, so the computing resources provided for other nodes to share will gradually decrease. Therefore, this node will automatically reduce the number of threads allocated to that node based on this information.
[0130] The above logic is planned to be executed within a time slice. After the end of this time slice, the distributed computing processing module A10 will plan the allocation of the next time slice for the reconstruction thread based on the computing resources of the local node and the nodes in the broadcast domain at that time. The length of the time slice can be set according to specific circumstances, usually between 500 milliseconds and 5000 milliseconds.
[0131] The above embodiments have different focuses, and there are no contradictions in their mutual integration. They can be selected and set according to needs in practical applications.
[0132] It should be noted that in the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the claims listing several means, several of these means can be embodied by one and the same item of hardware. The use of the terms first, second, third, etc. is for convenience of expression only and does not denote any order. These terms can be construed as part of the element name.
[0133] In addition, it should be noted that in the description of this specification, the description of terms such as "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0134] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications after learning the basic creative concept. Therefore, the claims should be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0135] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention should also include these modifications and variations.
Claims
1. A distributed reconstruction resource management method based on cloud devices, characterized in that, The described distributed reconstruction resource management method is applied to a distributed reconstruction system for medical image reconstruction in a medical imaging device. Multiple nodes implementing a data reconstruction program are configured in the distributed reconstruction system. All nodes are located in a broadcast domain of a distributed reconstruction system and periodically broadcast available resource information in a broadcast manner. The method includes: S10. The first node periodically obtains the available resource information within the first node and broadcasts it to other nodes in the broadcast domain, and receives the available resource information broadcast by each node in the broadcast domain; the domain composed of all nodes is decentralized; S20. The first node determines whether it has received a local reconstruction task sent by the imaging device connected to the first node; each node in the broadcast domain corresponds to one or more imaging devices installed locally in the hospital, and is used to receive the local reconstruction tasks uploaded by the imaging devices for implementing reconstruction; the node is a cloud workstation; S30. If the first node receives a local reconstruction task, it performs data segmentation on the local reconstruction task and creates threads for the data after segmentation according to the resource allocation strategy, the available resource information of the current first node, and the available resource information of other nodes in the broadcast domain, so as to complete the local reconstruction task distributively; The resource allocation strategy is used to determine one or more nodes capable of processing the local reconstruction task according to the processing information of the local reconstruction task, and segment the local reconstruction task according to the selected nodes, the processing efficiency of the nodes, and the available resource information of the nodes; Each node maintains a data structure of the available resource information of each node in the broadcast domain and a node efficiency information table of each node; Among them, the information in the node efficiency information table of each node stored in the first node is the processing efficiency of each node calculated by the first node based on the data packets processed by each node within a specified duration; The node processing efficiency of each node includes: the average processing time and average processing speed of the node for a data packet; Formula (1) is the average processing time, Formula (2) is the average processing speed; A is a node representation, n is the total number of data packets to be processed, and T k is the processing time of the k-th data packet; D k is the number of bytes of the k-th data packet.
2. The method according to claim 1, wherein The method further includes: After all the nodes processing the local reconstruction task have completed processing the data allocated to them, update the available resource information inside the node and broadcast it; Among them, when the first node determines one or more nodes capable of processing the local reconstruction task, the priority of the first node is the highest priority.
3. The method according to claim 1, wherein SOCKET communication mode is used for communication between any two nodes; The available resource information includes one or more of the following: Total number of CPU cores, current available cores; memory information of the node, free memory information, response speed, supported algorithms.
4. A server, characterized in that, The server belongs to a node of a distributed reconstruction system for medical image reconstruction in a medical imaging device. The server is located in a broadcast domain and broadcasts the available resource information within the current node to other nodes in the distributed reconstruction system in a broadcast manner, and receives the available resource information broadcast by each node in the broadcast domain. The server includes: A distributed computing processing module is used to communicate with other nodes in the distributed reconstruction system, receive local reconstruction tasks transmitted by imaging devices connected to the server, and perform data segmentation on the local reconstruction tasks and create threads and / or processes for the segmented data based on a resource allocation strategy, the available resource information of the current node, and the available resource information of other nodes in the broadcast domain; A reconstruction module is used to process the segmented local reconstruction tasks according to the threads and / or processes of at least one node determined by the distributed computing processing module through the distributed service interface in the distributed reconstruction system to complete the reconstruction; The resource allocation strategy is used to determine more than one node capable of processing the local reconstruction task according to the processing information of the local reconstruction task, and segment the local reconstruction task according to the selected nodes, the processing efficiency of these nodes, and the available resource information of these nodes; Each node in the broadcast domain corresponds to more than one imaging device installed locally in the hospital and is used to receive local reconstruction tasks uploaded by the imaging devices for realizing reconstruction; the nodes are cloud workstations, and the domain composed of all nodes is decentralized; Each node maintains a data structure of the available resource information of each node in the broadcast domain and a node efficiency information table of each node; Among them, the information in the node efficiency information table of each node stored in the first node is the processing efficiency of each node calculated by the first node based on the data packets processed by each node within a specified time period; The node processing efficiency of each node includes: the average processing time and average processing speed of the node for a data packet; Formula (1) is the average processing time, Formula (2) is the average processing speed; A is a node representation, n is the total amount of data packets to be processed, and T k is the processing time of the k-th data packet; D k is the number of bytes of the k-th data packet.
5. The server according to claim 4, wherein The distributed computing processing module includes: A network communication unit for communicating with other nodes in the distributed reconstruction system; A computing resource management unit for periodically obtaining the available resource information of other nodes in the distributed reconstruction system and maintaining a data structure of the available resource information of each node in the broadcast domain and a node efficiency information table of each node stored inside the server; A computing resource dynamic utilization unit for determining more than one node capable of processing the local reconstruction task based on a resource allocation strategy, the available resource information of the current node, and the available resource information of other nodes in the broadcast domain; A data segmentation and merging unit for segmenting the local reconstruction task according to all the selected nodes, the processing efficiency and available resource information of each node among all the selected nodes, and distributing it to the selected nodes for reconstruction processing.
6. The server according to claim 5, wherein The network communication unit is used to communicate by using the SOCKET communication method; or communicate by using the HTTPSOCKET communication method; The computing resource management unit is used to send the available resource information and reconstruction task information of the local node in the form of a broadcast to the broadcast domain; and receive broadcast packets sent by other nodes in the broadcast domain; The reconstruction task information includes the information of the local reconstruction task or the information of the reconstruction task distributed by other nodes; A computing resource dynamic utilization unit, which is configured to, after receiving a local reconstruction task and determining a reconstruction algorithm, obtain the available resource information that can be used by the local node; and obtain the available computing resource information of other nodes in the broadcast domain, and determine the available resource information of the local server and the available resource information of other nodes in the broadcast domain according to the local reconstruction concurrency limit and the reconstruction concurrency limit within the broadcast domain, select all nodes for processing the local reconstruction task, and create multiple threads and / or processes for processing the local reconstruction task; A data segmentation and merging unit, which is configured to receive the threads and / or processes sent by the computing resource dynamic utilization unit, the available resource information of all the selected nodes, and the processing efficiency of each of the selected nodes, segment the local reconstruction task and determine the data size to be allocated to each of the selected nodes for distribution; Part of the threads and / or processes for processing the segmented data runs on the local node, and part runs on other nodes in the broadcast domain.
7. The server according to claim 6, characterized in that, It further includes: The computing resource dynamic utilization unit creates the maximum number of threads for processing the local reconstruction task based on the available resource information of all the selected nodes.
8. A distributed reconstruction system based on cloud devices, characterized in that It includes multiple servers as described in any one of claims 4 to 7. Each server forms a decentralized broadcast domain as a node, and each server corresponds to one or more imaging devices located in a hospital.
Citation Information
Patent Citations
Task processing method and device
CN110728363A
Distributed computing task processing method and equipment, and electronic equipment
CN111381948A