A medical image processing method, device, system and storage medium
By performing three-dimensional reconstruction and gridding of medical image data, digital gene sequences are generated, solving the problems of tracing the source and accurately locating image data, and achieving efficient image matching and secure analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ZHONGWEI INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2026-05-11
- Publication Date
- 2026-07-31
AI Technical Summary
There are difficulties in tracing the source of existing medical imaging data, and microscopic lesions are difficult to accurately locate and track throughout their entire life cycle.
By performing three-dimensional reconstruction and adaptive discrete meshing on two-dimensional slice medical digital image sequences, a digital gene sequence containing three-dimensional spatial data, hierarchical data, modal data, and data features is generated. A global three-dimensional coordinate system is established, and the location and characteristics of the data block are uniquely represented by the digital gene sequence.
It achieves precise positioning and full lifecycle tracking of image data, solves the problem of limited image matching accuracy across levels, times, and devices, improves data security and compliance, reduces AI training costs, and achieves privacy protection for data analysis.
Smart Images

Figure CN122492932A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital imaging technology, specifically to a method for processing medical images. The invention also relates to a medical image processing apparatus, system, and storage medium. Background Technology
[0002] Medical imaging refers to the acquisition of image data of the internal structure and tissue state of the human body through physical principles such as X-rays, magnetic resonance imaging, and ultrasound. It includes modalities such as computed tomography (CT), magnetic resonance imaging (MRI), X-ray radiography (XR), and ultrasound imaging (US). Medical imaging can present human anatomical structures and pathological changes in a non-invasive manner, and is a core basis for clinical diagnosis, treatment planning, and efficacy evaluation. Statistics show that over 70% of clinical decisions rely on medical imaging information, giving it an irreplaceable and fundamental position in the modern medical system.
[0003] With the development of digital technology, medical image processing has evolved from the traditional film-based model to a digital model. Currently, the mainstream medical image processing technologies mainly include:
[0004] DICOM (Digital Imaging and Communications in Medicine) is an internationally recognized standard for medical imaging data, defining the file format and communication protocol for images. Each DICOM file consists of two parts: pixel data (recording the grayscale or CT value of each pixel) and metadata tags (recording patient ID, examination parameters, equipment information, etc.). The DICOM standard ensures that different imaging devices can output data in a uniform format.
[0005] PACS (Picture Archiving and Communication System) is a comprehensive information system used for the acquisition, storage, transmission, display, and management of medical images. PACS connects various imaging devices through the DICOM protocol, enabling centralized image storage, hospital-wide access, and remote consultation. It also integrates with the Hospital Information System (HIS) and Radiology Information System (RIS) to support the digital management of imaging diagnostic workflows.
[0006] The widespread application of these technologies has essentially transformed medical images from "physical films" to "digital assets," significantly improving the efficiency of image storage, transmission, and retrieval. However, these technologies still store images in a pixel matrix format, which is a low-level visual signal, and significant obstacles remain in the interpretation and application of medical images. Summary of the Invention
[0007] This invention provides a medical image processing method to solve the technical bottleneck of difficulty in tracing the source of existing medical image data processing and management, which makes it difficult to accurately locate microscopic lesions and track them throughout their entire life cycle.
[0008] This invention provides a medical digital image processing method, comprising:
[0009] Acquire two-dimensional slice medical digital image sequences for the same medical examination;
[0010] The two-dimensional slice medical digital image sequence is reconstructed in three dimensions to obtain the three-dimensional spatial data, hierarchical data and modal data of the two-dimensional slice medical digital image sequence in three-dimensional space;
[0011] Based on the three-dimensional spatial data and the hierarchical data, the two-dimensional segmented medical digital images in the two-dimensional slice medical digital image sequence are adaptively divided into discrete grids to obtain a set of gridded medical digital image data blocks; the data blocks
[0012] The collection includes multiple data blocks that have their own 3D spatial data, hierarchical data, modal data, and data characteristics in 3D space;
[0013] Based on the three-dimensional spatial data, hierarchical data, modal data, and data features of the data block, a digital gene sequence is generated to uniquely characterize the data features and location of each data block in the gridded medical digital image data block set.
[0014] In some embodiments, the step of performing three-dimensional reconstruction of the two-dimensional slice medical digital image sequence to obtain the three-dimensional spatial data, hierarchical data, and modal data of the two-dimensional slice medical digital image sequence in three-dimensional space includes:
[0015] For each two-dimensional slice medical digital image in the two-dimensional slice medical digital image sequence, a unified spatial reference origin and coordinate axis are defined;
[0016] A global three-dimensional coordinate system is established based on the spatial reference origin and the coordinate axes.
[0017] Based on the global three-dimensional coordinate system, all two-dimensional slice medical images are aligned to complete the three-dimensional reconstruction;
[0018] Based on the results of the three-dimensional reconstruction, the spatial mapping relationship of the two-dimensional slice medical digital image sequence is analyzed to obtain the three-dimensional spatial data, hierarchical data and modal data of each two-dimensional slice medical digital image in the three-dimensional space.
[0019] In some embodiments, the step of adaptively discretely dividing the two-dimensional slice medical digital images in the two-dimensional slice medical digital image sequence into a gridded medical digital image data block set based on the three-dimensional spatial data and the hierarchical data includes:
[0020] Based on the three-dimensional spatial data, determine the three-dimensional spatial boundary of the two-dimensional slice medical digital image;
[0021] The two-dimensional slice medical digital image sequence is scanned layer by layer;
[0022] Based on the adaptive segmentation size, the two-dimensional slice medical digital image is segmented stepwise to obtain data blocks divided on the two-dimensional slice medical digital image;
[0023] Extract the data features of the data block, and establish a mapping relationship between the three-dimensional spatial data to which the data block belongs, the hierarchical data to which the data block belongs, and the data features of the data block and the data block;
[0024] The set of data blocks with the aforementioned mapping relationship is determined as the set of gridded medical digital image data blocks for the two-dimensional slice medical digital image sequence.
[0025] In some embodiments, generating a digital gene sequence to uniquely characterize the data features and location of each data block in the gridded medical digital image data block set, based on the data block's three-dimensional spatial data, hierarchical data, modal data, and data features, includes:
[0026] Extract the three-dimensional spatial data, hierarchical data, modal data, and data features of the data block;
[0027] A string is constructed by concatenating at least two types of data selected arbitrarily from the three-dimensional spatial data, the hierarchical data, the modal data, and the data features.
[0028] The digital digest generated based on the string is determined as the check code;
[0029] Based on the string, as well as the data and checksums constructed outside the string, a digital gene sequence is generated to uniquely characterize the data features and location of each data block in the gridded medical digital image data block set.
[0030] In some embodiments, it also includes:
[0031] The digital gene sequence is stored in a database.
[0032] This application also provides a medical digital image processing device, comprising:
[0033] The acquisition unit is used to acquire a sequence of two-dimensional slice medical digital images for the same medical examination.
[0034] The reconstruction unit is used to perform three-dimensional reconstruction on the two-dimensional slice medical digital image sequence to obtain the three-dimensional spatial data, hierarchical data and modal data of the two-dimensional slice medical digital image sequence in three-dimensional space.
[0035] A partitioning unit is used to adaptively perform discrete grid partitioning on the two-dimensional segmented medical digital images in the two-dimensional slice medical digital image sequence based on the three-dimensional spatial data and the hierarchical data, to obtain a set of gridded medical digital image data blocks; the set of data blocks includes multiple data blocks that have their own three-dimensional spatial data, hierarchical data, modal data and data features in three-dimensional space.
[0036] The generation unit is used to generate a digital gene sequence that uniquely represents the data characteristics and location of each data block in the gridded medical digital image data block set, based on the three-dimensional spatial data, hierarchical data, modal data, and data characteristics of the data block.
[0037] This application also provides a medical digital image processing system, including: a data acquisition module, a data processing module, and a data storage module;
[0038] The data acquisition module is used to acquire two-dimensional slice medical digital image sequences for the same medical examination;
[0039] The data processing module includes a 3D reconstruction submodule, a partitioning submodule, and a generation submodule. The 3D reconstruction submodule is used to perform 3D reconstruction on the 2D slice medical digital image sequence to obtain the 3D spatial data, hierarchical data, and modal data of the 2D slice medical digital image sequence in 3D space. The partitioning submodule is used to adaptively partition the 2D slice medical digital images in the 2D slice medical digital image sequence into a discrete grid based on the 3D spatial data and the hierarchical data to obtain a set of gridded medical digital image data blocks. The data block set includes multiple data blocks with their respective 3D spatial data, hierarchical data, modal data, and data features in 3D space. The generation submodule is used to generate a digital gene sequence that uniquely represents the data features and location of each data block in the gridded medical digital image data block set, based on the data block's respective 3D spatial data, hierarchical data, modal data, and data features.
[0040] The data storage module is used to store the digital gene sequence and the data block corresponding to the digital gene sequence, and to establish a mapping relationship between the digital gene sequence and the data block.
[0041] In some embodiments, the data storage module includes: a basic information storage partition, a digital gene storage partition, an image data storage partition, and a composite index storage partition;
[0042] The basic information storage partition is used to store attribute data related to medical examinations in the two-dimensional slice medical digital image sequence;
[0043] The digital gene storage partition is used to store the digital gene sequence; the identifier in the digital gene sequence is associated with the patient identifier in the medical examination-related attribute data stored in the basic information storage partition;
[0044] The image data storage partition is used to store the data features of the data block, and to establish a mapping relationship between the hash value of the data block and the digital gene sequence stored in the digital gene storage partition;
[0045] The composite index storage partition is used to store a multidimensional index constructed based on pathological features and anatomical locations, and the digital gene storage partition is retrieved through the multidimensional index.
[0046] In some embodiments, it further includes: an application service module, configured to respond to a query request, locate the corresponding digital gene sequence based on the digital gene identifier carried in the query request, and provide one or more services such as data location, recombinant imaging, and report generation based on the digital gene sequence.
[0047] This application also provides a computer storage medium including a computer program that, when run on an electronic device, causes the electronic device to perform the relevant steps of the medical digital image processing method described above.
[0048] Compared with the prior art, the present invention has the following advantages:
[0049] This invention first decomposes the traditional file-level whole image into a set of microscopic gridded medical digital image data blocks by performing three-dimensional reconstruction and adaptive discrete mesh partitioning on the two-dimensional slice medical digital image sequence. This processing breaks the storage limitation of treating the image as a black box in the prior art, transforming the image data into a structured asset composed of multiple data blocks carrying independent attributes (three-dimensional spatial data, hierarchical data, modal data, and data features). This achieves a fine leap in data granularity from macro to micro, supporting the independent extraction and analysis of local lesion areas, and meeting the high standard requirements of precision medicine for the analysis of microscopic pathological features. Secondly, a unified global three-dimensional coordinate system is established through three-dimensional reconstruction, and the three-dimensional spatial data and hierarchical data of the data blocks in three-dimensional space are obtained. Compared with the prior art that only relies on slice number for logical positioning, this invention assigns each data block precise physical spatial coordinates (A, B, C axes) and hierarchical index. This eliminates the positional deviation caused by different devices or scanning parameters, enabling heterogeneous image data to be aligned under the same spatial reference. This effectively solves the problem of limited image matching accuracy across levels, times, and devices, providing a precise spatial positioning foundation for constructing a full-lifecycle lesion evolution atlas. Next, the generated digital gene sequence contains the three-dimensional spatial data, hierarchical data, modal data, and data features of the data block, and has a unique representation function. This digital gene sequence, as a unique digital identity card, not only records the location and content characteristics of the data block but also establishes a bidirectional mapping channel between the diagnostic conclusion and the original data block. By parsing this sequence, specific spatial coordinates can be instantly located; by verifying its feature information, the integrity of the data can be verified. This forms a closed-loop verification path where diagnosis is the result and the result is the source, compensating for the lack of evidence chain protection in traditional processes and greatly improving the security and compliance of medical data. Finally, the digital gene sequence embeds image feature data (data features) and spatial location information, equivalent to an automated preliminary annotation of the original image, making the gridded data block set a high-quality structured training corpus. Simultaneously, the data block-based retrieval mechanism enables on-demand loading, avoiding the transmission of the entire data. This significantly reduces the costs of data cleaning and manual annotation required for training large AI models. On the other hand, it enables data analysis by extracting only features and grid content, thus achieving data desensitization and solving the privacy protection bottleneck in the development of medical AI. Attached Figure Description
[0050] Figure 1 This is a flowchart of a medical digital image processing method provided by the present invention.
[0051] Figure 2 This is a flowchart of an embodiment of the three-dimensional reconstruction method provided by the present invention.
[0052] Figure 3 This is a schematic diagram of an embodiment of adaptive discrete mesh partitioning in the processing method provided by the present invention.
[0053] Figure 4 This is a flowchart of an embodiment of adaptive discrete mesh partitioning in the processing method provided by the present invention.
[0054] Figure 5 This is a flowchart of an embodiment of the processing method for generating digital gene sequences provided by the present invention.
[0055] Figure 6-1 This is a schematic diagram of one embodiment of the processing method provided by the present invention for generating digital gene sequences.
[0056] Figure 6-2 This is a schematic diagram of another embodiment of the processing method provided by the present invention for generating digital gene sequences.
[0057] Figure 7 This is a schematic diagram of the structure of a medical digital image processing device provided by the present invention.
[0058] Figure 8 This is a schematic diagram of the structure of a medical digital image processing system provided by the present invention.
[0059] Figure 9 This is a schematic diagram of the structure of an embodiment of an electronic device provided by the present invention. Detailed Implementation
[0060] Numerous specific details are set forth in the following description to provide a full understanding of the invention. However, the invention can be practiced in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0061] The terminology used in this invention is for the purpose of describing particular embodiments only and is not intended to limit the invention. The descriptive terms used in this invention and the appended claims, such as "a," "first," and "second," are not intended to limit quantity or sequence, but rather to distinguish information of the same type from one another.
[0062] As described above in the background section, the technical concept of this invention stems from the technical deficiencies in the processing and management of medical digital images within the medical industry. Specifically, medical imaging, as an auxiliary means of medical diagnosis, has become a crucial source of data. The Medical Digital Imaging and Communication Standard (DICOM), as a universal protocol in the global medical imaging field, has completely solved the interoperability problem between different imaging devices (such as CT, MRI, and ultrasound) and workstations and printers. The DICOM standard not only defines the file storage format for medical images, encapsulating basic patient information, examination parameters, and pixel data within a standard file structure, but also standardizes network communication protocols, enabling seamless transmission of image data between heterogeneous systems. The widespread adoption of this standard has laid the foundation for modern digital hospitals.
[0063] Supported by the DICOM standard, image archiving and communication systems have emerged and rapidly become widespread. PACS systems, through efficient network transmission and high-capacity storage technology, enable centralized storage, archiving management, and clinical distribution of massive amounts of medical image data, completely eliminating the traditional film-based management model. The application of PACS has greatly improved the efficiency of image data flow, enabling radiologists to quickly access historical images via terminals, providing a crucial data infrastructure for clinical diagnosis.
[0064] With the rapid development of precision medicine and AI-assisted diagnostic technologies, the clinical demand for image data has shifted from basic "storage and retrieval" to "in-depth mining and refined analysis." However, the existing technical architecture, which uses DICOM files as the atomic unit and PACS systems as the management core, is gradually revealing deep-seated technical bottlenecks when dealing with sophisticated medical scenarios.
[0065] 1) Lack of identification system and weak traceability mechanism
[0066] In existing technologies, DICOM files only contain file-level unique identifiers (UIDs), lacking a standardized identification mechanism for microscopic regions within images (such as specific anatomical structures or lesion areas). Data units within images are "anonymous," unable to carry unique identity information like biological genes. This leads to a often "one-way" processing flow for image data, lacking effective data lineage records. Diagnostic conclusions are difficult to accurately trace back to specific original image areas, making it impossible to construct a closed-loop verification path from macroscopic diagnosis to microscopic features, severely limiting the interpretability and credibility of the diagnostic and treatment process.
[0067] (2) Coarse management granularity, lack of micro-management
[0068] Existing PACS systems typically use a single two-dimensional slice image (i.e., a DICOM file) as the smallest unit for storage and retrieval. While this "coarse-grained" management model is efficient and convenient at the file transfer level, it severs the logical connections within the image at the data content level. In the context of precision medicine, doctors often need to analyze tiny lesions (such as those a few millimeters in diameter).
[0069] The existing technology cannot independently extract, mark features, and manage the microscopic regions inside the images, making it difficult to meet the need for quantitative analysis of minute pathological changes.
[0070] (3) Spatial location is discrete, making cross-sequence association difficult.
[0071] Image data from different scan sequences and at different levels are spatially discrete and lack inherent spatial topological relationships. This limits the accuracy of image matching across levels and sequences.
[0072] (4) Data silo effect, feature correlation break
[0073] Image data from different scanning time points or different examination sequences are often stored as independent files, lacking an inherent correlation based on image features, forming "data silos." This "data silo" phenomenon makes it difficult for the system to automatically identify the evolution trajectory of the same lesion at different times, and it is impossible to construct a continuous and dynamic pathological evolution feature map, resulting in the inability to fully explore and utilize the value of massive historical data.
[0074] In view of the deficiencies of the prior art, the present invention provides a medical image processing method to overcome the aforementioned problems and deficiencies, such as... Figure 1 As shown, the method includes:
[0075] Step S101: Obtain a two-dimensional slice medical digital image sequence for the same medical examination;
[0076] Step S102: Perform three-dimensional reconstruction on the two-dimensional slice medical digital image sequence to obtain the three-dimensional spatial data, hierarchical data and modal data of the two-dimensional slice medical digital image sequence in three-dimensional space;
[0077] Step S103: Based on the three-dimensional spatial data and the hierarchical data, adaptive discrete grid division is performed on the two-dimensional segmented medical digital images in the two-dimensional slice medical digital image sequence to obtain a set of gridded medical digital image data blocks; the set of data blocks includes multiple data blocks that have their own three-dimensional spatial data, hierarchical data, modal data and data features in three-dimensional space;
[0078] Step S104: Based on the three-dimensional spatial data, hierarchical data, modal data, and data features of the data block, generate a digital gene sequence that uniquely represents the data features and location of each data block in the gridded medical digital image data block set.
[0079] The steps S101 to S104 described above are described in detail below.
[0080] Regarding step S101: Obtain a two-dimensional slice medical digital image sequence for the same medical examination.
[0081] The two-dimensional slice medical digital image sequence in this step can also be called a two-dimensional slice medical digital image sequence. It consists of N two-dimensional tomographic images arranged at equal or unequal intervals along the human body. Each image has a regular pixel grid (e.g., 512×512, depending on the device), and each pixel carries a numerical value reflecting the physical properties of the tissues in the human body (e.g., CT value, MRI signal intensity). Each image records its absolute position and orientation in three-dimensional space through the ImagePosition and ImageOrientation tags in the DICOM standard. DICOM (Digital Imaging and Communications in Medicine) is an international standard protocol in the field of medical imaging for the storage, transmission, printing, and sharing of medical images. Its core purpose is to solve the interoperability problem between imaging equipment (such as CT, MRI, ultrasound equipment, etc.) from different manufacturers and terminal devices such as workstations, servers, and printers, ensuring the seamless exchange and interoperability of medical image data. It mainly includes two parts: a file header and pixel data. The file header contains metadata organized in the form of "data elements". These metadata records detailed patient information (such as name and patient ID), examination information (such as examination time and examination site), imaging parameters (such as slice thickness, window width and level, and equipment model), and image pixel description information. Each data element consists of a label, a value description, and a value length, following strict encoding rules. The pixel data includes actual...
[0082] The medical image information consists of image grayscale or color data represented by a numerical matrix. This data is encoded according to the imaging parameters in the file header and can be reconstructed into a visualized medical image using workstations and other terminal devices.
[0083] In other words, a two-dimensional slice medical digital image sequence refers to a set of multiple tomographic images (such as 200 chest CT slices) arranged in spatial order, generated by equipment such as CT and MRI. Each image is a two-dimensional digital image with spatial coordinates. It can be input data for three-dimensional reconstruction or one of the underlying data sources of the technical architecture.
[0084] Step 101 acquires a two-dimensional slice medical digital image sequence for the same medical examination, which can also be understood as a medical image DICOM file sequence for the same medical examination. In this embodiment, the same medical examination can refer to an examination of the same part of the same patient. The examination time can be unlimited, or the time range for acquiring the two-dimensional slice medical digital image sequence can be set according to the needs of the specific application scenario.
[0085] It is understood that this embodiment uses DICOM as an example of a two-dimensional slice medical digital image sequence for illustration, and the image sequence composed of two-dimensional slice medical digital images is applicable both theoretically and practically. In this embodiment, using DICOM as an example of a two-dimensional slice medical digital image sequence, step S101 may include:
[0086] Step S101-1: Obtain a set of medical image DICOM files from the same medical examination;
[0087] Step S101-2: Determine the DICOM file set as a two-dimensional slice medical digital image sequence.
[0088] As for how to obtain two-dimensional slice medical digital image sequences, they can be generated during medical examinations using medical equipment and obtained with authorization.
[0089] Regarding step S102: Perform three-dimensional reconstruction on the two-dimensional slice medical digital image sequence to obtain the three-dimensional spatial data, hierarchical data and modal data of the two-dimensional slice medical digital image sequence in three-dimensional space.
[0090] In step S102, the three-dimensional spatial data can be understood as the coordinates (x, y, z) of each slice image (two-dimensional slice medical digital image) in the three-dimensional middle of the two-dimensional slice medical digital image sequence.
[0091] Layer data can be understood as the layer in a three-dimensional medical digital image sequence where each slice image (two-dimensional slice medical digital image) is located. Because a two-dimensional medical digital image sequence is a stack of multiple slice images, each slice image is located at a different layer. For example, it corresponds to the slice index in a CT / MRI / PET image sequence (such as layer 1 to layer 100).
[0092] Coordinates (a, b, c): can be understood as the discretized coordinate values of the data block in the cross section, longitudinal section and depth direction.
[0093] Modality data can be understood as a type identifier for medical imaging equipment, such as CT, MRI, and PET. CT stands for Computed Tomography, which uses X-rays to penetrate tissues of different densities to create tomographic images, which are then reconstructed by a computer. It is mainly used for examining the brain, heart, chest and abdomen, and limbs, for example, to detect brain atrophy, developmental delays, brain tumors, cerebral vascular malformations, and bone and calcification problems. RI stands for Magnetic Resonance Imaging, which uses a magnetic field to change the rotational alignment of hydrogen atoms, causing the nuclei to release energy and electromagnetic signals. These signals are then analyzed by a computer to reconstruct images. Among CT, MRI, PET, PET-CT, and PET-MRI, it is the only radiation-free detection method. PET stands for Positron Emission Tomography, which involves injecting a positron-emitting drug such as glucose fluoroquinolones (FDG) into a vein, collecting signals, and then using a computer to reconstruct images of the distribution of positron isotopes in human tissues or organs. Its main applications are in the study of brain metabolism, as well as the detection of lung cancer, melanoma, colon cancer, lymphoma, esophageal cancer, head and neck cancer (including thyroid cancer), and breast cancer.
[0094] Based on the above step S101, it can be seen that a medical digital image sequence, such as an original DICOM slice, is composed of multiple independent two-dimensional images.
[0095] Success. Although each image contains pixel coordinates, the images are isolated from each other. The system can know that a certain image is the 1st image and another image is the 100th image, but it does not know the distance between the 1st and 100th images in physical space (or three-dimensional space), nor their relative positional relationship in three-dimensional space. Therefore, step S102 requires three-dimensional reconstruction of the two-dimensional slice medical digital image sequence to align all slice images. Therefore, as... Figure 2 As shown, the specific implementation process of step S102 includes:
[0096] Step S102-1: Define a unified three-dimensional spatial reference origin and coordinate axis for each two-dimensional slice medical digital image in the two-dimensional slice medical digital image sequence. In this embodiment, the spatial reference origin can be defined as the coordinate origin (0, 0, 0) of the two-dimensional slice medical digital image sequence in three-dimensional space. The coordinate axis can be defined as the x-axis, y-axis, and z-axis of the plane of the two-dimensional slice medical digital image in the sequence (e.g., the depth data of the current slice image in the overall slice data or the distance between the current slice image and the top image of the slice). The three-dimensional spatial structure can be obtained through the x-axis, y-axis, z-axis, and coordinate origin.
[0097] Step S102-2: Establish a global three-dimensional coordinate system based on the spatial reference origin and the coordinate axes;
[0098] Step S102-3: Based on the global three-dimensional coordinate system, align all two-dimensional slice medical images to complete the three-dimensional reconstruction.
[0099] In this embodiment, although the slice images (two-dimensional slice medical digital images) in the two-dimensional slice medical digital image sequence are for the same medical examination, they may come from different devices or scanning parameters. For example, a patient may have undergone multiple medical examinations targeting the lungs within one or more different time periods, but the source may be the same device with different scanning parameters, or different devices, etc. Therefore, by establishing a global three-dimensional coordinate system, the positional deviation caused by different devices or scanning parameters can be eliminated. Continuing with the above example, by parsing the spatial position information in the DICOM file header, the coordinates of each two-dimensional slice medical digital image in physical space are obtained. Based on the thinnest point, all two-dimensional slice medical digital images are mapped to the global three-dimensional coordinate system for spatial alignment, thereby completing the three-dimensional reconstruction.
[0100] Step S102-4: Based on the 3D reconstruction results, analyze the spatial mapping relationship of the 2D slice medical digital image sequence to obtain the 3D spatial data, hierarchical data, and modal data of each 2D slice medical digital image in the 3D slice medical digital image sequence. In this embodiment, based on the established global 3D coordinate system, determine the specific coordinate position (3D spatial data) of each 2D slice medical digital image in 3D space and its layer depth position (hierarchical data), and extract the imaging modal information (modal data) from the DICOM file header.
[0101] Regarding step S103: Based on the three-dimensional spatial data and the hierarchical data, the two-dimensional segmented medical digital images in the two-dimensional slice medical digital image sequence are adaptively divided into discrete grids to obtain a set of gridded medical digital image data blocks; the set of data blocks includes multiple data blocks that have their own three-dimensional spatial data, hierarchical data, modal data, and data features in three-dimensional space.
[0102] like Figure 3As shown, step S103 can employ adaptive discrete grid partitioning for two-dimensional segmented medical digital images. Grid partitioning refers to dividing the two-dimensional segmented medical digital image into small squares along the X and Y directions, aiming to decompose a whole into multiple units for easier fine-grained management. "Discrete" means that the grids are not continuous and are independent; each data block is an independent entity, and operations on one data block do not affect other data blocks. "Adaptive" means that the partitioning can be automatically adjusted according to the specific application scenario, and is not fixed. That is, the grid size can be automatically adjusted based on the image content. For example, for lung medical digital images, key areas (such as the lungs) can be partitioned using adaptive discrete grid partitioning.
[0103] For regions where the grid size is automatically reduced, the subdivisions become finer and the details clearer; for non-critical regions (such as background air), the grid size is automatically increased to reduce data volume and save storage. Therefore, in this embodiment, the adaptive discrete grid division refers to dynamically adjusting the size or distribution of grid cells based on the content characteristics of medical digital images (such as lesion size and texture complexity) or preset accuracy requirements, decomposing continuous image data into several spatially discrete and logically independent sets of data blocks.
[0104] like Figure 4 As shown in the embodiment of this application, the specific implementation process of step S103 includes:
[0105] Step S103-1: Determine the three-dimensional spatial boundary of the two-dimensional slice medical digital image based on the three-dimensional spatial data;
[0106] Step S103-2: Perform layer-by-layer scanning on the two-dimensional slice medical digital image sequence;
[0107] Step S103-3: According to the adaptive segmentation size, the two-dimensional slice medical digital image is segmented stepwise to obtain data blocks divided on the two-dimensional slice medical digital image;
[0108] Step S103-4: Extract the data features of the data block, and establish a mapping relationship between the three-dimensional spatial data to which the data block belongs, the hierarchical data to which it belongs, and the data features of the data block and the data block;
[0109] Step S103-5: Determine the set of data blocks with the mapping relationship as the set of gridded medical digital image data blocks of the two-dimensional slice medical digital image sequence.
[0110] The purpose of step S103-1 is to determine the range. Before segmentation, the effective area of the two-dimensional slice medical digital image needs to be known. The so-called three-dimensional spatial boundary refers to the coverage area of the two-dimensional slice medical digital image (or slice image) in the three-dimensional global coordinate system. For example: X-axis boundary: the coordinate range from the left pixel to the right pixel of the image (assuming it is from X=10mm to X=200mm). Y-axis boundary: the coordinate range from the upper pixel to the lower pixel of the image. Z-axis boundary: the layer depth position of the slice image. These positions can be determined by reading the pixel matrix size (Rows, Columns) and pixel spacing of the slice image to calculate the three-dimensional spatial range occupied by the slice image.
[0111] The function of step S103-2 is to perform traversal. The slice images are stacked layer by layer and need to be traversed and scanned in the order of their stacking. In this embodiment, each slice image can be retrieved and scanned sequentially according to the hierarchical order (such as L1, L2, L3Ln).
[0112] Step S103-3 is used to perform a segmentation operation. The adaptive segmentation size means that the segmentation of the slice image can be dynamically adjusted based on the complexity of the image content or a preset strategy. For example, areas with dense lesions can be segmented smaller, while background areas can be segmented larger. The so-called step-by-step segmentation can be understood as using a segmentation algorithm, for example, starting from the upper left corner of the segmentation window, moving step by step to the right and down according to a fixed or dynamic step size, cutting a block at each position. Of course, this is just an example and does not limit the segmentation method. Finally, the complete slice image is segmented into physically adjacent but logically independent blocks.
[0113] Steps S103-4 are used to assign attributes to the segmented blocks, making them data blocks. For example, feature vectors such as the grayscale mean and texture gradient of pixels within a block can be extracted as data features of the data block. A mapping relationship is established, for example: data block ↔ coordinates (A, B, C) ↔ feature value (ΔG) ↔ level (L). Through the mapping relationship, the corresponding data block can be found instantly by coordinates or by feature value.
[0114] The purpose of step S103-5 is to aggregate the segmented data blocks that already have mapping relationships, thereby obtaining a three-dimensional data block that is structured, indexable, and includes rich semantic information.
[0115] Regarding step S104: Based on the three-dimensional spatial data, hierarchical data, and modal data of the data block,
[0116] In addition to the data features, a digital gene sequence is generated to uniquely characterize the data features and location of each data block in the gridded medical digital image data block set.
[0117] The purpose of step S104 is to assign a unique, semantic, and tamper-proof "digital ID card" to each data block generated by the aforementioned gridding, wherein the data block may be in an unordered state.
[0118] This step is based on a technical logic combining "content addressing" and "multidimensional attribute encoding." Instead of relying on traditional serial numbers or random number generation for identifiers, it uses the data block's own "identity attributes" (e.g., location, type) and "content attributes" (appearance, feature values) as raw materials. Through a specific encoding algorithm, this multidimensional discrete information is compressed and mapped into a standardized string sequence. Specifically, this step processes at least four types of key information: the three-dimensional spatial data to which it belongs (definition gene), the hierarchical data to which it belongs (depth gene), the modal data to which it belongs (source gene), and data features (trait gene).
[0119] The three-dimensional spatial data refers to the coordinates of the corresponding data block in the global three-dimensional coordinate system, specifically the horizontal (A), vertical (B), and C-axis coordinates. Its function is to establish the spatial physical location of the data block within the human anatomical structure, that is, the precise location of the data block within the current slice image plane. This is the foundation for achieving forward rapid indexing, solving the problems of spatial positioning modules in existing technologies, and overcoming the shortcomings of existing technologies that rely solely on visual observation or pixel coordinates for precise cross-device positioning. Specifically, the horizontal and vertical axes determine the physical location of the data block within the slice plane; the C-axis determines the physical depth of the data block in three-dimensional space, for example, by parsing DICOM tags or calculating the interlayer spacing. The A, B, and C axes ensure seamless positioning in three-dimensional space, providing a depth benchmark for accurate lesion tracing across sequences and time points.
[0120] The data at the corresponding level refers to the slice layer data (e.g., L) where the corresponding data block is located. Its purpose is to enable rapid location of data within the sequence, facilitating quick retrieval of the target slice from massive data sequences.
[0121] The modal data refers to the type of image acquisition equipment (e.g., CT, MRI). Its function is to record the source attributes of the data, distinguish image features under different imaging principles, and avoid confusion between different modal data. The data features are quantified feature values extracted from within the data block, such as gradients and texture features. Their function is to characterize the image content features within the data block, providing key data for content-based image retrieval and lesion-assisted diagnosis, solving the problems of crude management and inability to identify microscopic content in existing technologies. In other words, the digital gene sequence generated in this step is not only a unique identifier at the technical level, but also a condensed data package carrying the location information and content features of the data block. This digital gene sequence transforms massive amounts of medical image data from unreadable pixel piles into computable, searchable, and traceable structured information assets, providing a solid data foundation for subsequent data loop verification and pathological evolution analysis.
[0122] like Figure 5 As shown, the specific implementation process of step S104 includes:
[0123] Step S104-1: Extract the three-dimensional spatial data, hierarchical data, modal data, and data features of the data block;
[0124] Step S104-2: Concatenate at least two types of data selected arbitrarily from the three-dimensional spatial data, the hierarchical data, the modal data, and the data features to construct a string;
[0125] Step S104-3: Determine the digital digest generated based on the string as the check code;
[0126] Step S104-4: Based on the string, and the data and check code constructed outside the string, generate a digital gene sequence to uniquely characterize the data features and location of each data block in the gridded medical digital image data block set.
[0127] The above process of generating digital gene sequences can be implemented in at least the following two ways:
[0128] Method 1:
[0129] like Figure 6-1As shown, step S104-1 can be understood as needing to collect data before encoding. This data comes from the three-dimensional spatial data, hierarchical data, modal data, and data features of the data block. Specifically, extracting the three-dimensional spatial data involves obtaining the specific coordinate values of the data block in the global coordinate system (e.g., A=10, B=20, C=5). This represents location information. Extracting the hierarchical data involves obtaining the slice layer number of the data block (e.g., L=1). This represents image hierarchical information. Extracting the modal data involves obtaining the image type (e.g., CT, MRI). This represents equipment source information. Extracting the data features involves obtaining the computational features of the data block (e.g., ΔG=120.5). This represents content information. The purpose is to transform specific image attributes in the physical world into numerical variables that can be processed by a computer.
[0130] Step S104-2 can be understood as integrating discrete numerical values into an information string with a unified format, i.e., serialization / encoding. Specifically, the data extracted in step S104-1 can be concatenated into a string (raw_code) according to preset format rules. For example, concatenating the extracted data according to the format "modality-level-horizontal coordinate-vertical coordinate-depth coordinate" yields: CT:L1:A10B20C5. This can be understood as constructing the plaintext part of the digital gene, making the string readable and allowing the basic attributes of the data block to be known.
[0131] To ensure security and the uniqueness of the string, step S104-3 generates a digital digest of the string, which serves as a checksum. The digital digest is a cryptographic algorithm (such as MD5, SHA-256, etc.) that maps a string of arbitrary length to a fixed-length random string. It possesses uniqueness and irreversibility. Uniqueness means that if the input string changes (e.g., the x-coordinate changes from 10 to 11), the generated digital digest will be completely different. Irreversibility means that the original data cannot be deduced from the digital digest, thus protecting privacy. Using the string CT:L50:A10B20C5 generated in step S104-2, a hash operation is performed on the string, and the first N characters (e.g., the first 8 characters) are extracted to obtain a short code, such as 8a2b3c4d. This step is to prevent the digital code from being tampered with and to fundamentally guarantee its uniqueness.
[0132] Step S104-4 is used to assemble / synthesize the string, data feature, and check code to generate a digital gene sequence. Specifically, the string from step S104-2 is concatenated with the data feature ΔG=120.5, and then combined with the check code from step S104-3 using a specific connecting symbol (such as a period). For example: CT:L1:A10B20C5:G120.CHK8a2b3c4d.
[0133] Method 2:
[0134] like Figure 6-2 As shown, step S104-1 can be understood as needing to collect data before encoding. This data comes from the three-dimensional spatial data, hierarchical data, modal data, and data features of the data block. Specifically, extracting the three-dimensional spatial data involves obtaining the specific coordinate values of the data block in the global coordinate system (e.g., A=10, B=20, C=5). This represents location information. Extracting the hierarchical data involves obtaining the slice layer number of the data block (e.g., L=1). This represents image hierarchical information. Extracting the modal data involves obtaining the image type (e.g., CT, MRI). This represents equipment source information. Extracting the data features involves obtaining the computational features of the data block (e.g., ΔG=120.5). This represents content information. The purpose is to transform specific image attributes in the physical world into numerical variables that can be processed by a computer.
[0135] Step S104-2 can be understood as integrating discrete numerical values into an information string with a unified format, i.e., serialization / encoding. Specifically, the data extracted in step S104-1 can be concatenated into a string (raw_code) according to preset format rules. For example, concatenating the extracted data according to the format "modality-level-horizontal coordinate-vertical coordinate-depth coordinate-feature value" yields: CT:L1:A10B20C5:G120. This can be understood as constructing the plaintext part of the digital gene, making the...
[0136] Strings are readable and can reveal the basic properties of data blocks.
[0137] To ensure security and the uniqueness of the string, step S104-3 generates a digital digest of the string, which serves as a checksum. The digital digest is a cryptographic algorithm (such as MD5, SHA-256, etc.) that maps a string of arbitrary length to a fixed-length random string. It possesses uniqueness and irreversibility. Uniqueness means that if the input string changes (e.g., the x-coordinate changes from 10 to 11), the generated digital digest will be completely different. Irreversibility means that the original data cannot be deduced from the digital digest, thus protecting privacy. Using the string CT:L50:A10B20C5:G120 generated in step S104-2, a hash operation is performed on the string, and the first N characters (e.g., the first 8 characters) are extracted to obtain a short code, such as 8a2b3c4d. This step is to prevent the digital code from being tampered with and to fundamentally guarantee its uniqueness.
[0138] Step S104-4 involves assembling / synthesizing the string and checksum to generate a digital gene sequence. Specifically, the string from step S104-2 and the checksum from step S104-3 are combined using a specific connecting symbol (such as a period). For example: CT:L1:A10B20C5:G120.CHK8a2b3c4d. Of course, the sequence representation can include various methods; the above is merely an example and not intended to limit the specific combination representation.
[0139] The difference between Method 1 and Method 2 lies in the different string compositions. This can be understood as referring to the hierarchical encoding or full-scale encoding during the digital gene sequence generation process. In the process of generating the digital gene sequence, the plaintext portion of this invention can be achieved by string concatenation using any combination or all of the data block's three-dimensional spatial data, hierarchical data, modal data, and data features. A checksum is then generated based on the concatenated string. The data outside the generated string (in cases where the string generation is not full-scale data), the string, and the checksum are combined according to a preset format to obtain the final digital gene sequence. One or more combination methods, such as concatenation and XOR, can be used during data combination. Other methods can also be used; this embodiment only uses concatenation as an example. In some application scenarios, such as when it is necessary to hide some information in the digital gene sequence (i.e., ciphertext), the plaintext portion and the ciphertext portion can be XORed to generate a digital gene sequence including both plaintext and ciphertext. For example, given CT:L1:A10B20C5 (plaintext) and G120 (the hidden part - ciphertext), we can use Hash(CT:L50:A10B20C5)XORG185 to generate CT:L50:A10B20C5.CHK8a2b3c4d. If the digital gene sequence does not contain the data that needs to be hidden, it can be generated by concatenation. For example, Hash(CT:L50:A10B20C5)+G185 generates CT:L1:A10B20C5:G120.CHK8a2b3c4d. Therefore, step S104 above can select string combinations based on different scenario requirements when generating digital gene sequences.
[0140] It should be noted that, Figure 6-1 and Figure 6-2 The Unicode string ATGCATGC represents biological DNA, where A, T, G, and C represent the four bases of DNA. In this invention, it is converted into a data-mapped encoding format. This represents the transformation of complex medical imaging data (DICOM) into a standardized, unique, and biologically similar gene-structure-like string sequence.
[0141] It is understood that the feature vector of the data block can be a high-dimensional vector, including feature vectors such as grayscale, texture, edge intensity, and difference from the neighborhood. The feature vector of the data block can be extracted according to the needs of the application scenario. If a high-dimensional vector is required in some application scenarios, the high-dimensional feature vector can be mapped and transformed into ΔG (which can be an integer), thereby facilitating concatenation and / or XOR operations. In this embodiment, ΔG can be understood as the aggregated value of the multi-dimensional feature vectors of the data block, converted into an integer for concatenation and / or XOR operations. Of course, the extraction of data features can be combined with the needs of specific application scenarios, such as combining lesion changes or anatomical structure features, to determine which data feature is more representative of the examination results. If it is necessary to understand the treatment results, edge intensity features can be extracted. Therefore, in this embodiment,
[0142] The generation of digital gene sequences is described using only the aggregated value of multidimensional feature vectors as an example, but it is not limited to this for data features.
[0143] The following describes the process of generating this digital gene sequence using specific examples.
[0144] Assume the following scenario: Patient Zhang San, examination site is the lungs, modal data is CT, target data block is a slice image located at layer 50, with spatial coordinates (A=128, B=45, C=5), containing a nodular lesion. First, the attribute values need to be read from the gridded data block:
[0145] 1. Modal data: Recognized as CT (Computed Tomography).
[0146] 2. Hierarchical data: This data block is located in the 50th slice of the image sequence and is recorded as L50.
[0147] 3. Three-dimensional spatial data: The grid coordinates of this data block in the slice plane are the 128th grid on the horizontal axis and the 45th grid on the vertical axis, with a depth of 5, recorded as A=128, B=45, C=5.
[0148] 4. Data characteristics: The pixel gradient changes within the data block were calculated, and its sharp edges were found (suspected nodules). The calculated feature value (ΔG) was 185.6.
[0149] The extraction result is:
[0150] Mode (M) = CT;
[0151] Level (L) = 50;
[0152] x-coordinate (A) = 128;
[0153] The vertical axis (B) is 45.
[0154] Depth coordinate (C) = 5;
[0155] Eigenvalue (ΔG) = 185.6.
[0156] Following the preset encoding rules, such as: modality: layer: x-coordinate, y-coordinate, depth: feature value; the extracted data is concatenated into the string: CT:L50:A128B45C5:G185.6 (example provided in method two above). This string represents the plaintext portion of the digital gene, describing the data block as a CT image, at layer 50, at location (A=128, B=45, C=5), and with feature value ΔG=185.6.
[0157] To ensure the uniqueness and tamper-proof nature of digital gene sequences, hash operations are required, such as the MD5 algorithm.
[0158] Input: CT:L50:A128B45C5:G185.6; Of course, the computer can convert the floating-point number G185.6 to an integer, i.e., CT:L50:A128B45C5:G185;
[0159] Algorithm processing: Execute MD5(CT:L50:A128B45C5:G185);
[0160] Calculation result: Generates a long hash value, such as e99a1803b0a9952b0c6c0a8a3c8c6c0a.
[0161] Extraction: To simplify storage, the system extracts the first 8 bits of the hash value as the final checksum, which is e99a1803. Combining the string and the checksum generates the digital gene sequence: CT:L50:A128B45C5:G185.CHKe99a1803. Based on the above, a unique digital gene sequence was generated for the data block containing nodules in Zhang San's lungs.
[0162] CT:L50:A128B45C5:G185.CHKe99a1803. When a doctor or system reads CT:L50:A128B45C5, they can determine that the CT image is on slice 50, with coordinates (128, 45, 5), enabling rapid localization and indexing. G185 records the lesion characteristics, and CHKe99a1803 ensures the data's authenticity. If the data is tampered with or the coordinates are entered incorrectly, the checksum will not match, thus achieving data traceability and anti-counterfeiting.
[0163] Furthermore, based on the three-dimensional spatial data (A=128, B=45, C=5) included in the digital gene sequence, it is possible to enough
[0164] Construct a pathological evolution map across time. The specific implementation method is as follows:
[0165] Suppose patient Zhang San underwent a follow-up examination of the same area three months later. Using the aforementioned 3D reconstruction to establish the same global coordinate system, a data block is located at the same spatial coordinates (A=128, B=45, C=5) in the lung. This new data block will then generate a new digital gene sequence, for example: CT:L80:A128B45C5:G210.CHK7f8e9d2a (Note: the slice level becomes L=80, and the feature value becomes ΔG=210). By comparing the digital gene sequences at the same coordinates (A=128, B=45, C=5) in the two examinations, the evolution of the lesion can be intuitively analyzed: although the slice numbers in the two scans (L=50 and L=80) may differ due to respiratory differences, the anchoring based on the global coordinates (A=128, B=45, C=5) ensures that the same anatomical location is being compared. By comparing the eigenvalue (G185 increasing to G210), the increased gradient at the nodule's margin was identified, suggesting potential lesion deterioration or increased density. Historical and current digital gene sequences at the same coordinate system were arranged along a timeline to automatically generate a digital pathological evolution map of the nodule. This map not only records changes in image morphology but also enables full-cycle traceability management through digital gene sequences, resolving the data silo problem caused by inconsistent slice layers in traditional techniques and providing precise quantitative data for subsequent processing.
[0166] In this embodiment, the digital gene sequence can be represented by the following data table, and is not limited to the above example or the example in the table below. This table is for illustrative purposes only.
[0167] Example table of data block encoding:
[0168] Level 1 10 20 5 CT 120 CT:L1:A10B20C5.G120 8a2b3c4d CT:L1:A10B20C5.G120.CHK8a2b3c4d 2nd floor 10 20 5 CT 120 CT:L2:A10B20C5.G120 9f1e2d3c CT:L2:A10B20C5.G120.CHK9f1e2d3c 2nd floor 11 20 5 CT 130 CT:L1:A11B20C5.G130 1b2c3d4e CT:L1:A11B20C5.G130.CHK1b2c3d4e 50th floor 50 50 10 CT 98 CT:L50:A50B50C10.G98 5e6f7a8b CT:L50:A50B50C10.G98.CHK5e6f7a8b
[0169] The above is just an example of digital gene sequence generation. Corresponding digital gene sequences can be generated for each data block's relevant attributes. The table above is for reference only.
[0170] Once the digital gene sequence is generated, it can be stored in a database for easy use by downstream applications and for traceability, thereby constructing a three-dimensional holographic digital index map. This transforms the originally disordered and cumbersome image files into ordered, lightweight, and computable data assets. The functions it can provide for downstream applications include:
[0171] Quick search and on-demand access:
[0172] Doctors only need to input the patient's relevant attribute information, such as name, medical record number, and consultation number, or a digital gene sequence, i.e., a digital gene ID (e.g., CT:L50:A128B45...), to quickly retrieve the corresponding data block (a few KB) from the underlying storage. This achieves pixel-level on-demand loading, significantly reducing network transmission bandwidth and terminal memory consumption. In this embodiment, there is a mapping relationship between the digital gene sequence and the patient's relevant attribute information, such as a mapping relationship with the name, medical record number, and consultation ID. Therefore, regardless of the type of patient attribute information or digital gene sequence used as query data, the ability to quickly retrieve relevant data from the database can be achieved.
[0173] Precise source tracing of lesions:
[0174] When AI or a doctor diagnoses a lesion, not only is the diagnosis result saved, but also the corresponding digital gene ID (digital gene sequence) is saved. In subsequent follow-up examinations, the coordinates of A, B, and C are parsed using the digital gene ID. Based on the coordinates, the system can automatically jump to the specific location in the original image and verify the data integrity (checksum). This establishes a reliable data link from the diagnostic conclusion to the original evidence, preventing data tampering.
[0175] Content-based intelligent screening:
[0176] The database stores digital gene sequences including feature values (such as G185 gradient values). If a doctor wants to find all sharp-edged nodules, they don't need to reanalyze all images; they can simply execute an SQL query in the database, for example: `SELECT * FROM Digital_Gene WHERE Gradient_Feature > 150`. This enables image-based screening within seconds, transforming historical data from inactive to valuable data, providing data support for scientific research and assisted diagnosis.
[0177] Analysis of cross-temporal evolution:
[0178] This feature is key to solving data silos. Even though CT scans taken at different times may have different files, as long as a unified global coordinate system is established, the spatial coordinates (A, B, C) of the same anatomical location are fixed. For example, a patient may have a lung CT scan a year ago and another a year later. Doctors who want to compare changes in lesions can automatically generate a pathological evolution atlas by comparing the digital gene sequences and their characteristic values (such as changes in nodule size and density) at the same coordinates in the two examinations. This breaks down the time barrier and enables full lifecycle management of lesions.
[0179] Therefore, in this embodiment, the digital gene sequence is stored in a database to construct a spatial index and feature index architecture for medical digital image data blocks, thereby enabling rapid spatial location retrieval, lesion source verification, content feature screening, and cross-temporal evolution analysis of image data.
[0180] The above describes a medical digital image processing method and its embodiments provided by the present invention. As can be seen from the above, the present invention, by introducing digital gene sequences, reconstructs the storage and management paradigm of medical image data, and includes advantages in at least the following three dimensions:
[0181] 1. Data processing paradigm shift: from "extensive overall approach" to "refined dynamic decomposition", i.e., positive data decomposition capability.
[0182] This invention breaks the limitation of traditional PACS systems that store images as a whole "black box" file, and achieves a fine leap in data granularity.
[0183] The processing of medical imaging data adopts adaptive discrete grid partitioning to have flexible network segmentation capabilities, supporting dynamic splitting according to clinical needs, such as organ-level, tissue-level, or even simulated cell-level granularity.
[0184] Through a "penetrating" data structure, image data is no longer just a single image, but becomes a structured data asset that can be computed, recombined, and deeply analyzed, providing ideal "atomic-level" data raw materials for AI-assisted data reading.
[0185] 2. Build a closed-loop traceability system across the entire supply chain to improve trust.
[0186] This invention solves the problem of the "broken chain" between diagnostic conclusions and original image data, and makes up for the lack of evidence chain protection in traditional processes.
[0187] By generating digital gene sequences and their embedded three-dimensional spatial coordinates (A, B, C) and check codes, a bidirectional channel is constructed: a forward channel and a reverse channel. Medical image data processing enables forward interpretation and reverse tracing, namely: efficient indexing and retrieval of data (the forward channel); and instantaneous tracing from diagnostic conclusions to specific spatial coordinates and original data blocks (the reverse channel). This closed-loop mechanism of "what is diagnosed is what is obtained, and what is obtained is what is sourced" greatly enhances the security, compliance, and credibility of medical data, providing a solid technical foundation for related medical data processing scenarios.
[0188] 3. Digital expression that gives data "life characteristics"
[0189] This invention transcends the superficial data processing of traditional computer science, which merely involves "labeling" data. It innovatively introduces biological thinking, particularly for the processing of medical digital images. Through a global coordinate system established by 3D reconstruction, and by mapping the location genes (spatial coordinates) and trait genes (image features) contained in digital gene sequences, it perfectly maps the "uniqueness" (ID anti-counterfeiting), "heritability" (trait inheritance), and "location" (spatial coordinates) of biological genes to the field of IT data management. This process is not merely a change in naming conventions, but an upgrade in data governance thinking. It will provide a solid foundation for downstream applications such as constructing a full-lifecycle patient pathological evolution map, achieving cross-hospital and cross-device data interoperability, and building a precision medicine "digital twin" model.
[0190] 4. Provide training data for the model
[0191] Current medical AI models (such as Med-PaLM) face significant challenges in training due to data scarcity and high annotation costs. This invention utilizes gridded data blocks and digital gene sequences to create pre-cleaned and pre-annotated corpora. For example, each data block's digital gene sequence includes location labels (A, B, C) and feature labels (G values). This means the digital gene sequences are structured, labeled data that can be directly fed into AI models for self-supervised learning or feature learning. Furthermore, AI models need to understand the relationship between images and text. In this invention, data blocks (image slices) and digital gene sequences (textual descriptions containing modalities, locations, and features) are mutually corresponding. This allows the model to learn more quickly what the corresponding image structure (data block) looks like when a feature value presents a certain form (digital gene sequence information). Moreover, training large models often involves privacy risks. This invention extracts only feature data and grid content without directly transmitting the original DICOM file and patient names. Data blocks themselves are difficult to trace back to specific individuals, thus achieving data anonymization and perfectly matching large model training. It is evident that the gridded medical digital image data block set and its corresponding digital gene sequences generated by this invention provide high-quality structured training materials for medical artificial intelligence. The three-dimensional spatial data and image feature data embedded in the digital gene sequences are equivalent to automated preliminary annotation of the original images, greatly reducing the data cleaning and manual annotation costs for AI model training.
[0192] The above is a detailed description of an embodiment of a medical digital image processing method provided by the present invention. Corresponding to the aforementioned embodiment of a medical digital image processing method, the present invention also discloses an embodiment of a medical digital image processing device. Please refer to [link / reference]. Figure 7Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The device embodiments described below are merely illustrative.
[0193] like Figure 7 As shown, the device includes: an acquisition unit 701, a reconstruction unit 702, a segmentation unit 703, and a generation unit 704. The acquisition unit 701 is used to acquire a sequence of two-dimensional slice medical digital images for the same medical examination.
[0194] The reconstruction unit 702 is used to perform three-dimensional reconstruction on the two-dimensional slice medical digital image sequence to obtain the three-dimensional spatial data, hierarchical data and modal data of the two-dimensional slice medical digital image sequence in three-dimensional space.
[0195] The partitioning unit 703 is used to adaptively discretely partition the two-dimensional segmented medical digital images in the two-dimensional slice medical digital image sequence according to the three-dimensional spatial data and the hierarchical data, to obtain a set of gridded medical digital image data blocks; the set of data blocks includes multiple data blocks that have their own three-dimensional spatial data, hierarchical data, modal data and data features in three-dimensional space.
[0196] The generation unit 704 is used to generate a digital gene sequence that uniquely represents the data characteristics and location of each data block in the gridded medical digital image data block set, based on the three-dimensional spatial data, hierarchical data, modal data, and data characteristics of the data block.
[0197] The specific content of the acquisition unit 701 can be found in the relevant description of step S101 above, and will not be detailed here.
[0198] The reconstruction unit 702 includes: a definition subunit, an establishment subunit, a reconstruction subunit, and a parsing subunit. The definition subunit defines a unified spatial reference origin and coordinate axes for each two-dimensional slice medical digital image in the two-dimensional slice medical digital image sequence. The establishment subunit establishes a global three-dimensional coordinate system based on the spatial reference origin and the coordinate axes. The reconstruction subunit aligns all two-dimensional slice medical images based on the global three-dimensional coordinate system to complete the three-dimensional reconstruction. The parsing subunit analyzes the spatial mapping relationship of the two-dimensional slice medical digital image sequence based on the result of the three-dimensional reconstruction, obtaining the three-dimensional spatial data, hierarchical data, and modal data of each two-dimensional slice medical digital image in the three-dimensional space. For details, please refer to the description of step S102 above; it will not be elaborated here.
[0199] The partitioning unit 703 includes: a first determining subunit, a scanning subunit, a segmenting subunit, an establishing subunit, and a second determining subunit. The first determining subunit is used to determine the three-dimensional spatial boundary of the two-dimensional slice medical digital image based on the three-dimensional spatial data. The scanning subunit is used to scan the two-dimensional slice medical digital image sequence layer by layer. The segmenting subunit is used to perform step-by-step segmentation of the two-dimensional slice medical digital image according to an adaptive segmentation size, obtaining data blocks divided on the two-dimensional slice medical digital image. The establishing subunit is used to extract the data features of the data blocks and establish a mapping relationship between the three-dimensional spatial data, the hierarchical data, and the data features of the data blocks and the data blocks. The second determining subunit is used to determine the set of data blocks with the mapping relationship as a set of gridded medical digital image data blocks for the two-dimensional slice medical digital image sequence. For details, please refer to the relevant description of step S103 above; it will not be elaborated here.
[0200] The generation unit 704 includes: an extraction subunit, a construction subunit, a determination subunit, and a generation subunit. The extraction subunit is used to extract the three-dimensional spatial data, hierarchical data, modal data, and data features of the data block. The construction subunit is used to construct a string by concatenating at least two types of data arbitrarily selected from the three-dimensional spatial data, hierarchical data, modal data, and data features. The determination subunit is used to determine the digital digest generated based on the string as a checksum. The generation subunit is used to generate a digital gene sequence that uniquely represents the data features and location of each data block in the gridded medical digital image data block set, based on the string, the data constructed outside the string, and the checksum. For details, please refer to the relevant description of step S104 above; it will not be elaborated here.
[0201] It may further include: a storage unit for storing the digital gene sequence in a database.
[0202] The above is a description of an embodiment of a medical digital image processing device provided in this application. For details of the device embodiment, please refer to the above method embodiment. The description will not be repeated here.
[0203] Based on the above, this application also provides a medical digital image processing system, such as... Figure 8 As shown, the system includes: a data acquisition module 801, a data processing module 802, and a data storage module 803.
[0204] The data acquisition module 801 is used to acquire two-dimensional slice medical digital image sequences for the same medical examination. For example, it parses and reads the file stream of a DICOM file, converting it into a processable sequence object. It also extracts metadata, such as modality data, from the DICOM file. Therefore, the data acquisition module may include a medical image data parsing submodule and a metadata extraction submodule. The medical image data parsing submodule is used to read the original medical image file and parse it into a standard two-dimensional slice medical digital image sequence. The metadata extraction submodule is used to extract key information from the file header of the medical image file, including but not limited to image modality data, spatial location information, and basic patient information. It should be noted that this embodiment uses a DICOM file as an example, but is not limited to this type of medical digital image; it may also include NIfTI format, Analyze format, NRRD format, etc., which will not be listed here.
[0205] The data processing module 802 includes a 3D reconstruction submodule, a partitioning submodule, and a generation submodule. The 3D reconstruction submodule is used to perform 3D reconstruction on the 2D slice medical digital image sequence to obtain the 3D spatial data, hierarchical data, and modal data of the 2D slice medical digital image sequence in 3D space. The partitioning submodule is used to adaptively partition the 2D segmented medical digital images in the 2D slice medical digital image sequence into a discrete grid based on the 3D spatial data and the hierarchical data to obtain a set of gridded medical digital image data blocks. The set of data blocks includes multiple data blocks with their respective 3D spatial data, hierarchical data, modal data, and data features in 3D space. The generation submodule is used to generate a digital gene sequence that uniquely represents the data features and location of each data block in the set of gridded medical digital image data blocks, based on the respective 3D spatial data, hierarchical data, modal data, and data features of the data blocks.
[0206] The data storage module 803 is used to store the digital gene sequence and the data block corresponding to the digital gene sequence, and to establish a mapping relationship between the digital gene sequence and the data block. The data storage module includes: a basic information storage partition, a digital gene storage partition, an image data storage partition, and a composite index storage partition. The basic information storage partition is used to store attribute data related to medical examinations in the two-dimensional slice medical digital image sequence. The digital gene storage partition is used to store the digital gene sequence; the identifier in the digital gene sequence is associated with the patient identifier in the medical examination-related attribute data stored in the basic information storage partition. The image data storage partition is used to store the data features of the data block, and the mapping relationship established between the hash value of the data block and the digital gene sequence stored in the digital gene storage partition. The composite index storage partition is used to store a multidimensional index constructed based on pathological features and anatomical location, and to retrieve data from the digital gene storage partition using the multidimensional index.
[0207] It may further include: an application service module 804, which, in response to a query request, locates the corresponding digital gene sequence based on the digital gene identifier carried in the query request, and provides one or more services, including data location, recombinant imaging, and report generation, based on the digital gene sequence.
[0208] For details regarding the system described above, please refer to the description of the method described above; further details will not be provided here.
[0209] Based on the above, this application also provides a computer storage medium including a computer program, which, when run on an electronic device, causes the electronic device to perform the relevant steps of the medical digital image processing method described above.
[0210] Based on the above, this application also provides an electronic device, such as... Figure 9 As shown, the electronic device includes: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface communicate with each other through the communication bus. The memory is used to store executable instructions, which cause the processor to execute the instruction buffer method of the aforementioned embodiment.
[0211] The processor can be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), a FPGA (Field Programmable Gate Array), or other programmable devices, transistor logic devices, hardware components, or any combination thereof. The processor can also be a combination that implements computational functions, such as a combination of one or more microprocessors, or a combination of a DSP and a microprocessor.
[0212] The communication bus may include a path for transmitting information between the memory and the communication interface. The communication bus may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 9 The symbol is represented by only one line, but this does not mean that there is only one bus or one type of bus.
[0213] The memory may be ROM (Read-Only Memory) or other types of static storage devices that can store static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it may be EEPROM (Electrically Erasable Programmable Read-Only Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, and optical data storage devices, etc.
[0214] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0215] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0216] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0217] These computer program instructions may also be stored in a computer-readable storage medium capable of directing a computer or other programmable data processing terminal device to operate in a predictive manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 One or more boxes
[0218] The functions specified in the document.
[0219] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0220] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.
[0221] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0222] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this invention are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0223] Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope defined in the claims of the present invention.
Claims
1. A medical digital image processing method, characterized by, include: Acquire two-dimensional slice medical digital image sequences for the same medical examination; The two-dimensional slice medical digital image sequence is reconstructed in three dimensions to obtain the three-dimensional spatial data, hierarchical data and modal data of the two-dimensional slice medical digital image sequence in three-dimensional space; Based on the three-dimensional spatial data and the hierarchical data, the two-dimensional segmented medical digital images in the two-dimensional slice medical digital image sequence are adaptively divided into discrete grids to obtain a set of gridded medical digital image data blocks; the set of data blocks includes multiple data blocks that have their own three-dimensional spatial data, hierarchical data, modal data and data features in three-dimensional space. Based on the three-dimensional spatial data, hierarchical data, modal data, and data features of the data block, a digital gene sequence is generated to uniquely characterize the data features and location of each data block in the gridded medical digital image data block set.
2. The method of claim 1, wherein, The step of reconstructing the two-dimensional slice medical digital image sequence into three dimensions to obtain the three-dimensional spatial data, hierarchical data, and modal data of the two-dimensional slice medical digital image sequence in three-dimensional space includes: For each two-dimensional slice medical digital image in the two-dimensional slice medical digital image sequence, a unified spatial reference origin and coordinate axis are defined; A global three-dimensional coordinate system is established based on the spatial reference origin and the coordinate axes. Based on the global three-dimensional coordinate system, all two-dimensional slice medical images are aligned to complete the three-dimensional reconstruction; Based on the results of the three-dimensional reconstruction, the spatial mapping relationship of the two-dimensional slice medical digital image sequence is analyzed to obtain the three-dimensional spatial data, hierarchical data and modal data of each two-dimensional slice medical digital image in the three-dimensional space.
3. The method of claim 1, wherein, The step of adaptively discretizing the two-dimensional slice medical digital images in the two-dimensional slice medical digital image sequence according to the three-dimensional spatial data and the hierarchical data to obtain a set of gridded medical digital image data blocks includes: Based on the three-dimensional spatial data, determine the three-dimensional spatial boundary of the two-dimensional slice medical digital image; The two-dimensional slice medical digital image sequence is scanned layer by layer; Based on the adaptive segmentation size, the two-dimensional slice medical digital image is segmented stepwise to obtain data blocks divided on the two-dimensional slice medical digital image; Extract the data features of the data block, and establish a mapping relationship between the three-dimensional spatial data to which the data block belongs, the hierarchical data to which the data block belongs, and the data features of the data block and the data block; The set of data blocks with the aforementioned mapping relationship is determined as the set of gridded medical digital image data blocks for the two-dimensional slice medical digital image sequence.
4. The method of claim 1, wherein, The step of generating a digital gene sequence to uniquely characterize the data features and location of each data block in the gridded medical digital image data block set, based on the data block's three-dimensional spatial data, hierarchical data, modal data, and data features, includes: Extract the three-dimensional spatial data, hierarchical data, modal data, and data features of the data block; A string is constructed by concatenating at least two types of data selected arbitrarily from the three-dimensional spatial data, the hierarchical data, the modal data, and the data features. The digital digest generated based on the string is determined as the check code; Based on the string, as well as the data and checksums constructed outside the string, a digital gene sequence is generated to uniquely characterize the data features and location of each data block in the gridded medical digital image data block set.
5. The method of claim 1, wherein, Also includes: The digital gene sequence is stored in a database.
6. A medical digital image processing apparatus characterized by comprising: include: The acquisition unit is used to acquire a sequence of two-dimensional slice medical digital images for the same medical examination. The reconstruction unit is used to perform three-dimensional reconstruction on the two-dimensional slice medical digital image sequence to obtain the three-dimensional spatial data, hierarchical data and modal data of the two-dimensional slice medical digital image sequence in three-dimensional space. A partitioning unit is used to adaptively perform discrete grid partitioning on the two-dimensional segmented medical digital images in the two-dimensional slice medical digital image sequence based on the three-dimensional spatial data and the hierarchical data, to obtain a set of gridded medical digital image data blocks; the set of data blocks includes multiple data blocks that have their own three-dimensional spatial data, hierarchical data, modal data and data features in three-dimensional space. The generation unit is used to generate a digital gene sequence that uniquely represents the data characteristics and location of each data block in the gridded medical digital image data block set, based on the three-dimensional spatial data, hierarchical data, modal data, and data characteristics of the data block.
7. A medical digital image processing system, characterized by include: Data acquisition module, data processing module, data storage module; The data acquisition module is used to acquire two-dimensional slice medical digital image sequences for the same medical examination; The data processing module includes a 3D reconstruction submodule, a partitioning submodule, and a generation submodule; The 3D reconstruction submodule is used to perform 3D reconstruction on the 2D slice medical digital image sequence to obtain the 3D spatial data, hierarchical data, and modal data of the 2D slice medical digital image sequence in 3D space; the partitioning submodule is used to adaptively partition the 2D slice medical digital images in the 2D slice medical digital image sequence into a discrete grid based on the 3D spatial data and the hierarchical data to obtain a set of gridded medical digital image data blocks; the data block set includes multiple data blocks that have their own 3D spatial data, hierarchical data, modal data, and data features in 3D space; The generation submodule is used to generate a digital gene sequence that uniquely represents the data characteristics and location of each data block in the gridded medical digital image data block set, based on the three-dimensional spatial data, hierarchical data, modal data, and data characteristics of the data block. The data storage module is used to store the digital gene sequence and the data block corresponding to the digital gene sequence, and to establish a mapping relationship between the digital gene sequence and the data block.
8. The system of claim 7, wherein, The data storage module includes: a basic information storage partition, a digital gene storage partition, an image data storage partition, and a composite index storage partition; The basic information storage partition is used to store attribute data related to medical examinations in the two-dimensional slice medical digital image sequence; The digital gene storage partition is used to store the digital gene sequence; the identifier in the digital gene sequence is associated with the patient identifier in the medical examination-related attribute data stored in the basic information storage partition; The image data storage partition is used to store the data features of the data block, and to establish a mapping relationship between the hash value of the data block and the digital gene sequence stored in the digital gene storage partition; The composite index storage partition is used to store a multidimensional index constructed based on pathological features and anatomical locations, and the digital gene storage partition is retrieved through the multidimensional index.
9. The system of claim 7, wherein, Also includes: The application service module is used to respond to a query request, locate the corresponding digital gene sequence based on the digital gene identifier carried in the query request, and provide one or more services such as data positioning, recombinant imaging, and report generation based on the digital gene sequence.
10. A computer storage medium, characterized in that, Includes a computer program that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1-5.