Reconstruction Method, Device, System and Medium for Diffuse Optical Tomography
By constructing reconstruction models of prior distribution coding networks, feature coding networks and decoding networks, the problem of discomfort qualitativeness of diffusion optical tomography technology in deep brain tissue detection is solved, and 3D image reconstruction with higher spatial resolution and accuracy is achieved, which improves the clinical applicability of brain functional imaging technology.
Patent Information
- Application Number
- CN202510345513.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-03-24
AI Technical Summary
The existing diffusion optical tomography technology has unfavorable qualities in the detection of deep brain tissue, resulting in high noise sensitivity and low spatial resolution of imaging results, making it difficult to achieve accurate 3D image reconstruction.
The reconstruction model of the prior distribution encoding network, feature encoding network and decoding network is adopted. The trained network generates prior hidden spatial parameter distribution information and structural characteristic implicit information, and the optical transmission parameter spatial distribution of the body to be tested is reconstructed after fusion, and the attention mechanism is used to improve the deep tissue detection ability.
It improves the credibility and accuracy of diffusion light tomography, realizes 3D image reconstruction with higher spatial resolution, especially in deep tissue detection, and promotes the application of brain function imaging technology in clinical practice.
Smart Images

Figure CN119863591B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of medical artificial intelligence, and particularly relates to a reconstruction method, device, system and medium for diffuse optical tomography. Background Art
[0002] Early diagnosis and treatment evaluation of brain diseases are particularly important. Brain functional imaging technologies include functional magnetic resonance imaging (fMRI), electrophysiological imaging (EEG), magnetoencephalography (MEG), functional near-infrared spectroscopy (fNIRS), etc., which can quantitatively image the spatio-temporal state of brain activities and provide important reference bases for early diagnosis and treatment of brain diseases.
[0003] Among them, as an emerging brain functional imaging technology, fNIRS has the advantages of low cost, high portability, non-invasiveness, no ionizing radiation, etc., and is suitable for scenarios such as large-scale clinical screening and daily status monitoring and evaluation of brain disease patients. The diffuse optical tomography (DOT) technology based on fNIRS provides technical possibilities for 3D reconstruction of internal images of brain tissues. However, as an inverse problem, DOT image reconstruction also has ill-posedness: on the one hand, the layered structure of the brain (scalp, skull, cerebrospinal fluid, gray matter, white matter) makes the detection of the cerebral cortex state vulnerable to physiological interference from superficial tissues; on the other hand, the strong scattering of internal brain tissues makes photons prone to random migration during the transmission process in tissues, resulting in higher noise sensitivity of DOT imaging results, especially in the detection of deep tissues, there are still significant deficiencies. Therefore, how to improve the credibility, accuracy and spatial resolution of the reconstruction results when performing 3D image reconstruction based on diffuse optical tomography measurement data is an urgent problem to be solved currently. Summary of the Invention
[0004] This application is provided to solve the above-mentioned defects in the prior art. There is a need for a reconstruction method, device, system and medium for diffuse optical tomography that can perform more credible, accurate and higher spatial resolution 3D image reconstruction based on diffuse optical tomography measurement data at a deeper depth, thereby promoting the clinical applicability of related technologies such as fNIRS brain functional imaging.
[0005] According to the first aspect of the present application, a reconstruction method for diffuse optical tomography is provided. The reconstruction method includes using a processor to reconstruct the spatial distribution of the optical parameters of the object to be measured based on the measurement data obtained by measuring the object to be measured using diffuse optical tomography technology, specifically including: based on the measurement data, using the trained prior distribution encoding network to generate prior latent space parameter distribution information corresponding to the measurement data, and generating latent space sampling based on the prior latent space parameter distribution information; based on the measurement data, using the trained feature encoding network to extract the structural feature latent information in the measurement data; performing fusion processing on the latent space sampling and the structural feature latent information to generate fused latent space sampling; based on the fused latent space sampling, using the decoding network to reconstruct the spatial distribution of the optical transmission parameters of the object to be measured for generating a reconstructed three-dimensional image of the object to be measured.
[0006] According to the second aspect of the present application, a reconstruction device for diffuse optical tomography is provided. The reconstruction device at least includes a processor and a memory. The memory stores computer-executable instructions, and the processor performs the operations of the reconstruction method for diffuse optical tomography as described in various embodiments of the present application when executing the computer-executable instructions.
[0007] According to the third aspect of the present application, a near-infrared brain functional imaging system is provided, including a near-infrared optical data acquisition device and a reconstruction device for diffuse optical tomography; wherein, the near-infrared optical data acquisition device is used to acquire near-infrared data of an object; the reconstruction device is used to perform the operations of the reconstruction method for diffuse optical tomography as described in various embodiments of the present application using the near-infrared data.
[0008] According to the fourth aspect of the present application, a non-transitory computer-readable storage medium storing a program is provided, and the program causes the processor to perform the operations of the reconstruction method for diffuse optical tomography as described in various embodiments of the present application.
[0009] The reconstruction methods, devices, systems and media for diffuse optical tomography provided by various embodiments of the present application construct a reconstruction model including a prior distribution encoding network, a feature encoding network and a decoding network. In the process of reconstructing the spatial distribution of the optical parameters of the object to be measured based on the diffuse optical measurement data of the object to be measured, it can not only accurately estimate the latent space parameter distribution information corresponding to the measurement data by using the pre-trained prior distribution encoding network, but also use the pre-trained feature encoding network to extract the potential structural feature information of the object to be measured from the measurement data and use it to constrain the latent space sampling data generated based on the latent space parameter distribution information, so that the decoding network has higher credibility and accuracy when predicting the spatial distribution of the optical transmission parameters of the object to be measured based on the constrained latent space sampling data. Especially the utilization of the structural feature information of the object to be measured can effectively reduce the ill-posedness of solving the inverse problem of reconstructing diffuse optical tomography based on the probability model and improve the detection ability of deep tissues. Therefore, it can achieve reliable, accurate, higher spatial resolution and more robust 3D image reconstruction of the object to be measured at a deeper depth, which is of great significance in promoting the early screening, diagnosis and prognosis evaluation of brain diseases and more extensive clinical applications of fNIRS technology. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0011] Figure 1 FIG. shows a schematic flowchart of a reconstruction method for diffuse optical tomography according to an embodiment of the present application.
[0012] Figure 2 FIG. shows a partial composition block diagram of a reconstruction model in the inference application stage according to an embodiment of the present application.
[0013] Figure 3 FIG. shows a schematic diagram of the network structure in the training stage of a reconstruction model according to an embodiment of the present application.
[0014] Figure 4 FIG. shows another schematic diagram of the network structure in the training stage of a reconstruction model according to an embodiment of the present application.
[0015] Figure 5 FIG. shows a schematic diagram of the composition structure of a feature encoding network according to an embodiment of the present application.
[0016] Figure 6Schematic diagrams showing the reconstruction results of the reconstruction method according to the present application under different types of media and numbers of targets.
[0017] Figure 7 Schematic diagrams comparing the multi-target image reconstruction performance of the reconstruction method according to embodiments of the present application under different signal-to-noise ratios. Detailed implementation manners
[0018] To enable those skilled in the art to better understand the technical solutions of the present application, the present application will be described in detail below in conjunction with the accompanying drawings and specific implementation manners. The embodiments of the present application will be further described in detail below in conjunction with the accompanying drawings and specific examples, but this is not a limitation to the present application.
[0019] The "first", "second" and similar terms used in the present application do not indicate any order, quantity or importance, but are only used for distinction. The terms such as "including" or "comprising" used in the present application mean that the elements before this word cover the elements listed after this word, and do not exclude the possibility of also covering other elements. In the present application, the arrows shown in the figures for each step are only examples of the execution order and not limitations. The technical solutions of the present application are not limited to the execution order described in the embodiments. Each step in the execution order can be executed jointly, can be decomposed, and can be reordered, as long as the logical relationship of the execution content is not affected.
[0020] All terms used in the present application (including technical terms or scientific terms) have the same meaning as understood by those of ordinary skill in the art to which the present application pertains, unless otherwise specifically defined. It should also be understood that terms defined in a general dictionary, such as, should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and should not be interpreted in an idealized or overly formal sense, unless specifically defined as such here. Technologies and devices known to those of ordinary skill in the relevant field may not be discussed in detail, but where appropriate, such technologies and devices should be regarded as part of the specification.
[0021] In the present application, the arrows shown in the figures for each step are only examples of the execution order and not limitations. The technical solutions of the present application are not limited to the execution order described in the embodiments. Each step in the execution order can be executed jointly, can be decomposed, and can be reordered, as long as the logical relationship of the execution content is not affected.
[0022] Figure 1 Schematic flow diagram showing a reconstruction method for diffuse optical tomography according to an embodiment of the present application. Figure 2 Schematic partial composition block diagram showing a reconstruction model in the inference application stage according to an embodiment of the present application.
[0023] As Figure 1As shown, the reconstruction method includes using a processor to reconstruct the spatial distribution of the optical parameters of the object to be measured based on the measurement data obtained by measuring the object to be measured using diffuse optical tomography technology, using a reconstruction model 200 including a prior distribution encoding network 201, a feature encoding network 202, and a decoding network 203 as shown in Figure 2 The object to be measured can be, for example, the whole body of a living being or a specific organ of a living being, such as the head of the subject, the abdomen of the subject, etc., as long as it is a thick tissue body, and under multi-point excitation such as near-infrared light, the time, space, and spectral distribution measurement information of the diffuse light on its surface can reflect the three-dimensional distribution of the optical characteristic parameters inside the tissue body, and no specific limitation is made here. In addition, the diffuse optical tomography technology includes irradiating the object to be measured with near-infrared light (600nm - 900nm), and the wavelength range of the near-infrared light used can be set as needed according to the characteristics of the object to be measured. The specific measurement method can, for example, arrange light sources and light source detectors on the surface of the object to be measured in a certain topological structure, and form effective detection channels between the light source detectors and at least some of the light sources. During measurement, incident light is emitted from the light source to the object to be measured, and the diffuse light in the outgoing light of the corresponding detection channels obtained from the light source detectors is used as measurement data. The topological structures such as the number, placement positions, and connection methods of the light sources and light source detectors can be set as needed, and the measurement values of each detection channel can be arranged as a one-dimensional array, a two-dimensional matrix, or in any other applicable data format in a predetermined order, which will not be elaborated here.
[0024] The above process can be specifically implemented by performing the following steps 101 - 104.
[0025] First, in step 101, based on the measurement data, use the trained prior distribution encoding network 201 to generate prior latent space parameter distribution information corresponding to the measurement data, and generate latent space sampling based on the prior latent space parameter distribution information.
[0026] In some embodiments, the prior distribution encoding network 201 can be implemented, for example, as a deep network constructed based on multiple continuously stacked convolutional layers to facilitate extracting multi-level feature representations layer by layer and learning the feature parameter distribution of the object to be measured in the latent space from the measurement data.
[0027] In step 102, based on the measurement data, use the trained feature encoding network 202 to extract the structural feature latent information in the measurement data. It should be noted that there is no sequential order between step 102 and step 101, and usually they can be executed synchronously and in parallel.
[0028] Then, in step 103, perform fusion processing on the latent space sampling and the structural feature latent information to generate fused latent space sampling.
[0029] Next, in step 104, based on the fused latent space sampling, the optical transmission parameter spatial distribution of the object to be measured is reconstructed using the decoding network 203 for generating a reconstructed three-dimensional image of the object to be measured.
[0030] According to the reconstruction method for diffuse optical tomography according to the embodiments of the present application, through pre-training, on the one hand, the trained prior distribution encoding network can accurately estimate the latent space parameter distribution information reflecting the structural and functional information of the object to be measured based on the measurement data and the probability model; on the other hand, the feature encoding network is further used to extract additional potential structural feature information of the object to be measured from the measurement data, and fuse it with the latent space sampling generated based on the latent space parameter distribution information, so that the fused latent space sampling can more accurately and with higher precision reflect the optical transmission parameter spatial distribution information of the object to be measured, especially can improve the deficiency of the diffuse optical tomography technology in deep tissue detection, thereby achieving reliable, accurate, higher spatial resolution and more robust 3D image reconstruction of the object to be measured at a deeper depth. The numerical simulation and physical model experiment results of various scenarios show that the reconstruction method according to the embodiments of the present application can reconstruct an object to be measured with a depth of 1.5 cm, and the spatial resolution of the reconstructed image (the minimum spatial distance that can be distinguished as different objects / absorbers during reconstruction, also called the effective resolution) can reach 1 cm. The performance superiority relative to other existing technologies will be specifically described in combination with the experimental results later.
[0031] In some embodiments, the feature encoding network can be implemented as any applicable type of neural network and includes a convolutional module. The reconstruction method can further include, using the processor: introducing an attention mechanism in the convolutional module of the feature encoding network and outputting an attention weight matrix associated with the tissue depth as the implicit structural feature information in the measurement data.
[0032] Considering that the DOT image reconstruction has deficiencies in aspects such as deep tissue detection ability and anti-interference ability, using the attention weight matrix carrying tissue depth-related information as a prior condition constraint for latent space sampling will help the decoding network selectively focus on regions where significant changes in optical properties occur. On the one hand, it can improve the sensitivity to changes in deep tissue parameter features, and at the same time can effectively remove the measurement data perturbations caused by spatial noise rather than real tissue structures, thereby enhancing the credibility, accuracy, precision and robustness of the DOT image reconstruction.
[0033] Figure 3 Shows a schematic diagram of the network structure in the training stage of the reconstruction model according to the embodiments of the present application. As Figure 3 shown, in the training stage of the reconstruction model, a different training network structure 300 will be adopted from the inference application stage.
[0034] Compared with the reconstruction model 200 in Figure 2 , a posterior distribution encoding network 301 is introduced into the training network structure 300. Using the training samples in the training sample set, based on the loss function composed of the weighted sum of the first loss function and the second loss function, the prior distribution encoding network 201, the feature encoding network 202, the decoding network 203 and the posterior distribution encoding network 301 are co-trained. Moreover, during the co-training process, the network parameters of the prior distribution encoding network 201, the feature encoding network 202 and the decoding network 203 are co-optimized until the pre-set training cut-off condition is met. For example, when the value of the loss function is less than the pre-set threshold, or the number of training rounds reaches the predetermined upper limit value, etc., it can be considered that the training is completed. The obtained prior distribution encoding network 201, the feature encoding network 202, and the decoding network 203 can be used as the trained reconstruction model for reconstructing the spatial distribution of the optical parameters of the object to be measured.
[0035] In the embodiments of the present application, both the prior distribution and the posterior distribution can be constructed and approximated by using a multivariate Gaussian distribution probability model with the same number of dimensions. Therefore, after the training is completed, the prior latent space parameter distribution information generated by the trained prior distribution encoding network 201 will follow the first multivariate Gaussian distribution, and the posterior latent space parameter distribution information generated by the trained posterior distribution encoding network 301 will follow the second multivariate Gaussian distribution, and the number of dimensions of the first multivariate Gaussian distribution and the second multivariate Gaussian distribution is the same.
[0036] In some embodiments, the training samples for training the network structure 300 can be obtained as follows: First, based on the anatomical information of the object to be measured for three-dimensional image reconstruction, such as the human head, etc., multiple simulation / simulation objects to be measured with the ground truth of the reconstructed three-dimensional image for training are constructed. The simulation / simulation object to be measured will be collectively referred to as the simulation object to be measured in this application. However, it should be noted that in fact, it can be a simulation object to be measured modeled in a simulation system, or a simulation object to be measured constructed in a physical or semi-physical simulation system. This application does not make specific restrictions on this, as long as the optical transmission forward calculation model adapted to the object to be measured is integrated in the simulation system / simulation system, and the transmission law of photons in the object to be measured can be simulated / simulated.
[0037] When modeling, each different simulation object to be measured can have the same or not completely the same external contour, and different internal tissue structures, for simulating different individuals of the same type of object to be measured. On this basis, each constructed simulation object to be measured is simulated and measured to generate corresponding simulation measurement data, and the paired simulation measurement data and the ground truth of the reconstructed three-dimensional image of each simulation object to be measured are used as training samples.
[0038] In some embodiments, the simulation measurement of each simulation object to be measured to generate corresponding simulation measurement data specifically includes: setting the topological structure composed of the light source and the light source detector of the diffuse light imaging technology on the surface of each simulation object to be measured, wherein the light source is used to emit incident light to each simulation object to be measured, the light source detector is used to detect the outgoing light, and a detection channel is formed between the light source detector and at least part of the light sources; in the case where there are absorbers in each simulation object to be measured, based on the change amount of the optical transmission parameters of each simulation object to be measured and the change amount of the outgoing light parameters of each detection channel arranged in a preset order, simulation measurement data is constructed, wherein the absorber has an absorption effect on the light emitted by the light source, and the voxel resolution of the simulation object to be measured is between 1 mm and 2 mm, which is higher than the 5 mm commonly used in the prior art. Generally, it can be considered that the higher the voxel resolution (that is, the smaller the size of a single voxel in the three-dimensional space), the finer the reconstructed image can be correspondingly (that is, the higher the spatial resolution). However, within the framework of the prior art, it will lead to an excessive amount of calculation, a too long solution time, and the ill-posedness of the image reconstruction inverse problem will be further aggravated, and the reconstruction effect, such as the spatial resolution, may not be able to be stably improved. However, in the case of using the reconstruction algorithm of the embodiments of the present application, since the feature encoding network can provide additional structural feature prior information for the latent space sampling generated by the prior distribution encoding network, the fused latent space sampling input to the decoding network has double-domain constraints in both the measurement domain and the reconstructed image domain, which can not only greatly reduce the ill-posedness of the DOT image reconstruction inverse problem, but also promote the rapid convergence during the training of the reconstruction model, shorten the training time, and after completing the training with the samples generated by simulation, it can robustly generate a three-dimensional DOT reconstruction image with a higher spatial resolution for the simulation object to be measured with a higher voxel resolution at an acceptable amount of computation and a shorter solution time.
[0039] Figure 4 Another schematic diagram of the network structure in the training stage of the reconstruction model according to an embodiment of the present application is shown. In other embodiments, as Figure 4 shown in the training network 400, before feeding the simulation measurement data of the training samples to the prior distribution encoding network 201, the feature encoding network 202, and the posterior distribution encoding network 301, the simulation measurement data can be upsampled by using the upsampling module 401 according to the voxel resolution of the three-dimensional image to be reconstructed, and feature data with dimension alignment matching the voxel resolution of the three-dimensional image to be reconstructed is output. Only as an example, the upsampling module 401 can be implemented as a fully connected network, so that while achieving dimension alignment, it can also achieve the effect of providing a rough reconstruction for sparse measurements, so that the subsequent encoding networks can perform more refined feature extraction based on the high-dimensional features containing richer semantics, and can alleviate the ill-posedness of the inverse problem solution to a certain extent.
[0040] As Figure 4 shown, assuming that the resolution of the three-dimensional image to be reconstructed is 16×26×26, where 16 is the number of layers of the reconstructed image in the depth dimension, and 26×26 is the resolution of each two-dimensional image layer. Then, the dimension of the dimension-aligned feature data output by the upsampling module 401 can be expressed as (16, 26, 26). In this case, the dimension of the attention weight matrix as the implicit information of the structural features will also match the voxel resolution of the three-dimensional image to be reconstructed, that is: the weight matrix is also a three-dimensional matrix with a dimension of (16, 26, 26).
[0041] Next, when fusing the latent space sampling and the implicit information of the structural features, different methods can be adopted. Assuming that the dimension of the first multivariate Gaussian distribution is 4, so the latent space sampling is a three-dimensional matrix with a dimension of (4, 1, 1) containing 4 random sampling values (such as s1, s2, s3, and s4). Then, first, according to the three dimensions of the implicit information of the structural features, the three-dimensional matrix can be extended in the other two dimensions outside the image layer (the number of layers is 16), that is, 26×26. The dimension of the extended latent space sampling three-dimensional matrix is (4, 26, 26). In the case of alignment in these two dimensions, for example, the extended latent space sampling information can be concatenated with the implicit information of the structural features in the image layer dimension. The dimension of the fused latent space sampling obtained after concatenation will be (20, 26, 26). And the ground truth of the reconstructed three-dimensional image of the simulated object to be measured can be concatenated with the dimension-aligned feature data in the image layer dimension. Therefore, the matrix dimensions of the image information and the fused data input to the posterior distribution encoding network 301 will be (32, 26, 26).
[0042] In addition, in the inference application stage of the reconstruction model after training, before the measurement data of the object to be measured is fed into the trained prior distribution encoding network 201 and the trained feature encoding network 202, the upsampling module 401 is also used to perform upsampling processing on the measurement data, which will not be elaborated here.
[0043] In a simulation / simulation system for generating simulated measurement data, only as an example, for instance, an extended form of the radiative transfer equation (RTE) of the first-order spherical harmonic function, that is, the diffusion equation, can be adopted as the theoretical model for photon propagation in tissue. The diffusion equation in continuous wave form is as shown in Equation (1):
[0044] (1)
[0045] Where represents the position of each voxel, is the absorption coefficient, is the diffusion term, defined as , where is the reduced scattering coefficient. represents the light source term and represents the position of the light intensity at, represents the light intensity distribution, represents the speed of light in vacuum.
[0046] The model shown in Equation (1) relates the variation of the optical absorption coefficient in the tissue to the logarithmic ratio of the measured values of each voxel or grid node to each source-detector pair relative to the baseline measurement value as shown in Equation (2):
[0047] (2)
[0048] where represents the position of the light source, represents the position of the detector, represents the optical flux of the light with wavelength generated by the light source at position and is formally expressed as the Green's function solution of the steady-state diffusion equation. Each measurement channel is approximated by the discrete sum over the voxels and is represented using the following notations of Equation (3) and Equation (4):
[0049] (3)
[0050] (4)
[0051] By combining Equation (2), Equation (3) and Equation (4), a linear forward model can be obtained as shown in Equation (5):
[0052] (5)
[0053] where, , , Y is the measurement domain of the boundary measurement, X is the image domain of the optical properties in the tissue, A is the sensitivity matrix, and each row of A represents the spatial sensitivity of the measured value at wavelength to the change in the absorption coefficient of each voxel in the tissue.
[0054] Based on the numerical simulation of the photon transport theory, by setting a high-density light source and a light source-detector topology on the surface of the simulated object to be measured, and setting one or more absorbers in different ways inside the simulated object to be measured (equivalent to generating multiple different ground truths of the reconstructed three-dimensional images ), Based on the forward model such as the above formula (5) and the theoretical value of A, realistic and rich simulated measurement data can be generated (that is, corresponding to each corresponding ), A large number of pairs of are used as training samples. After supervised training of the reconstruction model, the reconstruction model will be able to acquire the high-precision representation ability of the inverse mapping of DOT measurement data, that is, the ability to solve F -1 (Y)→X the inverse problem. And through the reasonable setting of training samples and the optimization of training methods, the generalization ability and robustness of the reconstruction model can be further enhanced.
[0055] Refer to Figure 4 to illustrate the construction methods of the first loss function and the second loss function. In some embodiments, the first loss function can be constructed in the following way, for example: Based on the dimension-aligned feature data, using the prior distribution encoding network 201, generate the prior latent space parameter distribution information corresponding to the dimension-aligned feature data; Based on the dimension-aligned feature data together with the reconstructed three-dimensional image ground truth, using the posterior distribution encoding network 301, generate the posterior latent space parameter distribution information corresponding to the dimension-aligned feature data; Construct the first loss function based on the prior latent space parameter distribution information and the posterior latent space parameter distribution information. Only as an example, the KL divergence of the above two latent space parameter distribution information can be used as the first loss function. As mentioned before, the prior distribution encoding network and the posterior distribution encoding network can be constructed according to the multivariate Gaussian distribution probability model. Specifically, both the prior distribution encoding network 201 and the posterior distribution encoding network 301 can include multiple stacked modules such as convolutional layers, batch normalization layers, and activation function layers, so as to map based on the input information to obtain the multivariate Gaussian distribution information with corresponding mean and standard deviation.
[0056] When the mean of the first multivariate Gaussian distribution is μ proir , and the standard deviation is σ proir , a sample generated by the prior distribution encoding network 201 based on the dimension-aligned feature data y can be expressed as the following formula (6):
[0057] (6)
[0058] When the mean of the second multivariate Gaussian distribution is μ post , and the standard deviation is σ post , since the posterior distribution encoding network 301 needs to use the dimension-aligned feature data together with the reconstructed three-dimensional image ground truth both, a sample generated by it can be expressed as the following formula (7):
[0059] (7)
[0060] At the beginning of training, for a specific training sample, the prior latent space parameter distribution information generated by the prior distribution encoding network and the posterior latent space parameter distribution information generated by the posterior distribution encoding network are different. Therefore, the training objective of the first loss function part in the loss function is to minimize the KL divergence of the above two multivariate Gaussian distributions. That is, the first loss function can be expressed as the following formula (8):
[0061] (8)
[0062] The smaller the first loss function, the closer the two multivariate Gaussian distributions will be, which means that the trained prior distribution encoding network can more accurately generate the latent space features related to image reconstruction hidden therein based on the input measurement data.
[0063] In some embodiments, the second loss function can be constructed based on the reconstructed three-dimensional image ground truth x and the mean-square error (MSE) or other similar error functions of the three-dimensional image of the simulated object to be measured reconstructed based on the spatial distribution of the optical parameters of the simulated object to be measured output by the decoding network. Taking MSE as an example only, the second loss function can be expressed as the following formula (9):
[0064] (9)
[0065] Where is the equivalent mapping function of the entire reconstruction network during the training process of the reconstruction model.
[0066] In some embodiments, the loss function can be defined as the weighted sum of the first loss function and the second loss function, for example, it can be expressed as the following formula (10):
[0067] (10)
[0068] Where represents the weight coefficient between the two loss functions, which is used to adjust the relative proportion of the two parts of the loss. The specific value can be set as needed considering factors such as the convergence speed and the performance of the model after training. This application does not limit this.
[0069] In addition, the specific number of dimensions of the first multivariate Gaussian distribution and the second multivariate Gaussian distribution can be determined through experiments to obtain appropriate values, so that the first multivariate Gaussian distribution and the second multivariate Gaussian distribution with a specific number of dimensions can adequately and without redundancy characterize the problem of the distribution of optical transmission parameters in a three-dimensional space, such as the position of the activation region inside the object to be measured, the shape of the activation region, the number of activation regions, the activation degree of the activation region, etc. It should be noted that since the samples generated based on the first multivariate Gaussian distribution and the second multivariate Gaussian distribution in the embodiments of the present application are all latent space samplings, the sample data of each dimension does not necessarily explicitly correspond to the above variables or accurately correspond to the actual values of these variables. Through the above analysis and comparative experiments on different numbers of dimensions, it is shown that setting the number of dimensions of the first multivariate Gaussian distribution and the second multivariate Gaussian distribution in the embodiments of the present application to be between 2 and 6 will make the reconstruction model have a wide range of applicability, avoiding both the poor reconstruction effect caused by too few dimensions and the problems of low model training efficiency, difficulty in convergence, and even degradation of model performance caused by too many dimensions, and achieving a good balance between the reconstruction effect and the training efficiency. In particular, when the number of dimensions is set to 4, the model performance is close to the optimal, and the training and application costs are acceptable.
[0070] Figure 5 FIG. shows a schematic structural diagram of the composition of a feature encoding network according to an embodiment of the present application. As Figure 5 shown, in the feature encoding network 202 into which the dimension-aligned feature data enters, a depth perception attention mechanism is integrated ( Figure 5Before the convolutional module of "Depth-aware attention" in [reference], it can first pass through two convolutional layers (Conv), one batch normalization layer (BN), and one activation function layer (ReLu). During the organization of depth-aware attention extraction, the max pooling layer (MaxPool) and average pooling layer (AvgPool) are used to aggregate information of the same depth. Subsequently, two fully connected layers are used to adjust the focus of the model on different depth layers, and the Sigmoid activation function is applied to obtain the normalized weights for feature map adjustment. In addition, residual connections are added to these layers to prevent performance degradation caused by excessive network depth. For the feature encoding network, the number of network layers is a key parameter affecting the inverse mapping representation of the reconstruction model. Generally, the more layers, the stronger the representation ability, but it does not always show a linear positive correlation trend. Too few or too many layers may not achieve the optimal performance. Especially when there are too many layers, it may lead to problems such as overfitting of the network model or vanishing gradients. The reduction ratio parameter of the features in the organization depth-aware attention mechanism reflects the scale of feature fusion across different depths. If its value is too low, it may cause the attention mechanism to be unable to effectively extract the key information to be concerned. If its value is too high, it may prevent the effective recovery of details during the organization depth-aware feature refinement process. The specific values of the above hyperparameters (such as the reduction ratio) can be jointly optimized through experiments and together with other parameters. This application does not limit this.
[0071] In the process of practicing this application, a numerical simulation platform and a physical model experiment platform were also synchronously constructed to verify the reconstruction method of the embodiments of this application, and performance comparison evaluations were carried out with two representative existing DOT image reconstruction models in multiple scenarios and multiple angles. One is a traditional DOT image reconstruction model (referred to as Model 1) based on inverse scattering theory that combines linear methods and regularization techniques to constrain the solution space, and the other is a DOT image reconstruction model based on deep learning (referred to as Model 2).
[0072] In the reconstruction method of the embodiments of this application, the model hyperparameter combination of the reconstruction network is adjusted to the following optimal values through experiments: the number of network layers in the feature encoding network is 10, the dimensions of the first multivariate Gaussian distribution / second multivariate Gaussian distribution are set to 4, and the reduction ratio of the features in the organization depth-aware attention mechanism is 4 to construct the final reconstruction network model.
[0073] Figure 6 Shows a schematic diagram of the reconstruction results of the reconstruction method according to this application under different types of media and the number of targets. In Figure 6In the experiment, the object to be measured was set as two cases of homogeneous medium and layered medium, and the numerical simulation experiment results of single-object, double-object and multi-object reconstructions using the reconstruction method of this application and two other existing technologies were compared. Specifically, the single-object reconstruction experiment was used to evaluate the ability of the model to reconstruct absorbers with different sizes and optical properties, the double-object reconstruction experiment was used to evaluate the ability of the model to distinguish different absorbers, and the multi-object reconstruction experiment was mainly used to test the generalization performance of the model. Figure 6 In (a1)-(a3), it corresponds to the single-object reconstruction experiment, (b1)-(b3) corresponds to the double-object reconstruction experiment, and (c1)-(c3) corresponds to the multi-object reconstruction experiment. In the three types of experiments, (a1), (b1), (c1) represent the xy cross-section at a depth of 15 mm, and the white dotted line represents the target shape contour at y = 25 mm; (a2), (b2), (c2) represent the reconstruction results of the corresponding items after applying the full width at half maximum threshold; (a3), (b3), (c3) represent the 3D reconstruction results of the corresponding items. The experiments normalized each reconstruction result based on the maximum absorption coefficient of the target absorber in the reconstructed image, and the normalized values are shown in the legend on the right in Figure 6 The larger the normalized value, the greater the absorption coefficient and the higher the activation level of the target absorber.
[0074] From Figure 6 the first row of image ground truth, it can be seen that in the single-object reconstruction experiment corresponding to (a1)-(a3), a cylindrical optical absorber with a radius of 7.5 mm and a height of 20 mm is located at the center of the imaging area, with a depth of 20 mm. In the double-object reconstruction experiment corresponding to (b1)-(b3), two cylindrical absorbers with a radius of 5 mm and a height of 20 mm are located on both sides of the center of the imaging area, with a depth of 20 mm, and the interval between them is 10 mm. In the multi-object reconstruction experiment corresponding to (c1)-(c3), multiple cylindrical absorbers with different radii and heights are randomly distributed throughout the imaging area.
[0075] As Figure 6 shown by the reconstruction results in the homogeneous medium, Model 1 cannot provide accurate reconstructions in single-object, double-object and multi-object experiments. The overall boundary of its reconstructed image is significantly different from the ground truth situation and tends to shift towards the shallow layer of the tissue. In contrast, Model 2 shows higher accuracy in reconstructing the shape and position of the absorber and can basically effectively distinguish the foreground and background. In the homogeneous medium and single-object reconstruction experiment, the reconstruction method of this application is comparable to Model 2 in performance. The reconstructed single template is highly consistent with the image ground truth in terms of shape, contour and tissue depth, indicating that both methods can effectively reconstruct a single relatively simple optical absorber with a simple geometric shape.
[0076] As the number of target absorbers increases, as shown in (b1)-(b3), both Model 2 and the reconstruction method of the present application can achieve dual-target reconstruction. However, comparatively speaking, there are artifacts in the foreground and background reconstructions of Model 2, which results in its performance in aspects such as the reconstruction of the optical characteristic parameters of the target absorber and the clarity of the target contour being lower than that of the reconstruction method of the present application. This indicates that the attention mechanism introduced in the reconstruction method of the present application enhances the ability of the reconstruction model to focus on details.
[0077] In the multi-target reconstruction experiment shown in (c1)-(c3), both Model 2 and the reconstruction method of the present application accurately locate the optical absorbers with significant changes. However, for the smaller absorbers with lower optical characteristics that coexist, Model 2 tends to underestimate them, resulting in ineffective reconstruction. It is speculated that this problem occurs because the use of 3D convolution in Model 2 tends to focus on the average level of overall reconstruction, and when the absorber is small, it is more vulnerable to the influence of the background. On the contrary, the reconstruction method of the present application combines the embedded features generated from the latent space with the attention mechanism, enabling the reconstruction model to have the ability to accurately capture the shape contour and reconstruct the target optical characteristic parameters under more complex target absorber conditions.
[0078] Compared with the homogeneous medium, such as Figure 6When performing reconstruction in a layered medium, the reconstructed images of the three methods all have problems of quality degradation in terms of contrast between the target absorber and the background and shape recovery. Specifically, Model 1 still fails to accurately reconstruct in all three experiments. The reconstructed boundary is significantly different from the ground truth and is significantly shifted towards the shallow layer. In addition, after applying the full width at half maximum threshold, it is still unable to filter out unimportant absorbers during reconstruction. The background contains a large number of artifacts, whose values are comparable to those of the absorbers being reconstructed, and the smoothness of the shape contour decreases. This indicates that the excessive dependence of Model 1 on the forward model limits its flexibility and adaptability during reconstruction. In contrast, Model 2 and the reconstruction method of the present application show significant advantages in reconstructing the shape and position of absorbers while effectively suppressing background noise. Further, for single-target and dual-target reconstructions, Model 2 is not as effective as the reconstruction method of the present application in distinguishing the foreground and background. There are a large number of artifacts at the absorber boundary in its reconstructed image, resulting in an increase in the fluctuations of the shape contour. In addition, compared with Model 2, the reconstruction method of the present application shows higher spatial resolution, clearly delineating the boundaries of dual-target absorbers. For multi-target reconstructions, Model 2 can only provide a rough localization of the absorbers and cannot accurately distinguish the boundaries between each absorber. Moreover, its sensitivity to smaller absorbers is significantly reduced, and there is a problem of overestimating the optical properties of larger absorbers. It is speculated that this phenomenon is due to the difficulty in effectively extracting absorber patterns when Model 2 is applied to layered medium data, resulting in performance limitations. On the contrary, the reconstruction method of the present application shows strong adaptability, has good stability in terms of shape recovery, absorber localization and spatial resolution, and is the least affected by changes in medium properties among the three methods. Further quantitative analysis of performance metrics also shows that the reconstruction method of the present application is less affected by the medium type on all three datasets, and its reconstruction metrics in the layered medium are even better than those of Model 2 in the homogeneous medium. These findings indicate that the reconstruction method of the present application not only performs well under homogeneous medium conditions, but also shows strong generality and generalization ability between different medium types, and has stronger practical value in more complex actual application scenarios.
[0079] In addition, considering the inevitable environmental light interference and sensor measurement noise in the diffusion optical tomography measurement data, the present application specifically conducts a comparative experiment on the performance of the reconstruction method under noise conditions.
[0080] Figure 7 Schematic diagram showing the comparison of multi-target image reconstruction performance of the reconstruction method according to the embodiments of the present application at different signal-to-noise ratios. At Figure 7In the experiment, the reconstruction method of this application is still considered for comparison with the aforementioned Model 1 and Model 2. Among them, the signal-to-noise ratios (SNRs) of the multi-objective datasets used are 25 dB, 15 dB, and 5 dB respectively. In the embodiments of this application, regarding the image reconstruction performance, an index system commonly used in the field of image processing and image quality assessment is adopted, including SSIM (Structural Similarity Index), PSNR (Peak Signal-to-Noise Ratio), CNR (Contrast-to-Noise Ratio), R value (Pearson Correlation Coefficient, also known as Pearson correlation coefficient), Jaccard index, and MAE (Mean Absolute Error). Their definitions, calculation methods, and the meanings of the index values are well-known to those skilled in the art and will not be elaborated here. From Figure 7 It can be seen that as the noise level increases, the overall reconstruction performance of the reconstruction method of this application shows a downward trend. However, compared with Model 1 and Model 2, the reconstruction method of this application exhibits higher SSIM, PSNR, CNR, R value, and Jaccard index, as well as lower MAE, indicating its superior anti-noise performance and robustness. In particular, even when the SNR is 5 dB, the reconstruction results of this application exceed the performance of Model 1 and Model 2 with SNRs of 15 dB and 25 dB, demonstrating that the reconstruction method of this application has more excellent anti-interference ability compared with the prior art.
[0081] In some embodiments of this application, a reconstruction device for diffuse optical tomography is also provided. The reconstruction device at least includes a processor and a memory. Computer-executable instructions are stored on the memory. When the processor executes the computer-executable instructions, it executes the reconstruction method for diffuse optical tomography described in each embodiment of this application.
[0082] The processor can be a processing device including more than one general-purpose processing device, such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), etc. More specifically, the processor can be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor running other instruction sets, or a processor running a combination of instruction sets. The processor can also be more than one dedicated processing device, such as an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), a system-on-chip (SoC), etc.
[0083] In some embodiments of the present application, a near-infrared brain functional imaging system is provided, including a near-infrared optical data acquisition device and a reconstruction device for diffuse optical tomography; wherein, the near-infrared optical data acquisition device is used to acquire near-infrared data of an object; the reconstruction device is used to execute the reconstruction method for diffuse optical tomography described in various embodiments of the present application by using the near-infrared data.
[0084] For the reconstruction method, device, system and medium for diffuse optical tomography provided in various embodiments of the present application, first, by using the high-density diffuse optical tomography training data generated based on the anatomical information of the object to be measured, with the assistance of the posterior distribution encoding network, the prior distribution encoding network constructed based on the multivariate Gaussian distribution probability model is enabled to acquire the ability to map the measurement data into the latent space parameter distribution information, and, the feature encoding network with an attention mechanism is enabled to acquire the ability to extract the structural feature information including depth information from the measurement data, and in the reconstruction model inference application stage, only the prior distribution encoding network, the feature encoding network and the decoding network are used to complete the high-precision reconstruction of the spatial distribution of the optical transmission parameters of the object to be measured, providing a new approach for high-spatial-resolution near-infrared brain functional imaging, and through experimental results, it is shown that the reconstruction performance is significantly better than a variety of existing DOT image reconstruction techniques, which has important significance and great application potential for promoting the fNIRS technology in the early screening, diagnosis and prognosis evaluation of brain diseases and other application scenarios.
[0085] The present application describes various operations or functions, which can be implemented as software code or instructions or defined as software code or instructions. Such content can be source code that can be directly executed or differential code ("incremental" or "patch" code) ("object" or "executable" form). The software code or instructions can be stored in a computer-readable storage medium, and when executed, can cause a machine to execute the described functions or operations, and include any mechanism for storing information in a form accessible by a machine (such as a computing device, an electronic system, etc.), such as a recordable or non-recordable medium (such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium, an optical storage medium, a flash device, etc.).
[0086] The exemplary methods described in the present application can be at least partially implemented by a machine or a computer. In some embodiments, a computer-readable storage medium stores computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the operations of the reconstruction method for diffuse optical tomography described in various embodiments of the present application.
[0087] Implementations of such methods may include software code, such as microcode, assembly language code, high-level language code, and the like. Various software programming techniques may be used to create various programs or program modules. For example, program portions or program modules may be designed using or with the aid of Java, Python, C, C++, assembly language, or any other known programming language. One or more of such software portions or modules may be integrated into a computer system and / or computer-readable media. Such software code may include computer-readable instructions for performing the various methods. The software code may form part of a computer program product or computer program module. Furthermore, in an example, the software code may be tangibly stored on one or more volatile, non-transitory, or non-volatile tangible computer-readable media, such as during execution or at other times. Examples of such tangible computer-readable media may include, but are not limited to, hard disks, removable disks, removable optical disks (e.g., optical disks and digital video disks), magnetic cassettes, memory cards or sticks, random access memory (RAM), read-only memory (ROM), and the like.
[0088] Furthermore, although exemplary embodiments have been described herein, the scope includes any and all embodiments based on the present application with equivalent elements, modifications, omissions, combinations (e.g., solutions that intersect various embodiments), adaptations, or changes. The elements of the claims are to be interpreted broadly based on the language employed in the claims and are not limited to the examples described in this specification or during the prosecution of the application, which examples are to be construed as non-exclusive. Therefore, it is intended that this specification and examples be considered merely as examples, with the true scope and spirit being indicated by the following claims and their full scope of equivalents.
[0089] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of their embodiments) may be used in combination with each other. For example, a person of ordinary skill in the art may use other embodiments when reading the above description. In addition, in the above detailed description, various features may be grouped together to simplify the application. This should not be interpreted as an intention that a disclosed feature that is not claimed for protection is essential to any claim. On the contrary, the subject matter of the present application may have less than all the features of a particular disclosed embodiment. Thus, the claims are incorporated herein into the detailed description as examples or embodiments, with each claim independently serving as a separate embodiment, and it is contemplated that these embodiments may be combined with each other in various combinations or arrangements. The scope of the present application should be determined with reference to the appended claims and the full scope of equivalents to which such claims are entitled.
[0090] The above embodiments are only exemplary embodiments of the present application and are not used to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements to the present application within the essence and protection scope of the present application, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present application.
Claims
1. A reconstruction method for diffuse light tomography, characterized in that: The reconstruction method includes using a processor to reconstruct the optical parameter spatial distribution of the object to be measured based on measurement data obtained after measuring the object to be measured using diffuse optical tomography technology, and using a reconstruction model including a prior distribution encoding network, a feature encoding network, and a decoding network, specifically including: Based on the measurement data, using a trained prior distribution encoding network, generating prior latent space parameter distribution information corresponding to the measurement data, and generating latent space samples based on the prior latent space parameter distribution information; Based on the measurement data, using a trained feature encoding network, extracting implicit information of structural features in the measurement data; The latent space samples and the structural feature latent information are fused, specifically comprising: expanding and dimensionally aligning the latent space samples in two dimensions other than the image layer according to the three dimensions of the structural feature latent information; and splicing the expanded and dimensionally aligned latent space sample information with the structural feature latent information in the image layer dimension to generate a fused latent space sample. Based on the fused latent space sampling, the decoding network is used to reconstruct the spatial distribution of the optical transmission parameters of the object to be measured, so as to generate a reconstructed three-dimensional image of the object to be measured.
2. The reconstruction method according to claim 1, characterized in that The feature encoding network includes a convolution module, and the reconstruction method further includes, using the processor: introducing an attention mechanism in the convolution module of the feature encoding network, and outputting an attention weight matrix associated with tissue depth as implicit information of structural features in the measurement data.
3. The reconstruction method according to claim 1, wherein: The reconstruction method further includes using the processor: introducing a posterior distribution encoding network in the training stage, using training samples in the training sample set, and collaboratively training the prior distribution encoding network, the feature encoding network, the decoding network, and the posterior distribution encoding network based on a loss function composed of a weighted sum of a first loss function and a second loss function.
4. The reconstruction method according to claim 3, characterized in that: The prior distribution and the posterior distribution are constructed using a multivariate Gaussian distribution probability model with the same number of dimensions. The prior latent space parameter distribution information generated by the trained prior distribution encoding network follows a first multivariate Gaussian distribution; the posterior latent space parameter distribution information generated by the trained posterior distribution encoding network follows a second multivariate Gaussian distribution; and the first multivariate Gaussian distribution and the second multivariate Gaussian distribution.
5. The reconstruction method according to claim 4, characterized in that: The number of dimensions of the first multivariate Gaussian distribution and the second multivariate Gaussian distribution is 4.
6. The reconstruction method according to claim 3, characterized in that: The training samples are obtained as follows: Based on the anatomical information of the object to be tested, a plurality of simulated objects to be tested with ground truth values of reconstructed three-dimensional images for training are constructed, and simulated measurements are performed on each simulated object to generate corresponding simulated measurement data; The simulated measurement data and the reconstructed 3D image ground truth of each simulated object are used as training samples.
7. The reconstruction method according to claim 6, characterized in that: The reconstruction method further includes upsampling the simulated measurement data of the training samples to generate dimensionally aligned feature data according to the voxel resolution of the three-dimensional image to be reconstructed before feeding the simulated measurement data to the prior distribution encoding network, the feature encoding network and the posterior distribution encoding network.
8. The reconstruction method according to claim 7, characterized in that: The first loss function and the second loss function are constructed as follows: Based on the dimensionally aligned feature data, using the prior distribution encoding network, generating prior latent space parameter distribution information corresponding to the simulated measurement data; Based on the dimensionally aligned feature data and the reconstructed three-dimensional image ground truth, using the posterior distribution encoding network, generating posterior latent space parameter distribution information corresponding to the simulated measurement data; Constructing the first loss function based on the KL divergence of the prior latent space parameter distribution information and the posterior latent space parameter distribution information; A second loss function is constructed based on the mean square error of the reconstructed three-dimensional image ground truth and the reconstructed three-dimensional image of the simulated object to be tested.
9. The reconstruction method according to claim 6, characterized in that: Performing simulation measurements on each simulated object to generate corresponding simulation measurement data specifically includes: Setting a topological structure consisting of a light source and a light source detector using diffuse light imaging technology on the surface of each simulated object to be tested, wherein the light source is used to emit incident light to each simulated object to be tested, and the light source detector is used to detect the outgoing light, and a detection channel is formed between the light source detector and at least part of the light source; In the case where there is an absorber in each simulated test body, simulated measurement data is constructed based on the change in the optical transmission parameters of each simulated test body and the change in the output light parameters of each detection channel arranged in a preset order, wherein the absorber has an absorption effect on the light emitted by the light source, and the voxel resolution of the simulated test body is between 1mm and 2mm.
10. The reconstruction method according to any one of claims 1 to 9, characterized in that: The diffuse light tomography technique involves the use of near-infrared light.
11. A reconstruction device for diffuse optical tomography, characterized in that: The reconstruction device comprises at least a processor and a memory, wherein the memory stores computer-executable instructions, and the processor executes the reconstruction method for diffuse optical tomography according to any one of claims 1 to 10 when executing the computer-executable instructions.
12. A near-infrared brain function imaging system, characterized in that: It includes a near-infrared optical data acquisition device and a reconstruction device for diffuse light tomography; wherein, The near-infrared optical data acquisition device is used to collect near-infrared data of the object; The reconstruction device is configured to execute the reconstruction method for diffuse optical tomography according to any one of claims 1 to 10 using the near-infrared data.
13. A non-transitory computer-readable storage medium storing a program, the program causing a processor to execute the operations of the reconstruction method for diffuse optical tomography according to any one of claims 1 to 10.
Citation Information
Patent Citations
Method and system for reconstructing mesoscopic fluorescent probe based on full convolution codec architecture
CN110772227A
Method for reconstructing three-dimensional scene based on variation fraction distillation and electroencephalogram encoder
CN118262045A