Data integration method and device for space transcriptome and space metabolome, electronic equipment and storage medium
By performing distance fitting and correlation between spatial transcriptome and metabolome data points, the compatibility problem of data integration under different sequencing chip structures was solved, and the reliability and accuracy of the integration results were improved.
Patent Information
- Application Number
- CN202510853498.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-23
AI Technical Summary
Existing spatial transcriptomics and spatial metabolomics data integration technologies are not compatible with different sequencing chip structures, resulting in poor reliability of the integration results and mismatches in the spatial structure distribution of the two omics data.
By obtaining the spatial transcriptome and metabolome data points after spatial coordinate registration, the nearest neighbor metabolome data point corresponding to each transcriptome data point is determined, and the distance between the two is used for fitting and correlation to achieve the unification of the data points in the same spatial coordinate system.
It improves the reliability and compatibility of data integration results, adapts to the data characteristics under different chip structures, and significantly improves the accuracy and consistency of integration results.
Smart Images

Figure CN120689376A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of bioinformatics analysis, and in particular to a method, device, electronic device and storage medium for integrating spatial transcriptome and spatial metabolome data. Background Art
[0002] Spatial metabolomics uses mass spectrometry technology to obtain high-throughput qualitative and quantitative information on small-molecule metabolites, while spatial transcriptomics uses sequencing technology to obtain high-throughput information on gene expression in samples. Transcribed mRNA is the direct product of gene expression and is processed, modified, and translated into protein. Metabolites are the products or intermediates of protein (enzyme)-catalyzed reactions in biological substances and are the ultimate outcome of gene expression and the material basis of an organism's phenotype. Spatial transcriptomics and spatial metabolomics respectively examine different stages of the central principle. Joint analysis of spatial transcriptomics and spatial metabolomics data can cross-validate the conclusions of the two groups, helping researchers explore the biological functional changes caused by changes in conditions (environment, disease, drugs, development) through differences in gene expression and metabolite phenotypes.
[0003] In the practical application of the joint analysis of spatial transcriptomics and spatial metabolomics, data integration is two key and indispensable technical links. However, in the process of realizing the present invention, the inventors found that the existing integration method has some defects in its specific implementation: the sequencing spots of the spatial transcriptomics sequencing chip are not closely arranged, and there are large gaps between each spot point, while the basic detection unit in the spatial metabolomics detection - the pixel point - is densely arranged, and there are significant differences in the spatial structure distribution between the two. This structural mismatch reduces the reliability of the integration results. In addition, there are obvious differences in the structure of spatial transcriptomics chips between different manufacturers, which is mainly reflected in the arrangement of sequencing spots on the chip. This difference makes the integration method of the two-omics data show a large difference. However, the existing integration technology fails to fully consider all chip structures in process design, and is only applicable to mainstream foreign chip structures and data structures. It is difficult to meet the diverse integration needs brought about by chip structures from different manufacturers.
[0004] Therefore, how to provide a reliable method for integrating spatial transcriptomics data and spatial metabolomics data that is compatible with different sequencing chip structures is a technical problem that needs to be solved. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method, device, electronic device and storage medium for integrating spatial transcriptome and spatial metabolome data, which are used to improve the reliability of data integration results and be compatible with spatial transcriptome and spatial metabolome data integration scenarios under different sequencing chip structures. In order to achieve the above purpose, the scheme adopted in the embodiment of the present invention is as follows: In a first aspect, the present invention provides a method for integrating spatial transcriptome and spatial metabolome data, the method comprising: obtaining spatial transcriptome data points and spatial metabolome data points after spatial coordinate alignment; determining the nearest neighbor spatial metabolome data point corresponding to each spatial transcriptome data point; using the distance between the spatial transcriptome data point and the nearest neighbor spatial metabolome data point, fitting the original data at the nearest neighbor spatial metabolome data point, and associating it with the spatial transcriptome data point.
[0006] In a second aspect, the present invention provides a data integration device for a spatial transcriptome and a spatial metabolome, comprising: an acquisition module, a search module, and an association module; the acquisition module is used to obtain spatial transcriptome data points and spatial metabolome data points after spatial coordinate alignment; the search module is used to determine the nearest neighbor spatial metabolome data point corresponding to each spatial transcriptome data point; the integration module is used to use the distance between the spatial transcriptome data point and the nearest neighbor spatial metabolome data point to fit the original data at the nearest neighbor spatial metabolome data point and associate it with the spatial transcriptome data point.
[0007] In a third aspect, the present invention provides an electronic device comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the method for integrating spatial transcriptome and spatial metabolome data as described in any of the aforementioned embodiments.
[0008] In a fourth aspect, the present invention provides a storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the method for integrating spatial transcriptome and spatial metabolome data as described in any of the aforementioned embodiments is implemented.
[0009] The data integration method, device, electronic device and storage medium of spatial transcriptome and spatial metabolome provided by the present invention, method comprises: obtaining spatial transcriptome data point and spatial metabolome data point after spatial coordinate alignment; determining the nearest neighbor spatial metabolome data point corresponding to each spatial transcriptome data point; utilizing the distance between the spatial transcriptome data point and the nearest neighbor spatial metabolome data point, fitting the raw data at the nearest neighbor spatial metabolome data point, and associating with the spatial transcriptome data point. The present invention first obtains the spatial transcriptome data point and spatial metabolome data point after spatial coordinate alignment, that is, it is possible to unify two different types of data points into the same spatial coordinate system, providing a reliability basis for subsequent integration, utilizing the distance between the spatial transcriptome data point and the nearest neighbor spatial metabolome data point, fitting the raw data at the nearest neighbor spatial metabolome data point, and associating with the spatial transcriptome data point, by this distance-based fitting method, it is possible to effectively make up for the data mismatch problem caused by spatial structure differences, thereby significantly improving the reliability of data integration results. At the same time, the flexibility of this method enables it to adapt to the data characteristics under different chip structures, further enhancing the compatibility of the scheme.
[0010] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without making any creative efforts.
[0012] Figure 1A A schematic diagram of the distribution of spots collected by a 15μm spatial transcriptome chip provided in an embodiment of the present invention; Figure 1B This is a schematic diagram of the distribution of basic detection unit pixels for 15 μm spatial metabolomics provided by an embodiment of the present invention; Figure 2 is a schematic flow chart of a method for integrating spatial transcriptome and spatial metabolome data provided by an embodiment of the present invention; Figure 3A This is a display diagram of the spatial transcriptome optical image provided by an embodiment of the present invention; Figure 3B This is an effect diagram of the preprocessing of the spatial transcriptome optical image provided by an embodiment of the present invention; Figure 4AThis is an effect diagram of the spatial metabolomics optical image preprocessing provided by an embodiment of the present invention; Figure 4B This is a display diagram of the spatial metabolome mass spectrometry image provided by an embodiment of the present invention; Figure 4C This is an effect diagram of the spatial metabolomics mass spectrometry image preprocessing provided by an embodiment of the present invention; Figure 5 A functional module diagram of the image registration model provided by an embodiment of the present invention; Figure 6 A schematic flow chart of an image registration process provided in an embodiment of the present invention; Figure 7 The superimposed image and the local magnified image of the registered spatial transcriptome optical image and the spatial metabolome optical image in an embodiment of the present invention are shown; Figure 8A This is a schematic diagram of the geometric relationship between the data points of the first spatial metabolome and spatial transcriptome provided by an embodiment of the present invention; Figure 8B Schematic diagram of the geometric relationship between the data points of the second spatial metabolome and spatial transcriptome provided by an embodiment of the present invention; Figure 8C Schematic diagram of the geometric relationship between the data points of the third spatial metabolome and spatial transcriptome provided by an embodiment of the present invention; Figure 9 A functional module diagram of a device for integrating spatial transcriptome and spatial metabolome data provided by an embodiment of the present invention; Figure 10 This is a structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0013] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0014] During the research, the inventors found that there are significant differences in the spatial structure distribution between the sequencing spots of spatial transcriptomics sequencing chips and the basic detection units in spatial metabolomics detection - pixels. Figure 1Aand Figure 1B As shown, Figure 1A A schematic diagram of the distribution of spots collected by a 15μm spatial transcriptome chip provided in an embodiment of the present invention; Figure 1B This is a schematic diagram of the distribution of basic detection unit pixels of 15 μm spatial metabolomics provided by an embodiment of the present invention. Figure 1A In the figure, the green squares represent the spots of spatial transcriptomics. Figure 1B In the figure, the red squares represent the pixels of the spatial metabolome.
[0015] Observations show that the pixels of the spatial metabolomics group exhibit a dense distribution in the image, while the spots of the spatial transcriptomics group exhibit a relatively sparse distribution, forming a sharp contrast with the pixel distribution of the spatial metabolomics group. This mismatch in spatial structure reduces the reliability of the integration results of the two omics data. In addition, the differences in the structure of spatial transcriptomics chips—that is, the arrangement of sequencing spots on the chip—also make existing integration technologies incompatible with different sequencing chip structures.
[0016] Considering the limitations of existing spatial transcriptomics and spatial metabolomics data integration technologies in terms of reliability and compatibility, the present invention provides a method for integrating spatial transcriptomics and spatial metabolomics data. Figure 2 , Figure 2 : is a schematic flow chart of a method for integrating spatial transcriptome and spatial metabolome data provided by an embodiment of the present invention. The method may be performed by an electronic device with data processing capabilities, and mainly includes steps S201 to S203, as described below: S201: obtaining spatial transcriptome data points and spatial metabolome data points after spatial coordinate registration; In the embodiment of the present invention, the data point of the spatial transcriptome image is the collection center of the spatial transcriptome data, that is, Figure 1A The spot point in the spatial metabolome image is the data point of the spatial metabolome data collection center, that is, Figure 1B Pixels in .
[0017] S202: Determine the nearest neighboring spatial metabolome data point corresponding to each spatial transcriptome data point.
[0018] S203: Using the distance between the spatial transcriptome data point and the nearest neighboring spatial metabolome data point, fitting the original data at the nearest neighboring spatial metabolome data point and associating it with the spatial transcriptome data point.
[0019] The data integration method of steps S101 to S103 provided by the present invention first obtains spatial transcriptome data points and spatial metabolome data points after spatial coordinate alignment. Through precise spatial coordinate alignment, the two different types of data points can be unified in the same spatial coordinate system, providing a reliability basis for subsequent data integration; then, the nearest neighbor spatial metabolome data point corresponding to each spatial transcriptome data point is determined. By finding the nearest neighbor data point, the relationship between the two data points can be more accurately matched. The distance between the spatial transcriptome data point and the nearest neighbor spatial metabolome data point is used to fit the original data at the nearest neighbor spatial metabolome data point and associate it with the spatial transcriptome data point. This integration process is carried out based on the spatial distance relationship between the spatial transcriptome and spatial metabolome data points at the same spatial coordinates, does not depend on a specific chip structure, effectively compensates for the data mismatch problem caused by spatial structure differences, and adapts to the arrangement of different sequencing chips, thereby significantly improving the reliability of the data integration results.
[0020] Next, the embodiment of the present invention will provide a detailed and clear description of the above steps S101 to S103 with reference to the relevant drawings.
[0021] In the embodiment of the present invention, with respect to step S101, in one implementation manner, existing data registration technology may be used to obtain spatial transcriptome data points and spatial metabolome data points after spatial coordinate registration.
[0022] For example, as an optional implementation, the spatial transcriptome image and the spatial metabolome image can be manually aligned, and the change matrix parameters of the two omics images, including rotation angle and scaling ratio, can be inversely calculated by evaluating the landmark offset, so as to unify the data points of the spatial transcriptome and the spatial metabolome into a unified coordinate system.
[0023] Considering that the accuracy of data registration will directly affect the accuracy of data integration results, the embodiment of the present invention also provides a new data registration method, which can improve the accuracy of data registration and provide a reliability basis for subsequent data integration methods.
[0024] In an optional implementation, the implementation of step S101 may include the following steps a1 to a3, which are described as follows: Step a1: Obtain the spatial transcriptome image and the spatial metabolome image to be registered; Step a2: perform image registration on the spatial transcriptome image to be registered and the spatial transcriptome image using an image registration model to obtain the optimal registration parameters; Step a3: Using the optimal registration parameters, the spatial coordinates of the registered spatial transcriptome image and the data points of the spatial transcriptome image are registered.
[0025] Next, the embodiments of the present invention will describe the implementation of the above steps a1 to a3 in detail and clearly with reference to the relevant drawings.
[0026] In step a1, the spatial transcriptome image and the spatial metabolome image to be registered include the following embodiments: (1) two images, a spatial transcriptome optical image and a spatial metabolome optical image; (2) a spatial transcriptome optical image and a spatial metabolome mass spectrometry image; (3) a spatial transcriptome optical image, a spatial metabolome optical image, and a spatial metabolome mass spectrometry image.
[0027] In the embodiments of the present invention, the first two implementations can directly use the two images for data registration. In the third implementation, the spatial metabolomics optical image will serve as a "bridge" for data registration, assisting in the completion of the spatial coordinate registration of the data points of the spatial transcriptomics optical image and the spatial metabolomics mass spectrometry image, achieving a more reliable and accurate registration effect.
[0028] To facilitate understanding of the process of obtaining the image to be registered in step a1, the third embodiment described above is used as an example for description below. It should be understood that the process of obtaining the image to be registered in the first and second embodiments is similar to the process of obtaining the image in the third embodiment, and will not be further described in detail in this embodiment of the present invention.
[0029] In the third embodiment, step a1 may include the following steps: Step a1-1: obtaining a spatial metabolome mass spectrometry image, a spatial metabolome optical image, and a spatial transcriptome optical image; Step a1-2: The spatial metabolome mass spectrometry image and the spatial metabolome optical image are used as a set of images to be registered, and the spatial metabolome optical image and the spatial transcriptome optical image are used as another set of images to be registered.
[0030] In step a1-1, the embodiment of the present invention can obtain the above three images based on the biological tissue slice sample, and obtain one image of each type. That is, step a1-1 can be implemented according to the following process: Step 1: Use a pathology section scanner to perform bright field scanning on two adjacent biological tissue sections, one stained and one unstained, to obtain spatial transcriptome optical images and spatial metabolome optical images; In an embodiment of the present invention, relevant personnel can select two biological tissue slice samples for slice preparation. One biological tissue slice sample is used to obtain a spatial transcriptome optical image, which is recorded as sample A; the other biological tissue slice sample is used to obtain a spatial metabolome optical image and a spatial metabolome mass spectrometry image, which is recorded as sample B.
[0031] In order to ensure that the two biological tissue slice samples have a high degree of similarity in morphology, sample A and sample B can preferably be two adjacent biological tissue slices.
[0032] The process for obtaining the spatial transcriptome optical image based on the two prepared samples is as follows: Sample A is stained and then scanned using a pathology slide scanner under brightfield conditions to acquire a high-quality optical image, namely the spatial transcriptome optical image. The process for obtaining the spatial metabolome optical image is: Sample B is mounted and dried, and then directly scanned using a pathology slide scanner under brightfield mode to acquire its optical image, thus obtaining the spatial metabolome optical image.
[0033] Step 2: Perform mass spectrometry imaging on unstained biological tissue sections to obtain spatial metabolome mass spectrometry images; In the embodiment of the present invention, based on the above-mentioned sample B, the spatial coordinate information in the spatial metabolomics data is converted into a spatial metabolomics mass spectrometry image in the form of a single-channel grayscale image at a ratio of 1:1.
[0034] Step 3: Preprocess the spatial transcriptome optical images, spatial metabolome optical images and spatial metabolome mass spectrometry images.
[0035] In the embodiment of the present invention, the pre-processing operations that can be performed on all three images include but are not limited to: cropping the blank areas around the image so that the sample in the image is located at the center of the image; and removing impurities and dirt from the surrounding area of the sample.
[0036] For the optical images of the spatial transcriptome and the spatial metabolome, the embodiment of the present invention can also convert them into single-channel grayscale images, and then downsample them according to the original image size. The downsampling multiple can be 2 times, 4 times, or 8 times, etc.
[0037] For the spatial metabolome mass spectrometry image, considering that its low resolution will affect the data registration effect, the embodiment of the present invention can also refer to the downsampling multiple of the spatial metabolome optical image to upsample the spatial metabolome mass spectrometry image. For example, if the downsampling multiple of the spatial metabolome optical image is 3 times, 5 times, 7 times, or 9 times, then the upsampling multiple of the spatial metabolome mass spectrometry image is correspondingly 3 times, 5 times, 7 times, or 9 times.
[0038] The benefits of the above preprocessing operations are that they can reduce the size of image data, improve image quality, and ensure improved operational efficiency and accuracy of the image registration model. Furthermore, considering that the image registration model in the embodiments of the present invention may fall into a "saddle point" during the optimization process, causing the model to mistakenly regard it as the optimal solution, resulting in registration accuracy failing to achieve the expected effect, the above preprocessing operations can effectively avoid this risk, ensuring that the image registration model can accurately achieve the expected accuracy.
[0039] In an optional embodiment, the various pre-processed images may be stored in, but not limited to, image formats such as png and tiff for subsequent registration.
[0040] In order to intuitively display the three images in the embodiment of the present invention, please refer to Figures 3A to 3B as well as Figures 4A to 4C , Figure 3A This is a display diagram of the spatial transcriptome optical image provided by an embodiment of the present invention; Figure 3B This is a diagram showing the effect of preprocessing the spatial transcriptome optical image provided by an embodiment of the present invention. Figure 4A This is the effect diagram after the spatial metabolomics optical image preprocessing provided by the embodiment of the present invention. Figure 4B is a display diagram of the spatial metabolome mass spectrometry image provided by an embodiment of the present invention, Figure 4C This is an effect diagram of the spatial metabolomics mass spectrometry image preprocessing provided by an embodiment of the present invention.
[0041] After obtaining the images to be registered through the above step a1-1, in step a1-2, the embodiment of the present invention uses the spatial metabolomics mass spectrometry image and the spatial metabolomics optical image as a group of images to be registered, and uses the spatial metabolomics optical image and the spatial transcriptomics optical image as another group of images to be registered.
[0042] Regarding step a2, before executing this step, embodiments of the present invention may also unify the resolutions of the spatial transcriptome image and the spatial metabolome image to be registered. This can be understood as follows: if the resolutions of the two images are different, the high-resolution image can be downsampled, or the low-resolution image can be upsampled, to bring the resolutions of the two images to the same level.
[0043] For example, for the two sets of images to be registered obtained in step a1-2, each set of images must first complete a unified resolution process. For the convenience of subsequent explanation, the embodiment of the present invention collectively refers to the spatial transcriptome optical image and the spatial metabolome optical image in step a1-2 as the "first image group", and the spatial metabolome optical image and the spatial metabolome mass spectrometry image as the "second image group". It should be made clear that this naming method is only for the sake of simplicity in language expression and is not a limitation on the relevant images or technical solutions.
[0044] In the first image set, if the resolution of the spatial transcriptome optical image is greater than the resolution of the spatial metabolome optical image, then the spatial transcriptome optical image is downsampled, or the spatial metabolome optical image is upsampled to make the resolution of the two equal. In the second image set, the resolution of the spatial metabolome mass spectrometry image is significantly lower than that of the spatial transcriptome optical image, so the spatial metabolome mass spectrometry image can be upsampled with reference to the resolution of the spatial metabolome optical image in the first image set.
[0045] In the image after sampling processing, the conversion formula of the data point coordinates can be expressed as: (1) Where t is the multiple of upsampling (or downsampling); ( , ) and (x, y) are the coordinates of the data points in the image after upsampling (or downsampling) and before upsampling (or downsampling), respectively.
[0046] Next, based on the images to be registered with unified resolution, the image registration model provided by the embodiment of the present invention can be used to automatically complete image registration and obtain optimal registration parameters, see step a2.
[0047] In step a2, the embodiment of the present invention automatically completes the image registration of the spatial transcriptome image and the spatial metabolome image through the image registration model to obtain the optimal registration parameters, including steps a2-1 to a2-3, as described below: Step a2-1: determining a transformed image and a fixed image in the spatial transcriptome image and the spatial metabolome image; In the embodiment of the present invention, a transformed image refers to an image that needs to be transformed, and a fixed image refers to an image that does not need to be transformed.
[0048] For ease of understanding, we continue to use the first and second groups of images above as examples. In the first group of images, the spatial transcriptome optical image is a fixed image, and the spatial metabolome optical image is a transformed image; in the second group of images, the spatial metabolome optical image is a fixed image, and the spatial metabolome mass spectrometry image is a transformed image. Through this configuration, the spatial metabolome optical image will serve as a "bridge" in the data alignment process, assisting in the completion of the spatial coordinate alignment of data points in the spatial transcriptome optical image and the spatial metabolome mass spectrometry image.
[0049] Step a2-2: performing multiple transformations on the transformed image using the image registration model; In an embodiment of the present invention, the transformation can be any one of the following and their combinations: translation transformation, rotation transformation, isotropic / anisotropic scaling, affine transformation, rigid transformation (Euclidean transformation), etc., to meet the diverse needs of image transformation in different scenarios.
[0050] For a transformed image, there is a clear equation relationship between the coordinates of the data points before and after the transformation. For example, as an example, this equation relationship can be expressed as: (2) in, is the coordinate matrix of the data points after transformation; x is the coordinate matrix of the data points before transformation; M is the transformation coefficient matrix; c is the coordinate matrix of the transformation center point; t is the initial translation vector, M, c and t constitute the transformation parameters in the embodiment of the present invention.
[0051] It should be understood that formula (2) and its corresponding transformation parameters are merely examples, and the specific transformation parameters output in the embodiment of the present invention depend on the specific transformation algorithm used.
[0052] Step a2-3: When it is determined that the similarity between the transformed image and the fixed image reaches a preset convergence threshold, the image registration is completed and the optimal registration parameters are output.
[0053] In the embodiment of the present invention, the image registration model can not only perform multiple transformations on the transformed image, but also evaluate the similarity between the fixed image and the transformed image after each transformation to determine whether the image registration is completed.
[0054] For example, continuing with the first image group obtained in step a1-2, after the spatial transcriptome optical image and the spatial metabolome optical image are input into the image registration model, the model performs multiple transformation operations on the spatial metabolome optical image. After each transformation, the model will evaluate the similarity between the spatial transcriptome optical image and the transformed spatial metabolome optical image, and determine whether a preset convergence threshold has been reached. If the convergence threshold is not reached, the model will transform the spatial metabolome optical image again, and iterate repeatedly until the convergence threshold is reached. At this point, the model stops performing the transformation on the transformed image and can output the current transformation parameters.
[0055] Alternatively, the convergence threshold can be flexibly set by relevant personnel based on experience. For example, when the convergence threshold is set to 0.8, if the similarity is less than 0.8, the transformation of the transformed image continues. Once the convergence threshold reaches 0.8, the transformation can be stopped and the current transformation parameters can be output.
[0056] Based on the inventive concept of the above-mentioned image registration process, the embodiment of the present invention also provides a functional structure diagram of an optional image registration model. Figure 5 , Figure 5 The functional structure diagram of the image registration model provided by the embodiment of the present invention includes a registrator, an evaluator and an optimizer. The functions of these three components are introduced in turn below.
[0057] The aligner has a built-in specific transformation algorithm, based on which the aligner can perform corresponding transformation operations on the transformed image.
[0058] The evaluator has a built-in similarity evaluation algorithm that can accurately evaluate the similarity between the fixed image and the transformed image after the transformation operation.
[0059] Optionally, the similarity evaluation algorithm may include, but is not limited to, a matrix similarity algorithm (Correlation), a least mean square (MeanSquares), a Demons algorithm, etc. Relevant personnel may flexibly select an algorithm based on specific application scenarios and image characteristics to achieve the best similarity evaluation effect.
[0060] The optimizer, a key iterative framework for the image registration model, incorporates a specific optimization algorithm. Based on pre-defined parameters such as the learning rate, convergence threshold, and number of iterations, it monitors the model's convergence in real time during each iteration. Specifically, it monitors whether the similarity reaches the preset convergence threshold. This ensures that the image registration model optimally aligns fixed and transformed images within a limited number of iterations, providing data support for the subsequent unification of data point spatial coordinates.
[0061] In the context of the optimizer, the "learning rate" refers to the magnitude or degree of transformation that the aligner applies to the transformed image. For example, when rotating an image, the learning rate might correspond to the angle of rotation; or when scaling an image, the learning rate might refer to the degree of scaling.
[0062] The optimization method built into the optimizer can be, but is not limited to, the gradient descent algorithm. Different optimization methods set different model parameters. The optimizer completes the next iteration by adjusting the model parameters after each iteration until the iteration ends.
[0063] For example, in the case of the gradient algorithm, the preset model parameters may include: initial learning rate, maximum iteration coefficient, convergence threshold, etc. These parameters can be set by relevant personnel based on actual experience until the image registration model automatically completes image registration. For example, suppose the initial value of the learning rate is set to 3.0, the maximum number of iterations is set to 300 times, and the convergence threshold is set to 0.8. During each round of iteration, the optimizer will continue to monitor whether the similarity output by the evaluator reaches 0.8. If not, the optimizer will adjust the initial value of the learning rate, and then the registrant will perform the corresponding transformation operation on the transformed image again based on the adjusted learning rate. Through continuous iterations until the similarity reaches 0.8, the image registration is completed, and the current transformation parameters are output as the optimal registration parameters.
[0064] In an optional embodiment, the optimal registration parameters can be output in the form of a file, which records in detail the optimal parameter combination such as the transformation coefficient matrix, transformation center, rotation angle, scaling ratio, etc. involved in formula (2). The specific parameter combination depends on the specific registration algorithm selected in the aligner.
[0065] based on Figure 5For an overall understanding of the working process of the image registration model provided by the embodiment of the present invention, please refer to the image registration model shown in FIG. Figure 6 , Figure 6 The workflow diagram of the image registration model provided in the embodiment of the present invention includes the following process: S1: Input the transformed image and the fixed image into the image registration model; S2: The register uses a preset transformation algorithm to transform the transformation image; S3: The evaluator calculates the similarity between the transformed image and the fixed image using a preset similarity evaluation algorithm; S4: The optimizer detects whether the similarity reaches the preset convergence threshold; If yes, execute S5; otherwise, adjust the preset learning rate and return to S2; S5: Outputting the optimal registration parameters, the transformed image under the optimal registration parameters, and the superimposed image of the transformed image under the optimal registration parameters and the fixed image.
[0066] In order to intuitively demonstrate the image registration effect in the embodiment of the present invention, the first set of images (spatial metabolome optical image and spatial transcriptome optical image) obtained in step a1-2 is still taken as an example. Figure 7 , Figure 7 The superimposed image of the spatial transcriptome optical image and the spatial metabolome optical image after image registration in the embodiment of the present invention and its local magnification are displayed. The superimposed image of the spatial transcriptome image and the spatial metabolome image can intuitively present the effect of image registration, providing relevant personnel with a convenient visualization analysis method.
[0067] Through the above step a2, the embodiment of the present invention can automatically complete the image registration of the spatial transcriptome image and the spatial metabolome image, and obtain the optimal registration parameters. Based on this information, the data points in the spatial transcriptome image and the spatial metabolome image can be quickly and accurately completed into the spatial coordinate system, see step a3.
[0068] In step a3, the data points of the spatial transcriptome image and the spatial metabolome image, after image registration, are spatially aligned using the optimal registration parameters. This involves converting the coordinates of the data points in one image to the coordinate system of the data points in the other image using the optimal registration parameters. For example, the coordinates of the data points in the spatial metabolome mass spectrometry image are converted to the coordinate system of the data points in the spatial metabolome optical image, thereby achieving spatial transcriptome and metabolome data point coordinate alignment.
[0069] For example, as one embodiment, for the first image group (spatial transcriptome optical image and spatial metabolome optical image) and the second image group (spatial metabolome optical image and spatial metabolome mass spectrometry image), the embodiment of step a3 may be: Step a3-1: performing spatial coordinate registration on the data points of the spatial metabolome mass spectrometry image and the data points of the spatial metabolome optical image according to the optimal registration parameters of the spatial metabolome mass spectrometry image and the spatial metabolome optical image; Step a3-2: performing spatial coordinate registration on the data points of the spatial metabolome mass spectrometry image and the spatial transcriptome optical image after spatial coordinate registration according to the optimal registration parameters corresponding to the spatial metabolome optical image and the spatial transcriptome optical image.
[0070] Through the above implementation, the data point coordinates of spatial metabolomics and spatial transcriptomics can be accurately unified into the same spatial plane coordinate system. At this point, the embodiment of the present invention completes the spatial coordinate alignment of the data points of spatial transcriptomics and spatial metabolomics.
[0071] It can be seen from the above steps a1 to a3 that the data registration method provided by the embodiment of the present invention has the following significant advantages: First, the image registration model designed by the embodiment of the present invention is used to realize automatic image registration, and no manual marking is required throughout the process, thereby greatly improving the registration efficiency; second, in the process of automatic image registration, the image registration model can accurately obtain the optimal transformation parameters through multiple rounds of iteration and optimization. Using these optimal image registration parameters for spatial position registration can significantly improve the registration accuracy. In addition, the registration algorithm adopted by the image registration model can be flexibly adjusted and can take into account the global and local features of the image. After the registration is completed, the model can automatically output the registration result image and key graphic transformation parameters. The entire registration process does not require manual intervention, and the registration results and effects can be quantitatively evaluated by similarity.
[0072] The data registration method provided by the embodiment of the present invention can quickly and accurately obtain the spatial transcriptome data points and the spatial metabolome data points after spatial coordinate registration and unify them into the same coordinate system.
[0073] To intuitively understand the data registration results in the embodiment of the present invention, see Figures 8A to 8C , these figures show the geometric relationship between spot points and pixel points after unifying them into the same spatial coordinate system. Figure 8A This is a first effect diagram of the geometric relationship diagram of the data points of the first spatial metabolome and spatial transcriptome provided by an embodiment of the present invention. Figure 8B : is a schematic diagram of the geometric relationship between the data points of the second spatial metabolome and spatial transcriptome provided by an embodiment of the present invention, Figure 8CSchematic diagram of the geometric relationship between the data points of the third spatial metabolome and spatial transcriptome provided by an embodiment of the present invention.
[0074] observe Figures 8A to 8C It can be seen that each spatial transcriptome spot (green square) is constantly surrounded by four spatial metabolome pixels (red squares), and these spatial metabolome pixels together form a quadrilateral structure. Figure 8A What is presented in the figure is that the spatial transcriptome spot points and the spatial metabolome pixel points are independent of each other and do not overlap. Figure 8B It depicts the scene where the spatial transcriptome spot coincides with one of the spatial metabolome pixels; Figure 8C It shows the special layout of the spatial transcriptome spot point located exactly on one edge of a regular polygon composed of multiple adjacent spatial metabolome pixel points.
[0075] It should be made clear that in Figure 8B In the figure, the red square drawn in the lower left corner overlaps with the green square and part of the red square exceeds the image boundary. This is only to more intuitively show the state of overlap of the two pixel positions, and is not a strict limitation on the relationship between the two pixel positions.
[0076] Based on the geometric relationship between the spots of the spatial transcriptome and the pixels of the spatial metabolome, the embodiment of the present invention proposes an implementation method of steps S202 to S203.
[0077] In step S202, Figures 8A to 8C As shown in the geometric relationship diagram, each spot point is surrounded by four nearest pixels. Therefore, in this embodiment of the present invention, these pixels are referred to as the "nearest neighboring spatial metabolome data points" of the spot point. In order to quickly find these nearest neighboring spatial metabolome data points, this embodiment of the present invention proposes an implementation method using a sliding window strategy to determine the nearest neighboring spatial metabolome data points corresponding to each spatial transcriptome data point, including the following process: Step b1: Determine the side length of the sliding window; In an embodiment of the present invention, the distance between the spatial transcriptome data points determines the sliding window side length. For example, as an example, the sliding window side length can be set to twice the spot point spacing. Such a setting can ensure that the pixel points corresponding to each spot point can be fully captured within the area covered by the sliding window, thereby improving the accuracy of data integration.
[0078] Step b2: traverse each spatial transcriptome data point according to the sliding window side length to obtain several candidate spatial metabolome pixel points within the sliding window side length; It can be understood that during the sliding window operation, as the sliding window gradually moves, each spot point can obtain a pixel point set, which includes several adjacent pixel points. The coordinates of all pixel points are within the range defined by the sliding window. Subsequently, the nearest neighbor pixel point of each spot point will be found from the pixel point set based on the distance.
[0079] Step b3: Calculate the distance between each candidate spatial metabolome data point and the spatial transcriptome data point; In the embodiment of the present invention, the distance may be, but is not limited to, the Euclidean distance. For example, the Euclidean distance between each candidate pixel and the spot point can be expressed by formula (3): (3) In formula (3), (x, y) is the coordinate of the spot point; ( , ) is the coordinate of the i-th candidate pixel.
[0080] Step b4: The candidate spatial metabolome data point with the smallest distance is used as the nearest neighbor spatial metabolome data point. The smallest distance here does not refer only to the spatial metabolome data point with the smallest distance to the spatial transcriptome data point, but rather to several spatial metabolome data points surrounding the spatial transcriptome. Thus, for a spatial metabolome data point with the smallest distance to a data point in a certain spatial transcriptome, the spatial metabolome data point is associated with the spatial transcriptome data point with the smallest distance. In this way, when spatial transcriptome data points are sparse and spatial metabolome data points are dense, a spatial transcriptome data point will generally be associated with multiple spatial metabolome data points.
[0081] It should be understood that using a sliding window strategy to find metabolome pixels in the target space that satisfy the aforementioned specific geometric distribution is only one example of many possible approaches. In practical applications, researchers can flexibly adopt a variety of other methods, such as clustering, image segmentation, or machine learning to achieve this effect, thereby accurately screening the nearest neighbor pixels corresponding to each spot.
[0082] Next, using the nearest neighbor pixel points corresponding to each spot point and the distances therebetween, the embodiment of the present invention can fit and correlate the original data at the nearest neighbor spatial metabolomics data points, ie, execute step S203.
[0083] In step S203 , for each nearest neighbor spatial metabolome data point, the embodiment of the present invention may use the distance between it and the spatial transcriptome data point to quantify the weight of the original data on the nearest neighbor spatial metabolome data point.
[0084] In the embodiment of the present invention, the raw data refers to all mass spectrometry response values. The greater the weight, the more important the mass spectrometry response value at the nearest neighbor spatial metabolome data point is, and the stronger the influence on the final integration result. Furthermore, the embodiment of the present invention uses the obtained weight coefficient to fit the raw data and complete data association. Therefore, the embodiment of the present invention provides the following implementation method for implementing step S203, which can include the following process: Step c1: Determine the weight coefficient of the original data using the distance corresponding to the metabolomics data points in the nearest neighbor space; In the embodiment of the present invention, the weight of the original data at each nearest neighbor pixel is determined by converting the distance into a Gaussian distance and then normalizing it. Specifically: In the embodiment of the present invention, the distances corresponding to the metabolome data points in the nearest neighbor space are first converted into Gaussian distance weights according to the following formula (4): (4) In formula (4), is the height of the normal distribution curve, the default value is 1; b is the offset of the center line of the normal distribution curve on the x-axis, the default value is 0; The half-peak width of the normal distribution curve is 2.0 by default.
[0085] Next, the Gaussian distance weights obtained by formula (4) are converted to relative values to obtain the weight coefficient. The calculation formula is as follows: (5) In formula (5), represents the Gaussian distance weight corresponding to the kth nearest neighbor spatial metabolome data point; K represents the number of nearest neighbor spatial metabolome data points.
[0086] Step c2: weighting the original data according to the weight coefficient; Step c3: Add the weighted raw data as the fitting result and associate it with the spatial transcriptome data points.
[0087] In the embodiment of the present invention, weighting the original data according to the weight coefficient can be understood as: multiplying the mass spectrometry response value at the nearest neighbor spatial metabolome data point by the weight coefficient, and then adding the weighted mass spectrometry response values to complete the fitting of the original data, and associating the fitting result with the coordinates of the spatial transcriptome data point to complete the data integration.
[0088] In summary, the data integration method proposed in the embodiment of the present invention has the following advantages: First, an embodiment of the present invention provides a data registration method, which can automatically complete the image registration of spatial transcriptome images and spatial metabolome images through an image registration model to obtain optimal registration parameters. Based on the optimal registration parameters, the spatial transcriptome data points and the spatial metabolome data points can be unified to the same coordinates to obtain the spatial transcriptome data points and spatial metabolome data points after spatial coordinate registration, which provides a reliability basis for subsequent data integration.
[0089] Secondly, the embodiment of the present invention is based on the geometric relationship between the spatial transcriptome data points and the spatial metabolome data points after spatial coordinate registration in the same spatial coordinate system, finds the nearest neighbor spatial metabolome data point corresponding to each spatial transcriptome data point, and then performs data fitting and association based on the distance between them. This nearest neighbor-based matching method can effectively reduce the error caused by differences in spatial structure distribution, further improving the reliability of data integration. At the same time, this method does not rely on a specific chip structure and can adapt to the arrangement of different sequencing chips, laying the foundation for compatibility with different chip structures.
[0090] In general, the data integration method provided by the embodiments of the present invention can not only effectively improve the reliability of data integration results, but also be compatible with scenarios with different sequencing chip structures, thereby meeting diverse integration needs.
[0091] Based on Figure 2 With the same inventive concept, the present invention also provides a data integration device 90 for spatial transcriptome and spatial metabolome. Figure 9 , Figure 9 The functional module diagram of the spatial transcriptome and spatial metabolome data integration device provided by the embodiment of the present invention includes: an acquisition module 901, a search module 902 and an integration module 903; An acquisition module 901 is used to obtain spatial transcriptome data points and spatial metabolome data points after spatial coordinate registration; Search module 902, for determining the nearest neighboring spatial metabolome data point corresponding to each spatial transcriptome data point; The integration module 903 is used to fit the original data at the nearest neighboring spatial metabolome data point using the distance between the spatial transcriptome data point and the nearest neighboring spatial metabolome pixel point, and associate it with the spatial transcriptome data point.
[0092] It is understandable that the acquisition module 901, the search module 902 and the integration module 903 can be executed in a coordinated manner. Figure 2 Each step in the process is performed to achieve the corresponding technical effects.
[0093] It should be noted that the data integration device of the spatial transcriptome and the spatial metabolome provided in the embodiment of the present invention can be specific hardware on the device or software or firmware installed on the device. The device provided in the embodiment of the present invention has the same implementation principle and technical effects as the aforementioned method embodiment. For the sake of brief description, for any part not mentioned in the device embodiment, reference can be made to the corresponding content in the aforementioned method embodiment. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can all refer to the corresponding processes in the aforementioned method embodiment, and will not be repeated here.
[0094] Optionally, the above modules can be stored in the form of software or firmware. Figure 10 The memory shown may be fixed in the operating system (OS) of the electronic device and can be executed by the processor. At the same time, the data and program codes required to execute the above modules may be stored in the memory.
[0095] See Figure 10 , Figure 10 An electronic device 100 provided in an embodiment of the present invention includes a memory 1001, a processor 1002, and a communication interface 1003. The memory 1001, processor 1002, and communication interface 1003 are electrically connected to each other, directly or indirectly, to enable data transmission or interaction. For example, these components can be electrically connected to each other via one or more communication buses or signal lines.
[0096] Optionally, the bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 10 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0097] In an embodiment of the present invention, the processor 1002 may be a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and may implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in conjunction with the embodiments of the present invention may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor. The software module may be located in the memory 1001, and the processor 1002 reads the program instructions in the memory 1001 and performs the steps of the above method in conjunction with its hardware.
[0098] In an embodiment of the present invention, the memory 1001 may be a non-volatile memory, such as a hard disk drive (HDD) or a solid-state drive (SSD), or a volatile memory (Volatile Memory), such as RAM. The memory may also be any other medium that can be used to carry or store the desired program executable code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory in an embodiment of the present invention may also be a circuit or any other device that can implement a storage function, for storing instructions and / or data.
[0099] The memory 1001 can be used to store software programs and modules, such as the instructions / modules of the spatial transcriptome and spatial metabolome data integration device 90 provided in an embodiment of the present invention. These instructions / modules can be stored in the memory 1001 in the form of software or firmware or embedded in the operating system (OS) of the electronic device 100. The processor 1002 executes the software programs and modules stored in the memory 1001 to perform various functional applications and data processing. The communication interface 1003 can be used for signaling or data communication with other node devices.
[0100] I understand. Figure 10 The structure shown is for illustration only. The electronic device 100 may also include Figure 10 More or fewer components than shown, or with Figure 10 Different configurations shown. Figure 10 The components shown may be implemented in hardware, software, or a combination thereof.
[0101] Based on the above embodiments, the present invention also provides a readable storage medium, which stores a computer program. When the computer program is executed by a computer, the computer executes the data integration method of the spatial transcriptome and spatial metabolome provided in the above embodiments. The specific implementation can be found in the method embodiment and will not be repeated here.
[0102] An embodiment of the present invention can also provide a computer program product for executing a method for integrating spatial transcriptome and spatial metabolome data, including a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method in the previous method embodiment. For specific implementation, please refer to the method embodiment and will not be repeated here.
[0103] In the embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interface, the indirect coupling or communication connection of the device or unit can be electrical, mechanical or other forms.
[0104] In addition, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0105] Furthermore, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0106] It should be noted that if the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.
Claims
1. A method for integrating spatial transcriptome and spatial metabolome data, characterized in that: The method comprises: Obtain spatial transcriptome data points and spatial metabolome data points after spatial coordinate registration; Determine the nearest neighboring spatial metabolome data point corresponding to each spatial transcriptome data point; The distance between the spatial transcriptome data point and the nearest neighbor spatial metabolome data point is used to fit the original data at the nearest neighbor spatial metabolome data point and associate it with the spatial transcriptome data point.
2. The method for integrating spatial transcriptome and spatial metabolome data according to claim 1, characterized in that: Using the distance between the spatial transcriptome data point and the nearest neighbor spatial metabolome data point, fitting the original data at the nearest neighbor spatial metabolome data point and associating it with the spatial transcriptome data point, comprising: Determining the weight coefficient of the original data using the distance corresponding to the nearest neighbor spatial metabolomics data point; weighting the original data according to the weight coefficient; The weighted raw data are added together as a fitting result, and associated with the spatial transcriptome data points.
3. The method for integrating spatial transcriptome and spatial metabolome data according to claim 2, characterized in that: Determining the weight coefficient of the original data using the distance corresponding to the metabolomics data point in the nearest neighbor space includes: Converting the distances corresponding to the metabolome data points in the nearest neighbor space into Gaussian distance weights; Perform relative value conversion on each of the Gaussian distance weights to obtain the weight coefficient.
4. The method for integrating spatial transcriptome and spatial metabolome data according to claim 2, characterized in that: The method further comprises: A sliding window strategy was used to determine the nearest neighboring spatial metabolome data point corresponding to each spatial transcriptome data point.
5. The method for integrating spatial transcriptome and spatial metabolome data according to claim 4, characterized in that: A sliding window strategy is used to determine the nearest neighboring spatial metabolome data points corresponding to each spatial transcriptome data point, including: Determine the side length of the sliding window; Traversing each of the spatial transcriptome data points according to the sliding window side length to obtain a plurality of candidate spatial metabolome pixel points located within the sliding window side length; Calculating the distance between each candidate spatial metabolome data point and the spatial transcriptome data point; The candidate spatial metabolome data point with the smallest distance is used as the nearest neighbor spatial metabolome data point.
6. The method for integrating spatial transcriptome and spatial metabolome data according to claim 5, characterized in that: Determine the sliding window side length, including: The sliding window side length is determined based on the spacing between the spatial transcriptome data points.
7. The method for integrating spatial transcriptome and spatial metabolome data according to any one of claims 1 to 6, characterized in that: Obtain spatial transcriptome data points and spatial metabolome data points after spatial coordinate registration, including: Obtaining the spatial transcriptome image and the spatial metabolome image to be registered; Performing image registration on the spatial transcriptome image to be registered and the spatial transcriptome image using an image registration model to obtain optimal registration parameters; The optimal registration parameters are used to perform spatial coordinate registration on the registered spatial transcriptome image and the data points of the spatial transcriptome image.
8. A data integration device for spatial transcriptome and spatial metabolome, characterized in that: include: Acquisition module, search module and integration module; The acquisition module is used to obtain spatial transcriptome data points and spatial metabolome data points after spatial coordinate registration; The search module is used to determine the nearest neighboring spatial metabolome data point corresponding to each spatial transcriptome data point; The integration module is used to fit the original data at the nearest neighboring spatial metabolome data point using the distance between the spatial transcriptome data point and the nearest neighboring spatial metabolome data point, and associate the data with the spatial transcriptome data point.
9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to execute the computer program to implement the method for integrating spatial transcriptome and spatial metabolome data according to any one of claims 1 to 7.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for integrating spatial transcriptome and spatial metabolome data according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Multi-modal integration analysis method based on spatial multi-omics data alignment
CN121258929A