Projection method for high-dimensional biological data, electronic device, storage medium

By calculating the difference and similarity of high-dimensional biological data, and combining low-dimensional structural data with cluster analysis, continuous projection data is generated, which solves the problem of insufficient temporal continuity in the projection of high-dimensional biological data and realizes the coherent analysis of data in time series.

CN120705787BActive Publication Date: 2025-11-11TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511205880.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-11-11
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Existing technologies struggle to maintain the temporal continuity of high-dimensional biological data, resulting in a lack of continuity between projected data.

Method used

By calculating the difference and similarity between high-dimensional static biological data, and combining low-dimensional structural data with cluster analysis, continuous projected data is generated. Distribution structure characteristics and difference constraints are introduced to ensure temporal continuity.

Benefits of technology

It achieves temporal coherence between high-dimensional biological data projection data, enabling more accurate analysis of data changes over time and reflecting the characteristics of the data itself.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705787B_ABST
    Figure CN120705787B_ABST
Patent Text Reader

Abstract

This invention provides a projection method, electronic device, and storage medium for high-dimensional biological data, which can be applied to the field of data analysis technology. The method includes: for the nth frame of high-dimensional static biological data in N frames of high-dimensional static biological data, obtaining the (n-1)th low-dimensional structural data based on the (n-1)th projection data; obtaining the nth high-dimensional difference degree based on the (n-1)th and nth frames of high-dimensional static biological data; obtaining the nth projection data based on the nth frame of high-dimensional static biological data, the (n-1)th low-dimensional structural data, and the nth high-dimensional difference degree; and obtaining the projection data of the high-dimensional biological data based on the first to the Nth projection data. This ensures that during the entire high-dimensional biological data projection process, the categories divided in the nth projection data remain continuous with the categories divided in the (n-1)th projection data, reflecting the characteristics of the nth frame of high-dimensional static biological data itself, thereby guaranteeing the temporal continuity of the obtained projection data of the high-dimensional biological data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data analysis technology, and more specifically to a projection method, electronic device, and storage medium for high-dimensional biological data. Background Technology

[0002] High-dimensional biological data contains rich and complex information, and mining high-dimensional biological data is of great significance in practical production.

[0003] In related technologies, projection methods for high-dimensional static biological data can be used to obtain projection data of high-dimensional biological data, but this method is difficult to maintain the temporal continuity between two adjacent frames of high-dimensional static biological data. Summary of the Invention

[0004] In view of the above problems, the present invention provides a projection method, electronic device, and storage medium for high-dimensional biological data.

[0005] According to a first aspect of the present invention, a projection method for high-dimensional biological data is provided, wherein the high-dimensional biological data includes N frames of high-dimensional static biological data, each frame of high-dimensional static biological data including high-dimensional data corresponding to a plurality of high-dimensional data points, and N is an integer greater than or equal to 1; for the nth frame of high-dimensional static biological data in the above N frames of high-dimensional static biological data, based on the (n-1)th projection data, the (n-1)th low-dimensional structural data is obtained, wherein the (n-1)th projection data represents the projection data of the (n-1)th frame of high-dimensional static biological data, the (n-1)th low-dimensional structural data represents the distribution structure characteristics of the (n-1)th projection data, and the first projection data represents the projection data of the first frame of high-dimensional static biological data, n∈{ 2, ..., N}; Based on the above-mentioned (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data, the nth high-dimensional difference degree is obtained, wherein the above-mentioned nth high-dimensional difference degree characterizes the degree of difference between the (n-1)th high-dimensional data point group and the nth high-dimensional data point group, the above-mentioned (n-1)th high-dimensional data point group includes high-dimensional data points in the above-mentioned (n-1)th frame of high-dimensional static biological data, and the above-mentioned nth high-dimensional data point group includes high-dimensional data points in the above-mentioned nth frame of high-dimensional static biological data; Based on the above-mentioned nth frame of high-dimensional static biological data, the above-mentioned (n-1)th low-dimensional structural data and the above-mentioned nth high-dimensional difference degree, the nth projection data is obtained; Based on the above-mentioned first projection data to the Nth projection data, the projection data of the above-mentioned high-dimensional biological data is obtained.

[0006] According to an embodiment of the present invention, the aforementioned nth high-dimensional difference includes at least one of the following: nth high-dimensional distance difference and nth similarity distinguishability; the aforementioned obtaining the nth high-dimensional difference based on the aforementioned (n-1)th frame of high-dimensional static biological data and the aforementioned nth frame of high-dimensional static biological data includes: obtaining the nth high-dimensional distance difference based on the aforementioned (n-1)th frame of high-dimensional static biological data and the aforementioned nth frame of high-dimensional static biological data, wherein the aforementioned nth high-dimensional distance difference characterizes every two high-dimensional data points in the aforementioned (n-1)th high-dimensional data point group. The degree of difference between the first high-dimensional distance between two data points and the second high-dimensional distance between two identical high-dimensional data points in the above-mentioned nth high-dimensional data point group; the nth similarity distinguishability is obtained based on the above-mentioned (n-1)th frame of high-dimensional static biological data and the above-mentioned nth frame of high-dimensional static biological data, wherein the above-mentioned nth similarity distinguishability characterizes the degree of difference between the first nearest neighbor high-dimensional data point group corresponding to each high-dimensional data point in the above-mentioned (n-1)th high-dimensional data point group and the second nearest neighbor high-dimensional data point group corresponding to each high-dimensional data point in the above-mentioned nth high-dimensional data point group.

[0007] According to an embodiment of the present invention, the method for obtaining the nth high-dimensional distance difference based on the (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data includes: for every two high-dimensional data points in the (n-1)th frame of high-dimensional static biological data, calculating the distance between each pair of high-dimensional data points based on their respective (n-1)th high-dimensional data to obtain the (n-1)th high-dimensional distance; for every two high-dimensional data points in the nth frame of high-dimensional static biological data, calculating the distance between each pair of high-dimensional data points based on their respective (n)th high-dimensional data to obtain the nth high-dimensional distance; obtaining the nth sub-high-dimensional distance difference based on the difference between the (n-1)th high-dimensional distance and the nth high-dimensional distance; and obtaining the nth high-dimensional distance difference based on multiple nth sub-high-dimensional distance differences.

[0008] According to an embodiment of the present invention, the above-mentioned similarity distinguishability based on the above-mentioned (n-1)th frame of high-dimensional static biological data and the above-mentioned nth frame of high-dimensional static biological data includes: for any high-dimensional data point among the multiple high-dimensional data points included in the above-mentioned (n-1)th frame of high-dimensional static biological data, based on the above-mentioned high-dimensional data point and the (n-1)th other high-dimensional data points respectively, calculating the first distance between the above-mentioned high-dimensional data point and the above-mentioned (n-1)th other high-dimensional data points, and obtaining the (n-1)th nearest neighbor high-dimensional data point group of the above-mentioned high-dimensional data point, wherein the first distance between the above-mentioned high-dimensional data point and the (n-1)th nearest neighbor high-dimensional data points included in the above-mentioned (n-1)th nearest neighbor high-dimensional data point group is one of I first distances, the above-mentioned I first distances are the first I obtained by arranging the multiple above-mentioned first distances in ascending order, I is an integer greater than 1, and the above-mentioned (n-1)th other high-dimensional data points are other high-dimensional data points in the above-mentioned (n-1)th frame of high-dimensional static biological data excluding the above-mentioned high-dimensional data point; for the above-mentioned... For each high-dimensional data point in n frames of high-dimensional static biological data, based on the nth high-dimensional data point and the nth other high-dimensional data points, a second distance is calculated between the high-dimensional data point and the nth other high-dimensional data points to obtain the nth nearest neighbor high-dimensional data point group. The second distance between the high-dimensional data point and the nth nearest neighbor high-dimensional data points included in the nth nearest neighbor high-dimensional data point group is one of I second distances. These I second distances are the first I obtained by arranging multiple second distances in ascending order. The nth other high-dimensional data points are the other high-dimensional data points in the nth frame of high-dimensional static biological data besides the high-dimensional data point. The difference neighbor high-dimensional data points between the (n-1)th nearest neighbor high-dimensional data point group and the nth nearest neighbor high-dimensional data point group of the high-dimensional data point are determined to obtain the nth sub-similarity discriminant degree of the high-dimensional data point. Based on the nth sub-similarity discriminant degrees of the multiple high-dimensional data points, the nth similarity discriminant degree is obtained.

[0009] According to an embodiment of the present invention, obtaining the (n-1)th low-dimensional structural data based on the (n-1)th projection data includes: clustering the (n-1)th projection data to generate K clusters, where K is an integer greater than or equal to 1; performing secondary clustering on any of the K clusters to generate M sub-clusters, where M is an integer greater than or equal to 1; and for each low-dimensional data point in the (n-1)th projection data, calculating the clustering with the K clusters based on the low-dimensional data of the low-dimensional data points. The distance to the cluster centers is used to obtain K (n-1)th first sub-low-dimensional structure data points; the distance to the cluster centers of the M sub-clusters is calculated based on the low-dimensional data of each of the above low-dimensional data points to obtain M (n-1)th second sub-low-dimensional structure data points; based on preset weights, the above K (n-1)th first sub-low-dimensional structure data points and the above M (n-1)th second sub-low-dimensional structure data points are concatenated to obtain the (n-1)th target sub-low-dimensional structure data points; based on the above (n-1)th target sub-low-dimensional structure data points, the above (n-1)th low-dimensional structure data points are obtained.

[0010] According to an embodiment of the present invention, obtaining the nth projection data based on the nth frame high-dimensional static biological data, the (n-1)th low-dimensional structural data, and the nth high-dimensional difference includes: stitching the (n-1)th low-dimensional structural data and the nth high-dimensional difference to obtain the nth spatial structural constraint, wherein the nth spatial structural constraint is used to constrain the temporal coherence of each projection structure between the (n-1)th projection data and the nth projection data; stitching the nth spatial structural constraint and the nth frame high-dimensional static biological data to obtain the target nth frame high-dimensional static biological data; and obtaining the nth projection data based on the target nth frame high-dimensional static biological data.

[0011] According to an embodiment of the present invention, the first projection data is obtained by the following method: for any high-dimensional data point among the plurality of high-dimensional data points included in the first frame of high-dimensional static biological data, the similarity between the high-dimensional data points is calculated based on the first high-dimensional data of each of the plurality of high-dimensional data points to obtain a high-dimensional similarity probability matrix; for each low-dimensional data point in the initial projection data, the similarity between the low-dimensional data points is calculated based on the low-dimensional data of each of the plurality of low-dimensional data points to obtain a low-dimensional similarity probability matrix; the relative entropy between the high-dimensional similarity probability matrix and the low-dimensional similarity probability matrix is ​​calculated; the low-dimensional data points in the initial projection data are adjusted based on the relative entropy until the relative entropy satisfies a first preset condition to obtain the first projection data.

[0012] According to an embodiment of the present invention, the first projection data is obtained by the following method: For each high-dimensional data point in the first frame of high-dimensional static biological data, based on the first high-dimensional data of the high-dimensional data point and the first other high-dimensional data points, the distance between the high-dimensional data point and the first other high-dimensional data points is calculated to obtain a first nearest neighbor high-dimensional data point group; based on the first high-dimensional data of each of the multiple high-dimensional data points included in the first nearest neighbor high-dimensional data point group, the fuzzy similarity weight between the high-dimensional data points is calculated to obtain a high-dimensional fuzzy similarity matrix; for each low-dimensional data point in the initial projection data, based on the low-dimensional data of each of the multiple low-dimensional data points, the fuzzy similarity between every two low-dimensional data points is calculated to obtain a low-dimensional fuzzy similarity matrix; the cross-entropy loss value between the high-dimensional fuzzy similarity matrix and the low-dimensional fuzzy similarity matrix is ​​calculated; the low-dimensional data points in the initial projection data are adjusted based on the cross-entropy loss value until the loss value satisfies a second preset condition to obtain the first projection data.

[0013] A second aspect of the present invention provides a projection device for high-dimensional biological data, wherein the high-dimensional biological data includes N frames of high-dimensional static biological data, each frame of high-dimensional static biological data including high-dimensional data corresponding to multiple high-dimensional data points, where N is an integer greater than or equal to 1. The device includes: a low-dimensional structure data determination module, used to obtain (n-1)th low-dimensional structure data based on (n-1)th projection data for the nth frame of high-dimensional static biological data among the N frames of high-dimensional static biological data, wherein the (n-1)th projection data represents the projection data of the (n-1)th frame of high-dimensional static biological data, the (n-1)th low-dimensional structure data represents the distribution structure characteristics of the (n-1)th projection data, and the first projection data represents the projection data of the first frame of high-dimensional static biological data, where n∈{2, ..., N}; high-dimensional... The difference determination module is used to obtain the nth high-dimensional difference degree based on the above-mentioned (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data, wherein the above-mentioned nth high-dimensional difference degree characterizes the degree of difference between the (n-1)th high-dimensional data point group and the nth high-dimensional data point group, the above-mentioned (n-1)th high-dimensional data point group includes high-dimensional data points in the above-mentioned (n-1)th frame of high-dimensional static biological data, and the above-mentioned nth high-dimensional data point group includes high-dimensional data points in the above-mentioned nth frame of high-dimensional static biological data; the projection data determination module is used to obtain the nth projection data based on the above-mentioned nth frame of high-dimensional static biological data, the above-mentioned (n-1)th low-dimensional structural data and the above-mentioned nth high-dimensional difference degree; the dynamic projection module is used to obtain the projection data of the above-mentioned high-dimensional biological data based on the above-mentioned first projection data to the Nth projection data.

[0014] A third aspect of the present invention provides an electronic device comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method described above.

[0015] A fourth aspect of the present invention also provides a computer-readable storage medium having a computer program or instructions stored thereon, wherein the computer program or instructions, when executed by a processor, implement the steps of the above-described method.

[0016] A fifth aspect of the present invention also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described method.

[0017] According to an embodiment of the present invention, for high-dimensional static biological data that is not in the first frame, low-dimensional structural data of the (n-1)th frame is obtained through the (n-1)th projection data, thereby introducing the distribution structure characteristics of the (n-1)th projection data into the generated nth projection data, thus constraining the nth projection data. The difference between the (n-1)th and nth frames of high-dimensional static biological data is used to introduce the degree of difference between the (n-1)th and nth high-dimensional data point groups, thereby providing constraints on the nth projection data that are more in line with the variation characteristics of high-dimensional biological data, thus reflecting the characteristics of the nth frame of high-dimensional static biological data itself, thereby ensuring the temporal continuity between the first projection data and the Nth projection data. Attached Figure Description

[0018] The above-described features, other objects, and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which:

[0019] Figure 1 The diagram illustrates an application scenario of a projection method and apparatus for high-dimensional biological data according to an embodiment of the present invention.

[0020] Figure 2 A flowchart of a projection method for high-dimensional biological data according to an embodiment of the present invention is shown;

[0021] Figure 3 A schematic diagram of multi-layer clustering according to an embodiment of the present invention is shown;

[0022] Figure 4 The first projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown;

[0023] Figure 5 The second projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown;

[0024] Figure 6The third projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown;

[0025] Figure 7 The fourth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown;

[0026] Figure 8 The fifth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown;

[0027] Figure 9 The sixth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown;

[0028] Figure 10 The seventh projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown;

[0029] Figure 11 The eighth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown;

[0030] Figure 12 The ninth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown;

[0031] Figure 13 The tenth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown;

[0032] Figure 14 This diagram illustrates the trajectory results based on frame number coloring from the first projection data to the tenth projection data according to an embodiment of the present invention.

[0033] Figure 15 A trajectory result diagram based on cluster label coloring from the first projection data to the tenth projection data is shown according to an embodiment of the present invention.

[0034] Figure 16 The image shows the trajectory results of the static projection method based on correlation technology, from the first projection data to the tenth projection data, colored based on the frame number.

[0035] Figure 17 The image shows the trajectory results of the static projection method based on correlation technology from the first projection data to the tenth projection data based on cluster label coloring;

[0036] Figure 18 A flowchart illustrating the projection process for high-dimensional biological data according to an embodiment of the present invention is shown;

[0037] Figure 19 A structural block diagram of a projection device for high-dimensional biological data according to an embodiment of the present invention is shown;

[0038] Figure 20 A block diagram of an electronic device suitable for implementing a projection method for high-dimensional biological data according to an embodiment of the present invention is shown. Detailed Implementation

[0039] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the invention. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the invention for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.

[0040] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0041] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.

[0042] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).

[0043] High-dimensional biological data contains rich and complex information. Mining high-dimensional biological data has significant implications for practical production. Visualization is a comprehensive discipline combining computer science, human-computer interaction, and psychology. It studies how to visually represent data in a graphical way that closely resembles human perception, enhancing human perception and revealing the hidden information and patterns within the data. Exploring the visualization of high-dimensional biological data helps people more quickly extract information from the data and obtain useful insights.

[0044] Projection methods are widely used in high-dimensional data visualization. Applying projection methods to the visualization of high-dimensional biological data has become one of the hot topics in the field of visualization.

[0045] In related technologies, projection methods for high-dimensional static biological data can be used to obtain projected data of high-dimensional biological data. The high-dimensional biological data may include N frames of high-dimensional static biological data. The projected data of the high-dimensional biological data may include first projected data to Nth projected data.

[0046] In the process of implementing this invention, it was found that the projection data obtained based on the above method is difficult to maintain the temporal continuity between the (n-1)th projection data and the nth projection data. The nth projection data is the projection data of the nth frame of high-dimensional static biological data. The (n-1)th projection data is the projection data of the (n-1)th frame of high-dimensional static biological data.

[0047] In view of this, embodiments of the present invention provide a projection method for high-dimensional biological data, wherein the high-dimensional biological data includes N frames of high-dimensional static biological data, each frame of high-dimensional static biological data includes high-dimensional data corresponding to multiple high-dimensional data points, and N is an integer greater than or equal to 1; for the nth frame of high-dimensional static biological data in the N frames of high-dimensional static biological data, based on the (n-1)th projection data, the (n-1)th low-dimensional structural data is obtained, wherein the (n-1)th projection data represents the projection data of the (n-1)th frame of high-dimensional static biological data, the (n-1)th low-dimensional structural data represents the distribution structure characteristics of the (n-1)th projection data, and the first projection data represents the projection of the first frame of high-dimensional static biological data. Image data, n∈{2,……,N}; Based on the (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data, the nth high-dimensional difference degree is obtained, where the nth high-dimensional difference degree characterizes the degree of difference between the (n-1)th high-dimensional data point group and the nth high-dimensional data point group. The (n-1)th high-dimensional data point group includes high-dimensional data points in the (n-1)th frame of high-dimensional static biological data, and the nth high-dimensional data point group includes high-dimensional data points in the nth frame of high-dimensional static biological data. Based on the nth frame of high-dimensional static biological data, the (n-1)th low-dimensional structural data, and the nth high-dimensional difference degree, the nth projection data is obtained; Based on the first projection data to the Nth projection data, the projection data of the high-dimensional biological data is obtained.

[0048] Figure 1 The diagram illustrates an application scenario of a projection method and apparatus for high-dimensional biological data according to an embodiment of the present invention.

[0049] like Figure 1 As shown, application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0050] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).

[0051] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0052] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0053] It should be noted that the projection method for high-dimensional biological data provided in this embodiment of the invention can generally be executed by server 105. Correspondingly, the projection device for high-dimensional biological data provided in this embodiment of the invention can generally be located in server 105. The projection method for high-dimensional biological data provided in this embodiment of the invention can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the projection device for high-dimensional biological data provided in this embodiment of the invention can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105.

[0054] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0055] The following will be based on Figure 1 The described scene, through Figures 2-19 The projection method for high-dimensional biological data according to the embodiments of the invention will be described in detail.

[0056] Figure 2A flowchart of a projection method for high-dimensional biological data according to an embodiment of the present invention is shown.

[0057] like Figure 2 As shown, the projection method for high-dimensional biological data in this embodiment includes operations S210 to S240.

[0058] In operation S210, for the nth frame of high-dimensional static biological data in N frames of high-dimensional static biological data, the (n-1)th low-dimensional structural data is obtained based on the (n-1)th projection data.

[0059] Wherein, the (n-1)th projection data represents the projection data of the (n-1)th frame of high-dimensional static biological data, the (n-1)th low-dimensional structural data represents the distribution structure characteristics of the (n-1)th projection data, and the first projection data represents the projection data of the first frame of high-dimensional static biological data, n∈{2,……,N}.

[0060] According to embodiments of the present invention, the distribution structure features can characterize the distribution characteristics of each low-dimensional data point in the (n-1)th projection data. These distribution characteristics can be, for example, the cluster to which the low-dimensional data points belong. The first projection data is obtained by projecting the first frame of high-dimensional static biological data using a static projection method. This static projection method can be, for example, t-SNE (t-Stochastic Neighbor Embedding) or UMAP (Uniform Manifold Approximation and Projection), or other static projection methods.

[0061] In operation S220, the high-dimensional difference is obtained based on the (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data.

[0062] Among them, the nth high-dimensional difference characterizes the degree of difference between the (n-1)th high-dimensional data point group and the nth high-dimensional data point group. The (n-1)th high-dimensional data point group includes high-dimensional data points in the (n-1)th frame of high-dimensional static biological data, and the nth high-dimensional data point group includes high-dimensional data points in the nth frame of high-dimensional static biological data.

[0063] In operation S230, the nth projection data is obtained based on the nth frame of high-dimensional static biological data, the (n-1)th low-dimensional structural data, and the nth high-dimensional difference.

[0064] In operation S240, projection data of high-dimensional biological data is obtained based on the first projection data to the Nth projection data.

[0065] According to an embodiment of the present invention, a continuous projection of high-dimensional biological data can be obtained by using the first projection data and the second to Nth projection data obtained by the above method.

[0066] According to an embodiment of the present invention, for high-dimensional static biological data that is not in the first frame, low-dimensional structural data of the (n-1)th frame is obtained through the (n-1)th projection data, thereby introducing the distribution structure characteristics of the (n-1)th projection data into the generated nth projection data, thus constraining the nth projection data. The difference between the (n-1)th and nth frames of high-dimensional static biological data is used to introduce the degree of difference between the (n-1)th and nth high-dimensional data point groups, thereby providing constraints on the nth projection data that are more in line with the variation characteristics of high-dimensional biological data, thus reflecting the characteristics of the nth frame of high-dimensional static biological data itself, thereby ensuring the temporal continuity between the first projection data and the Nth projection data.

[0067] According to embodiments of the present invention, the nth high-dimensional difference degree includes at least one of the following: nth high-dimensional distance difference degree and nth similarity distinguishability degree; the nth high-dimensional difference degree is obtained based on the (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data, including: obtaining the nth high-dimensional distance difference degree based on the (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data, wherein the nth high-dimensional distance difference degree characterizes the degree of difference between the first high-dimensional distance between every two high-dimensional data points in the (n-1)th high-dimensional data point group and the second high-dimensional distance between the same two high-dimensional data points in the nth high-dimensional data point group; the nth similarity distinguishability degree is obtained based on the (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data, wherein the nth similarity distinguishability degree characterizes the degree of difference between the first nearest neighbor high-dimensional data point group corresponding to each high-dimensional data point in the (n-1)th high-dimensional data point group and the second nearest neighbor high-dimensional data point group corresponding to each high-dimensional data point in the nth high-dimensional data point group.

[0068] According to embodiments of the present invention, by calculating the nth high-dimensional distance difference and the nth similarity difference, the changes of high-dimensional static biological data between consecutive frames can be presented in a quantitative manner, thereby enabling more accurate analysis of the evolution of high-dimensional biological data in time series and making the changes of high-dimensional biological data more intuitive.

[0069] According to an embodiment of the present invention, the nth high-dimensional distance difference is obtained based on the (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data, including: for every two high-dimensional data points in the (n-1)th frame of high-dimensional static biological data, the distance between each two high-dimensional data points is calculated based on their respective (n-1)th high-dimensional data to obtain the (n-1)th high-dimensional distance; for every two high-dimensional data points in the nth frame of high-dimensional static biological data, the distance between each two high-dimensional data points is calculated based on their respective (n)th high-dimensional data to obtain the nth high-dimensional distance; the nth sub-high-dimensional distance difference is obtained based on the difference between the (n-1)th high-dimensional distance and the nth high-dimensional distance; and the nth high-dimensional distance difference is obtained based on multiple nth sub-high-dimensional distance differences.

[0070] According to an embodiment of the present invention, the distance between every two high-dimensional data points can be Euclidean distance.

[0071] For example, there are three high-dimensional data points A, B, and C. The high-dimensional data corresponding to each of the three high-dimensional data points includes location information. Then, for high-dimensional data point A, the first high-dimensional distance is the distance ab between high-dimensional data point A and high-dimensional data point B and the distance ac between high-dimensional data point A and high-dimensional data point C in the (n-1)th frame of high-dimensional static biological data. The second high-dimensional distance is the distance ab' between high-dimensional data point A and high-dimensional data point B and the distance ac' between high-dimensional data point A and high-dimensional data point C in the nth frame of high-dimensional static biological data. The difference in the nth high-dimensional distance can be expressed as (ab-ab')+(ac-ac'), where ab-ab' is the difference in the nth sub-high-dimensional distance between high-dimensional data point A and high-dimensional data point B, and ac-ac' is the difference in the nth sub-high-dimensional distance between high-dimensional data point A and high-dimensional data point B.

[0072] According to an embodiment of the present invention, the (n-1)th high-dimensional distance is obtained by calculating the distance between every two high-dimensional data points in the (n-1)th frame of high-dimensional static biological data, and the nth high-dimensional distance is obtained by calculating the distance between every two high-dimensional data points in the nth frame of high-dimensional static biological data. The difference between the (n-1)th and nth high-dimensional distances is then used to obtain the nth sub-high-dimensional distance difference, and thus the nth high-dimensional distance difference is obtained. This can capture and quantify the changes in high-dimensional biological data in the time dimension, thereby constraining the nth projection data based on the nth high-dimensional distance difference between two adjacent frames of high-dimensional static biological data, thereby maintaining the consistency of the projection data between two adjacent frames.

[0073] According to an embodiment of the present invention, based on the (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data, the nth similarity distinguishability is obtained, including: for any high-dimensional data point among the multiple high-dimensional data points included in the (n-1)th frame of high-dimensional static biological data, based on the (n-1)th high-dimensional data of the high-dimensional data point and the (n-1)th other high-dimensional data points respectively, calculating the first distance between the high-dimensional data point and the (n-1)th other high-dimensional data points, obtaining the (n-1)th nearest neighbor high-dimensional data point group of the high-dimensional data point, wherein the first distance between the high-dimensional data point and the (n-1)th nearest neighbor high-dimensional data points included in the (n-1)th nearest neighbor high-dimensional data point group is one of I first distances, I first distances are the first I obtained by arranging multiple first distances in ascending order, I is an integer greater than 1, and the (n-1)th other high-dimensional data points are the other high-dimensional data points in the (n-1)th frame of high-dimensional static biological data excluding the high-dimensional data point; for the nth frame of high-dimensional static biological data, the first distance between the high-dimensional data point and the (n-1)th nearest neighbor high-dimensional data point group is obtained, and the first distance between the high-dimensional data point and the (n-1)th nearest neighbor high-dimensional data point group is obtained, and the first distance between the high-dimensional data point and the (n-1)th other high-dimensional data points are the other high-dimensional data points in the (n-1)th frame of high-dimensional static biological data excluding the high-dimensional data point; for the (n-1)th frame of high-dimensional static biological data, the first distance between the high-dimensional data point and the (n-1)th other high-dimensional data points are obtained, and the first distance between the high-dimensional data point and the (n-1)th other high-dimensional data points are obtained, and the first distance between the high-dimensional For each high-dimensional data point in n frames of high-dimensional static biological data, based on the nth high-dimensional data point and the nth other high-dimensional data points, the second distance between the high-dimensional data point and the nth other high-dimensional data points is calculated to obtain the nth nearest neighbor high-dimensional data point group of the high-dimensional data point. The second distance between the high-dimensional data point and the nth nearest neighbor high-dimensional data points included in the nth nearest neighbor high-dimensional data point group is one of I second distances, which are the first I obtained by arranging multiple second distances in ascending order. The nth other high-dimensional data points are the other high-dimensional data points in the nth frame of high-dimensional static biological data besides the high-dimensional data point. The (n-1)th nearest neighbor high-dimensional data point group and the nth nearest neighbor high-dimensional data point group of the high-dimensional data point are determined to be the difference nearest neighbor high-dimensional data points, thus obtaining the nth sub-similarity discriminant degree of the high-dimensional data point. Based on the nth sub-similarity discriminant degrees of multiple high-dimensional data points, the nth similarity discriminant degree is obtained.

[0074] According to an embodiment of the present invention, the k-nearest neighbor method can be used to calculate the I (n-1)th nearest neighbor high-dimensional data points and the I nth nearest neighbor high-dimensional data points in the (n-1)th frame of high-dimensional static biological data for the same high-dimensional data point, and calculate the k-nearest neighbor similarity between the I (n-1)th nearest neighbor high-dimensional data points and the I nth nearest neighbor high-dimensional data points to obtain the nth sub-similarity discriminability for the high-dimensional data point, and then obtain the nth similarity discriminability.

[0075] For example, there are six high-dimensional data points A, B, C, D, E, and F. For demonstration purposes, only the nearest high-dimensional data point groups of A and E are calculated here. In practice, the calculation should be based on the nearest high-dimensional data point groups of A, B, C, D, E, and F respectively. In the (n-1)th frame of high-dimensional static biological data, for high-dimensional data point A, the (n-1)th nearest high-dimensional data point group is [B, C], and for high-dimensional data point E, the (n-1)th nearest high-dimensional data point group is [D, F]. In the nth frame of high-dimensional static biological data... For a high-dimensional data point A, the nth nearest neighbor high-dimensional data point group is [B,D], and for a high-dimensional data point E, the nth nearest neighbor high-dimensional data point group is [D,F]. Then, for high-dimensional data point A, the calculated sub-similarity discriminant is 1 (as long as the (n-1)th nearest neighbor high-dimensional data point group and the nth nearest neighbor high-dimensional data point group are different, it is counted as 1). For high-dimensional data point E, the calculated sub-similarity discriminant is 0. Further, the calculated nth similarity discriminant is (1+0) / 2=0.5.

[0076] According to an embodiment of the present invention, obtaining the (n-1)th low-dimensional structural data based on the (n-1)th projection data includes: clustering the (n-1)th projection data to generate K clusters, where K is an integer greater than or equal to 1; performing secondary clustering on any cluster among the K clusters to generate M sub-clusters, where M is an integer greater than or equal to 1; and for each low-dimensional data point in the (n-1)th projection data, calculating the clustering mean of the low-dimensional data points and the K clusters. The distance to the center is used to obtain K (n-1)th first sub-low-dimensional structure data; the distance to the cluster center of M sub-clusters is calculated based on the low-dimensional data of each of the multiple low-dimensional data points to obtain M (n-1)th second sub-low-dimensional structure data; based on preset weights, the K (n-1)th first sub-low-dimensional structure data and the M (n-1)th second sub-low-dimensional structure data are concatenated to obtain the (n-1)th target sub-low-dimensional structure data; based on the multiple (n-1)th target sub-low-dimensional structure data, the (n-1)th low-dimensional structure data is obtained.

[0077] According to an embodiment of the present invention, the (n-1)th projected data can be clustered using the K-Means method to obtain K clusters. Then, the Euclidean distance from each low-dimensional data point in the projected data to the centers of the K clusters is calculated. For each low-dimensional data point in the projected data, there is a corresponding K-dimensional vector representing the distance between each low-dimensional data point and the centers of the K clusters. The K-Means method requires specifying the number of clusters K, which can be determined using the elbow method. The elbow method is a heuristic method that determines the optimal value of K by plotting the relationship between the sum of squared errors and the number of clusters K. In this embodiment of the invention, K can be set to 5. When the data scale increases, the value of K can be appropriately increased; the present invention does not impose a limitation on the value of K.

[0078] According to embodiments of the present invention, the value of M can be the same as the value of K, which can both be set to 5 in this invention, or it can be different from the value of K. For any of the K clusters, secondary clustering is performed to generate M sub-clusters. This hierarchical clustering operation is inspired by the spatial pyramid concept. The spatial pyramid is a feature representation method widely used in image processing and computer vision. Its core idea is to capture spatial features at different scales by refining the spatial division of the image through multiple levels, and to fuse these features to enhance expressive power. Specifically, it recursively divides the image by progressively reducing the region size, capturing information at different levels from global to local. It can further cluster the sub-clusters again. However, as the number of iterations increases, the gain for temporal coherence decreases, and as the number of pyramid layers increases, the gain for capturing image structure also decreases. Therefore, the number of pyramid layers generally does not exceed 3 layers. In this invention, 2-3 clustering operations are generally sufficient.

[0079] Figure 3 A schematic diagram of multi-layer clustering according to an embodiment of the present invention is shown.

[0080] like Figure 3 As shown, for a low-dimensional data point yj in the projection data, five clusters are obtained in the first-level clustering: cluster 1, cluster 2, cluster 3, cluster 4, and cluster 5. The distances from yj to the center points of the five clusters are calculated to obtain five (n-1)th first sub-low-dimensional structure data. In the second-level clustering, a second clustering is performed in cluster 1 to obtain five sub-clusters. At the same time, the distances from yj to the center points of the five sub-clusters are calculated to obtain five (n-1)th second sub-low-dimensional structure data. The five (n-1)th first sub-low-dimensional structure data and the five (n-1)th second sub-low-dimensional structure data are weighted and concatenated to obtain the (n-1)th target sub-low-dimensional structure data. The (n-1)th target sub-low-dimensional structure data corresponding to the low-dimensional data of each low-dimensional data point in the aforementioned (n-1)th projection data constitute the aforementioned (n-1)th low-dimensional structure data.

[0081] According to an embodiment of the present invention, by clustering the projected data into K clusters, the distribution characteristics in the projected data can be quickly identified. Furthermore, each cluster is further subdivided into M sub-clusters, which allows for a more detailed analysis of the internal structure and differences within each cluster. By calculating the distances to the cluster centers of the K clusters and the M sub-clusters using the low-dimensional data points, K (n-1)th first sub-low-dimensional structural data and M (n-1)th second sub-low-dimensional structural data are obtained, which can simultaneously retain global and local feature information. Furthermore, the K (n-1)th first sub-low-dimensional structural data and M (n-1)th second sub-low-dimensional structural data are weighted and concatenated according to preset weights, which can adjust the importance between different layers or features according to timing requirements, making the generated (n-1)th low-dimensional structural data more targeted.

[0082] According to an embodiment of the present invention, the nth projection data is obtained based on the nth frame of high-dimensional static biological data, the (n-1)th low-dimensional structural data, and the nth high-dimensional difference, including: stitching the (n-1)th low-dimensional structural data and the nth high-dimensional difference to obtain the nth spatial structural constraint, wherein the nth spatial structural constraint is used to constrain the temporal coherence of each projection structure between the (n-1)th projection data and the nth projection data; stitching the nth spatial structural constraint and the nth frame of high-dimensional static biological data to obtain the target nth frame of high-dimensional static biological data; and obtaining the nth projection data based on the target nth frame of high-dimensional static biological data.

[0083] According to an embodiment of the present invention, the (n-1)th low-dimensional structural data and the nth high-dimensional difference can be multiplied by Hadamard to obtain the nth spatial structural constraint. The projection structure can characterize the spatial distribution characteristics of low-dimensional data points in each cluster. The temporal coherence of the above projection structure indicates that for each cluster in the projection data, the distribution of its corresponding low-dimensional data points remains within a certain range from the first projection data to the Nth projection data. The following will illustrate this through... Figures 4-13 This explains the temporal coherence of the above projection structure.

[0084] Figure 4 The first projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown.

[0085] Figure 5 The second projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown.

[0086] Figure 6 The third projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown.

[0087] Figure 7The fourth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown.

[0088] Figure 8 The fifth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown.

[0089] Figure 9 The sixth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown.

[0090] Figure 10 The seventh projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown.

[0091] Figure 11 The eighth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown.

[0092] Figure 12 The ninth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown.

[0093] Figure 13 The tenth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown.

[0094] like Figures 4-13 As shown, n represents the number of frames, and each color represents a cluster label. In the projection data from frame 1 to frame 10, taking the yellow cluster in the figure as an example, the distribution in the projection data is basically maintained within a spatial range. That is, the projection structure is continuous in both time and space, and the cluster radius of the cluster decreases accordingly as the number of frames increases.

[0095] Figure 14 A trajectory result diagram based on frame number coloring from the first projection data to the tenth projection data is shown according to an embodiment of the present invention.

[0096] Figure 15 The image shows the trajectory results of frame number coloring from the first projection data to the tenth projection data using a static projection method based on correlation techniques.

[0097] Figure 16 The diagram shows the trajectory results of cluster label coloring from the first projection data to the tenth projection data according to an embodiment of the present invention.

[0098] Figure 17 The diagram shows the trajectory results of the static projection method based on correlation techniques, from the first projection data to the tenth projection data, based on cluster label coloring.

[0099] like Figure 14 and Figure 15As shown, the trajectory result diagram based on frame number coloring, which is drawn by projecting high-dimensional biological data using the method of the embodiment of the present invention, is significantly different from the trajectory result diagram based on frame number coloring drawn by the static projection method in related technologies. The evolution law of high-dimensional biological data is obvious, and it can be seen that the cluster radius of the clusters gradually decreases as the number of frames increases.

[0100] like Figure 16 and Figure 17 As shown, the trajectory result map based on cluster label coloring for projecting high-dimensional biological data according to the method of the present invention, compared with the trajectory effect map based on cluster label coloring for projecting high-dimensional biological data by static projection method in related art, shows that the spatial position of high-dimensional biological data changes within a certain range in the projected data as the number of frames increases.

[0101] Figure 18 A flowchart illustrating the projection process for high-dimensional biological data according to an embodiment of the present invention is shown.

[0102] like Figure 18 As shown, for the (n-1)th frame of high-dimensional static biological data, there are J high-dimensional data points, each with d feature dimensions. Projecting the (n-1)th frame of high-dimensional static biological data yields the (n-1)th projected data. For each of the J high-dimensional data points, there is an l×k low-dimensional structure data, where k can be the clustering parameter of the K-Means clustering method, and l is the number of layers in the low-dimensional structure data. For each high-dimensional data point, there is a scalar nth sub-high-dimensional dissimilarity. The nth sub-high-dimensional distance dissimilarity and the nth sub-similarity discriminability can be weighted and summed to obtain the aforementioned nth sub-high-dimensional dissimilarity. The J nth sub-high-dimensional dissimilarity constitutes the nth high-dimensional data dissimilarity. The (n-1)th low-dimensional structure data and the nth high-dimensional dissimilarity are concatenated using Hadamard multiplication to obtain the nth spatial structure constraint. The nth spatial structure constraint is concatenated with the nth frame of high-dimensional static biological data to obtain the target nth frame of high-dimensional static biological data. Then, the target nth frame of high-dimensional static biological data is projected to obtain the nth projection data.

[0103] According to an embodiment of the present invention, the first projection data is obtained by the following method: for any high-dimensional data point among multiple high-dimensional data points included in the first frame of high-dimensional static biological data, the similarity between each high-dimensional data point is calculated based on the first high-dimensional data of each of the multiple high-dimensional data points to obtain a high-dimensional similarity probability matrix; for each low-dimensional data point in the initial projection data, the similarity between each low-dimensional data point is calculated based on the low-dimensional data of each of the multiple low-dimensional data points to obtain a low-dimensional similarity probability matrix; the relative entropy between the high-dimensional similarity probability matrix and the low-dimensional similarity probability matrix is ​​calculated; the low-dimensional data points in the initial projection data are adjusted based on the relative entropy until the relative entropy meets a first preset condition to obtain the first projection data.

[0104] For the high-dimensional data point x in the first high-dimensional data i′ x j′ x m′ and x n′ x can be calculated using the following formula (1). i′ and x j′ The similarity between the corresponding high-dimensional data.

[0105] (1);

[0106] Where, p i′j′ Represents a high-dimensional data point x i′ and x j′ The similarity between them, σ i′ It is related to high-dimensional data points x i′ The standard deviation of the relevant Gaussian distribution is used to control the range of similarity calculation, σ. m′ It is related to high-dimensional data points x m′ The standard deviation of the relevant Gaussian distribution, given that the first high-dimensional data includes S high-dimensional data points, is x i′ x j′ x m′ and x n′ Let S be the i′, j′, m′, and n′ high-dimensional data points, where i′, j′, m′, and n′ are all integers greater than 1 and less than or equal to S. S is an integer greater than or equal to 1.

[0107] According to an embodiment of the present invention, high-dimensional data points are mapped to a low-dimensional space to obtain initial projection data. For the low-dimensional data points y in the initial projection data... i′ y j′ y m′ and y n′ The similarity between low-dimensional data points can be calculated using the following formula (2).

[0108] (2);

[0109] Where, q i′j′ Represents low-dimensional data points y i′ and y j′ The similarity between them, assuming the first high-dimensional data includes S high-dimensional data points, is y i′ y j′ y m′ and y n′ Let i be the i′, j′, m′, and n′ high-dimensional data points, where i′, j′, m′, and n′ are all integers greater than 1 and less than or equal to S.

[0110] According to an embodiment of the present invention, the above formula (2) adopts a t-distribution, which has a wider tail and can better handle the data distribution in low-dimensional space.

[0111] The relative entropy between the high-dimensional similarity probability matrix and the low-dimensional similarity probability matrix can be calculated using the following formula (3).

[0112] (3);

[0113] in, It is a high-dimensional data point x i′ and x j′ Similarity between them and low-dimensional data points y i′ and y j′ The relative entropy of the similarity between them, p i′j′ It is a high-dimensional data point x in a high-dimensional space. i′ and x j′ The similarity between them, q i′j′ It is a low-dimensional data point y in a low-dimensional space. i′ and y j′ The similarity between them.

[0114] According to an embodiment of the present invention, the position of low-dimensional data points can be updated by calculating the gradient of relative entropy until the relative entropy satisfies a first preset condition. The first preset condition may be, for example, the relative entropy converging to a small value or reaching a preset maximum number of iterations.

[0115] According to an embodiment of the present invention, the first projection data is obtained by the following method: For each high-dimensional data point in the first frame of high-dimensional static biological data, based on the first high-dimensional data of the high-dimensional data point and the first high-dimensional data of other high-dimensional data points, the distance between the high-dimensional data point and the first high-dimensional data points is calculated to obtain the first nearest neighbor high-dimensional data point group of the high-dimensional data point; based on the first high-dimensional data of each of the multiple high-dimensional data points included in the first nearest neighbor high-dimensional data point group, the fuzzy similarity weight between each high-dimensional data point is calculated to obtain the high-dimensional fuzzy similarity matrix; for each low-dimensional data point in the initial projection data, based on the low-dimensional data of each of the multiple low-dimensional data points, the fuzzy similarity between every two low-dimensional data points is calculated to obtain the low-dimensional fuzzy similarity matrix; the cross-entropy loss value between the high-dimensional fuzzy similarity matrix and the low-dimensional fuzzy similarity matrix is ​​calculated; the low-dimensional data points in the initial projection data are adjusted based on the cross-entropy loss value until the loss value meets the second preset condition to obtain the first projection data.

[0116] According to an embodiment of the present invention, for each high-dimensional data point x i′ The k-nearest neighbor method can be used to find the first nearest neighbor high-dimensional data point group kNN(x). i′ x can be calculated using the following formula (4). i′ and x j′ The high-dimensional fuzzy similarity weights between them.

[0117] (4);

[0118] Where, p′ i′j′ x represents i′ and x j′ The high-dimensional fuzzy similarity weights between them, σ i′ It is related to high-dimensional data points x i′ The standard deviation of the relevant Gaussian distribution is determined by binary search.

[0119] According to embodiments of the present invention, high-dimensional data points can be randomly initialized or reduced using other dimensionality reduction methods, such as PCA (Principal Component Analysis).

[0120] According to an embodiment of the present invention, the low-dimensional data point y can be calculated using the above formula (2). i′ and y j′ The low-dimensional fuzzy similarity weights q′ between them i′j′ .

[0121] The cross-entropy loss between the high-dimensional fuzzy matrix and the low-dimensional fuzzy matrix can be calculated using the following formula (5).

[0122] (5);

[0123] Where H(P, Q) represents the high-dimensional data point x i′ and x j′ The high-dimensional fuzzy similarity weight p′ between them i′j′ With low-dimensional data points y i′ and y j′ The low-dimensional fuzzy similarity weights q′ between them i′j′ The cross-entropy between them, where P represents the high-dimensional fuzzy similarity matrix, Q represents the low-dimensional fuzzy similarity matrix, and p′ i′j′ Represents a high-dimensional data point x i′ and x j′ The high-dimensional fuzzy similarity weights between them, q′ i′j′ Represents low-dimensional data points y i′ and y j′ Low-dimensional fuzzy similarity weights between them.

[0124] According to embodiments of the present invention, the above-described projection method for high-dimensional biological data can be applied to the determination of microbial community metagenomics, specifically including: acquiring metagenomic sequencing data of a microbial community, the metagenomic sequencing data including N frames of metagenomic sequencing static data, the N frames of metagenomic sequencing static data representing the metagenomic abundance on day n; passing the metagenomic sequencing data through the above-described projection method for high-dimensional biological data to obtain projection data of the metagenomic sequencing data of the microbial community, wherein the nth high-dimensional dissimilarity can be the species abundance vector of the microbial community; determining the anomalous drift data of antibiotic perturbation on day n based on the projection data of the metagenomic sequencing data, and determining the recovery time of the microbial community based on the anomalous drift data of antibiotic perturbation on day n.

[0125] Based on the above-described projection method for high-dimensional biological data, this invention also provides a projection device for high-dimensional biological data. The following will be combined with... Figure 19 The device is described in detail.

[0126] Figure 19 A structural block diagram of a projection device for high-dimensional biological data according to an embodiment of the present invention is shown.

[0127] like Figure 19 As shown, the projection device 1900 for high-dimensional biological data in this embodiment includes a low-dimensional structure data determination module 1910, a high-dimensional difference determination module 1920, a projection data determination module 1930, and a dynamic projection module 1940.

[0128] The low-dimensional structure data determination module 1910 is used to determine the (n-1)th low-dimensional structure data based on the (n-1)th projection data of the nth high-dimensional static biological data in N frames of high-dimensional static biological data. The (n-1)th projection data represents the projection data of the (n-1)th high-dimensional static biological data, the (n-1)th low-dimensional structure data represents the distribution structure characteristics of the (n-1)th projection data, and the first projection data represents the projection data of the first frame of high-dimensional static biological data, where n ∈ {2, ..., N}. In one embodiment, the low-dimensional structure data determination module 1910 can be used to perform the operation S210 described above, which will not be repeated here.

[0129] The high-dimensional difference determination module 1920 is used to obtain the nth high-dimensional difference based on the (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data. The nth high-dimensional difference characterizes the degree of difference between the (n-1)th high-dimensional data point group and the nth high-dimensional data point group. The (n-1)th high-dimensional data point group includes high-dimensional data points from the (n-1)th frame of high-dimensional static biological data, and the nth high-dimensional data point group includes high-dimensional data points from the nth frame of high-dimensional static biological data. In one embodiment, the high-dimensional difference determination module 1920 can be used to perform the operation S220 described above, which will not be repeated here.

[0130] The projection data determination module 1930 is used to obtain the nth projection data based on the nth frame of high-dimensional static biological data, the (n-1)th low-dimensional structural data, and the nth high-dimensional difference. In one embodiment, the projection data determination module 1930 can be used to perform the operation S230 described above, which will not be repeated here.

[0131] The dynamic projection module 1940 is used to obtain projection data of high-dimensional biological data based on the first projection data to the Nth projection data. In one embodiment, the dynamic projection module 1940 can be used to perform the operation S240 described above, which will not be repeated here.

[0132] According to embodiments of the present invention, any plurality of modules among the low-dimensional structure data determination module 1910, high-dimensional difference determination module 1920, projection data determination module 1930, and dynamic projection module 1940 can be combined into one module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in one module. According to embodiments of the present invention, at least one of the low-dimensional structure data determination module 1910, high-dimensional difference determination module 1920, projection data determination module 1930, and dynamic projection module 1940 can be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging the circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the low-dimensional structure data determination module 1910, the high-dimensional difference determination module 1920, the projection data determination module 1930, and the dynamic projection module 1940 can be at least partially implemented as a computer program module, which can perform corresponding functions when the computer program module is run.

[0133] Figure 20 A block diagram of an electronic device suitable for implementing a projection method for high-dimensional biological data according to an embodiment of the present invention is shown.

[0134] like Figure 20 As shown, an electronic device 2000 according to an embodiment of the present invention includes a processor 2001, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 2002 or a program loaded from a storage portion 2008 into a random access memory (RAM) 2003. The processor 2001 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 2001 may also include onboard memory for caching purposes. The processor 2001 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0135] RAM 2003 stores various programs and data required for the operation of electronic device 2000. Processor 2001, ROM 2002, and RAM 2003 are interconnected via bus 2004. Processor 2001 executes various operations of the method flow according to embodiments of the present invention by executing programs in ROM 2002 and / or RAM 2003. It should be noted that the programs may also be stored in one or more memories other than ROM 2002 and RAM 2003. Processor 2001 may also execute various operations of the method flow according to embodiments of the present invention by executing programs stored in said one or more memories.

[0136] According to embodiments of the present invention, the electronic device 2000 may further include an input / output (I / O) interface 2005, which is also connected to a bus 2004. The electronic device 2000 may also include one or more of the following components connected to the input / output (I / O) interface 2005: an input section 2006 including a keyboard, mouse, etc.; an output section 2007 including a cathode ray tube (CRN), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 2008 including a hard disk, etc.; and a communication section 2009 including a network interface card such as a LAN card, modem, etc. The communication section 2009 performs communication processing via a network such as the Internet. A drive 2010 is also connected to the input / output (I / O) interface 2005 as needed. A removable medium 2011, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 2010 as needed so that computer programs read from it can be installed into the storage section 2008 as needed.

[0137] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.

[0138] According to embodiments of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of the present invention, a computer-readable storage medium may include ROM 2002 and / or RAM 2003 and / or one or more memories other than ROM 2002 and RAM 2003 described above.

[0139] Embodiments of the present invention also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code enables the computer system to implement the projection method for high-dimensional biological data provided in the embodiments of the present invention.

[0140] When the computer program is executed by the processor 2001, it performs the functions defined in the system / apparatus of this invention. According to embodiments of the invention, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0141] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 2009, and / or installed from a removable medium 2011. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.

[0142] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 2009, and / or installed from the removable medium 2011. When the computer program is executed by the processor 2001, it performs the functions defined in the system of this embodiment of the invention. According to embodiments of the invention, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.

[0143] According to embodiments of the present invention, program code for executing the computer programs provided in the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C", or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0144] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0145] Those skilled in the art will understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention can be combined and / or combined in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.

[0146] The embodiments of the present invention have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of the invention. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of the invention, and all such substitutions and modifications should fall within the scope of the invention.

Claims

1. A projection method for high-dimensional biological data, characterized in that, The high-dimensional biological data includes N frames of high-dimensional static biological data, each frame of high-dimensional static biological data includes high-dimensional data corresponding to multiple high-dimensional data points, where N is an integer greater than or equal to 1. For the nth frame of high-dimensional static biological data in the N frames of high-dimensional static biological data, the (n-1)th low-dimensional structural data is obtained based on the (n-1)th projection data. The (n-1)th projection data represents the projection data of the (n-1)th frame of high-dimensional static biological data, the (n-1)th low-dimensional structural data represents the distribution structure characteristics of the (n-1)th projection data, and the first projection data represents the projection data of the first frame of high-dimensional static biological data, n∈{2,……,N}; Based on the (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data, the nth high-dimensional difference is obtained, wherein the nth high-dimensional difference characterizes the degree of difference between the (n-1)th high-dimensional data point group and the nth high-dimensional data point group, the (n-1)th high-dimensional data point group includes high-dimensional data points in the (n-1)th frame of high-dimensional static biological data, and the nth high-dimensional data point group includes high-dimensional data points in the nth frame of high-dimensional static biological data; Based on the nth frame of high-dimensional static biological data, the (n-1)th low-dimensional structural data, and the nth high-dimensional difference, the nth projection data is obtained; Based on the first projection data to the Nth projection data, the projection data of the high-dimensional biological data is obtained; The process of obtaining the nth projection data based on the nth frame of high-dimensional static biological data, the (n-1)th low-dimensional structural data, and the nth high-dimensional difference includes: The (n-1)th low-dimensional structural data and the nth high-dimensional difference are concatenated to obtain the nth spatial structural constraint, wherein the nth spatial structural constraint is used to constrain the temporal coherence of each projection structure between the (n-1)th projection data and the nth projection data; The nth spatial structure constraint and the nth frame of high-dimensional static biological data are spliced ​​together to obtain the target nth frame of high-dimensional static biological data; Based on the target's nth frame of high-dimensional static biological data, the nth projection data is obtained.

2. The method according to claim 1, characterized in that, The nth high-dimensional difference includes at least one of the following: nth high-dimensional distance difference, nth similarity difference; The process of obtaining the nth high-dimensional difference based on the (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data includes: Based on the (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data, the nth high-dimensional distance difference is obtained, wherein the nth high-dimensional distance difference characterizes the degree of difference between the first high-dimensional distance between every two high-dimensional data points in the (n-1)th high-dimensional data point group and the second high-dimensional distance between the same two high-dimensional data points in the nth high-dimensional data point group. The nth similarity distinguishability is obtained based on the (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data, wherein the nth similarity distinguishability characterizes the degree of difference between the first nearest neighbor high-dimensional data point group corresponding to each high-dimensional data point in the (n-1)th high-dimensional data point group and the second nearest neighbor high-dimensional data point group corresponding to each high-dimensional data point in the nth high-dimensional data point group.

3. The method according to claim 2, characterized in that, The process of obtaining the nth high-dimensional distance difference degree based on the (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data includes: For every two high-dimensional data points in the (n-1)th frame of high-dimensional static biological data, the distance between each pair of high-dimensional data points is calculated based on their respective (n-1)th high-dimensional data to obtain the (n-1)th high-dimensional distance. For every two high-dimensional data points in the nth frame of high-dimensional static biological data, the distance between each two high-dimensional data points is calculated based on their respective nth high-dimensional data to obtain the nth high-dimensional distance; Based on the difference between the (n-1)th high-dimensional distance and the nth high-dimensional distance, the nth sub-high-dimensional distance difference degree of each pair of high-dimensional data points is obtained; The nth high-dimensional distance difference is obtained based on multiple nth sub-high-dimensional distance difference degrees.

4. The method according to claim 2 or 3, characterized in that, The process of obtaining the nth similarity distinguishability based on the (n-1)th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data includes: For any high-dimensional data point among the multiple high-dimensional data points included in the (n-1)th frame of high-dimensional static biological data, based on the (n-1)th high-dimensional data of the high-dimensional data point and the (n-1)th other high-dimensional data points respectively, the first distance between the high-dimensional data point and the (n-1)th other high-dimensional data points is calculated to obtain the (n-1)th nearest neighbor high-dimensional data point group of the high-dimensional data point. The first distance between the high-dimensional data point and the (n-1)th nearest neighbor high-dimensional data points included in the (n-1)th nearest neighbor high-dimensional data point group is one of I first distances. The I first distances are the first I obtained by arranging multiple first distances in ascending order, where I is an integer greater than 1. The (n-1)th other high-dimensional data points are the other high-dimensional data points in the (n-1)th frame of high-dimensional static biological data excluding the high-dimensional data point. For each high-dimensional data point in the nth frame of high-dimensional static biological data, based on the nth high-dimensional data point and the nth other high-dimensional data points respectively, the second distance between the high-dimensional data point and the nth other high-dimensional data points is calculated to obtain the nth nearest neighbor high-dimensional data point group of the high-dimensional data point. The second distance between the high-dimensional data point and the nth nearest neighbor high-dimensional data points included in the nth nearest neighbor high-dimensional data point group is one of I second distances. The I second distances are the first I obtained by arranging multiple second distances in ascending order. The nth other high-dimensional data points are other high-dimensional data points in the nth frame of high-dimensional static biological data excluding the high-dimensional data point. Determine the difference neighbor high-dimensional data points between the (n-1)th nearest neighbor high-dimensional data point group and the nth nearest neighbor high-dimensional data point group of the high-dimensional data point, and obtain the nth sub-similarity discriminability of the high-dimensional data point; The nth similarity distinguishability is obtained based on the nth sub-similarity distinguishability of each of the multiple high-dimensional data points.

5. The method according to any one of claims 1 to 3, characterized in that, The process of obtaining the (n-1)th low-dimensional structural data based on the (n-1)th projection data includes: Cluster the (n-1)th projection data to generate K clusters, where K is an integer greater than or equal to 1; For any of the K clusters, perform secondary clustering on the cluster to generate M sub-clusters, where M is an integer greater than or equal to 1; For each low-dimensional data point in the (n-1)th projection data, the distance to the cluster center of the K clusters is calculated based on the low-dimensional data of the low-dimensional data point to obtain the K (n-1)th first sub-low-dimensional structure data. Based on the low-dimensional data of each of the multiple low-dimensional data points, the distance to the cluster center of the M sub-clusters is calculated to obtain the M (n-1)th second sub-low-dimensional structure data. Based on preset weights, the K (n-1)th first sub-low-dimensional structure data and the M (n-1)th second sub-low-dimensional structure data are concatenated to obtain the (n-1)th target sub-low-dimensional structure data. The (n-1)th low-dimensional structural data is obtained based on multiple sets of the low-dimensional structural data of the (n-1)th target sub-sub ...

6. The method according to any one of claims 1 to 3, characterized in that, The first projection data was obtained through the following method: For any high-dimensional data point among the multiple high-dimensional data points included in the first frame of high-dimensional static biological data, the similarity between the multiple high-dimensional data points is calculated based on the first high-dimensional data of each of the multiple high-dimensional data points to obtain a high-dimensional similarity probability matrix. For each low-dimensional data point in the initial projection data, the similarity between the low-dimensional data points is calculated based on the low-dimensional data of each low-dimensional data point to obtain a low-dimensional similarity probability matrix. Calculate the relative entropy between the high-dimensional similarity probability matrix and the low-dimensional similarity probability matrix; The low-dimensional data points in the initial projection data are adjusted based on the relative entropy until the relative entropy meets the first preset condition, thereby obtaining the first projection data.

7. The method according to any one of claims 1 to 3, characterized in that, The first projection data was obtained through the following method: For each high-dimensional data point in the first frame of high-dimensional static biological data, based on the high-dimensional data point and the first high-dimensional data of each of the other high-dimensional data points, the distance between the high-dimensional data point and the other high-dimensional data points is calculated to obtain the first nearest neighbor high-dimensional data point group of the high-dimensional data point. Based on the first high-dimensional data of each of the multiple high-dimensional data points included in the first nearest high-dimensional data point group, the fuzzy similarity weight between the high-dimensional data points is calculated to obtain a high-dimensional fuzzy similarity matrix. For each low-dimensional data point in the initial projection data, based on the low-dimensional data of each of the multiple low-dimensional data points, the fuzzy similarity between every two low-dimensional data points is calculated to obtain a low-dimensional fuzzy similarity matrix; Calculate the cross-entropy loss value between the high-dimensional fuzzy similarity matrix and the low-dimensional fuzzy similarity matrix; The low-dimensional data points in the initial projection data are adjusted based on the cross-entropy loss value until the loss value meets the second preset condition, thereby obtaining the first projection data.

8. An electronic device, comprising: One or more processors; Memory, used to store one or more computer programs. The characteristic feature is that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method for identifying underground coal mine personnel based on gait identification

    CN109241870A

  • High-dimensional data clustering method, device and equipment, medium and product

    CN119128566A