Projection method for high-dimensional biological data, electronic equipment and storage medium
By calculating the difference and similarity distinction of high-dimensional biological data, combining low-dimensional structural data and cluster analysis, continuous projection data is generated, which solves the problem of insufficient temporal continuity of high-dimensional biological data and realizes efficient data change analysis and coherence projection.
Patent Information
- Application Number
- CN202511205880.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing technologies have difficulty maintaining the temporal coherence of high-dimensional biological data, resulting in insufficient continuity of projection data between adjacent frames.
By calculating the difference and similarity between high-dimensional static biological data, combining low-dimensional structural data and cluster analysis, continuous projection data is generated, and distribution structure characteristics and difference degree constraints are introduced to ensure temporal continuity.
It realizes the continuous projection of high-dimensional biological data on time series, which can more accurately analyze data changes and maintain the temporal continuity and data characteristic reflection of the projected data.
Smart Images

Figure CN120705787A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis technology, and more specifically to a projection method, electronic device, and storage medium for high-dimensional biological data. Background Art
[0002] High-dimensional biological data contains rich and complex information. Mining high-dimensional biological data is of great significance in actual production practice.
[0003] In the related art, a projection method for high-dimensional static biological data can be used to obtain projection data of the high-dimensional biological data. However, this method is difficult to maintain the temporal continuity between two adjacent frames of high-dimensional static biological data. Summary of the Invention
[0004] In view of the above problems, the present invention provides a projection method, electronic device, and storage medium for high-dimensional biological data.
[0005] According to a first aspect of the present invention, a projection method for high-dimensional biological data is provided, wherein the high-dimensional biological data includes N frames of high-dimensional static biological data, each frame of high-dimensional static biological data includes high-dimensional data corresponding to multiple high-dimensional data points, and N is an integer greater than or equal to 1; for the n-th frame of high-dimensional static biological data in the N frames of high-dimensional static biological data, based on the n-1th projection data, the n-1th low-dimensional structural data is obtained, wherein the n-1th projection data represents the projection data of the n-1th frame of high-dimensional static biological data, the n-1th low-dimensional structural data represents the distribution structure characteristics of the n-1th projection data, and the first projection data represents the projection data of the first frame of high-dimensional static biological data, n∈{ 2, ..., N}; based on the above-mentioned n-1th frame high-dimensional static biological data and the nth frame high-dimensional static biological data, obtain the nth high-dimensional difference, wherein the above-mentioned nth high-dimensional difference characterizes the degree of difference between the n-1th high-dimensional data point group and the nth high-dimensional data point group, the above-mentioned n-1th high-dimensional data point group includes the high-dimensional data points in the above-mentioned n-1th frame high-dimensional static biological data, and the above-mentioned nth high-dimensional data point group includes the high-dimensional data points in the above-mentioned n-1th frame high-dimensional static biological data; based on the above-mentioned nth frame high-dimensional static biological data, the above-mentioned n-1th low-dimensional structural data and the above-mentioned nth high-dimensional difference, obtain the nth projection data; based on the above-mentioned first projection data to the Nth projection data, obtain the projection data of the above-mentioned high-dimensional biological data.
[0006] According to an embodiment of the present invention, the above-mentioned n-th high-dimensional difference includes at least one of the following: n-th high-dimensional distance difference, n-th similarity difference; the above-mentioned n-th high-dimensional difference based on the above-mentioned n-1-th frame high-dimensional static biological data and the n-th frame high-dimensional static biological data, includes: based on the above-mentioned n-1-th frame high-dimensional static biological data and the above-mentioned n-th frame high-dimensional static biological data, the n-th high-dimensional distance difference is obtained, wherein the above-mentioned n-th high-dimensional distance difference represents the difference between every two high-dimensional data points in the above-mentioned n-1-th high-dimensional data point group. The degree of difference between the first high-dimensional distance between the two high-dimensional data points and the second high-dimensional distance between the same two high-dimensional data points in the above-mentioned n-th high-dimensional data point group; the n-th similarity distinction is obtained based on the above-mentioned n-1-th frame high-dimensional static biological data and the above-mentioned n-th frame high-dimensional static biological data, wherein the above-mentioned n-th similarity distinction represents the degree of difference between the first neighboring high-dimensional data point group corresponding to each high-dimensional data point in the above-mentioned n-1-th high-dimensional data point group and the second neighboring high-dimensional data point group corresponding to each high-dimensional data point in the above-mentioned n-th high-dimensional data point group.
[0007] According to an embodiment of the present invention, the above-mentioned n-th high-dimensional distance difference is obtained based on the above-mentioned n-1th frame high-dimensional static biological data and the above-mentioned n-th frame high-dimensional static biological data, including: for every two high-dimensional data points in the above-mentioned n-1th frame high-dimensional static biological data, based on the n-1th high-dimensional data of each of the above-mentioned every two high-dimensional data points, the distance between the above-mentioned every two high-dimensional data points is calculated to obtain the n-1th high-dimensional distance; for every two high-dimensional data points in the above-mentioned n-th frame high-dimensional static biological data, based on the n-th high-dimensional data of each of the above-mentioned every two high-dimensional data points, the distance between the above-mentioned every two high-dimensional data points is calculated to obtain the n-th high-dimensional distance; based on the difference between the above-mentioned n-1th high-dimensional distance and the above-mentioned n-th high-dimensional distance, the n-th sub-high-dimensional distance difference of the above-mentioned every two high-dimensional data points is obtained; based on multiple above-mentioned n-th sub-high-dimensional distance differences, the above-mentioned n-th high-dimensional distance difference is obtained.
[0008] According to an embodiment of the present invention, the above-mentioned method of obtaining the nth similarity distinction based on the above-mentioned n-1th frame high-dimensional static biological data and the above-mentioned n-1th frame high-dimensional static biological data includes: for any high-dimensional data point among the multiple high-dimensional data points included in the above-mentioned n-1th frame high-dimensional static biological data, based on the n-1th high-dimensional data of each of the above-mentioned high-dimensional data point and the n-1th other high-dimensional data point, calculating the first distance between the above-mentioned high-dimensional data point and the above-mentioned n-1th other high-dimensional data point, and obtaining the n-1th nearest neighbor high-dimensional data point group of the above-mentioned high-dimensional data point, wherein the first distance between the above-mentioned high-dimensional data point and the n-1th nearest neighbor high-dimensional data point included in the above-mentioned n-1th nearest neighbor high-dimensional data point group is one of I first distances, and the above-mentioned I first distances are the first I obtained by arranging the multiple first distances in ascending order, I is an integer greater than 1, and the above-mentioned n-1th other high-dimensional data point is the other high-dimensional data point in the n-1th frame high-dimensional static biological data except the above-mentioned high-dimensional data point; for the above-mentioned For each high-dimensional data point in n frames of high-dimensional static biological data, based on the nth high-dimensional data of each of the above high-dimensional data point and the nth other high-dimensional data point, the second distance between the above high-dimensional data point and the above nth other high-dimensional data point is calculated to obtain the nth nearest neighbor high-dimensional data point group of the above high-dimensional data point, wherein the second distance between the above high-dimensional data point and the nth nearest neighbor high-dimensional data point included in the above nth nearest neighbor high-dimensional data point group is one of I second distances, the above I second distances are the first I obtained by arranging multiple above second distances in ascending order, and the above nth other high-dimensional data point is other high-dimensional data points in the nth frame of high-dimensional static biological data except the above high-dimensional data point; determine the difference nearest neighbor high-dimensional data point between the above n-1th nearest neighbor high-dimensional data point group and the above nth nearest neighbor high-dimensional data point group of the above high-dimensional data point to obtain the nth sub-similarity distinction of the above high-dimensional data point; obtain the above nth similarity distinction based on the nth sub-similarity distinction of each of the above multiple high-dimensional data points.
[0009] According to an embodiment of the present invention, the method of obtaining the n-1th low-dimensional structure data based on the n-1th projection data includes: clustering the n-1th projection data to generate K clusters, where K is an integer greater than or equal to 1; for any cluster in the K clusters, performing secondary clustering on the cluster to generate M sub-clusters, where M is an integer greater than or equal to 1; for each low-dimensional data point in the n-1th projection data, calculating the clustering of the K clusters based on the low-dimensional data of the low-dimensional data point. The method further comprises the following steps: calculating the distance between the n-1th first sub-low-dimensional structure data and the cluster center based on the low-dimensional data of each of the above-mentioned low-dimensional data points, and obtaining K n-1th first sub-low-dimensional structure data; calculating the distance between the low-dimensional data points and the cluster center of the above-mentioned M sub-cluster clusters based on the low-dimensional data of each of the above-mentioned low-dimensional data points, and obtaining M n-1th second sub-low-dimensional structure data; based on the preset weights, splicing the above-mentioned K n-1th first sub-low-dimensional structure data and the above-mentioned M n-1th second sub-low-dimensional structure data to obtain the n-1th target sub-low-dimensional structure data; and obtaining the above-mentioned n-1th low-dimensional structure data based on the above-mentioned multiple n-1th target sub-low-dimensional structure data.
[0010] According to an embodiment of the present invention, the above-mentioned nth projection data is obtained based on the above-mentioned nth frame high-dimensional static biological data, the above-mentioned n-1th low-dimensional structural data and the above-mentioned nth high-dimensional difference, including: splicing the above-mentioned n-1th low-dimensional structural data and the above-mentioned nth high-dimensional difference to obtain the nth spatial structure constraint, wherein the above-mentioned nth spatial structure constraint is used to constrain the temporal continuity of each projection structure between the above-mentioned n-1th projection data and the above-mentioned nth projection data; splicing the above-mentioned nth spatial structure constraint and the above-mentioned nth frame high-dimensional static biological data to obtain the target nth frame high-dimensional static biological data; and obtaining the above-mentioned nth projection data based on the above-mentioned target nth frame high-dimensional static biological data.
[0011] According to an embodiment of the present invention, the above-mentioned first projection data is obtained by the following method: for any high-dimensional data point among the multiple high-dimensional data points included in the above-mentioned first frame of high-dimensional static biological data, based on the first high-dimensional data of each of the above-mentioned multiple high-dimensional data points, the similarity between the above-mentioned high-dimensional data points is calculated to obtain a high-dimensional similarity probability matrix; for each low-dimensional data point in the initial projection data, based on the low-dimensional data of each of the above-mentioned low-dimensional data points, the similarity between the above-mentioned low-dimensional data points is calculated to obtain a low-dimensional similarity probability matrix; the relative entropy between the above-mentioned high-dimensional similarity probability matrix and the above-mentioned low-dimensional similarity probability matrix is calculated; based on the above-mentioned relative entropy, the low-dimensional data points in the above-mentioned initial projection data are adjusted until the above-mentioned relative entropy meets the first preset condition to obtain the above-mentioned first projection data.
[0012] According to an embodiment of the present invention, the above-mentioned first projection data is obtained by the following method: for each high-dimensional data point in the above-mentioned first frame of high-dimensional static biological data, based on the first high-dimensional data of each of the above-mentioned high-dimensional data point and the first other high-dimensional data point, the distance between the above-mentioned high-dimensional data point and the above-mentioned first other high-dimensional data point is calculated to obtain the first neighboring high-dimensional data point group of the above-mentioned high-dimensional data point; based on the first high-dimensional data of each of the multiple high-dimensional data points included in the above-mentioned first neighboring high-dimensional data point group, the fuzzy similarity weights between the above-mentioned high-dimensional data points are calculated to obtain a high-dimensional fuzzy similarity matrix; for each low-dimensional data point in the initial projection data, based on the low-dimensional data of each of the multiple low-dimensional data points, the fuzzy similarity between each two of the above-mentioned low-dimensional data points is calculated to obtain a low-dimensional fuzzy similarity matrix; the cross-entropy loss value between the above-mentioned high-dimensional fuzzy similarity matrix and the above-mentioned low-dimensional fuzzy similarity matrix is calculated; based on the above-mentioned cross-entropy loss value, the low-dimensional data points in the above-mentioned initial projection data are adjusted until the above-mentioned loss value meets the second preset condition to obtain the above-mentioned first projection data.
[0013] A second aspect of the present invention provides a projection device for high-dimensional biological data, wherein the high-dimensional biological data includes N frames of high-dimensional static biological data, each frame of high-dimensional static biological data includes high-dimensional data corresponding to multiple high-dimensional data points, N is an integer greater than or equal to 1, and includes: a low-dimensional structural data determination module for obtaining n-1th low-dimensional structural data based on n-1th projection data for the n-th frame of high-dimensional static biological data in the N frames of high-dimensional static biological data, wherein the n-1th projection data represents the projection data of the n-1th frame of high-dimensional static biological data, the n-1th low-dimensional structural data represents the distribution structure characteristics of the n-1th projection data, and the first projection data represents the projection data of the first frame of high-dimensional static biological data, n∈{2,…,N}; high-dimensional A difference determination module is used to obtain the nth high-dimensional difference based on the above-mentioned n-1th frame high-dimensional static biological data and the nth frame high-dimensional static biological data, wherein the above-mentioned nth high-dimensional difference represents the degree of difference between the n-1th high-dimensional data point group and the nth high-dimensional data point group, the above-mentioned n-1th high-dimensional data point group includes the high-dimensional data points in the above-mentioned n-1th frame high-dimensional static biological data, and the above-mentioned nth high-dimensional data point group includes the high-dimensional data points in the above-mentioned n-1th frame high-dimensional static biological data; a projection data determination module is used to obtain the nth projection data based on the above-mentioned nth frame high-dimensional static biological data, the above-mentioned n-1th low-dimensional structural data and the above-mentioned nth high-dimensional difference; a dynamic projection module is used to obtain the projection data of the above-mentioned high-dimensional biological data based on the above-mentioned first projection data to the Nth projection data.
[0014] A third aspect of the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.
[0015] The fourth aspect of the present invention further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.
[0016] The fifth aspect of the present invention further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor.
[0017] According to an embodiment of the present invention, for non-first frame high-dimensional static biological data, the n-1th low-dimensional structural data is obtained through the n-1th projection data, thereby introducing the distribution structural characteristics of the n-1th projection data into the generated nth projection data, and then constraining the nth projection data. The nth high-dimensional difference between the n-1th frame high-dimensional static biological data and the nth frame high-dimensional static biological data is used to introduce the degree of difference between the n-1th high-dimensional data point group and the nth high-dimensional data point group, so as to provide the nth projection data with a constraint that is more in line with the change characteristics of the high-dimensional biological data, and then reflect the characteristics of the nth frame high-dimensional static biological data itself, thereby ensuring the temporal continuity between the first projection data to the Nth projection data. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which:
[0019] Figure 1 A diagram showing an application scenario of a projection method and device for high-dimensional biological data according to an embodiment of the present invention is shown;
[0020] Figure 2 A flowchart of a projection method for high-dimensional biological data according to an embodiment of the present invention is shown;
[0021] Figure 3 A schematic diagram of multi-layer clustering according to an embodiment of the present invention is shown;
[0022] Figure 4 shows first projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention;
[0023] Figure 5 shows second projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention;
[0024] Figure 6shows third projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention;
[0025] Figure 7 shows fourth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention;
[0026] Figure 8 shows fifth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention;
[0027] Figure 9 shows sixth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention;
[0028] Figure 10 shows seventh projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention;
[0029] Figure 11 shows eighth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention;
[0030] Figure 12 shows ninth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention;
[0031] Figure 13 shows tenth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention;
[0032] Figure 14 A trajectory result diagram based on frame number coloring from the first projection data to the tenth projection data according to an embodiment of the present invention is shown;
[0033] Figure 15 shows a trajectory result diagram from the first projection data to the tenth projection data colored based on cluster labels according to an embodiment of the present invention;
[0034] Figure 16 A trajectory result diagram based on frame number coloring from the first projection data to the tenth projection data using a static projection method based on related art is shown;
[0035] Figure 17 The figure shows the trajectory result diagram from the first projection data to the tenth projection data colored based on cluster labels using a static projection method based on related technologies;
[0036] Figure 18 A projection flow chart for high-dimensional biological data according to an embodiment of the present invention is shown;
[0037] Figure 19 It shows a structural block diagram of a projection device for high-dimensional biological data according to an embodiment of the present invention;
[0038] Figure 20 A block diagram of an electronic device suitable for implementing a projection method for high-dimensional biological data according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0039] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.
[0040] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise", "include", etc. used herein indicate the presence of the features, steps, operations and / or components, but do not exclude the presence or addition of one or more other features, steps, operations or components.
[0041] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0042] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0043] High-dimensional biological data contains rich and complex information. Mining this data is of great significance in practical production. Visualization, a comprehensive discipline combining computer science, human-computer interaction, and psychology, studies how to visualize data using graphical representations that are close to human perception, enhancing human perception and revealing implicit information and patterns within the data. Visual exploration of high-dimensional biological data helps people more quickly uncover information within the data and obtain valuable insights.
[0044] Projection methods are widely used in the visualization of high-dimensional data. Applying projection methods to the visualization of high-dimensional biological data has become one of the hot issues in the field of visualization.
[0045] In related art, a projection method for high-dimensional static biological data can be used to obtain projection data of the high-dimensional biological data. The high-dimensional biological data can include N frames of high-dimensional static biological data. The projection data of the high-dimensional biological data can include first projection data to N-th projection data.
[0046] During the implementation of the present invention, it was found that it was difficult to maintain temporal continuity between the n-1th projection data and the n-th projection data obtained using the above method. The n-th projection data is the projection data of the n-th frame of high-dimensional static biological data. The n-1th projection data is the projection data of the n-1th frame of high-dimensional static biological data.
[0047] In view of this, an embodiment of the present invention provides a projection method for high-dimensional biological data, wherein the high-dimensional biological data includes N frames of high-dimensional static biological data, each frame of high-dimensional static biological data includes high-dimensional data corresponding to multiple high-dimensional data points, and N is an integer greater than or equal to 1; for the n-th frame of high-dimensional static biological data in the N frames of high-dimensional static biological data, based on the n-1th projection data, the n-1th low-dimensional structural data is obtained, wherein the n-1th projection data represents the projection data of the n-1th frame of high-dimensional static biological data, the n-1th low-dimensional structural data represents the distribution structure characteristics of the n-1th projection data, and the first projection data represents the projection data of the first frame of high-dimensional static biological data. projection data, n∈{2,…,N}; based on the n-1th frame high-dimensional static biological data and the nth frame high-dimensional static biological data, obtain the nth high-dimensional difference, wherein the nth high-dimensional difference characterizes the degree of difference between the n-1th high-dimensional data point group and the nth high-dimensional data point group, the n-1th high-dimensional data point group includes the high-dimensional data points in the n-1th frame high-dimensional static biological data, and the nth high-dimensional data point group includes the high-dimensional data points in the nth frame high-dimensional static biological data. Based on the nth frame high-dimensional static biological data, the n-1th low-dimensional structural data and the nth high-dimensional difference, obtain the nth projection data; based on the first projection data to the Nth projection data, obtain the projection data of the high-dimensional biological data.
[0048] Figure 1 A diagram showing an application scenario of a projection method and device for high-dimensional biological data according to an embodiment of the present invention is shown.
[0049] like Figure 1 As shown, the application scenario 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or optical fiber cables.
[0050] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0051] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0052] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.
[0053] It should be noted that the projection method for high-dimensional biological data provided in the embodiments of the present invention can generally be executed by the server 105. Accordingly, the projection device for high-dimensional biological data provided in the embodiments of the present invention can generally be located in the server 105. The projection method for high-dimensional biological data provided in the embodiments of the present invention can also be executed by a server or server cluster that is different from the server 105 and that is capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Accordingly, the projection device for high-dimensional biological data provided in the embodiments of the present invention can also be located in a server or server cluster that is different from the server 105 and that is capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0054] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0055] The following will be based on Figure 1 The scene described by Figures 2 to 19 The projection method for high-dimensional biological data according to an embodiment of the present invention is described in detail.
[0056] Figure 2A flowchart of a projection method for high-dimensional biological data according to an embodiment of the present invention is shown.
[0057] like Figure 2 As shown, the projection method for high-dimensional biological data in this embodiment includes operations S210 to S240.
[0058] In operation S210 , for an n-th frame of high-dimensional static biological data among N frames of high-dimensional static biological data, an n-1-th low-dimensional structure data is obtained based on an n-1-th projection data.
[0059] Among them, the n-1th projection data represents the projection data of the n-1th frame of high-dimensional static biological data, the n-1th low-dimensional structure data represents the distribution structure characteristics of the n-1th projection data, and the first projection data represents the projection data of the first frame of high-dimensional static biological data, n∈{2,...,N}.
[0060] According to an embodiment of the present invention, the distribution structure feature can characterize the distribution characteristics of each low-dimensional data point in the (n-1)th projection data. The distribution characteristics can be, for example, the cluster to which the low-dimensional data point belongs. The first projection data is obtained by projecting the first frame of high-dimensional static biological data using a static projection method, such as the t-SNE (t-Stochastic Neighbor Embedding) method or the UMAP (Uniform Manifold Approximation and Projection) method.
[0061] In operation S220, an nth high-dimensional difference is obtained based on the n-1th frame of high-dimensional static biometric data and the nth frame of high-dimensional static biometric data.
[0062] Among them, the nth high-dimensional difference represents the degree of difference between the n-1th high-dimensional data point group and the nth high-dimensional data point group, the n-1th high-dimensional data point group includes the high-dimensional data points in the n-1th frame of high-dimensional static biological data, and the nth high-dimensional data point group includes the high-dimensional data points in the nth frame of high-dimensional static biological data.
[0063] In operation S230 , nth projection data is obtained based on the nth frame of high-dimensional static biological data, the (n−1)th low-dimensional structural data, and the nth high-dimensional difference.
[0064] In operation S240 , projection data of high-dimensional biological data is obtained based on the first to N-th projection data.
[0065] According to an embodiment of the present invention, continuous projections for high-dimensional biological data can be obtained through the first projection data and the second projection data to the Nth projection data obtained by the above method.
[0066] According to an embodiment of the present invention, for non-first frame high-dimensional static biological data, the n-1th low-dimensional structural data is obtained through the n-1th projection data, thereby introducing the distribution structural characteristics of the n-1th projection data into the generated nth projection data, and then constraining the nth projection data. The nth high-dimensional difference between the n-1th frame high-dimensional static biological data and the nth frame high-dimensional static biological data is used to introduce the degree of difference between the n-1th high-dimensional data point group and the nth high-dimensional data point group, so as to provide the nth projection data with a constraint that is more in line with the change characteristics of the high-dimensional biological data, and then reflect the characteristics of the nth frame high-dimensional static biological data itself, thereby ensuring the temporal continuity between the first projection data to the Nth projection data.
[0067] According to an embodiment of the present invention, the nth high-dimensional difference includes at least one of the following: the nth high-dimensional distance difference, the nth similarity distinction; the nth high-dimensional difference is obtained based on the n-1th frame high-dimensional static biological data and the nth frame high-dimensional static biological data, including: the nth high-dimensional distance difference is obtained based on the n-1th frame high-dimensional static biological data and the nth frame high-dimensional static biological data, wherein the nth high-dimensional distance difference represents the degree of difference between the first high-dimensional distance between each two high-dimensional data points in the n-1th high-dimensional data point group and the second high-dimensional distance between the same two high-dimensional data points in the nth high-dimensional data point group; the nth similarity distinction is obtained based on the n-1th frame high-dimensional static biological data and the nth frame high-dimensional static biological data, wherein the nth similarity distinction represents the degree of difference between the first neighboring high-dimensional data point group corresponding to each high-dimensional data point in the n-1th high-dimensional data point group and the second neighboring high-dimensional data point group corresponding to each high-dimensional data point in the nth high-dimensional data point group.
[0068] According to an embodiment of the present invention, by calculating the nth high-dimensional distance difference and the nth similarity distinction, the changes in high-dimensional static biological data between consecutive frames can be presented in a quantitative manner, thereby being able to more accurately analyze the evolution of high-dimensional biological data in a time series, making the changes in high-dimensional biological data more intuitive.
[0069] According to an embodiment of the present invention, based on the n-1th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data, an nth high-dimensional distance difference is obtained, including: for every two high-dimensional data points in the n-1th frame of high-dimensional static biological data, based on the n-1th high-dimensional data of each two high-dimensional data points, the distance between each two high-dimensional data points is calculated to obtain the n-1th high-dimensional distance; for every two high-dimensional data points in the nth frame of high-dimensional static biological data, based on the nth high-dimensional data of each two high-dimensional data points, the distance between each two high-dimensional data points is calculated to obtain the nth high-dimensional distance; based on the difference between the n-1th high-dimensional distance and the nth high-dimensional distance, the nth sub-high-dimensional distance difference of each two high-dimensional data points is obtained; based on multiple nth sub-high-dimensional distance differences, the nth high-dimensional distance difference is obtained.
[0070] According to an embodiment of the present invention, the distance between each two high-dimensional data points may be a Euclidean distance.
[0071] Exemplarily, there are three high-dimensional data points A, B, and C, and the high-dimensional data corresponding to the three high-dimensional data points all include position information. Then, for the high-dimensional data point A, the first high-dimensional distance is the distance ab between the high-dimensional data point A and the high-dimensional data point B in the n-1th frame of high-dimensional static biological data, and the distance ac between the high-dimensional data point A and the high-dimensional data point C; the second high-dimensional distance is the distance ab' between the high-dimensional data point A and the high-dimensional data point B in the nth frame of high-dimensional static biological data, and the distance ac' between the high-dimensional data point A and the high-dimensional data point C. The nth high-dimensional distance difference can be expressed as (ab-ab') + (ac-ac'), where ab-ab' is the nth sub-high-dimensional distance difference between the high-dimensional data point A and the high-dimensional data point B, and ac-ac' is the nth sub-high-dimensional distance difference between the high-dimensional data point A and the high-dimensional data point B.
[0072] According to an embodiment of the present invention, the n-1th high-dimensional distance is obtained by calculating the distance between every two high-dimensional data points in the n-1th frame of high-dimensional static biological data, and the nth high-dimensional distance is obtained by calculating the distance between every two high-dimensional data points in the nth frame of high-dimensional static biological data, and the nth high-dimensional distance is obtained based on the difference between the n-1th high-dimensional distance and the nth high-dimensional distance, and then the nth high-dimensional distance difference is obtained. This can capture and quantify changes in high-dimensional biological data in the time dimension, thereby constraining the nth projection data based on the nth high-dimensional distance difference between two adjacent frames of high-dimensional static biological data, thereby maintaining the consistency of the projection data of two adjacent frames.
[0073] According to an embodiment of the present invention, based on the n-1th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data, an nth similarity distinction is obtained, including: for any high-dimensional data point among the multiple high-dimensional data points included in the n-1th frame of high-dimensional static biological data, based on the n-1th high-dimensional data of the high-dimensional data point and the n-1th other high-dimensional data point, a first distance between the high-dimensional data point and the n-1th other high-dimensional data point is calculated to obtain the n-1th nearest neighbor high-dimensional data point group of the high-dimensional data point, wherein the first distance between the high-dimensional data point and the n-1th nearest neighbor high-dimensional data point included in the n-1th nearest neighbor high-dimensional data point group is one of I first distances, the I first distances are the first I obtained by arranging the multiple first distances in ascending order, I is an integer greater than 1, and the n-1th other high-dimensional data point is the other high-dimensional data point in the n-1th frame of high-dimensional static biological data except the high-dimensional data point; for the For each high-dimensional data point in n frames of high-dimensional static biological data, based on the nth high-dimensional data of each of the high-dimensional data point and the nth other high-dimensional data point, the second distance between the high-dimensional data point and the nth other high-dimensional data point is calculated to obtain the nth nearest neighbor high-dimensional data point group of the high-dimensional data point, wherein the second distance between the high-dimensional data point and the nth nearest neighbor high-dimensional data point included in the nth nearest neighbor high-dimensional data point group is one of I second distances, the I second distances are the first I obtained by arranging multiple second distances in ascending order, and the nth other high-dimensional data point is other high-dimensional data points in the nth frame of high-dimensional static biological data except the high-dimensional data point; determine the difference between the n-1th nearest neighbor high-dimensional data point group and the nth nearest neighbor high-dimensional data point group of the high-dimensional data point to obtain the nth sub-similarity distinction of the high-dimensional data point; and obtain the nth similarity distinction based on the nth sub-similarity distinction of each of the multiple high-dimensional data points.
[0074] According to an embodiment of the present invention, the k-nearest neighbor method can be used to calculate the I n-1th nearest neighbor high-dimensional data points and the I nth nearest neighbor high-dimensional data points in the n-1th frame of high-dimensional static biological data for the same high-dimensional data point, and calculate the k-nearest neighbor similarity of the I n-1th nearest neighbor high-dimensional data points and the I nth nearest neighbor high-dimensional data points to obtain the nth sub-similarity distinction for the high-dimensional data point, and then obtain the nth similarity distinction.
[0075] For example, there are six high-dimensional data points A, B, C, D, E, and F. For the sake of demonstration, only the nearest high-dimensional data point groups of A and E are calculated here. In practice, they should be calculated based on the nearest high-dimensional data point groups of A, B, C, D, E, and F respectively. In the n-1th frame of high-dimensional static biological data, for high-dimensional data point A, the n-1th nearest high-dimensional data point group is [B, C], and for high-dimensional data point E, the n-1th nearest high-dimensional data point group is [D, F]; in the nth frame of high-dimensional static biological data, For the high-dimensional data point A, the nth nearest neighbor high-dimensional data point group is [B, D], and for the high-dimensional data point E, the nth nearest neighbor high-dimensional data point group is [D, F]. For the high-dimensional data point A, the calculated nth sub-similarity discrimination is 1 (as long as the n-1th nearest neighbor high-dimensional data point group and the nth nearest neighbor high-dimensional data point group are different, they are both counted as 1). For the high-dimensional data point E, the calculated nth sub-similarity discrimination is 0, and the further calculated nth similarity discrimination is (1+0) / 2=0.5.
[0076] According to an embodiment of the present invention, based on the n-1th projection data, obtaining the n-1th low-dimensional structure data includes: clustering the n-1th projection data to generate K clusters, where K is an integer greater than or equal to 1; for any cluster in the K clusters, performing secondary clustering on the cluster to generate M sub-clusters, where M is an integer greater than or equal to 1; for each low-dimensional data point in the n-1th projection data, calculating the low-dimensional data of the low-dimensional data point and the clustering of the K clusters based on the low-dimensional data of the low-dimensional data point. The method comprises the following steps: calculating the distance between the n-1th first sub-low-dimensional structure data and the cluster center of the M sub-cluster clusters based on the low-dimensional data of each of the multiple low-dimensional data points, and obtaining the n-1th second sub-low-dimensional structure data; calculating the distance between the n-1th first sub-low-dimensional structure data and the cluster center of the M sub-cluster clusters based on the low-dimensional data of each of the multiple low-dimensional data points, and obtaining the n-1th second sub-low-dimensional structure data; splicing the K n-1th first sub-low-dimensional structure data and the M n-1th second sub-low-dimensional structure data based on the preset weights, and obtaining the n-1th target sub-low-dimensional structure data; and obtaining the n-1th low-dimensional structure data based on multiple n-1th target sub-low-dimensional structure data.
[0077] According to an embodiment of the present invention, the n-1th projection data can be clustered using the K-Means method to obtain K clusters. The Euclidean distance from the low-dimensional data points in the projection data to the centers of the K clusters is then calculated. For each low-dimensional data point in the projection data, a K-dimensional vector corresponds to it, representing the distance between each low-dimensional data point and the centers of the K clusters. The K-Means method requires specifying the number of clusters K, which can be determined using the elbow method. The elbow method is a heuristic method that determines the optimal K value by plotting the relationship between the sum of squared errors and the number of clusters K. In an embodiment of the present invention, K can be set to 5. When the data scale becomes larger, the K value can be appropriately increased. The present invention does not limit the K value.
[0078] According to an embodiment of the present invention, the M value can be the same as the K value, and in the present invention, both can be set to 5, or different from the K value. For any of the K clusters, secondary clustering is performed to generate M sub-clusters. This hierarchical clustering operation is inspired by the idea of spatial pyramid. Spatial pyramid is a feature representation method widely used in image processing and computer vision. Its core idea is to capture spatial features of different scales by refining the spatial partitioning of the image at multiple levels, and to fuse these features to enhance the expressive power. Specifically, it recursively divides the image by reducing the region size layer by layer, capturing information at different levels from the global to the local, and can further cluster the sub-clusters. However, as the number of iterations increases, the gain in temporal coherence decreases. As the number of pyramid layers increases, the gain in capturing image structure also decreases. Therefore, the number of pyramid layers generally does not exceed 3. In the present invention, 2-3 clustering operations are generally sufficient.
[0079] Figure 3 A schematic diagram of multi-layer clustering according to an embodiment of the present invention is shown.
[0080] like Figure 3 As shown, for the low-dimensional data point yj in the projection data, 5 clusters are obtained in the first layer of clustering, namely cluster 1, cluster 2, cluster 3, cluster 4 and cluster 5, and the distance from yj to the center points of the 5 clusters is calculated to obtain 5 n-1th first sub-low-dimensional structure data. In the second layer of clustering, secondary clustering is performed in cluster 1 to obtain 5 sub-cluster clusters. At the same time, the distance from yj to the center points of the 5 sub-cluster clusters is calculated to obtain 5 n-1th second-sub-low-dimensional structure data. The 5 n-1th first-sub-low-dimensional structure data and the 5 n-1th second-sub-low-dimensional structure data are weightedly spliced to obtain the n-1th target sub-low-dimensional structure data. The n-1th target sub-low-dimensional structure data corresponding to the low-dimensional data of all low-dimensional data points in the aforementioned n-1th projection data constitute the above-mentioned n-1th low-dimensional structure data.
[0081] According to an embodiment of the present invention, by clustering the projection data and dividing it into K clusters, the distribution characteristics in the projection data can be quickly identified. Furthermore, each cluster is further subdivided into M sub-clusters, and the internal structure and differences in each cluster can be analyzed in more detail. The low-dimensional data of the low-dimensional data points are used to calculate the distances to the cluster centers of the K clusters and the distances to the cluster centers of the M sub-clusters, respectively, to obtain K n-1th first sub-low-dimensional structure data and M n-1th second sub-low-dimensional structure data, which can retain global and local feature information at the same time. The K n-1th first sub-low-dimensional structure data and the M n-1th second sub-low-dimensional structure data are further weighted and spliced according to preset weights. The importance of different layers or features can be adjusted according to timing requirements, so that the generated n-1th low-dimensional structure data is more targeted.
[0082] According to an embodiment of the present invention, based on the nth frame high-dimensional static biological data, the n-1th low-dimensional structural data and the nth high-dimensional difference, the nth projection data is obtained, including: splicing the n-1th low-dimensional structural data and the nth high-dimensional difference to obtain the nth spatial structure constraint, wherein the nth spatial structure constraint is used to constrain the temporal continuity of each projection structure between the n-1th projection data and the nth projection data; splicing the nth spatial structure constraint and the nth frame high-dimensional static biological data to obtain the target nth frame high-dimensional static biological data; and obtaining the nth projection data based on the target nth frame high-dimensional static biological data.
[0083] According to an embodiment of the present invention, the Hadamard product operation can be performed on the n-1th low-dimensional structure data and the nth high-dimensional difference to obtain the nth spatial structure constraint. The projection structure can characterize the spatial distribution characteristics of the low-dimensional data points in each cluster. The temporal coherence of the above projection structure means that for each cluster in the projection data, the distribution of its corresponding low-dimensional data points from the first projection data to the Nth projection data remains within a certain range. Figures 4 to 13 To explain the temporal coherence of the above projection structure.
[0084] Figure 4 It shows first projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention.
[0085] Figure 5 The diagram shows second projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention.
[0086] Figure 6 The third projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown.
[0087] Figure 7The fourth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown.
[0088] Figure 8 The fifth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown.
[0089] Figure 9 Shown is sixth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention.
[0090] Figure 10 The seventh projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown.
[0091] Figure 11 FIG. 8 shows eighth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention.
[0092] Figure 12 The ninth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention is shown.
[0093] Figure 13 The figure shows tenth projection data obtained by projecting high-dimensional biological data according to an embodiment of the present invention.
[0094] like Figures 4 to 13 As shown in the figure, n represents the number of frames, and each color represents a cluster label. In the projection data from the 1st frame to the 10th frame, taking the yellow cluster in the figure as an example, the distribution in the projection data is basically maintained within a spatial range, that is, the projection structure is continuous in time and space, and as the number of frames increases, the cluster radius of the cluster decreases accordingly.
[0095] Figure 14 A trajectory result diagram of frame number-based coloring from the first projection data to the tenth projection data according to an embodiment of the present invention is shown.
[0096] Figure 15 A trajectory result diagram of the static projection method based on the related art from the first projection data to the tenth projection data colored based on the frame number is shown.
[0097] Figure 16 A trajectory result diagram based on cluster label coloring from the first projection data to the tenth projection data according to an embodiment of the present invention is shown.
[0098] Figure 17 The figure shows the trajectory result of the static projection method based on the related technology from the first projection data to the tenth projection data colored based on the cluster cluster labels.
[0099] like Figure 14 and Figure 15As shown, the trajectory result diagram based on frame number coloring drawn by projecting high-dimensional biological data according to the method of an embodiment of the present invention is compared with the trajectory effect diagram based on frame number coloring drawn by projecting high-dimensional biological data using the static projection method in the related art. The evolution law of high-dimensional biological data is obvious, and it can be seen that the clustering radius of the cluster cluster gradually decreases with the increase of the frame number.
[0100] like Figure 16 and Figure 17 As shown, the trajectory result diagram based on clustering label coloring drawn by projecting high-dimensional biological data according to the method of an embodiment of the present invention is compared with the trajectory effect diagram based on clustering label coloring drawn by projecting high-dimensional biological data using the static projection method in the related art. The spatial position of the high-dimensional biological data in the projected data changes within a certain range as the number of frames increases.
[0101] Figure 18 A projection flow chart for high-dimensional biological data according to an embodiment of the present invention is shown.
[0102] like Figure 18 As shown, for the n-1th frame of high-dimensional static biological data, it has J high-dimensional data points, each high-dimensional data point has d feature dimensions, and the n-1th frame of high-dimensional static biological data is projected to obtain the n-1th projection data. For each of the J high-dimensional data points, it has an l×k low-dimensional structural data, where k can be the clustering parameter of the clustering method K-Means, and l is the number of layers of the low-dimensional structural data. For each high-dimensional data point, it has a scalar n-th sub-high-dimensional difference. The n-th sub-high-dimensional distance difference and the n-th sub-similarity difference can be weighted and summed to obtain the above-mentioned n-th sub-high-dimensional difference. The J n-th sub-high-dimensional differences constitute the n-th high-dimensional data difference. The n-1th low-dimensional structural data and the n-th high-dimensional difference are spliced by Hadamard multiplication to obtain the n-th spatial structure constraint. The nth spatial structure constraint is column-concatenated with the nth frame high-dimensional static biological data to obtain target nth frame high-dimensional static biological data, thereby projecting the target nth frame high-dimensional static biological data to obtain nth projection data.
[0103] According to an embodiment of the present invention, the first projection data is obtained by the following method: for any high-dimensional data point among the multiple high-dimensional data points included in the first frame of high-dimensional static biological data, the similarity between each high-dimensional data point is calculated based on the first high-dimensional data of each of the multiple high-dimensional data points to obtain a high-dimensional similarity probability matrix; for each low-dimensional data point in the initial projection data, the similarity between each low-dimensional data point is calculated based on the low-dimensional data of each of the multiple low-dimensional data points to obtain a low-dimensional similarity probability matrix; the relative entropy between the high-dimensional similarity probability matrix and the low-dimensional similarity probability matrix is calculated; the low-dimensional data points in the initial projection data are adjusted based on the relative entropy until the relative entropy meets the first preset condition to obtain the first projection data.
[0104] For the high-dimensional data point x in the first high-dimensional data i′ 、x j′ 、x m′ and x n′ , x can be calculated by the following formula (1) i′ and x j′ The similarity between the corresponding high-dimensional data.
[0105] (1);
[0106] Among them, p i′j′ Represents a high-dimensional data point x i′ and x j′ The similarity between i′ is related to the high-dimensional data point x i′ The standard deviation of the associated Gaussian distribution, used to control the range of similarity calculation, σ m′ is related to the high-dimensional data point x m′ The standard deviation of the related Gaussian distribution, in the case where the first high-dimensional data includes S high-dimensional data points, x i′ 、x j′ 、x m′ and x n′ are the i′, j′, m′, and n′th high-dimensional data points, i′, j′, m′, and n′ are all integers greater than 1 and less than or equal to S. S is an integer greater than or equal to 1.
[0107] According to an embodiment of the present invention, high-dimensional data points are mapped to low-dimensional space to obtain initial projection data. For the low-dimensional data point y in the initial projection data, i′ 、y j′ 、y m′ and y n′ , the similarity between low-dimensional data points can be calculated by the following formula (2).
[0108] (2);
[0109] Among them, q i′j′ Represents a low-dimensional data point y i′ and y j′ The similarity between them, when the first high-dimensional data includes S high-dimensional data points, y i′ 、y j′ 、y m′ and y n′ are the i′, j′, m′, and n′th high-dimensional data points, i′, j′, m′, n′ are all integers greater than 1 and less than or equal to S.
[0110] According to an embodiment of the present invention, the above formula (2) adopts the t distribution, which has a wider tail and can better handle data distribution in low-dimensional space.
[0111] The relative entropy between the high-dimensional similarity probability matrix and the low-dimensional similarity probability matrix can be calculated using the following formula (3).
[0112] (3);
[0113] in, is a high-dimensional data point x i′ and x j′ The similarity between the two and the low-dimensional data points y i′ and y j′ The relative entropy of similarity between i′j′ is a high-dimensional data point x in a high-dimensional space i′ and x j′ The similarity between i′j′ is a low-dimensional data point y in the low-dimensional space i′ and y j′ The similarities between them.
[0114] According to an embodiment of the present invention, the position of the low-dimensional data point can be updated by calculating the gradient of the relative entropy until the relative entropy meets a first preset condition. The first preset condition may be, for example, that the relative entropy converges to a smaller value or reaches a preset maximum number of iterations.
[0115] According to an embodiment of the present invention, the above-mentioned first projection data is obtained by the following method: for each high-dimensional data point in the first frame of high-dimensional static biological data, based on the first high-dimensional data of each high-dimensional data point and the first other high-dimensional data point, the distance between the high-dimensional data point and the first other high-dimensional data point is calculated to obtain the first neighboring high-dimensional data point group of the high-dimensional data point; based on the first high-dimensional data of each of the multiple high-dimensional data points included in the first neighboring high-dimensional data point group, the fuzzy similarity weights between each high-dimensional data point are calculated to obtain a high-dimensional fuzzy similarity matrix; for each low-dimensional data point in the initial projection data, based on the low-dimensional data of each of the multiple low-dimensional data points, the fuzzy similarity between each two low-dimensional data points is calculated to obtain a low-dimensional fuzzy similarity matrix; the cross-entropy loss value between the high-dimensional fuzzy similarity matrix and the low-dimensional fuzzy similarity matrix is calculated; based on the cross-entropy loss value, the low-dimensional data points in the initial projection data are adjusted until the loss value meets the second preset condition to obtain the first projection data.
[0116] According to an embodiment of the present invention, for each high-dimensional data point x i′ , we can find the first nearest neighbor high-dimensional data point group kNN (x i′ ), x can be calculated by the following formula (4) i′ and x j′ The high-dimensional fuzzy similarity weight between them.
[0117] (4);
[0118] Among them, p′ i′j′ Represents x i′ and x j′ The high-dimensional fuzzy similarity weight between i′ is related to the high-dimensional data point x i′ The standard deviation of the associated Gaussian distribution, determined using a binary search.
[0119] According to an embodiment of the present invention, high-dimensional data points may be randomly initialized or implemented through other dimensionality reduction methods, such as PCA (Principal Component Analysis).
[0120] According to an embodiment of the present invention, the low-dimensional data point y can be calculated using the above formula (2) i′ and y j′ The low-dimensional fuzzy similarity weight q′ between i′j′ .
[0121] The cross entropy loss value between the high-dimensional fuzziness matrix and the low-dimensional fuzziness matrix can be calculated by the following formula (5).
[0122] (5);
[0123] Among them, H(P,Q) represents the high-dimensional data point x i′ and x j′ The high-dimensional fuzzy similarity weight p′ between i′j′ With low-dimensional data points y i′ and y j′ The low-dimensional fuzzy similarity weight q′ between i′j′ The cross entropy between them, P represents the high-dimensional fuzzy similarity matrix, Q represents the low-dimensional fuzzy similarity matrix, p′ i′j′ Represents a high-dimensional data point x i′ and x j′ The high-dimensional fuzzy similarity weight between i′j′ Represents a low-dimensional data point y i′ and y j′ The low-dimensional fuzzy similarity weight between them.
[0124] According to an embodiment of the present invention, the above-mentioned projection method for high-dimensional biological data can be applied to the determination of the metagenome of a microbial community, specifically including: obtaining the metagenome sequencing data of the microbial community, the metagenome sequencing data including N frames of metagenome sequencing static data, the N frames of metagenome sequencing static data representing the metagenome abundance on the nth day; subjecting the metagenome sequencing data to the above-mentioned projection method for high-dimensional biological data to obtain the projection data of the metagenome sequencing data of the microbial community, wherein the nth high-dimensional difference can be the species abundance vector of the microbial community; determining the abnormal drift data of the antibiotic disturbance on the nth day based on the projection data of the metagenome sequencing data, and determining the recovery time of the microbial community based on the abnormal drift data of the antibiotic disturbance on the nth day.
[0125] Based on the above projection method for high-dimensional biological data, the present invention also provides a projection device for high-dimensional biological data. Figure 19 The device is described in detail.
[0126] Figure 19 A structural block diagram of a projection device for high-dimensional biological data according to an embodiment of the present invention is shown.
[0127] like Figure 19 As shown, the projection device 1900 for high-dimensional biological data of this embodiment includes a low-dimensional structural data determination module 1910 , a high-dimensional difference determination module 1920 , a projection data determination module 1930 and a dynamic projection module 1940 .
[0128] Low-dimensional structure data determination module 1910 is configured to obtain, based on the n-1th projection data, the n-1th low-dimensional structure data for the n-th frame of high-dimensional static biological data among N frames of high-dimensional static biological data, where the n-1th projection data represents the projection data of the n-1th frame of high-dimensional static biological data, the n-1th low-dimensional structure data represents the distribution structure characteristics of the n-1th projection data, and the first projection data represents the projection data of the first frame of high-dimensional static biological data, where n∈{2,…,N}. In one embodiment, low-dimensional structure data determination module 1910 can be configured to perform operation S210 described above and will not be further described herein.
[0129] High-dimensional difference determination module 1920 is configured to obtain an nth high-dimensional difference based on the n-1th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data, wherein the nth high-dimensional difference represents the degree of difference between the n-1th high-dimensional data point group and the nth high-dimensional data point group, the n-1th high-dimensional data point group including the high-dimensional data points in the n-1th frame of high-dimensional static biological data, and the nth high-dimensional data point group including the high-dimensional data points in the nth frame of high-dimensional static biological data. In one embodiment, high-dimensional difference determination module 1920 can be configured to perform operation S220 described above and will not be further described here.
[0130] The projection data determination module 1930 is configured to obtain nth projection data based on the nth frame of high-dimensional static biological data, the n-1th low-dimensional structural data, and the nth high-dimensional difference. In one embodiment, the projection data determination module 1930 may be configured to perform the operation S230 described above, which will not be described in detail here.
[0131] The dynamic projection module 1940 is configured to obtain projection data of the high-dimensional biological data based on the first projection data to the Nth projection data. In one embodiment, the dynamic projection module 1940 may be configured to perform the operation S240 described above, which will not be described in detail here.
[0132] According to embodiments of the present invention, any multiple modules among the low-dimensional structure data determination module 1910, the high-dimensional difference determination module 1920, the projection data determination module 1930, and the dynamic projection module 1940 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present invention, at least one of the low-dimensional structure data determination module 1910, the high-dimensional difference determination module 1920, the projection data determination module 1930, and the dynamic projection module 1940 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or implemented in any one of software, hardware, and firmware, or any suitable combination of these. Alternatively, at least one of the low-dimensional structure data determination module 1910, the high-dimensional difference determination module 1920, the projection data determination module 1930 and the dynamic projection module 1940 can be at least partially implemented as a computer program module, which can perform the corresponding function when it is executed.
[0133] Figure 20 A block diagram of an electronic device suitable for implementing a projection method for high-dimensional biological data according to an embodiment of the present invention is shown.
[0134] like Figure 20 As shown, electronic device 2000 according to an embodiment of the present invention includes a processor 2001, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 2002 or programs loaded from storage 2008 into random access memory (RAM) 2003. Processor 2001 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or related chipsets and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. Processor 2001 may also include onboard memory for caching purposes. Processor 2001 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0135] The RAM 2003 stores various programs and data required for the operation of the electronic device 2000. The processor 2001, ROM 2002, and RAM 2003 are connected to each other via a bus 2004. The processor 2001 executes the programs in the ROM 2002 and / or RAM 2003 to perform the various operations of the method flow according to the embodiment of the present invention. It should be noted that the programs may also be stored in one or more memories other than the ROM 2002 and RAM 2003. The processor 2001 may also execute the programs stored in the one or more memories to perform the various operations of the method flow according to the embodiment of the present invention.
[0136] According to an embodiment of the present invention, electronic device 2000 may further include an input / output (I / O) interface 2005, which is also connected to bus 2004. Electronic device 2000 may also include one or more of the following components connected to I / O interface 2005: an input unit 2006 including a keyboard, mouse, etc.; an output unit 2007 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage unit 2008 including a hard disk; and a communication unit 2009 including a network interface card such as a LAN card or modem. Communication unit 2009 performs communication processing via a network such as the Internet. A drive 2010 is also connected to I / O interface 2005 as needed. Removable media 2011, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 2010 as needed, so that computer programs read from the removable media can be installed into storage unit 2008 as needed.
[0137] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the method according to the embodiments of the present invention.
[0138] According to an embodiment of the present invention, a computer-readable storage medium may be a non-volatile computer-readable storage medium, and may include, for example, but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, a computer-readable storage medium may include the ROM 2002 and / or RAM 2003 described above, and / or one or more memories other than ROM 2002 and RAM 2003.
[0139] Embodiments of the present invention also include a computer program product comprising a computer program containing program code for executing the method shown in the flowchart. When the computer program product is executed in a computer system, the program code causes the computer system to implement the projection method for high-dimensional biological data provided by the embodiments of the present invention.
[0140] The computer program executes the above functions defined in the system / device of the embodiment of the present invention when executed by the processor 2001. According to the embodiment of the present invention, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0141] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 2009, and / or installed from a removable medium 2011. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0142] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 2009 and / or installed from the removable medium 2011. When the computer program is executed by the processor 2001, the above-described functions defined in the system of the embodiment of the present invention are performed. According to the embodiment of the present invention, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0143] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0144] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0145] It will be understood by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention may be combined and / or coupled in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or couplings fall within the scope of the present invention.
[0146] The above describes embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.
Claims
1. A projection method for high-dimensional biological data, characterized in that: The high-dimensional biological data includes N frames of high-dimensional static biological data, each frame of high-dimensional static biological data includes high-dimensional data corresponding to a plurality of high-dimensional data points, and N is an integer greater than or equal to 1; For the nth frame of high-dimensional static biological data among the N frames of high-dimensional static biological data, obtaining n-1th low-dimensional structural data based on the n-1th projection data, wherein the n-1th projection data represents the projection data of the n-1th frame of high-dimensional static biological data, the n-1th low-dimensional structural data represents the distribution structure characteristics of the n-1th projection data, and the first projection data represents the projection data of the first frame of high-dimensional static biological data, n∈{2, ..., N}; Based on the n-1th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data, obtaining an nth high-dimensional difference, wherein the nth high-dimensional difference represents a degree of difference between the n-1th high-dimensional data point group and the nth high-dimensional data point group, the n-1th high-dimensional data point group includes high-dimensional data points in the n-1th frame of high-dimensional static biological data, and the nth high-dimensional data point group includes high-dimensional data points in the nth frame of high-dimensional static biological data; Obtaining nth projection data based on the nth frame of high-dimensional static biological data, the n-1th low-dimensional structural data, and the nth high-dimensional difference; Projection data of the high-dimensional biological data is obtained based on the first projection data to the Nth projection data.
2. The method according to claim 1, characterized in that The nth high-dimensional difference includes at least one of the following: nth high-dimensional distance difference, nth similarity difference; The obtaining of the n-th high-dimensional difference based on the n-1-th frame of high-dimensional static biological data and the n-th frame of high-dimensional static biological data includes: Obtaining an nth high-dimensional distance difference based on the n-1th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data, wherein the nth high-dimensional distance difference represents a degree of difference between a first high-dimensional distance between every two high-dimensional data points in the n-1th high-dimensional data point group and a second high-dimensional distance between the same two high-dimensional data points in the nth high-dimensional data point group; An nth similarity distinction is obtained based on the n-1th frame high-dimensional static biological data and the nth frame high-dimensional static biological data, wherein the nth similarity distinction represents the degree of difference between a first neighboring high-dimensional data point group corresponding to each high-dimensional data point in the n-1th high-dimensional data point group and a second neighboring high-dimensional data point group corresponding to each high-dimensional data point in the nth high-dimensional data point group.
3. The method according to claim 2, characterized in that The obtaining of the nth high-dimensional distance difference based on the n-1th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data includes: For every two high-dimensional data points in the n-1th frame of high-dimensional static biological data, calculating the distance between the every two high-dimensional data points based on the n-1th high-dimensional data of each of the every two high-dimensional data points to obtain the n-1th high-dimensional distance; For each two high-dimensional data points in the n-th frame of high-dimensional static biological data, calculating the distance between each two high-dimensional data points based on the n-th high-dimensional data of each of the two high-dimensional data points to obtain an n-th high-dimensional distance; Obtaining the nth sub-high-dimensional distance difference between each two high-dimensional data points based on the difference between the n-1th high-dimensional distance and the nth high-dimensional distance; The nth high-dimensional distance difference is obtained based on a plurality of the nth sub-high-dimensional distance differences.
4. The method according to claim 2 or 3, characterized in that The obtaining of an nth similarity distinction based on the n-1th frame of high-dimensional static biological data and the nth frame of high-dimensional static biological data includes: For any high-dimensional data point among the multiple high-dimensional data points included in the n-1th frame of high-dimensional static biological data, based on the n-1th high-dimensional data of each of the high-dimensional data point and the n-1th other high-dimensional data point, calculate the first distance between the high-dimensional data point and the n-1th other high-dimensional data point to obtain the n-1th nearest neighbor high-dimensional data point group of the high-dimensional data point, wherein the first distance between the high-dimensional data point and the n-1th nearest neighbor high-dimensional data point included in the n-1th nearest neighbor high-dimensional data point group is one of I first distances, the I first distances are the first I obtained by arranging the multiple first distances in ascending order, I is an integer greater than 1, and the n-1th other high-dimensional data point is the other high-dimensional data point in the n-1th frame of high-dimensional static biological data except the high-dimensional data point; For each high-dimensional data point in the n-th frame of high-dimensional static biological data, based on the n-th high-dimensional data of each of the high-dimensional data point and the n-th other high-dimensional data point, a second distance between the high-dimensional data point and the n-th other high-dimensional data point is calculated to obtain an n-th nearest neighbor high-dimensional data point group of the high-dimensional data point, wherein the second distance between the high-dimensional data point and the n-th nearest neighbor high-dimensional data point included in the n-th nearest neighbor high-dimensional data point group is one of I second distances, the I second distances being the first I of a plurality of second distances arranged in ascending order, and the n-th other high-dimensional data point being other high-dimensional data points in the n-th frame of high-dimensional static biological data except the high-dimensional data point; Determining the difference neighboring high-dimensional data points between the n-1th neighboring high-dimensional data point group and the nth neighboring high-dimensional data point group of the high-dimensional data point, and obtaining an nth sub-similarity distinctiveness of the high-dimensional data point; The nth similarity distinction is obtained based on the nth sub-similarity distinctions of the multiple high-dimensional data points.
5. The method according to any one of claims 1 to 3, characterized in that The step of obtaining the n-1th low-dimensional structure data based on the n-1th projection data includes: Clustering the n-1th projection data to generate K clusters, where K is an integer greater than or equal to 1; For any cluster among the K clusters, perform secondary clustering on the cluster to generate M sub-clusters, where M is an integer greater than or equal to 1; For each low-dimensional data point in the n-1th projection data, calculating the distance from the cluster center of the K clusters based on the low-dimensional data of the low-dimensional data point to obtain K n-1th first sub-low-dimensional structure data; Calculating the distances from the cluster centers of the M sub-clusters based on the low-dimensional data of each of the plurality of low-dimensional data points to obtain M (n-1) second sub-low-dimensional structure data; Based on preset weights, the K n-1th first sub-low-dimensional structure data and the M n-1th second sub-low-dimensional structure data are concatenated to obtain the n-1th target sub-low-dimensional structure data; Based on the plurality of the n-1th target sub-low-dimensional structure data, the n-1th low-dimensional structure data is obtained.
6. The method according to any one of claims 1 to 3, characterized in that The obtaining of nth projection data based on the nth frame of high-dimensional static biological data, the n-1th low-dimensional structural data, and the nth high-dimensional difference comprises: splicing the n-1th low-dimensional structure data and the nth high-dimensional difference to obtain an nth spatial structure constraint, wherein the nth spatial structure constraint is used to constrain the temporal coherence of each projection structure between the n-1th projection data and the nth projection data; splicing the nth spatial structure constraint and the nth frame of high-dimensional static biological data to obtain target nth frame of high-dimensional static biological data; The nth projection data is obtained based on the nth frame of high-dimensional static biological data of the target.
7. The method according to any one of claims 1 to 3, characterized in that The first projection data is obtained by the following method: For any high-dimensional data point among the plurality of high-dimensional data points included in the first frame of high-dimensional static biological data, calculating similarities between the plurality of high-dimensional data points based on the first high-dimensional data of each of the plurality of high-dimensional data points to obtain a high-dimensional similarity probability matrix; For each low-dimensional data point in the initial projection data, based on the low-dimensional data of each of the plurality of low-dimensional data points, calculating the similarity between the low-dimensional data points to obtain a low-dimensional similarity probability matrix; Calculating the relative entropy between the high-dimensional similarity probability matrix and the low-dimensional similarity probability matrix; Low-dimensional data points in the initial projection data are adjusted based on the relative entropy until the relative entropy satisfies a first preset condition, thereby obtaining the first projection data.
8. The method according to any one of claims 1 to 3, characterized in that The first projection data is obtained by the following method: For each high-dimensional data point in the first frame of high-dimensional static biological data, calculating a distance between the high-dimensional data point and the first other high-dimensional data point based on the first high-dimensional data of each high-dimensional data point, to obtain a first group of neighboring high-dimensional data points of the high-dimensional data point; Calculating the fuzzy similarity weights between the high-dimensional data points based on the first high-dimensional data of each of the plurality of high-dimensional data points included in the first group of neighboring high-dimensional data points to obtain a high-dimensional fuzzy similarity matrix; For each low-dimensional data point in the initial projection data, based on the low-dimensional data of each of the plurality of low-dimensional data points, calculating the fuzzy similarity between each two of the low-dimensional data points to obtain a low-dimensional fuzzy similarity matrix; Calculating a cross entropy loss value between the high-dimensional fuzzy similarity matrix and the low-dimensional fuzzy similarity matrix; Low-dimensional data points in the initial projection data are adjusted based on the cross entropy loss value until the loss value satisfies a second preset condition, thereby obtaining the first projection data.
9. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Method for identifying underground coal mine personnel based on gait identification
CN109241870A
High-dimensional data clustering method, device and equipment, medium and product
CN119128566A
Distributed similarity learning for high-dimensional image features
US20150146973A1