Method and device for generating a map representative of an n-dimensional dataset, and use of the map to analyse biological drift in a subject
By generating a representative mapping of N-dimensional biological data using a graph-based method with convex hulls, the method addresses the limitations of linear modeling in health diagnostics, enhancing diagnostic precision and early detection of abnormalities.
Patent Information
- Application Number
- PCT/EP2025/058596
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-28
- Filing Date
- 2025-03-28
- Publication Date
- 2025-10-02
AI Technical Summary
Existing data analysis methods, particularly in health diagnostics, often rely on linear modeling which is inadequate for nonlinear data spaces, failing to account for correlations between biological parameters and leading to inaccurate or erroneous results.
A method and device for generating a representative mapping of an N-dimensional data set using a graph with vertices and edges, constructing convex hulls around neighborhoods of points, and forming a mapping by the union of these hulls, allowing for the analysis of biological drift by comparing a subject's parameters to a reference population.
This approach provides a more precise analysis of biological parameters, enabling earlier detection of abnormal health states and accounting for correlations, thus improving diagnostic accuracy.
Smart Images

Figure EP2025058596_02102025_PF_FP_ABST
Abstract
Description
[0001] DESCRIPTION
[0002] Method and device for generating a representative map of an N-dimensional data set, and use of said map for analyzing biological drift in a subject
[0003] The present invention relates to a method and a device for generating a representative mapping of an N-dimensional data set, and an associated computer program.
[0004] The invention also relates to uses of the mapping thus generated for the analysis of the biological drift of at least one of a number N greater than or equal to two of quantitative biological parameters in a subject.
[0005] In many application areas, for example in the monitoring and predictive maintenance of industrial systems or diagnostics based on physiological variables, values of many parameters are measured.
[0006] The exploitation of parameter values, by data analysis methods, particularly for prevention and diagnostic assistance applications, is a rapidly growing field.
[0007] However, most known models are based, explicitly or implicitly, on assumptions of linear modeling of the data space. This is the case, for example, when data analysis uses linear regression.
[0008] However, in many application areas, the data to be analyzed, for example from measurements of various physical or biological parameters, actually belong to a nonlinear subspace. In such a case, classical linear models and linear metrics are not applicable, or they provide an inaccurate or even erroneous result.
[0009] In particular, in the field of health, the analysis of biological parameters (or characteristics) in a subject is an essential tool to aid in the diagnosis of a pathology.
[0010] Indeed, there are more or less direct links between certain biological parameters and given pathologies, for example between blood sugar and diabetes, iron levels and anemia, blood pressure and hypertension, the presence of a pathogen and an infection linked to this pathogen, sedimentation rate and an infection or inflammation, etc. The measurement of certain biological parameters thus makes it possible to guide or confirm a diagnosis. It also makes it possible to monitor a patient's response to treatment, by following the evolution of one or more parameters linked to the disease concerned.
[0011] The analysis of a biological parameter generally consists of measuring the value of this biological parameter and then determining whether the value obtained is within a range of reference values defined by a minimum reference threshold value and a maximum reference threshold value. Thus, the results of a biological analysis are generally classified into three categories: (i) below the minimum reference threshold value, (ii) between the minimum and maximum reference threshold values or (iii) above the maximum reference threshold value. A value falling within a range of reference threshold values will be considered "normal". The "normality" of a value that falls outside a range of reference threshold values is left to the judgment of the practitioner, in particular depending on the deviation of this value from one of the reference threshold values.
[0012] The analysis of biological parameters can also concern several biological parameters. Conventionally, each biological parameter has associated minimum reference threshold values and maximum reference threshold values, and the analysis of several biological parameters is carried out by comparing each of the biological parameters considered to the reference threshold values associated with it. If we consider N biological parameters, the estimation of normality amounts to checking whether the values of the N biological parameters for a subject belong to the N-dimensional hyper-cube formed from the reference threshold values for each of the parameters. However, such an analysis does not allow for taking into account a possible correlation between the biological parameters considered.
[0013] The invention aims to remedy the aforementioned drawbacks by providing a method and device for modeling an underlying subspace representative of an N-dimensional data set of a group of subjects.
[0014] A particular application is the application to the in vitro analysis of the biological drift of one or more quantitative biological parameters among N quantitative biological parameters considered jointly, to provide more precise biological analysis methods and / or allowing earlier detection of one or more biological parameters whose values correspond to an abnormal or potentially abnormal state of health.
[0015] For this purpose, the subject of the invention is a method for generating a mapping representative of an N-dimensional data set, N being an integer greater than or equal to 2, each dimension being associated with a given physical or biological parameter of a subject, the method being implemented by a processor of an electronic computing device, the method comprising an acquisition of said N-dimensional data set of a set of P subjects, in the form of P sets of N values, forming coordinates of P points in an N-dimensional subspace, P being an integer greater than or equal to two. This method further comprises steps, implemented by the computing processor, of: calculating a graph representing said P points, the graph comprising vertices and edges, each edge connecting two vertices together, in which for a given vertex, the set of vertices connected to said vertex by an edge forms a neighborhood of said given vertex in the graph,wherein each vertex of the graph corresponds to a first point of said subspace, and its neighborhood comprises vertices corresponding to second points of said subspace, the second points being the k nearest neighbors of the first point in the sense of a predetermined metric, k being a chosen positive integer; for each point of said subspace having an associated vertex in said graph, construction of an associated convex hull, encompassing the points corresponding to the vertices of the neighborhood of said associated vertex, said mapping being formed by the union of the convex hulls associated with the points of the subspace.,
[0016] Advantageously, the method makes it possible to faithfully map a reference population of P subjects according to N distinct biological parameters.
[0017] According to other advantageous aspects of the invention, the method for generating a mapping representative of an N-dimensional data set comprises one or more of the following characteristics, taken in isolation or in all technically possible combinations.
[0018] The method further comprises a calculation of at least one numerical value associated with each convex hull, the or each numerical value being calculated from an attribute of the convex hull.
[0019] The attribute of the convex hull is representative of a local density of points belonging to said convex hull.
[0020] For a given attribute, the numeric value is a quantile of a value of the attribute in a statistical distribution of the values of that attribute.
[0021] The method comprises an association of a weight to each edge of the graph, as a function of the distance, according to said metric, between the points corresponding to the vertices of the graph connected by said edge.
[0022] The metric is for example the Euclidean distance in N dimensions. The graph calculation step includes a sub-step of connecting unconnected components, each component being a sub-graph, the connection step comprising an addition of an edge between vertices each belonging to a distinct unconnected component, and an addition of a connecting neighbor in the neighborhood of the points corresponding to said vertices.
[0023] The graph calculation step includes a sub-step of symmetrization of the neighborhoods associated with the P points.
[0024] The method further comprises an identification and storage of extreme points of the union of the convex envelopes, each extreme point being an external point of one of said convex envelopes and which does not belong to another of the convex envelopes.
[0025] The method further comprises an update of the mapping formed from the first N-dimensional data of a first set of P subjects, forming P first points in the N-dimensional subspace, using second additional N-dimensional data relating to a second set of Q subjects, forming Q second points in the N-dimensional subspace, the update comprising a modification of the edges of the graph according to the second additional N-dimensional data making it possible to obtain a modified graph, a symmetrization of the edges of the modified graph and a reconstruction of convex hulls of the first points and of the second points whose neighborhood is modified by the modification and symmetrization steps.
[0026] The method further comprises a step of enriching the mapping by adding at least one set of synthetic data, by random sampling in at least one of the convex hulls of said mapping.
[0027] The invention also relates to a device for generating a mapping representative of an N-dimensional data set, N being an integer greater than or equal to 2, each dimension being associated with a given physical or biological parameter of a subject, the device being configured to acquire the N-dimensional data set in the form of P sets of N values, forming coordinates of P points in an N-dimensional subspace, P being an integer greater than or equal to two, the device comprising a calculation processor configured to implement: a module for calculating a graph representing said P points, the graph comprising vertices and edges, each edge connecting two vertices together, in which for a given vertex, the set of vertices connected to said vertex by an edge forms a neighborhood of said given vertex in the graph, in which each vertex of the graph corresponds to a first point of said subspace,and its neighborhood comprises vertices corresponding to second points of said subspace, the second points being the k nearest neighbors of the first point in the sense of a predetermined metric, k being a chosen positive integer; a module for constructing convex hulls configured, for each point of said subspace having an associated vertex in said graph, to construct an associated convex hull, encompassing the points corresponding to the vertices of the neighborhood of said associated vertex, a module for forming said mapping, said mapping being formed by the union of the convex hulls associated with the points of the subspace.,
[0028] The device for generating a mapping representative of an N-dimensional data set is configured to implement the method for generating a mapping representative of an N-dimensional data set as described above, according to all its embodiments.
[0029] The invention also relates to a computer program comprising software instructions which, when executed by a computer, implement a method for generating a mapping representative of an N-dimensional data set as defined above.
[0030] According to another aspect, the invention relates to a method for analyzing a biological drift of at least one of a number N greater than or equal to two of quantitative biological parameters in a subject, the method comprising steps, implemented by a calculation processor, of:
[0031] -for a reference population formed of a group of P subjects, obtain values of said N quantitative biological parameters forming an N-dimensional data set,
[0032] -generate a mapping representative of the N-dimensional data set, the generation of a mapping comprising steps of: calculating a graph representing said P points, the graph comprising vertices and edges, each edge connecting two vertices together, in which for a given vertex, the set of vertices connected to said vertex by an edge forms a neighborhood of said given vertex in the graph, in which each vertex of the graph corresponds to a first point of said subspace, and its neighborhood comprises vertices corresponding to second points of said subspace, the second points being the k closest neighbors of the first point in the sense of a predetermined metric, k being a chosen positive integer;for each point of said subspace having an associated vertex in said graph, construction of an associated convex hull, encompassing the points corresponding to the vertices of the neighborhood of said associated vertex, said mapping being formed by the union of the convex hulls associated with the points of the subspace of the N-dimensional space, said mapping being a reference mapping associated with the reference population, the analysis method further comprising:;
[0033] - obtaining values of the N quantitative biological parameters of said subject, - determining whether a point relative to said subject in the N-dimensional space of coordinates equal to the values of the N quantitative biological parameters of said subject is located inside the union of convex envelopes,
[0034] - if the said point relating to the said subject is not located within the union of convex envelopes, conclude that the combination of the said N biological parameters of the subject is potentially abnormal in view of the combinations of the said N parameters observed in the reference subjects.
[0035] Advantageously, the proposed method allows a subject to be compared more precisely with a reference population, according to N quantitative biological parameters.
[0036] According to other advantageous aspects of the invention, the method for analyzing a biological drift comprises one or more of the following characteristics, taken in isolation or in all technically possible combinations.
[0037] The generation of a mapping further comprises a calculation of at least one numerical value associated with each convex hull, the or each numerical value being calculated from an attribute of the convex hull.
[0038] The attribute of the convex hull is representative of a local density of points belonging to said convex hull.
[0039] For a given attribute, the numeric value is a quantile of a value of the attribute in a statistical distribution of the values of that attribute.
[0040] The generation of a mapping further includes an association of a weight to each edge of the graph, as a function of the distance, according to said metric, between the points corresponding to the vertices of the graph connected by said edge.
[0041] According to one embodiment, the metric is the Euclidean distance in N dimensions.
[0042] In the mapping generation, the graph calculation step comprises a sub-step of connecting unconnected components, each component being a sub-graph, the connection step comprising an addition of an edge between vertices each belonging to a distinct unconnected component, and an addition of a connection neighbor in the neighborhood of the points corresponding to said vertices.
[0043] In mapping generation, the graph computation step includes a sub-step of symmetrization of the neighborhoods associated with the P points.
[0044] The mapping generation step further comprises an identification and storage of extreme points of the union of the convex hulls, each extreme point being an external point of one of said convex hulls and which does not belong to another of the convex hulls.
[0045] The method further comprises an update of the mapping formed from the first N-dimensional data of a first set of P subjects, forming P first points in the N-dimensional subspace, using second additional N-dimensional data relating to a second set of Q subjects, forming Q second points in the N-dimensional subspace, the update comprising a modification of the edges of the graph according to the second additional N-dimensional data making it possible to obtain a modified graph, a symmetrization of the edges of the modified graph and a reconstruction of convex hulls of the first points and of the second points whose neighborhood is modified by the modification and symmetrization steps.
[0046] The method further comprises a step of enriching the mapping by adding at least one set of synthetic data, by random sampling in at least one of the convex hulls of said mapping.
[0047] Obtaining the values of the N quantitative biological parameters comprises obtaining measured values for Nm quantitative biological parameters among the N quantitative biological parameters, m being an integer, and an estimation, for the m missing quantitative biological parameters, according to an intersection between a geometric hyperplane associated with the subject, said geometric hyperplane associated with the subject being by said measured values of the Nm quantitative biological parameters in the N-dimensional space, and the calculated mapping.
[0048] The method of analyzing a biological drift includes calculating a drift value associated with the subject from the k closest neighbors of the point relating to said subject among the points of the mapping.
[0049] According to a variant, the method for analyzing a biological drift comprises a generation of a plurality of maps, each map being representative of a distinct reference population, associated with a pathology or a specific characteristic, and a calculation, for said subject, of a drift value in relation to each map, and a calculation of an overall drift value for said subject by combining the calculated drift values.
[0050] According to other advantageous aspects of the invention, each convex hull of said mapping has an associated numerical value, the numerical value being calculated from an attribute of the convex hull, the method for analyzing a biological drift further comprising, if said point relating to said subject is located inside the union of convex hulls, identifying the convex hull(s) containing the point relating to said subject, and deducing therefrom a level of drift as a function of the numerical value(s) associated with the hull(s) containing the point relating to said subject.
[0051] According to other advantageous aspects of the invention, the attribute is representative of a density of points of the convex hull, and a division of the drift levels is carried out according to quantile intervals of said attribute.
[0052] According to other advantageous aspects, the method for analyzing a biological drift further comprises a repetition of the calculation of a biological drift value for successive time instants, and a calculation of a temporal evolution of the biological drift of the subject.
[0053] According to other advantageous aspects, the method for analyzing a biological drift further comprises a transmission of the calculated biological drift values to an alert system.
[0054] According to another aspect, the invention relates to a device for analyzing a biological drift of at least one of a number N greater than or equal to two of quantitative biological parameters in a subject, the analysis device comprising a calculation processor, configured to implement modules configured to:
[0055] -for a reference population formed of a group of P subjects, obtain values of said N quantitative biological parameters forming an N-dimensional data set,
[0056] -generate a mapping representative of the N-dimensional data set, comprising a calculation of a graph representing said P points, the graph comprising vertices and edges, each edge connecting two vertices together, in which for a given vertex, the set of vertices connected to said vertex by an edge forms a neighborhood of said given vertex in the graph, in which each vertex of the graph corresponds to a first point of said subspace, and its neighborhood comprises vertices corresponding to second points of said subspace, the second points being the k closest neighbors of the first point in the sense of a predetermined metric, k being a chosen positive integer;for each point of said subspace having an associated vertex in said graph, constructing an associated convex hull, encompassing the points corresponding to the vertices of the neighborhood of said associated vertex, said mapping (30) being formed by the union of the convex hulls associated with the points of the subspace of the two-dimensional space, said mapping being a reference mapping associated with the reference population;
[0057] - obtain values of the N quantitative biological parameters of said subject,
[0058] -determine whether a point relative to said subject in the N-dimensional space of coordinates equal to the values of the N quantitative biological parameters of said subject is located inside the union of convex envelopes,
[0059] - if the said point relating to the said subject is not located inside the union of convex envelopes, conclude that the combination of the said N biological parameters of the subject is potentially abnormal in view of the combinations of the said N parameters observed in the reference subjects.
[0060] The device for analyzing a biological drift is configured to implement the method for analyzing a biological drift as described above according to all its embodiments.
[0061] The invention further relates to a computer program comprising software instructions which, when executed by a computer, implement a method of analyzing a biological drift as defined above.
[0062] The invention will appear more clearly on reading the description which follows, given solely by way of non-limiting example, and made with reference to the drawings in which: Figure 1 is a block diagram of a device for generating a mapping representative of an N-dimensional data set and of a device for analyzing a biological drift implementing the mapping; Figure 2 is a flowchart of the main steps of a method for generating a mapping representative of an N-dimensional data set according to one embodiment; Figure 3 represents an example of data mapping in a two-dimensional subspace; Figure 4 is a flowchart of the main steps of a method for analyzing a biological drift of at least one of a number N greater than or equal to two of quantitative biological parameters in a subject.
[0063] The invention relates to the generation of representative mapping of an N-dimensional data set, N being an integer greater than or equal to 2, in the context of data representative of physical or biological parameters of a subject, human or animal.
[0064] The term mapping refers to a spatial modeling of the underlying two-dimensional subspace, also called a "manifold" in mathematics or a "manifold" in English, by providing an approximation of its shape and boundaries. As will be described in more detail below, the mapping is formed by envelopes associated with points in the N-dimensional subspace.
[0065] In the field of health, the invention finds applications in the in vitro characterization of the biological drift of a subject from at least two quantitative biological parameters.
[0066] In the context of such applications, the invention comprises a phase of generating a mapping representative of N-dimensional data in an N-dimensional subspace, each dimension being associated with a physical or physiological parameter.
[0067] For example, in the health field, mapping represents a reference population.
[0068] By "reference population", also called "reference group", we mean a population of individuals (or subjects) in which the biological parameters of interest are measured.
[0069] In a particular embodiment, the reference population is for example made up of healthy individuals, i.e. individuals with a good general state of health.
[0070] In another embodiment, the reference population is for example made up of pathological individuals, having been diagnosed as suffering from a known pathology or from a plurality of known pathologies.
[0071] According to another variant, the reference population is, for example, made up of healthy individuals with exceptional abilities, for example a population of high-level athletes.
[0072] The reference population comprises a number P of subjects, preferably at least 50 subjects, preferably at least 60 subjects, more preferably at least 70 subjects, at least 80 subjects, at least 90 subjects, more preferably still at least
[0073] 100 subjects, at least 110 subjects, for example at least 120 subjects.
[0074] By "biological parameter" we mean here any parameter, quantitative or qualitative, allowing, directly or indirectly, the assessment of the state of health of a subject.
[0075] The biological parameter is preferably selected from the group consisting of a serum parameter, an infectious parameter, a genetic parameter and a non-serum parameter (such as a clinical parameter or a lifestyle parameter).
[0076] Lifestyle-related parameters are preferably qualitative.
[0077] Serum, genetic, clinical, non-serum and infectious parameters can be quantitative or qualitative parameters.
[0078] Biological parameter, especially serum, infectious, genetic and some non-serum parameters, can be measured in a biological sample of a subject.
[0079] The biological sample can be a blood sample, urine sample, cerebrospinal fluid sample, stool sample or a biopsy.
[0080] The biological sample may have undergone one or more pre-processing steps prior to analysis (such as purification, centrifugation, filtration, precipitation, PCR and / or RT-PCT).
[0081] By "serum parameter" is meant here a biological parameter measured in a blood sample or a sample obtained from a blood sample (for example, in plasma).
[0082] The serum parameter can, for example, be selected from the group consisting of:
[0083] - the number, concentration and / or proportion of given blood cells or blood cell populations: for example, the number, concentration and / or proportion of red blood cells, leukocytes, polyneutrophils, polyeosinophils, polybasophils, monocytes, lymphocytes or platelets,
[0084] - blood count (CBC),
[0085] - hemoglobin concentration,
[0086] - hematocrit, i.e. the volume occupied by red blood cells in a given volume of whole blood,
[0087] - the platelets,
[0088] - the sedimentation rate,
[0089] - VGM (Mean Cell Volume),
[0090] - CGMH (Mean Blood Hemoglobin Concentration),
[0091] - TCMH (Mean Corpuscular Hemoglobin Content), - the quantity, concentration and / or proportion of a given enzyme, for example lipase, the quantity, concentration and / or proportion of Alkaline Phosphatase, Gamma GT, transaminase TgO or transaminase TgP,
[0092] - the quantity, concentration and / or proportion of lipids, in particular total cholesterol, HDL cholesterol, LDL cholesterol or triglycerides,
[0093] - blood sugar,
[0094] - the quantity, concentration and / or proportion of chlorine, uric acid, urea, or creatinine,
[0095] - protein electrophoresis, in particular the quantity, concentration and / or proportion of total protein, albumin, alpha 1 globulin, alpha 2 globulin, beta 1 globulin, beta 2 globulin and / or gamma globulin,
[0096] - the protein profile, in particular the quantity, concentration and / or proportion of serum proteins, such as IgA, IgG, IgM and / or IgE,
[0097] - an ionogram, in particular the quantity, concentration and / or proportion of one or more ions, such as sodium, potassium, calcium, magnesium, chlorine, and / or bicarbonate, for example in blood plasma,
[0098] - the quantity, concentration and / or proportion of one or more hormones, or
[0099] - blood type (e.g. ABO blood type).
[0100] The inventors noted that when a plurality of quantitative biological parameters are taken into consideration, at least some of the parameters being correlated, the evaluation of normality by analysis of the value taken by each of the biological parameters taken independently, in relation to the associated minimum and maximum reference threshold values, is not satisfactory.
[0101] The generation of a mapping of an N-dimensional data set makes it possible to represent the values of the biological parameters of a reference population and to generalize "normality" by the membership or not of a set of values of the N biological parameters considered of a subject to the mapping represented in the form of a union of convex envelopes.
[0102] The method for generating a map is described below for generating a map for a reference population, it being understood that the method applies in a similar manner for generating a map per reference population, thus allowing several maps to be obtained for the same biological parameters or for distinct biological parameters. Figure 1 represents a device 2 for generating a map representative of an N-dimensional data set and for analyzing a biological drift of a subject.
[0103] Optionally, the device 2 comprises or is connected to an alert system, which generates diagnostic alerts based on an analysis of the biological drift of a subject.
[0104] The device 2 is a programmable electronic device, for example a computer or a plurality of computers connected to each other.
[0105] The device 2 comprises a user interface 4, comprising an input / output interface 6 and a display screen 8, making it possible to present results and receive commands and data from a user.
[0106] Alternatively, the user interface 4 is a touch screen which then combines the input / output interface and display screen functionalities.
[0107] The device 2 further comprises a calculation processor 10, an electronic memory unit 12, and a communication module 16 configured to connect to a communication network using a communication protocol.
[0108] In particular, the communication module 16 is configured to communicate data, for example biological drift values, to an alert system 18, which generates diagnostic alerts based on an analysis of a subject's biological drift.
[0109] According to one variant, the alert system 18 is integrated into the device 2, the diagnostic alerts being for example displayed on the display screen 8.
[0110] The various elements 4, 10, 12, 16 are configured to communicate with each other via a data communication bus 15.
[0111] The computing processor 10 is configured to implement:
[0112] - a module 20 for acquiring an N-dimensional data set forming coordinates of P points in an N-dimensional subspace,
[0113] - a module 22 for calculating a graph representing the P points,
[0114] - a module 24 for constructing a convex hull associated with each of the P points,
[0115] - a module 26 for forming and representing a mapping 30 of the N-dimensional data set, the mapping 30 of the N-dimensional data set being formed by the union of the convex envelopes thus constructed;
[0116] -optionally, a module 28 for using one or more maps 30 for analyzing the biological drift of a subject.
[0117] A representation of the mapping 30 is stored in the electronic memory unit 12. Optionally, the module 26 also performs a calculation and an association of at least one numerical value with each convex hull. Preferably, the numerical values associated with the convex hulls are also stored in the electronic memory unit 12.
[0118] A programmable electronic device 2 comprising a module 28 for using the mapping 30 for analyzing the biological drift of a subject is a device for analyzing the biological drift of at least one of a number N greater than 2 of quantitative biological parameters in a subject.
[0119] In one embodiment, the modules 20, 22, 24, 26 are implemented in the form of software instructions forming a computer program, which, when executed by a programmable electronic device, implements a method for generating a mapping representative of an N-dimensional data set as described.
[0120] In one embodiment, the modules 20, 22, 24, 26, 28 are produced in the form of software instructions forming a computer program, which, when executed by a programmable electronic device, implements a method for analyzing a biological drift of at least one of a number N greater than or equal to two of quantitative biological parameters in a subject as described.
[0121] In a variant not shown, the modules 20, 22, 24, 26, 28 are each produced in the form of programmable logic components, such as FPGAs (Field Programmable Gate Array), microprocessors, GPGPU components (General-purpose processing on graphics processing), or even dedicated integrated circuits, such as ASICs (Application Specific Integrated Circuit).
[0122] The computer program comprising software instructions is further capable of being recorded on a non-transitory, computer-readable information recording medium. This computer-readable medium is, for example, a medium capable of storing electronic instructions and of being coupled to a bus of a computer system. For example, this medium is an optical disk, a magneto-optical disk, a ROM memory, a RAM memory, any type of non-volatile memory (for example EPROM, EEPROM, FLASH, NVRAM), a magnetic card or an optical card.
[0123] With reference to Figure 2, the main steps for generating a representative map of an N-dimensional data set are described, N being an integer greater than or equal to 2, each dimension being associated with a physical or biological parameter of a subject (human or animal), previously defined according to the intended use. The method comprises a step 40 of acquiring P sets of N values, P being an integer greater than or equal to 2.
[0124] For example, P is greater than or equal to 50, each set of N values being representative of the values of the biological parameters considered for a subject of a reference population.
[0125] Each set of N values forms coordinates of a point of the N-dimensional subspace considered, hereinafter called the initial point.
[0126] The method further comprises a step 42 of calculating a graph representing the P initial points.
[0127] A graph is a data structure consisting of vertices and edges, each edge connecting two vertices.
[0128] Step 42 of calculating a graph includes several sub-steps described below.
[0129] Step 42 includes a sub-step 44 of initializing the graph, each initial point being associated with a vertex of the graph.
[0130] It also comprises a sub-step 46 of determining, for each initial point, called first point, a neighborhood formed of the k second points, called neighbors or neighboring points, closest to the first point in the sense of a predetermined metric, also called first metric, for example the Euclidean metric in the IR' space. V, and connecting by an edge between the vertex associated with the first point and each vertex associated with a second point in the neighborhood. The number k is a predetermined integer, for example greater than or equal to 1, preferably greater than or equal to 10, k being less than or equal to P. In an example implementation, the number k is between 5 and 20, while the number P is greater than 100, for example equal to 120.
[0131] Additionally, a weight is associated with each edge (step 48), the weight being equal to the distance according to the first metric (eg Euclidean distance) between the two points associated with the vertices connected by the edge.
[0132] A second metric is defined to estimate the distance between points of the N-dimensional subspace, which is a metric associated with the graph: the distance between two points according to this second metric is equal to the smallest sum of the weights associated with the edges connecting the vertices corresponding to the two points in the graph.
[0133] Step 42 then comprises a symmetrization sub-step 50, consisting of adding connecting edges, if necessary, so as to ensure that if a point A is part of the neighborhood of a point B, point B is also part of the neighborhood of point A.
[0134] Thus, where appropriate, one or more additional points are added to the neighborhood of a first point given in the symmetrization step. For each point, at the end of the symmetrization sub-step 50, the associated neighborhood is stored, this neighborhood comprising the “natural” neighbors (the k nearest neighbors), as well as, where appropriate, the neighbors, called “forced” neighbors, obtained by symmetry.
[0135] Finally, step 42 includes a step 52 of connecting, if applicable, subgraphs which are not connected by edges.
[0136] If the resulting graph is composed of C components, each component being a graph of points connected by edges, identifiable by a depth-first traversal algorithm.
[0137] In other words, a component C is a subset of points of the graph such that it is possible to connect any two points of C by a path connecting one or more edges.
[0138] If C is strictly greater than 1, then the formed graph is an unconnected graph, and the following process implemented at the connection step 52 makes it possible to connect the C components by minimizing the sum of the weights of the added edges:
[0139] • S1- association with each component of an identifier being an integer;
[0140] • S2-calculation and storage of the distance, according to the first predefined metric (eg the Euclidean distance), of each pair of points located in different components and for each pair of different components, storage of the pair of points with the smallest distance;
[0141] • S3 - connection of the two components having the minimum distance, via the edge linking the points of the pair of points having the smallest distance;
[0142] • S4 - update of the component identifiers so that the two components that have just been connected have the same identifier;
[0143] • S5 - repeat steps 3 and 4 until all components have the same identifier. A total of C-1 iterations are required to obtain a single remaining connected component (i.e. connected graph).
[0144] Thus, at the end of connection step 52, a fully connected graph is obtained, and some of the neighborhoods are enriched with an additional neighbor called a connecting neighbor.
[0145] In the end, the constructed graph represents the set of P initial points, and each vertex of the graph is connected to all the other vertices of the graph by traversing the edges of the graph. In other words, all the vertices are connected via one or more edges of the graph. In other words, at the end of step 42, each initial point called first point is associated with a first vertex of the graph, and with second neighboring points, the second neighboring points being the initial points associated with second vertices, each second vertex being connected to the first vertex by an edge.
[0146] The neighborhood of a first point, associated with a vertex of the graph, is formed by the second points, whose associated vertices are connected directly (by a single edge) to the first vertex. This neighborhood is memorized for each first point.
[0147] The neighborhood of a given first point includes at least the points which are its k nearest neighbors (called natural neighbors), but can also include one or more points resulting from the symmetrization (called forced neighbors), and one or more points resulting from the connection step (called connecting neighbors).
[0148] The method then comprises a step 54 of calculating (or constructing), for each initial point having an associated vertex in the graph, an associated convex hull, constructed from the points corresponding to the vertices of the natural neighbors of said associated vertex. The representative mapping of the initial data set is then formed by the union of the calculated convex hulls.
[0149] According to a first variant, the convex hull is constructed on the natural neighbors.
[0150] According to a second variant, the convex hull is constructed on the natural neighbors and the forced neighbors.
[0151] According to a third variant, the convex hull is built on natural neighbors, forced neighbors and connecting neighbors.
[0152] In one embodiment, the computation of the convex hulls of a set of points is performed via an implementation of the algorithm known as the “QUICKHULL ALGORITHM,” described in particular in the article “The Quickhull Algorithm for Convex Hulls” by C. Bradford Barber et al., published in ACM Transactions on Mathematical Software, vol. 22, pp. 469-483 and available at the following address: https: / / dl.acm.org / doi / 10.1145 / 235815.235821. For example, programmed functions for processing convex hulls are used.
[0153] The method involves, for each convex hull, the memorization of the different faces of the convex hull (which forms a polytope) and the R points, among the points of the neighborhood, located on the surface of the convex hull.
[0154] As is known, to check whether a point lies inside a convex hull, it is necessary to check whether the given point can be expressed as a linear combination of the S vertices {pi,...,p s} of the convex hull, such that the coefficients of the linear combination are all positive and their sum is equal to one.
[0155] In other words: the elements belonging to a convex hull whose vertices are a set of S points {pi,...,ps} are the points x which can be written in the form: s
[0156] X = ^ / iPi i = 1
[0157] Where the coefficients Àj are positive real numbers such that:
[0158] ± , = 1 i = 1
[0159] The union of the computed convex hulls forms a mapping of the N-dimensional data space.
[0160] The method further comprises an identification 56 of the extreme points, that is to say on the surface of the union of the convex envelopes.
[0161] A point belonging to a convex hull can be located inside it, or on the surface of this convex hull. A point is on the surface of a convex hull if it belongs to one of the faces of this convex hull.
[0162] A point is extreme, and therefore considered on the surface of the union of convex hulls, if it is on the surface of all the convex hulls to which it belongs.
[0163] The extreme points are stored, which allows the contour of the calculated map to be characterized.
[0164] In addition, the method includes a calculation 58 and a storage of at least one numerical value associated with each convex hull.
[0165] For example, the numeric value is a statistical quantile of an attribute associated with the convex hull in the distribution of that attribute.
[0166] The attribute is for example the volume of the convex hull, calculated by an algorithm such as "QUICKHULL ALGORITHM" mentioned above.
[0167] Alternatively, the attribute is the average distance according to the first metric (e.g. Euclidean distance) between the initial point for which the convex hull was calculated and its k nearest neighbors.
[0168] Advantageously, these attributes reflect a local density of points.
[0169] In one embodiment, a value of the selected attribute is calculated for each convex hull.
[0170] The distribution of the values obtained makes it possible to associate a statistical rank with each of them and then a quantile by dividing this rank by the number of values in the distribution. In statistics, the rank q of a statistical sample is equal to the q-th smallest value. The rank of a value relative to a set of values corresponds to its index in an increasingly ordered sequence of these values. Low quantiles of convex hull volume, or average distances to the k-nearest neighbors will then be associated with a high density and high quantiles with low densities.
[0171] Preferably, following the construction of the map, the following are stored:
[0172] • the set of convex hulls, and for each convex hull: o the faces constituting each convex hull; o the vertices of each convex hull; o one or more associated numerical values, ie the values of different attributes (volume, distance to the nearest neighbor, local density of points, surface area of the hull or any attribute relating to the hull),
[0173] • the neighbors of each point: o natural neighbors (k nearest neighbors) o neighbors forced by symmetrization, o connecting neighbors
[0174] • the nearest neighbor graph as described above: o its vertices o its edges and their associated weights
[0175] • the set of extreme points, following the calculation of the extreme points.
[0176] • the distribution and quantiles of the attribute associated with each convex hull.
[0177] Optionally, the method for generating a mapping comprises an update of the mapping formed from the first N-dimensional data of a first set of P subjects, forming P first points in the two-dimensional subspace, using second additional N-dimensional data relating to a second set of Q subjects, forming Q second points in the N-dimensional subspace, the update comprising a modification of the edges of the graph according to the second additional N-dimensional data making it possible to obtain a modified graph, a symmetrization of the edges of the modified graph and a reconstruction of convex hulls of the first points and of the second points whose neighborhood is modified by the modification and symmetrization steps.
[0178] Optionally, the method for generating a map further comprises a step of enriching the map by adding at least one set of synthetic data, by random sampling in at least one of the convex hulls of said map. In other words, this makes it possible to obtain one or more synthetic points, and therefore to increase the number of points of the map to obtain a more complete map.
[0179] With reference to Figure 3, as a simplified, non-limiting example, a 2-dimensional data map (N=2) is illustrated.
[0180] The points indicated by circles correspond to the initial points of the reference population and form a cloud of points, represented in the two-dimensional space of the parameters "hemoglobin" (in g / L) and "albumin" (in mmol / L). Each of the parameters has associated minimum and maximum reference threshold values, which are illustrated in the figure, namely V1_m and V1_M for albumin, and V2_m and V2_M for hemoglobin.
[0181] The obtained Carto mapping is illustrated schematically, delimited by a dotted line outline in Figure 3. In addition, three convex hulls EC_1, EC_2, EC_3, associated with respective initial points A, B, C are represented.
[0182] Points S1, S2, S3 are points corresponding to measurements taken for separate subjects for which the possible biological drift of the “albumin” and “hemoglobin” parameters is to be estimated.
[0183] Considering biological parameters independently amounts to deciding that from the moment the values of a subject's biological parameters are within the rectangle corresponding to the respective threshold values associated with these two parameters (rectangle shown in Figure 3), the subject is "normal" in relation to the reference population.
[0184] Point S1 corresponds to an albumin value between V1_m and V1_M, but a hemoglobin value greater than V2_M.
[0185] Points S2 and S3 both correspond to albumin values between V1_m and V1_M and hemoglobin values between V2_m and V2_M.
[0186] According to the independent interpretation of the parameters (i.e. hemoglobin and albumin parameters taken separately), point S1 is not considered "normal", but points S2 and S3 are normal.
[0187] However, given the distribution of the initial points, using the Carto mapping generated by the described method, we will consider that point S2 is not "normal", or in other words is outside the reference threshold values, because the values of the parameters for this subject S2 are outside the union of convex hulls calculated from the initial points, even if it is located in the admissible range of each of the reference threshold values. Point S3 is inside one of the calculated convex hulls, so it is considered "normal" in view of the union of the convex hulls, therefore of the global mapping representative of the reference population. A drift may nevertheless be observed for point S3, and a level of drift calculated according to the numerical value associated with the convex hull to which point S3 belongs, as explained in more detail below.
[0188] It is worth noting that thanks to the use of mapping, the representation of the parameters of the reference population is much more accurate than when considering the biological parameters independently.
[0189] A method for analyzing a biological drift of at least one of N quantitative biological parameters in a subject, the number N being greater than or equal to two, is described with reference to Figure 4.
[0190] The method comprises a step 60 of generating the mapping representation of a reference population for the N quantitative biological parameters considered, in the form of a union of convex envelopes, by implementing the method described above.
[0191] The number N varies depending on the applications, preferably between 2 and 10,000.
[0192] For example, such a mapping is generated during a prior step and stored in a memory of the programmable electronic device implementing the method or in a database stored on a remote server.
[0193] In the case where the number of subjects P in the reference population is insufficient, i.e. less than a threshold number P o predetermined, for example less than P o =12O, then a synthetic data enrichment process is implemented.
[0194] This process consists of adding P' points, with P' > (P o - P), sample P' convex hulls of the mapping with possible repetitions, randomly. For each of these P' hulls, choose the Àj with i = {1 , ... ,S}, S being the number of vertices in the convex hull, and 0 <Àj < 1 such that:
[0195] Then the associated convex synthetic point x is then: The method also includes obtaining 62 of the values of the N biological parameters for a subject considered at a given time instant Ti: v, = These values are provided to the programmable electronic device.
[0196] In the case where m of the N biological parameters for a subject under consideration are not available, m being an integer greater than or equal to 1, a preliminary step of imputation of missing data by estimation of the missing biological parameters is implemented. To do this, we seek the intersection between the hyperplane drawn by the (Nm) biological parameters in the N-dimensional space, with the generated mapping. We seek a convex hull of the mapping that admits at least one solution for the system of equations {equation of the hyperplane; equation of the convex hull}. If there is no solution, the following convex hull is considered. If no convex hull of the mapping is intersected by the hyperplane, then we can already deduce that the subject under consideration is outside the mapping.Otherwise, for the first convex hull admitting one or more solutions, random sampling among the solutions makes it possible to complete the m missing dimensions (corresponding to the m biological parameters missing for the subject considered), in a manner compatible with the biological information captured by the mapping.
[0197] It should also be noted that the mapping can be updated once constructed if new data becomes available, this update being useful in order to aggregate different data sources iteratively, in the context of federated learning for example. Consider a cartograph Cp constructed from P reference subjects forming a first set SP.
[0198] The update includes a modification of the edges of the graph based on the second additional N-dimensional data to obtain a modified graph, a symmetrization of the edges of the modified graph and a reconstruction of convex hulls of the first points and the second points whose neighborhood is modified by the modification and symmetrization steps.
[0199] In other words, the update includes a modification of the natural and forced neighborhoods of the points corresponding to the subjects of the first set SP, based on a second set of points SQ' associated with Q' subjects.
[0200] To perform the update, we iterate the following steps for each subject in the second set SQ', described by N biological parameters. Subsequently, we process points in the N-dimensional space, each point being associated with a subject.
[0201] Consider a first point p of the second set SQ'.
[0202] • for any point s of SP that is one of the k nearest neighbors of p among the set {SP ; p} o if p is also one of the k nearest neighbors of s among {SP ; p}, then the subject associated with point p is one of the k nearest neighbors of the subject associated with point s and the former k-th neighbor of s noted n k (s) is removed from the vicinity of s;
[0203] ■ if s is one of the k nearest neighbors of n k (s) among {SP ; p}, then n k (s) becomes a forced neighbor of point s;
[0204] ■ if s was a forced neighbor of n k (s) among SP, then s is no longer a forced neighbor of n k (s) among {SP ; p} o if p is not one of the nearest neighbors of s among {SP ; p}, then the point p becomes a forced neighbor of the point s
[0205] • for any point s of SP not being part of the k nearest neighbors of p among the set {SP;p}, o if p is part of the nearest neighbors of s, then p is added to the forced neighbors of s and the former k-th nearest neighbor of s among SP denoted n k (s) is no longer one of the k nearest neighbors of s among {SP;p}
[0206] ■ if s is one of the k nearest neighbors of n k (s) among {SP;p}, then n k (s) becomes a forced neighbor of s
[0207] ■ otherwise
[0208] • if s was a forced neighbor of n k (s) among SP, then s withdrew from the forced neighbors of n k (s) o if p is not one of s's nearest neighbors, then do not change anything.
[0209] Each point whose natural or forced neighborhood is modified during the update is called a point impacted by the addition of point p. Then the connections to be added are identified if the graph resulting from the neighborhood relations is not a connected graph, then the convex hulls of the points impacted by the addition of point p are recalculated. Advantageously, this update process is more parsimonious than a complete recalculation of the mapping with the set of points {SP ; SQ'} whose algorithmic complexity is quadratic in total number of subjects.
[0210] Then, it is determined (step 64) whether a point p(Vj) of the N-dimensional space with coordinates equal to the values of the N quantitative biological parameters of the subject {v n ,...,v i;v} lies inside the union of convex hulls.
[0211] In one embodiment, in step 64 it is checked whether there is a convex hull that contains the point p(Vi). If one or more convex hulls contain the point p(Vi), the point is considered to have normal values relative to the reference population.
[0212] Step 64 is then followed by a step 70 during which the convex hull C m to which the point p(Vi) belongs is identified and the numerical value Val(C m ) associated with the convex hull C m is obtained. If p(Vi) is contained by several convex hulls, p(Vi) is identified with the smallest value associated with the convex hulls obtained.
[0213] The numerical value Val(C m ) associated with point p(Vi) allows it to be assigned a drift level (step 72).
[0214] The more the point is located in a low density area, the more it is considered to be drifting.
[0215] In one embodiment, a slicing of drift levels is performed according to quantile intervals of convex hull volumes or average distance to k nearest neighbors, two attributes inversely correlated with density.
[0216] The optimum (or absence of drift) is defined for attribute values associated with a quantile located in a first interval [0; X1], with for example X1 = 68.26. A drift (or first level of drift) is defined for attribute values associated with a quantile located in a second interval [X1; X2], with for example X2 = 95.44. A significant drift (or second level of drift) is defined for attribute values associated with a quantile located in the interval [X2; X3] with X3 = 99.7. A very significant drift (or third level of drift) is defined for attribute values associated with a quantile located in the interval [X3; 100].
[0217] For example, referring to Figure 3, assuming that the numerical value of convex hulls is associated with their volume, if p(Vi) is located in the convex hull of point A, whose convex hull associated with a low volume, and therefore an associated quantile located in the interval [0; 68.26], p(Vi) is considered to be at the optimum (absence of drift).
[0218] If p(Vi) is located in the convex hull of point B whose convex hull is associated with a volume whose quantile is in the interval [68.26; 95.44], it is considered to have a first level drift.
[0219] If p(Vi) is located in the convex hull of point C whose convex hull associated with a larger volume (greater than the respective volumes of the convex hulls associated with points A and B) and therefore an associated quantile in the interval [95.44; 99.7]; the point p(Vi) is considered to have a significant drift (second level of drift), because the quantile of the volume is very high. Thus, when the numerical value associated with each convex hull is representative of the statistical distribution of the points of the reference population, the more the point p(Vi) is located in a dense area, the more it is considered "normal" (in other words the level of drift is low, or even zero and therefore the point is at the optimum), and the more it is located in a low density area, the higher the level of drift.
[0220] If none of the convex hulls contains the point p(Vi), then the point is considered potentially abnormal given the reference population approximated by the union of the convex hulls, and therefore probably pathological.
[0221] The method comprises a step of calculating a drift value associated with the subject at k nearest neighbors of the point associated with said subject among the points of the mapping.
[0222] In one embodiment, a convex hull is associated with the point p(Vi), this hull is constructed from its k nearest neighbors among the P points of the reference population. The numerical value calculated from an attribute of the convex hull is then characterized and the rank of this value is calculated among the P numerical values of the convex hulls of the mapping of the reference population in order to determine its level of drift.
[0223] If none of the convex hulls contains the point p(Vi) and the drift level associated with the subject is a significant drift or more, typically above a predetermined drift threshold, then the point is considered abnormal given the reference population approximated by the union of the convex hulls, and therefore probably pathological.
[0224] It is further provided to calculate a distance between the point p(Vi) and the union of the convex hulls and to memorize this distance (step 66), whether the point p(Vi) is outside or inside the mapping.
[0225] This distance is based on the weights of the edges of the mapping graph. The distance between point p(Vi) and the nearest point in the mapping is calculated by the first distance (e.g., Euclidean distance). Then, a distance relative to the mapping between this nearest point and another point in the mapping is calculated using the shortest path (or geodesic) connecting these two points in the mapping. It is then possible to measure a distance between p(Vi) and any point in the mapping.
[0226] According to one embodiment, the calculated distance is a drift value of the subject.
[0227] According to one embodiment, the drift of a subject associated with a point p(Vi) may be characterized by any combination of calculations characterizing p(Vi) as a function of the mapping. For example, the drift is characterized by the rank of the density as described previously, or the average distance between p(Vi) and all or a subset of the points of the mapping.
[0228] Optionally, an alert is raised (step 68) to indicate the presence of parameter values outside the reference population mapping and / or the calculated drift level. The distance value calculated in step 66 is optionally transmitted and / or displayed when the alert is raised.
[0229] Optionally, a diagnosis is also provided in connection with the alert.
[0230] According to one embodiment, it is envisaged to repeat the method for the same subject in order to follow the evolution over time of the values of the biological parameters considered, or in other words, the trajectory of the successive points whose coordinates are the values of the biological parameters considered for the subject at time instants Ti which are shifted, for example shifted by a predetermined time step.
[0231] This then makes it possible to add a time dimension, and to calculate a trend over time.
[0232] The use of mapping the N biological parameters of a reference population, healthy or pathological, makes it possible to carry out an in vitro diagnosis of a disease based on the membership of the values of the biological parameters considered for a subject, at a time instant or a succession of shifted time instants, in an N-dimensional subspace representative of the reference population. When a potential drift in the values of the biological parameters considered for the subject is observed, the diagnosis of a disease or a risk of suffering from a disease associated with the N biological parameters is made.
[0233] According to another variant, it is envisaged to construct several maps among multiple reference populations of patients characterized by a label and then to evaluate the biological drift of a subject for each of the maps of the set of maps, by calculating a drift value, according to one of the methods described above. It is then possible to combine these multiple drift values into a score (or global drift value) making it possible to quantify a meta-drift, or in other words a drift with respect to the set of maps. This combination can be a ratio between a drift with respect to a chosen map and the sum of the drifts or an average of all the drifts. Other mathematical operations on the calculated drift values are conceivable.Of course, it is also possible to calculate drift values for a plurality of maps, and their combination to calculate an overall drift value repeatedly over time, so as to obtain a temporal evolution of the overall drift value.
[0234] For example, it is envisaged to consider L reference populations, L being an integer greater than or equal to 1, associated with groups of subjects or patients characterized by a label composed of P1,..,PL subjects, each label being associated with a pathology or a specific characteristic for example, and for each of the L reference populations, obtain said N quantitative biological parameters forming an N-dimensional data set. The numbers P1, P2,...,PL can be equal or distinct.
[0235] For example, the reference population L1 can be associated with P1 patients diagnosed with diabetes, the reference population L2 can be associated with P2 patients diagnosed with rheumatoid arthritis, the reference population L3 can be associated with P3 high-level sports patients, etc.
[0236] For each of the L reference populations, the generation of a representative mapping of the N-dimensional data set according to the method described above is implemented. Each mapping Ci obtained is a reference mapping associated with the reference population Li, and is represented in the form of the union of convex hulls of the N-dimensional space.
[0237] For a given subject to be tested, the point s in the corresponding N-dimensional space is obtained as a function of the values of the subject's N quantitative biological parameters.
[0238] For each mapping of the L cartographies, the method comprises: determining whether the point s relating to said subject of the N-dimensional space of coordinates equal to the values of the N quantitative biological parameters is located inside the union of convex envelopes of the mapping and the quantitative level of drift associated with the numerical value of the envelope containing it, or formed from said subject and its k nearest neighbors, and calculating a score (or overall drift value) from the different drifts reflecting the proximity of a point with the L cartographies.
[0239] This score can be a ratio between the drift with respect to a mapping and the sum of the drifts of all the mappings. Then, in an application case, a temporal monitoring of the evolution of the biological drift of a subject among a set of L mappings of labeled patients is carried out:
[0240] - for T time steps, obtain values of the N quantitative biological parameters of said subject,
[0241] - for the T combinations of N biological parameters, calculate, for each of the L maps, calculate the associated biological drift.
[0242] The analysis of the evolution of biological drifts over time makes it possible to deduce the evolution of the subject according to characteristics or pathologies represented by the L maps.
[0243] In one embodiment, the temporal evolutions of biological drifts Der_i,t, i=1 to L, t being the time step index, calculated for the L maps for a subject are transmitted to an alert system, which generates diagnostic alerts associated with all the drifts among the L populations.
[0244] When the combination of the N biological parameters of the subject is considered abnormal compared to the mapping or mappings, the method described above may comprise a subsequent step of implementing a method, preferably in vitro and / or ex vivo, for diagnosing a disease or the risk of suffering from a disease, said disease being associated with the biological drift values calculated compared to the reference population.
[0245] The method according to the invention has the advantage of allowing earlier diagnosis of a disease.
[0246] The invention also relates to a method for preventing and / or treating a disease in a subject, linked to the belonging of the values of the biological parameters considered for a subject, at a time instant or at a succession of shifted time instants, to an N-dimensional subspace representative of the reference population comprising:
[0247] - an implementation of the method for analyzing a biological drift of at least one of a number N greater than or equal to two of quantitative biological parameters in a subject as described above,
[0248] - if the combination of said N biological parameters of the subject is abnormal in view of the combinations of said N parameters observed in the reference subjects, administration of a treatment appropriate to said disease. The treatment is for example administration of a drug chosen from the group comprising: an analgesic, an immunosuppressant, a corticosteroid, an anti-inflammatory, a vitamin, a blood treatment, an antibiotic.
Claims
CLAIMS 1. Method for generating a mapping representative of an N-dimensional data set, N being an integer greater than or equal to 2, each dimension being associated with a given physical or biological parameter of a subject, the method being implemented by a processor of an electronic computing device, the method comprising an acquisition of said N-dimensional data set of a set of P subjects, in the form of P sets of N values, forming coordinates of P points in an N-dimensional subspace, P being an integer greater than or equal to two, the method being characterized in that it further comprises steps, implemented by the computing processor, of: calculating (42) a graph representing said P points, the graph comprising vertices and edges, each edge connecting two vertices together, in which for a given vertex, the set of vertices connected to said vertex by an edge forms a neighborhood of said given vertex in the graph,in which each vertex of the graph corresponds to a first point of said subspace, and its neighborhood comprises vertices corresponding to second points of said subspace, the second points being the k nearest neighbors of the first point in the sense of a predetermined metric, k being a chosen positive integer; for each point of said subspace having an associated vertex in said graph, construction (54) of an associated convex hull, encompassing the points corresponding to the vertices of the neighborhood of said associated vertex, said mapping (30) being formed by the union of the convex hulls associated with the points of the subspace., 2. Method according to claim 1, further comprising a calculation (58) of at least one numerical value associated with each convex hull, the or each numerical value being calculated from an attribute of the convex hull.
3. Method according to claim 2, wherein said attribute of the convex hull is representative of a local density of points belonging to said convex hull.
4. Method according to one of claims 2 or 3, in which, for a given attribute, the numerical value is a quantile of a value of the attribute in a statistical distribution of the values of said attribute.
5. Method according to any one of claims 1 to 4, comprising an association (48) of a weight to each edge of the graph, as a function of the distance, according to said metric, between the points corresponding to the vertices of the graph connected by said edge.
6. Method according to any one of claims 1 to 5, wherein said metric is the Euclidean distance in N dimensions.
7. Method according to any one of claims 1 to 6, in which the step of calculating (42) the graph comprises a sub-step of connecting (52) non-connected components, each component being a sub-graph, the connecting step (52) comprising an addition of an edge between vertices each belonging to a distinct non-connected component, and an addition of a connecting neighbor in the neighborhood of the points corresponding to said vertices.
8. Method according to one of claims 1 to 7, in which the step of calculating the graph comprises a sub-step of symmetrization (50) of the neighborhoods associated with the P points.
9. Method according to any one of claims 1 to 8, further comprising an identification (56) and storage of extreme points of the union of the convex envelopes, each extreme point being an external point of one of said convex envelopes and which does not belong to another of the convex envelopes.
10. Method according to any one of claims 1 to 9, further comprising an update of the mapping formed from the first N-dimensional data of a first set of P subjects, forming P first points in the N-dimensional subspace, using second additional N-dimensional data relating to a second set of Q subjects, forming Q second points in the N-dimensional subspace, the update comprising a modification of the edges of the graph according to the second additional N-dimensional data making it possible to obtain a modified graph, a symmetrization of the edges of the modified graph and a reconstruction of convex hulls of the first points and of the second points whose neighborhood is modified by the modification and symmetrization steps.
11. Method according to any one of claims 1 to 10, further comprising a step of enriching the mapping by adding at least one set of synthetic data, by random sampling in at least one of the convex hulls of said mapping.
12. Device for generating a representative map of an N-dimensional data set, N being an integer greater than or equal to 2, each dimension being associated with a given physical or biological parameter of a subject, the device being configured to acquire the N-dimensional data set in the form of P sets of N values, forming coordinates of P points in an N-dimensional subspace, P being an integer greater than or equal to two, the device comprising a calculation processor configured to implement: a module (22) for calculating a graph representing said P points, the graph comprising vertices and edges, each edge connecting two vertices together, in which for a given vertex, the set of vertices connected to said vertex by an edge forms a neighborhood of said given vertex in the graph, in which each vertex of the graph corresponds to a first point of said subspace, and its neighborhood comprises vertices corresponding to second points of said subspace, the second points being the k closest neighbors of the first point in the sense of a predetermined metric, k being a chosen positive integer;a module (26) for constructing convex hulls configured, for each point of said subspace having an associated vertex in said graph, to construct an associated convex hull, encompassing the points corresponding to the vertices of the neighborhood of said associated vertex, a module (26) for forming said mapping (30), said mapping (30) being formed by the union of the convex hulls associated with the points of the subspace.; 13. Computer program comprising software instructions which, when executed by a programmable electronic device, implement a method for generating a mapping representative of an N-dimensional data set according to claims 1 to 11.
14. Method for analyzing a biological drift of at least one of a number N greater than or equal to two of quantitative biological parameters in a subject, the method being characterized in that it comprises steps, implemented by a calculation processor, of: -for a reference population formed of a group of P subjects, obtain values of said N quantitative biological parameters forming an N-dimensional data set, -generate (60) a mapping representative of the N-dimensional data set by a method according to claims 1 to 11, said mapping being a reference mapping associated with the reference population, and being represented in the form of the union of convex envelopes of the N-dimensional space, - obtain (62) values of the N quantitative biological parameters of said subject, -determine (64) whether a point relative to said subject in the N-dimensional space of coordinates equal to the values of the N quantitative biological parameters of said subject is located inside the union of convex envelopes, - if the said point relating to the said subject is not located within the union of convex envelopes, conclude that the combination of the said N biological parameters of the subject is potentially abnormal in view of the combinations of the said N parameters observed in the reference subjects.
15. Method for analyzing a biological drift according to the preceding claim, wherein obtaining the values of the N quantitative biological parameters comprises obtaining measured values for Nm quantitative biological parameters among the N quantitative biological parameters, m being an integer, and an estimation, for the m missing quantitative biological parameters, according to an intersection between a geometric hyperplane associated with the subject, said geometric hyperplane associated with the subject being by said measured values of the Nm quantitative biological parameters in the N-dimensional space, and the calculated mapping.
16. Method for analyzing a biological drift according to claim 14 or 15, comprising a calculation of a drift value associated with the subject from the k closest neighbors of the point associated with said subject among the points of the mapping.
17. Method for analyzing a biological drift according to any one of claims 14 to 16, comprising a generation of a plurality of maps, each map being representative of a distinct reference population, associated with a pathology or a specific characteristic, and a calculation, for said subject, of a drift value in relation to each map, and a calculation of an overall drift value for said subject by combining the calculated drift values.
18. Method for analyzing a biological drift according to any one of claims 16 to 18, wherein each convex hull of said mapping has an associated numerical value, the numerical value being calculated from an attribute of the convex hull, the method further comprising, if said point relating to said subject is located inside the union of convex hulls, identifying (70) the convex hull(s) containing the point relating to said subject, and deducing therefrom (72) a level of drift as a function of the numerical value(s) associated with the hull(s) containing the point relating to said subject.
19. Method for analyzing a biological drift according to the preceding claim, in which said attribute is representative of a density of points of the convex hull, and in which a division of the drift levels is carried out according to quantile intervals of said attribute.
20. Method for analyzing a biological drift according to any one of claims 16 to 19, further comprising a repetition of the calculation of a biological drift value for successive time instants, and a calculation of a temporal evolution of the biological drift of the subject.
21. Method for analyzing a biological drift according to any one of claims 16 to 20, further comprising a transmission of the calculated biological drift values to an alert system.
22. Device for analyzing a biological drift of at least one of a number N greater than or equal to two of quantitative biological parameters in a subject, the analysis device comprising a calculation processor, configured to implement modules configured to: -for a reference population formed of a group of P subjects, obtain values of said N quantitative biological parameters forming an N-dimensional data set, -generate a mapping representative of the N-dimensional data set by a method according to claims 1 to 11, said mapping being a reference mapping associated with the reference population, and being represented in the form of the union of convex envelopes of the N-dimensional space, - obtain values of the N quantitative biological parameters of said subject, -determine whether a point relative to said subject in the N-dimensional space of coordinates equal to the values of the N quantitative biological parameters of said subject is located inside the union of convex envelopes, - if the said point relating to the said subject is not located within the union of convex envelopes, conclude that the combination of the said N biological parameters of the subject is potentially abnormal in view of the combinations of the said N parameters observed in the reference subjects.