Gestational diabetes data analysis method and system based on serum metabolism markers

By constructing and standardizing serum metabolic marker data during pregnancy and applying machine learning, the problem of insufficient accuracy and generalization in gestational diabetes data analysis is solved, and more accurate and universal analysis results are achieved.

CN120221108AInactive Publication Date: 2025-06-27PEOPLES HOSPITAL OF INNER MONGOLIA AUTONOMOUS REGION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510242800.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems of insufficient accuracy and general use in the data analysis of gestational diabetes. Insufficient sensitivity and specificity of blood sugar tests lead to misdiagnosis or misdiagnosis, and the data is too single, so the general use of the analysis results is not high.

Method used

Using a data analysis method based on serum metabolic markers, the original pregnancy data table set was constructed, and the pregnancy duration group was extracted and balanced, metabolic marker analysis was performed, the data was standardized, and the predicted pregnancy data table was generated through machine learning.

Benefits of technology

The accuracy and generalization of gestational diabetes data analysis are improved, and the shortcomings of single data analysis are made up for the analysis of single data through multi-dimensional serum metabolic markers, and more accurate and universal diagnosis and prediction results are provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120221108A_ABST
    Figure CN120221108A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis, in particular to a gestational diabetes data analysis method and system based on a serum metabolism marker, and the method comprises the steps: constructing an original pregnancy data table set, extracting an original pregnancy duration group in each original pregnancy data table from the original pregnancy data table set, and obtaining an original pregnancy duration group set; performing pregnancy duration balance on the original pregnancy duration set to obtain an effective duration set, extracting a plurality of serum metabolism sets from the original pregnancy data table set, performing metabolism marker analysis according to the plurality of serum metabolism sets to obtain an effective metabolism marker set, and performing pregnancy duration balance on the basis of the effective duration set and the effective metabolism marker set according to the original pregnancy data table set. And performing machine learning based on the standard pregnancy data table set to obtain a predicted pregnancy data table. The gestational diabetes mellitus data analysis method can improve the accuracy and universality of gestational diabetes mellitus data analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and in particular, to a method and system for analyzing data of gestational diabetes mellitus based on serum metabolic markers. Background Art

[0002] Gestational diabetes mellitus is a disorder of glucose metabolism that occurs during pregnancy and poses a certain threat to the health of the mother and fetus. Therefore, it is crucial to analyze the clinical data of pregnant women with gestational diabetes mellitus to find timely and accurate methods for preventing and diagnosing gestational diabetes mellitus.

[0003] Currently, the data analysis of gestational diabetes mellitus mainly relies on frequent blood glucose tests on pregnant women. Blood glucose tests may be less sensitive or specific, resulting in missed or misdiagnosed cases, thereby reducing the accuracy of data analysis of gestational diabetes mellitus. At the same time, frequent tests on individual pregnant women also make the obtained data too single, resulting in low generality of the data analysis results. Summary of the Invention

[0004] The present invention provides a method and system for analyzing data of gestational diabetes mellitus based on serum metabolic markers, and its main purpose is to improve the accuracy and generality of data analysis of gestational diabetes mellitus.

[0005] To achieve the above object, a method for analyzing data of gestational diabetes mellitus based on serum metabolic markers provided by the present invention includes:

[0006] Construct an original pregnancy data table set, wherein the original pregnancy data table set contains the serum metabolomes of different pregnant women at different pregnancy durations, and the original pregnancy data table set contains multiple pregnancy duration groups and multiple serum metabolome sets, and one pregnancy duration group corresponds to one serum metabolome set;

[0007] Extract the original pregnancy duration groups in each original pregnancy data table in the original pregnancy data table set to obtain an original pregnancy duration group set, and balance the pregnancy durations of the original pregnancy duration group set to obtain an effective duration group;

[0008] Extract multiple serum metabolome sets in the original pregnancy data table set, wherein one original pregnancy data table corresponds to one serum metabolome set;

[0009] According to the multiple serum metabolome sets, perform metabolic marker analysis to obtain an effective metabolic marker group, wherein the effective metabolic marker group contains the names of serum metabolic markers and there are no duplicate names of serum metabolic markers;

[0010] Based on the effective duration group and the effective metabolic marker group, perform data standardization on the original pregnancy data table set to obtain a standard pregnancy data table set;

[0011] Based on the standard pregnancy data set, machine learning is performed to obtain a predicted pregnancy data table, and based on the predicted pregnancy data table, data analysis of gestational diabetes based on serum metabolic markers is completed.

[0012] Optionally, the construction of the original pregnancy data set includes:

[0013] Obtain a pregnancy database, where the pregnancy database contains serum metabolic information and pregnancy duration information of pregnant women during pregnancy;

[0014] Successively retrieve the pregnant woman identification in the pregnancy database, and obtain the original pregnancy data set corresponding to the pregnant woman identification in the pregnancy database, where the original pregnancy data set includes: a pregnancy duration set and a corresponding serum metabolome set, and one pregnancy duration in the pregnancy duration set corresponds to one serum metabolome in the serum metabolome set;

[0015] Based on the original pregnancy data set, create an original pregnancy data table for the pregnant woman identification;

[0016] When the original pregnancy data table is equal to the preset standard number of tables, stop the step of retrieving the pregnant woman identification;

[0017] Summarize the original pregnancy data tables to obtain an original pregnancy data set.

[0018] Optionally, the pregnancy duration balance of the original pregnancy duration set to obtain an effective duration set includes:

[0019] Obtain a reference pregnancy duration;

[0020] According to the reference pregnancy duration and the preset number of reference durations, construct multiple original duration distribution intervals, where the interval length of each original duration distribution interval is the same, and the number of original duration distribution intervals is equal to the number of reference durations plus 1;

[0021] Use the original pregnancy duration set to supplement the multiple original duration distribution intervals to obtain multiple visible duration distribution intervals;

[0022] Perform interval adjustment on the multiple visible duration distribution intervals to obtain multiple target duration distribution intervals;

[0023] Calculate the median duration of each target duration distribution interval in the multiple target duration distribution intervals to obtain a median duration set, and record the median duration as the effective duration set, where the number of median durations in the median duration set is equal to the number of target duration distribution intervals, and the median duration is expressed as:

[0024]

[0025] Among them, t z represents the median duration, and T max represents the maximum target duration of the target duration distribution interval, and T min represents the minimum target duration of the target duration distribution interval.

[0026] Optionally, the adjusting the multiple visible duration distribution intervals to obtain multiple target duration distribution intervals includes:

[0027] Counting the multiple interval durations of the multiple visible duration distribution intervals, and calculating an interval dispersion factor according to the multiple interval durations, where the interval dispersion factor is expressed as:

[0028]

[0029] Among them, α represents the interval dispersion factor, n represents the number of interval durations among the multiple interval durations, represents the i-th interval duration, represents the j-th interval duration;

[0030] Judging whether the interval dispersion factor is greater than a preset standard dispersion factor;

[0031] If the interval dispersion factor is greater than the standard dispersion factor, adjusting the multiple visible duration distribution intervals to obtain multiple adjusted duration distribution intervals;

[0032] Updating the multiple visible duration distribution intervals by using the multiple adjusted duration distribution intervals, and returning to the step of counting the multiple interval durations of the multiple visible duration distribution intervals;

[0033] If the interval dispersion factor is not greater than the standard dispersion factor, recording the multiple visible duration distribution intervals as multiple target duration distribution intervals.

[0034] Optionally, the adjusting the multiple visible duration distribution intervals to obtain multiple adjusted duration distribution intervals includes:

[0035] Calculating the average duration of the multiple interval durations, and sequentially extracting visible duration distribution intervals from the multiple visible duration distribution intervals;

[0036] Obtaining the number of visible durations and the interval length within the visible duration distribution interval, and judging whether the visible duration distribution interval is a preset majority duration distribution interval, where the number of visible durations within the majority duration distribution interval is greater than the average duration;

[0037] If the visible duration distribution interval is not a majority duration distribution interval, calculating an allocable length based on the standard dispersion factor, the number of visible durations, and the interval length, where the allocable length is expressed as:

[0038]

[0039] Among them, ΔL represents the allocable length, and α b represents the standard discrete factor, and t s represents the number of visible duration, and L q represents the interval length, and down(*) represents the floor function;

[0040] Identify the nearest duration distribution interval of the visible duration distribution interval, and determine whether the nearest duration distribution interval is the majority duration distribution interval;

[0041] If the nearest duration distribution interval is not the majority duration distribution interval, then based on the allocable length, use the visible duration distribution interval to supplement the nearest duration distribution interval to obtain a first allocated duration interval and a second allocated duration interval, where the first allocated duration interval is the allocated visible duration distribution interval, and the second allocated duration interval is the allocated nearest duration distribution interval;

[0042] Summarize the first allocated duration interval and the second allocated duration interval to obtain multiple adjusted duration distribution intervals.

[0043] Optionally, the performing metabolic marker analysis according to the multiple serum metabolome sets to obtain an effective metabolic marker group includes:

[0044] Successively obtain the number of metabolic markers in each serum metabolome of the multiple serum metabolome sets to obtain a metabolic marker number set;

[0045] Set a minimum marker number, and identify multiple low-frequency marker numbers less than the minimum marker number in the metabolic marker number set;

[0046] Remove the multiple low-frequency marker numbers from the metabolic marker number set to obtain an effective marker number set, and extract the corresponding multiple effective serum metabolome sets of the effective marker number set from the multiple serum metabolome sets;

[0047] Calculate the average marker number of the effective marker number set, and count the metabolite frequency group set of the multiple effective serum metabolome sets, where the metabolite frequency group contains a metabolic marker and the marker frequency of the metabolic marker appearing in the multiple effective serum metabolome sets;

[0048] Sort the metabolite frequency group set to obtain a metabolite frequency group sequence set, where the metabolite frequency group with a higher marker frequency is sorted in the front;

[0049] Based on the average number of markers, high-order extraction is performed on the metabolite frequency group sequence set to obtain an effective metabolite marker group, where the number of effective metabolite markers in the effective metabolite marker group is equal to the average number of markers.

[0050] Optionally, based on the effective duration group and the effective metabolite marker group, data standardization is performed on the original pregnancy data table set to obtain a standard pregnancy data table set, including:

[0051] Obtain the number of effective durations in the effective duration group, and identify the corresponding effective pregnancy data table set of the multiple effective serum metabolite group sets in the original pregnancy data table set;

[0052] Perform the following operations on each effective pregnancy data table in the effective pregnancy data table set:

[0053] Extract the effective pregnancy duration group and the effective pregnancy metabolite group set of the effective pregnancy data table, obtain the number of effective pregnancy durations in the effective pregnancy duration group, and calculate the effective difference between the number of effective pregnancy durations and the number of effective durations;

[0054] Judge whether the number of effective pregnancy durations and the effective difference are respectively less than the number of effective durations and the preset standard deviation number;

[0055] If the number of effective pregnancy durations and the effective difference are not respectively less than the number of effective durations and the standard deviation number, then perform pregnancy data standardization on the effective pregnancy metabolite group set according to the effective duration group and the effective metabolite marker group to obtain a standard pregnancy duration group and a standard pregnancy metabolite group set;

[0056] Use the standard pregnancy duration group and the standard pregnancy metabolite group set to update the effective pregnancy data table to obtain a standard pregnancy data table;

[0057] Summarize the standard pregnancy data tables to obtain a standard pregnancy data table set.

[0058] Optionally, the performing pregnancy data standardization on the effective pregnancy metabolite group set according to the effective duration group and the effective metabolite marker group to obtain a standard pregnancy duration group and a standard pregnancy metabolite group set includes:

[0059] Successively extract the effective durations in the effective duration group, and based on the effective durations, judge whether there are the same pregnancy durations in the effective pregnancy duration group;

[0060] If there are the same pregnancy durations in the effective pregnancy duration group, record the same pregnancy durations as the standard pregnancy duration, and extract the standard pregnancy metabolite corresponding to the standard pregnancy duration in the effective pregnancy metabolite group set;

[0061] If there is no same gestational duration within the group of effective gestational durations, record the effective duration as the standard gestational duration, and calculate the linear weight number of the group of effective gestational durations according to a preset regulation factor and a regulation value function, where the linear weight number is expressed as:

[0062]

[0063] where K represents the linear weight number, up(*) represents the ceiling function, f(a - b) represents the regulation value function, k represents the regulation factor, N1 is the number of effective gestational durations in the group of effective gestational durations, and N2 is the number of effective durations in the group of effective durations;

[0064] Based on the linear weight number, identify the linear gestational duration group of the effective duration in the group of effective gestational durations, where the linear gestational duration group includes multiple effective gestational durations close to the effective duration, and the number of linear gestational durations in the linear gestational duration group is the linear weight number;

[0065] Calculate the standard gestational metabolism group according to the linear gestational duration group, the standard gestational duration, and the group of effective metabolic markers;

[0066] Summarize the standard gestational duration and the standard gestational metabolism group respectively to obtain the group of standard gestational durations and the set of standard gestational metabolism groups.

[0067] Optionally, the calculating the standard gestational metabolism group according to the linear gestational duration group, the standard gestational duration, and the group of effective metabolic markers includes:

[0068] Identify the set of linear gestational metabolism groups corresponding to the linear gestational duration group, and perform data integration on the set of linear gestational metabolism groups based on the group of effective metabolic markers to obtain the set of linear metabolic content groups;

[0069] Successively extract the linear metabolic content groups in the set of linear metabolic content groups, and calculate the set of linear metabolic factors based on the linear gestational duration group and the linear metabolic content groups, where the set of linear metabolic factors is expressed as:

[0070] G = {…, G k ,…}

[0071]

[0072] where G represents the set of linear metabolic factors, G k represents the k-th linear metabolic factor, represents the (k + 1)-th linear metabolic content in the linear metabolic content group, represents the k-th linear metabolic content, represents the (k + 1)-th linear gestational duration in the linear gestational duration group, represents the k-th linear gestational duration;

[0073] Calculate the average metabolic factor of the set of linear metabolic factors, identify the nearest gestational duration to the standard gestational duration in the linear gestational duration group, and identify the nearest metabolic content corresponding to the nearest gestational duration in the linear metabolic content group;

[0074] Based on the average metabolic factor, the standard gestational duration, the nearest gestational duration, and the nearest metabolic content, calculate the standard gestational metabolic content, where the standard gestational metabolic content is expressed as:

[0075]

[0076] where W b represents the standard gestational metabolic content, represents the average metabolic factor, T b represents the standard gestational duration, T z represents the nearest gestational duration, W z represents the nearest metabolic content, and |*| represents the absolute value symbol;

[0077] Summarize the standard gestational metabolic content to obtain the standard gestational metabolic group.

[0078] To achieve the above object, the present invention also provides a data analysis system for gestational diabetes based on serum metabolic markers, including:

[0079] A pregnancy table construction module for constructing an original pregnancy data table set, where the original pregnancy data table set contains the serum metabolome of different pregnant women at different gestational durations, and the original pregnancy data table set contains multiple gestational duration groups and multiple serum metabolome sets, and one gestational duration group corresponds to one serum metabolome set;

[0080] A gestational duration balancing module for extracting the original gestational duration groups in each original pregnancy data table in the original pregnancy data table set to obtain an original gestational duration group set, and performing gestational duration balancing on the original gestational duration group set to obtain an effective duration group;

[0081] A gestational metabolism analysis module for extracting multiple serum metabolome sets in the original pregnancy data table set, where one original pregnancy data table corresponds to one serum metabolome set, and performing metabolic marker analysis based on the multiple serum metabolome sets to obtain an effective metabolic marker group, where the effective metabolic marker group contains the names of serum metabolic markers and there are no duplicate names of serum metabolic markers;

[0082] A pregnancy data prediction module for normalizing the original pregnancy data table set based on the effective duration group and the effective metabolic marker group to obtain a standard pregnancy data table set, and performing machine learning based on the standard pregnancy data table set to obtain a predicted pregnancy data table.

[0083] In order to solve the above problem, the present invention further provides an electronic device, the electronic device comprising:

[0084] a memory storing at least one instruction; and

[0085] The processor executes the instructions stored in the memory to implement the above-mentioned gestational diabetes data analysis method based on serum metabolic markers.

[0086] In order to solve the above problems, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored. The at least one instruction is executed by a processor in an electronic device to implement the above-mentioned gestational diabetes data analysis method based on serum metabolic markers.

[0087] In order to solve the problems described in the background technology, the present invention first constructs an original pregnancy data table set, which collects a large number of pregnancy serum information of pregnant women from different online databases, provides a large amount of data support for subsequent data analysis, and improves the accuracy of data analysis. At the same time, information from different sources also increases the versatility of data analysis. Then, the original pregnancy duration group set is balanced to obtain an effective duration group. The acquisition of the effective duration group provides an important analysis basis for subsequent data analysis. At the same time, this step also normalizes the data, improves the accuracy of data analysis, and performs metabolic marker analysis based on multiple serum metabolic group sets to obtain an effective metabolic marker group. This The first step is to standardize the metabolic markers of different quantities and different gestational periods, so that the information about the content of metabolic markers is more intuitive, and the missing of some data is compensated, which improves the accuracy of data analysis. Then, based on the effective duration group and the effective metabolic marker group, the original pregnancy data table set is standardized to obtain a standard pregnancy data table set. This step completes the construction of a data table of standard specifications, paving the way for the subsequent deep learning. Finally, based on the standard pregnancy data table set, machine learning is performed to obtain a predicted pregnancy data table. The standard pregnancy data table set is converted into a prediction table that can be used to guide practical applications by deep learning, which greatly improves the versatility of this data analysis. Therefore, the present invention can improve the accuracy and versatility of gestational diabetes data analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] Figure 1 A schematic diagram of a process for analyzing gestational diabetes data based on serum metabolic markers provided by one embodiment of the present invention;

[0089] Figure 2 A functional module diagram of a gestational diabetes mellitus data analysis system based on serum metabolic markers provided by one embodiment of the present invention;

[0090] Figure 3 This is a schematic structural diagram of an electronic device for implementing the method for analyzing data of gestational diabetes based on serum metabolic markers provided in an embodiment of the present invention.

[0091] Explanation of reference numerals:

[0092] 1. Electronic device; 10. Processor; 11. Memory; 12. Bus.

[0093] The implementation, functional characteristics, and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the drawings. Specific embodiments

[0094] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0095] An embodiment of the present application provides a method for analyzing data of gestational diabetes based on serum metabolic markers. The execution subject of the method for analyzing data of gestational diabetes based on serum metabolic markers includes, but is not limited to, at least one of an electronic device such as a server, a terminal, etc. that can be configured to execute the method provided in the embodiment of the present application. In other words, the method for analyzing data of gestational diabetes based on serum metabolic markers can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc.

[0096] Refer to Figure 1 As shown, it is a schematic flowchart of a method for analyzing data of gestational diabetes based on serum metabolic markers provided in an embodiment of the present invention. In this embodiment, the method for analyzing data of gestational diabetes based on serum metabolic markers includes:

[0097] S1. Construct an original pregnancy data table set, where the original pregnancy data table set contains the serum metabolomes of different pregnant women at different pregnancy durations, and the original pregnancy data table set contains multiple pregnancy duration groups and multiple serum metabolome sets, and one pregnancy duration group corresponds to one serum metabolome set.

[0098] Specifically, the construction of the original pregnancy data table set includes:

[0099] Obtain a pregnancy database, where the pregnancy database contains the serum metabolic information and pregnancy duration information of pregnant women during pregnancy;

[0100] Retrieve the pregnant woman identity identifier in the pregnancy database in sequence, and obtain the original pregnancy data set corresponding to the pregnant woman identity identifier in the pregnancy database. The original pregnancy data set includes: a pregnancy duration set and a corresponding serum metabolome set, and one pregnancy duration in the pregnancy duration set corresponds to one serum metabolome in the serum metabolome set;

[0101] Based on the original pregnancy data set, create an original pregnancy data table for the pregnant woman identity identifier;

[0102] When the original pregnancy data table is equal to the preset standard number of tables, stop the step of retrieving the pregnant woman identity identifier;

[0103] Summarize the original pregnancy data table to obtain an original pregnancy data table set.

[0104] It is understandable that the pregnancy database refers to a database with a large amount of serum metabolism information of pregnant women during pregnancy. For example: a database of serum samples collected by medical institutions during routine pregnancy examinations, a clinical sample library owned by research institutions or some universities, academic papers or related literature on published serum metabolism data, an online database maintained by the public health department containing pregnant women's health information, some online platforms such as Dryad, Figshare, etc. The pregnant woman identity identifier refers to the information used to identify pregnant women when querying the database. For example: in some literatures, the pregnant woman identity is marked as "Pregnant Woman 1" or "Pregnant Woman a", or in some medical institutions, there are specific pregnant women's names or medical record numbers. The original pregnancy data set refers to all the information about the pregnant woman identity identifier queried in the pregnancy database. The original pregnancy data table is a relationship table between pregnancy duration and serum metabolome, where one pregnancy duration corresponds to one serum metabolome, and the standard number of tables refers to a constant set by humans.

[0105] It should be explained that since the serum metabolome in the collected original pregnancy data may not be complete, the number of serum metabolisms included in different pregnant woman identity identifiers or different serum metabolomes of the same pregnant woman identity identifier may not be the same. When initially retrieving the original pregnancy data set, the differences in the above serum metabolomes are not considered, that is, pregnant woman identity identifiers with clear pregnancy duration and serum metabolome records in the pregnancy database can all be recorded.

[0106] S2. Extract the original pregnancy duration groups in each original pregnancy data table in the original pregnancy data table set to obtain an original pregnancy duration group set, and perform pregnancy duration balancing on the original pregnancy duration group set to obtain an effective duration group.

[0107] Specifically, the performing pregnancy duration balancing on the original pregnancy duration group set to obtain an effective duration group includes:

[0108] Obtain a reference pregnancy duration;

[0109] Construct a plurality of original duration distribution intervals according to the reference pregnancy duration and the preset number of reference durations. Among them, the interval length of each original duration distribution interval is the same, and the number of original duration distribution intervals is equal to the number of reference durations plus 1;

[0110] Use the set of original pregnancy durations to supplement the plurality of original duration distribution intervals to obtain a plurality of visible duration distribution intervals;

[0111] Adjust the intervals of the plurality of visible duration distribution intervals to obtain a plurality of target duration distribution intervals;

[0112] Calculate the median duration of each target duration distribution interval in the plurality of target duration distribution intervals to obtain a set of median durations, and record the median durations as a set of effective durations. Among them, the number of median durations in the set of median durations is equal to the number of target duration distribution intervals, and the median duration is expressed as:

[0113]

[0114] where t z represents the median duration, T max represents the maximum target duration of the target duration distribution interval, and T min represents the minimum target duration of the target duration distribution interval.

[0115] It can be understood that the reference pregnancy duration refers to the total duration of pregnancy. This reference pregnancy duration is usually 40 weeks. The number of reference durations refers to a constant set by humans and is used to determine the number of original duration distribution intervals. The original duration distribution interval refers to the duration range of pregnancy duration. For example, the original duration distribution interval can be set as: 0 weeks to 5 weeks. The visible duration distribution interval refers to the original duration distribution interval after digital filling. For example, A is an original duration distribution interval, which represents the pregnancy duration range from 0 weeks to 5 weeks. A certain set of original pregnancy durations contains 20 pregnancy durations from 0 weeks to 5 weeks. Now, use this set of original pregnancy durations to supplement this original duration distribution interval to obtain a visible duration distribution interval B = {20}. The median duration refers to the median of the target distribution interval.

[0116] It should be noted that the set of effective durations is a combination of multiple durations. For example: (5 weeks, 10 weeks, 15 weeks, 20 weeks...).

[0117] Specifically, the adjusting the intervals of the plurality of visible duration distribution intervals to obtain a plurality of target duration distribution intervals includes:

[0118] Count the number of interval durations of the plurality of visible duration distribution intervals, and calculate the interval dispersion factor according to the number of interval durations. The interval dispersion factor is expressed as:

[0119]

[0120] Wherein, α represents the interval discrete factor, n represents the number of interval durations among multiple interval durations, represents the i-th interval duration, represents the j-th interval duration;

[0121] Determine whether the interval discrete factor is greater than a preset standard discrete factor;

[0122] If the interval discrete factor is greater than the standard discrete factor, adjust the multiple visible duration distribution intervals to obtain multiple adjusted duration distribution intervals;

[0123] Update the multiple visible duration distribution intervals using the multiple adjusted duration distribution intervals, and return the step of counting the multiple interval durations of the multiple visible duration distribution intervals;

[0124] If the interval discrete factor is not greater than the standard discrete factor, record the multiple visible duration distribution intervals as multiple target duration distribution intervals.

[0125] It can be understood that the interval duration refers to the number of pregnancy durations in the visible duration distribution interval, the interval discrete factor refers to a value used to represent the discrete degree of the pregnancy duration distribution in the multiple visible duration distribution intervals, the standard discrete factor refers to a constant set by humans, and the adjusted duration distribution interval refers to the visible duration distribution interval after adjustment. When the interval discrete factor is greater than the standard discrete factor, it indicates that the current division of the duration intervals is not uniform. If data analysis is continued according to the current duration intervals, a large error will occur. Therefore, it is necessary to adjust the visible duration distribution intervals.

[0126] Specifically, the adjusting the multiple visible duration distribution intervals to obtain multiple adjusted duration distribution intervals includes:

[0127] Calculate the average duration of the multiple interval durations, and sequentially extract the visible duration distribution intervals in the multiple visible duration distribution intervals;

[0128] Obtain the visible duration number and the interval length within the visible duration distribution interval, and determine whether the visible duration distribution interval is a preset majority duration distribution interval, wherein the visible duration number within the majority duration distribution interval is greater than the average duration;

[0129] If the visible duration distribution interval is not the majority duration distribution interval, calculate the allocable length based on the standard discrete factor, the visible duration number, and the interval length, where the allocable length is expressed as:

[0130]

[0131] Among them, ΔL represents the allocable length, and α b represents the standard discrete factor, and t s represents the number of visible duration, and L q represents the interval length, and down(*) represents the floor function;

[0132] Identify the nearest duration distribution interval of the visible duration distribution interval, and determine whether the nearest duration distribution interval is the majority duration distribution interval;

[0133] If the nearest duration distribution interval is not the majority duration distribution interval, then based on the allocable length, use the visible duration distribution interval to supplement the nearest duration distribution interval to obtain the first allocated duration interval and the second allocated duration interval. Among them, the first allocated duration interval is the visible duration distribution interval after allocation, and the second allocated duration interval is the nearest duration distribution interval after allocation;

[0134] Summarize the first allocated duration interval and the second allocated duration interval to obtain multiple adjusted duration distribution intervals.

[0135] It should be explained that the average duration number refers to the average of the duration numbers of multiple intervals, the visible duration number refers to the number of pregnancy durations within the visible duration distribution interval, and the interval length refers to the length of the visible duration distribution interval.

[0136] Exemplarily, a certain visible duration distribution interval is: B = {30}, and the pregnancy duration range represented by this visible duration distribution interval is: 20 weeks to 25 weeks. Then the visible duration number is 30, and the interval length is 25 weeks - 20 weeks = 5 weeks.

[0137] Furthermore, the allocable length refers to the duration length that can be used to allocate duration to the minority duration distribution intervals in the visible duration distribution interval. The detailed steps of using the visible duration distribution interval to supplement the nearest duration distribution interval are as follows: Intercept a unit duration interval adjacent to the nearest duration distribution interval in the visible duration distribution interval, where the length of the unit duration interval is the allocable length, and supplement this unit duration interval to the nearest duration distribution interval.

[0138] Exemplarily, the pregnancy duration range represented by a certain visible duration distribution interval is: 30 weeks to 35 weeks, and for this visible duration distribution interval, the pregnancy duration range represented by the nearest duration distribution interval is: 35 weeks to 40 weeks. It is calculated that the allocable length of this visible duration distribution interval is 1 week. Then the duration range of the visible duration distribution interval is changed to: 30 weeks to (35 weeks - 1 week = 34 weeks), and the duration range of the nearest duration distribution interval is changed to: (35 weeks - 1 week = 34 weeks) to 40 weeks.

[0139] S3. Extract multiple serum metabolome sets from the original pregnancy data set, where one original pregnancy data table corresponds to one serum metabolome set.

[0140] It should be explained that the serum metabolome set refers to the set of all serum metabolomes in the original pregnancy data set.

[0141] S4. According to the multiple serum metabolome sets, perform metabolic marker analysis to obtain an effective metabolic marker group, where the effective metabolic marker group contains the names of serum metabolic markers and there are no duplicate names of serum metabolic markers.

[0142] It can be understood that the effective metabolic marker group refers to a combination without duplicate marker names, and its detailed acquisition method will be given later.

[0143] Specifically, the performing metabolic marker analysis according to the multiple serum metabolome sets to obtain an effective metabolic marker group includes:

[0144] Successively obtain the number of metabolic markers in each serum metabolome in the multiple serum metabolome sets to obtain a set of metabolic marker numbers;

[0145] Set a minimum marker number, and identify multiple low-frequency marker numbers less than the minimum marker number in the set of metabolic marker numbers;

[0146] Remove the multiple low-frequency marker numbers from the set of metabolic marker numbers to obtain a set of effective marker numbers, and extract the corresponding multiple effective serum metabolome sets from the multiple serum metabolome sets;

[0147] Calculate the average marker number of the set of effective marker numbers, and count the metabolite frequency group set of the multiple effective serum metabolome sets, where the metabolite frequency group contains a metabolic marker and the marker frequency of the metabolic marker appearing in the multiple effective serum metabolome sets;

[0148] Sort the metabolite frequency group set to obtain a metabolite frequency group order set, where the metabolite frequency group with a higher marker frequency is sorted in the front;

[0149] Based on the average marker number, perform high-order extraction in the metabolite frequency group order set to obtain an effective metabolic marker group, where the number of effective metabolic markers in the effective metabolic marker group is equal to the average marker number.

[0150] It is understandable that the number of metabolic markers refers to the number of serum metabolic markers in the serum metabolome. For example, for a certain serum metabolome: (insulin 30 mIU / L, glycated hemoglobin 3.5%, triglyceride 1.2 mmol / L, serum creatinine 60 μmol / L), the number of metabolic markers is 4. The minimum marker number refers to a constant number of markers set artificially. The low-frequency marker number refers to the number of metabolic markers less than the minimum marker number. The effective marker number set refers to the set of metabolic marker numbers after deleting the low-frequency marker numbers. The average marker number refers to the average in the effective marker number set.

[0151] It should be explained that the metabolite frequency group refers to the combination of a certain effective serum metabolic marker that appears in multiple effective serum metabolome sets and the number of times it appears. For example, if insulin appears 50 times in multiple effective serum metabolome sets, the metabolite frequency group corresponding to insulin is: (insulin: 50). The effective metabolic marker group refers to the metabolic markers in the metabolite frequency groups ranked in the top a positions in the metabolite frequency group sequence set, where a represents the average marker number.

[0152] It should be noted that the effective metabolic marker group is a combination composed of various serum metabolic markers. For example: (blood glucose, insulin, glycated hemoglobin, blood lipid, serum creatinine...). This serum metabolic marker group only contains the names of metabolic markers, but does not limit the specific serum metabolic marker data.

[0153] Specifically, based on the effective duration group and the effective metabolic marker group, data standardization is performed on the original pregnancy data dataset to obtain a standard pregnancy data dataset, including:

[0154] Obtain the effective duration number of the effective duration group, and identify the effective pregnancy data dataset corresponding to the multiple effective serum metabolome sets in the original pregnancy data dataset;

[0155] Perform the following operations on each effective pregnancy data table in the effective pregnancy data dataset:

[0156] Extract the effective pregnancy duration group and the effective pregnancy metabolome set of the effective pregnancy data table, obtain the effective pregnancy duration number of the effective pregnancy duration group, and calculate the effective difference number between the effective pregnancy duration number and the effective duration number;

[0157] Judge whether the effective pregnancy duration number and the effective difference number are respectively less than the effective duration number and a preset standard deviation number;

[0158] If the effective pregnancy duration number and the effective difference number are not respectively less than the effective duration number and the standard deviation number, then perform pregnancy data standardization on the effective pregnancy metabolome set according to the effective duration group and the effective metabolic marker group to obtain a standard pregnancy duration group and a standard pregnancy metabolome set;

[0159] Using the standard pregnancy duration group and the standard pregnancy metabolic group set, the effective pregnancy data table is updated to obtain the standard pregnancy data table;

[0160] The standard pregnancy data tables are summarized to obtain a set of standard pregnancy data tables.

[0161] It can be understood that the number of effective durations refers to the number of effective durations in the effective duration group, the effective pregnancy duration group and the effective pregnancy metabolic group set respectively refer to the pregnancy duration group and the serum metabolic group set in the effective pregnancy data table, the number of effective pregnancy durations refers to the number of effective pregnancy durations in the effective pregnancy duration group, the standard pregnancy duration group refers to the effective duration group after standardization, wherein the standard pregnancy duration group is numerically the same as the effective duration group, the standard pregnancy metabolic group refers to the effective pregnancy metabolic group in the effective pregnancy metabolic group set that corresponds to the standard pregnancy duration, and the detailed acquisition steps of the above-mentioned standard pregnancy duration group and standard pregnancy metabolic group set will be given later.

[0162] It should be explained that the purpose of the step of identifying the valid pregnancy data table set corresponding to the multiple valid serum metabolic group sets in the original pregnancy data table set is to delete the original pregnancy data tables with a smaller number of serum metabolic markers, and the number of serum metabolic markers in the remaining valid pregnancy data tables is greater than the minimum number of markers.

[0163] Furthermore, when the number of effective pregnancy durations in the effective pregnancy duration group is less than the effective duration number, it means that the metabolic information represented by the effective pregnancy duration group is less. At the same time, if the difference between the effective pregnancy duration number and the effective duration number at this time is less than the standard deviation number, it means that the metabolic information represented in the effective pregnancy duration group is very little. If it is brought into the subsequent data analysis, it may cause a large error and needs to be discarded, that is, only the effective pregnancy duration group when the effective pregnancy duration number is greater than the effective duration number, or the effective pregnancy duration number is less than the effective duration number but the corresponding effective difference number is greater than the standard deviation number is retained.

[0164] Specifically, the effective pregnancy duration group and the effective metabolic marker group are used to standardize the pregnancy data of the effective pregnancy metabolic group set to obtain the standard pregnancy duration group and the standard pregnancy metabolic group set, including:

[0165] Extracting effective durations in the effective duration groups in turn, and judging whether there is the same gestational duration in the effective gestational duration groups based on the effective durations;

[0166] If the same pregnancy duration exists in the effective pregnancy duration group, the same pregnancy duration is recorded as the standard pregnancy duration, and the standard pregnancy metabolic group corresponding to the standard pregnancy duration is extracted from the effective pregnancy metabolic group set;

[0167] If there is no same gestational age duration within the effective gestational age duration group, record the effective duration as the standard gestational age duration, and calculate the linear weight number of the effective gestational age duration group according to a preset regulation factor and a regulation value function, where the linear weight number is expressed as:

[0168]

[0169] where K represents the linear weight number, up(*) represents the ceiling function, f(a - b) represents the regulation value function, k represents the regulation factor, N1 is the number of effective gestational age durations in the effective gestational age duration group, and N2 is the number of effective durations in the effective duration group;

[0170] Based on the linear weight number, identify the linear gestational age duration group of the effective duration in the effective gestational age duration group, where the linear gestational age duration group includes multiple effective gestational age durations close to the effective duration, and the number of linear gestational age durations in the linear gestational age duration group is the linear weight number;

[0171] Calculate the standard gestational metabolic group according to the linear gestational age duration group, the standard gestational age duration, and the effective metabolic marker group;

[0172] Summarize the standard gestational age duration and the standard gestational metabolic group respectively to obtain the standard gestational age duration group and the standard gestational metabolic group set.

[0173] It can be understood that the same gestational age duration refers to the effective gestational age duration with the same value as the effective duration that appears in the effective gestational age duration group. The regulation factor refers to a constant set artificially and is used to regulate the result when calculating the linear weight number later. For example, the regulation factor can be set to 1, 2, etc. The regulation value function will present different functional relationships according to the relationship between different input variables and the regulation factor. The linear weight coefficient refers to the number of linear gestational age durations in the linear gestational age duration group. The linear gestational age duration group refers to K effective gestational age durations close to the effective duration. For example, if an effective gestational age duration group is: (10 weeks, 15 weeks, 20 weeks, 25 weeks, 30 weeks, 35 weeks, 40 weeks), the effective duration is 28 weeks, and the linear weight number is 4, then the linear gestational age duration group corresponding to this effective duration is: (20 weeks, 25 weeks, 30 weeks, 35 weeks).

[0174] S5. Based on the effective duration group and the effective metabolic marker group, perform data standardization on the original pregnancy data table set to obtain the standard pregnancy data table set.

[0175] Specifically, the calculating the standard gestational metabolic group according to the linear gestational age duration group, the standard gestational age duration, and the effective metabolic marker group includes:

[0176] Identify the linear pregnancy metabolome set corresponding to the linear pregnancy duration group, and based on the effective metabolite marker set, perform data integration on the linear pregnancy metabolome set to obtain a linear metabolite content set;

[0177] Successively extract linear metabolite content groups from the linear metabolite content set, and calculate a linear metabolite factor set based on the linear pregnancy duration group and the linear metabolite content group, where the linear metabolite factor set is expressed as:

[0178] G = {…, G k ,…}

[0179]

[0180] where G represents the linear metabolite factor set, and G k represents the k-th linear metabolite factor, represents the (k + 1)-th linear metabolite content in the linear metabolite content group, represents the k-th linear metabolite content, represents the (k + 1)-th linear pregnancy duration in the linear pregnancy duration group, represents the k-th linear pregnancy duration;

[0181] Calculate the average metabolite factor of the linear metabolite factor set, identify the nearest pregnancy duration closest to the standard pregnancy duration in the linear pregnancy duration group, and identify the nearest metabolite content corresponding to the nearest pregnancy duration in the linear metabolite content group;

[0182] Calculate the standard pregnancy metabolite content based on the average metabolite factor, the standard pregnancy duration, the nearest pregnancy duration, and the nearest metabolite content, where the standard pregnancy metabolite content is expressed as:

[0183]

[0184] where W b represents the standard pregnancy metabolite content, represents the average metabolite factor, T b represents the standard pregnancy duration, T z represents the nearest pregnancy duration, W z represents the nearest metabolite content, and |*| represents the absolute value symbol;

[0185] Summarize the standard pregnancy metabolite content to obtain the standard pregnancy metabolome.

[0186] It is understandable that the linear pregnancy metabolome set refers to the linear pregnancy metabolome set corresponding to the linear pregnancy duration group in the effective pregnancy data table. The linear metabolite content set refers to the content set containing the effective metabolite marker groups in the linear pregnancy metabolome set, which represents the contents of different serum metabolite markers at the same pregnancy duration in the linear pregnancy metabolome, while the linear metabolite content group represents the contents at different pregnancy durations of the same serum metabolite marker.

[0187] It should be explained that the step of integrating data for the linear pregnancy metabolome set is explained as follows: The linear pregnancy metabolome corresponds one-to-one with the linear pregnancy duration, that is, multiple metabolite marker contents correspond to one linear pregnancy duration. For example, when the linear pregnancy duration is 20 weeks, at this pregnancy duration, the linear pregnancy metabolome of the pregnant woman is: (insulin 30 mIU / L, glycated hemoglobin 3.5%, triglyceride 1.2 mmol / L, serum creatinine 60 μmol / L), and the linear metabolite content group corresponds one-to-one with the effective metabolite marker, that is, multiple contents corresponding to the effective metabolite marker correspond to one effective metabolite marker. For example, if the effective metabolite marker is insulin, the corresponding linear metabolite content group is: (content at 5 weeks is 30 mIU / L, content at 10 weeks is 29.5 mIU / L, content at 15 weeks is 29 mIU / L, content at 20 weeks is 30.5 mIU / L...).

[0188] Furthermore, the linear metabolite factor refers to the slope between two adjacent linear metabolite contents, the average metabolite factor refers to the average of all linear metabolite factors in the linear metabolite factor set, and the nearest metabolite content refers to the content of the metabolite marker corresponding to the nearest pregnancy duration.

[0189] Exemplarily, the linear metabolite content group of a certain insulin is: (content at 5 weeks is 30 mIU / L, content at 10 weeks is 29.5 mIU / L, content at 15 weeks is 29 mIU / L, content at 20 weeks is 30.5 mIU / L...). If the nearest pregnancy duration is 15 weeks, then the nearest metabolite content corresponding to this nearest pregnancy duration is: 29 mIU / L.

[0190] S6. Based on the standard pregnancy data table set, perform machine learning to obtain a predicted pregnancy data table, and complete the data analysis of gestational diabetes based on serum metabolite markers based on the predicted pregnancy data table.

[0191] It is understandable that the steps of performing machine learning are as follows: Select a machine learning model, such as: MLP, DNN, CNN. Store the standard pregnancy data table as a csv file, and train multiple table files on the selected machine learning model to obtain a predicted pregnancy data table. The predicted pregnancy data table contains a predicted pregnancy duration group and multiple predicted pregnancy metabolome sets, where the predicted pregnancy duration group is the same as the standard pregnancy duration group.

[0192] To solve the problems described in the background art, the present invention first constructs an original pregnancy data table set, which collects a large amount of pregnant women's pregnancy serum information from different online databases, providing a large amount of data support for subsequent data analysis, improving the accuracy of data analysis. At the same time, the information from different sources also increases the versatility of data analysis. Then, the pregnancy duration balance is carried out on the original pregnancy duration set to obtain an effective duration set. The acquisition of this effective duration set provides an important analysis basis for subsequent data analysis. At the same time, this step also standardizes the data, improving the accuracy of data analysis. Through the analysis of multiple serum metabolome sets, an effective metabolite marker set is obtained. This step standardizes the metabolite markers with different quantities and different pregnancy durations, making the information on the content of metabolite markers more intuitive and compensating for the missing part of the data, improving the accuracy of data analysis. Then, based on the effective duration set and the effective metabolite marker set, the original pregnancy data table set is standardized to obtain a standard pregnancy data table set. This step completes the construction of a data table with a standard specification, laying a foundation for the subsequent deep learning. Finally, based on the standard pregnancy data table set, machine learning is carried out to obtain a predicted pregnancy data table. By means of deep learning, the standard pregnancy data table set is transformed into a prediction table that can be used to guide practical applications, which greatly improves the versatility of this data analysis. Therefore, the present invention can improve the accuracy and versatility of gestational diabetes data analysis.

[0193] As Figure 2 shown, it is a functional module diagram of a gestational diabetes data analysis system based on serum metabolite markers provided by an embodiment of the present invention.

[0194] The gestational diabetes data analysis system 100 based on serum metabolite markers of the present invention can be installed in an electronic device. According to the implemented functions, the gestational diabetes data analysis system 100 based on serum metabolite markers can include a pregnancy table construction module 101, a pregnancy duration balance module 102, a pregnancy metabolism analysis module 103, and a pregnancy data prediction module 104. The modules of the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0195] The pregnancy table construction module 101 is used to construct an original pregnancy data table set, wherein the original pregnancy data table set includes the serum metabolomes of different pregnant women at different pregnancy durations, and the original pregnancy data table set includes multiple pregnancy duration sets and multiple serum metabolome sets, and one pregnancy duration set corresponds to one serum metabolome set;

[0196] The pregnancy duration balancing module 102 is configured to extract the original pregnancy duration groups from each original pregnancy data table in the original pregnancy data table set, obtain the original pregnancy duration group set, and perform pregnancy duration balancing on the original pregnancy duration group set to obtain the effective duration group;

[0197] The pregnancy metabolism analysis module 103 is configured to extract multiple serum metabolome sets from the original pregnancy data table set, where one original pregnancy data table corresponds to one serum metabolome set, and perform metabolite marker analysis based on the multiple serum metabolome sets to obtain the effective metabolite marker group, where the effective metabolite marker group includes the names of serum metabolite markers and there are no duplicate names of serum metabolite markers;

[0198] The pregnancy data prediction module 104 is configured to perform data standardization on the original pregnancy data table set based on the effective duration group and the effective metabolite marker group to obtain the standard pregnancy data table set, and perform machine learning based on the standard pregnancy data table set to obtain the predicted pregnancy data table.

[0199] Specifically, each module in the gestational diabetes data analysis system 100 based on serum metabolite markers in the embodiments of the present invention adopts the same technical means as those in the above Figure 1 The technical means described in the gestational diabetes data analysis method based on serum metabolite markers, and can produce the same technical effects, which will not be elaborated here.

[0200] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing the gestational diabetes data analysis method based on serum metabolite markers provided by an embodiment of the present invention.

[0201] The electronic device 1 may include a processor 10, a memory 11, and a bus 12, and may further include a computer program stored in the memory 11 and executable on the processor 10, such as a gestational diabetes data analysis method program based on serum metabolite markers.

[0202] Among them, the memory 11 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, mobile hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), magnetic memory, magnetic disk, optical disc, etc. The memory 11 can be an internal storage unit of the electronic device 1 in some embodiments, such as the mobile hard disk of the electronic device 1. The memory 11 can also be an external storage device of the electronic device 1 in some other embodiments, such as a plug-in mobile hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 1. Further, the memory 11 also includes the internal storage unit of the electronic device 1 and also includes an external storage device. The memory 11 can be used not only to store application software installed in the electronic device 1 and various types of data, such as the code of the data analysis method program for gestational diabetes based on serum metabolic markers, etc., but also to temporarily store data that has been output or will be output.

[0203] The processor 10 can be composed of integrated circuits in some embodiments. For example, it can be composed of a single packaged integrated circuit, or can also be composed of multiple integrated circuits with the same or different functions packaged, including a combination of one or more Central Processing Units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc. The processor 10 is the control core (Control Unit) of the electronic device, connecting various components of the entire electronic device through various interfaces and lines, and by running or executing programs or modules (such as the data analysis method program for gestational diabetes based on serum metabolic markers, etc.) stored in the memory 11, and calling the data stored in the memory 11, to perform various functions of the electronic device 1 and process data.

[0204] The bus 12 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus 12 can be divided into an address bus, a data bus, a control bus, etc. The bus 12 is set to realize the connection and communication between the memory 11 and at least one processor 10, etc.

[0205] Figure 3 Only the electronic device with components is shown. Those skilled in the art can understand that, Figure 3The shown structure does not constitute a limitation on the electronic device 1, and it may include fewer or more components than shown, or combine certain components, or have a different component arrangement.

[0206] For example, although not shown, the electronic device 1 may further include a power source (such as a battery) for supplying power to each component. Preferably, the power source can be logically connected to the at least one processor 10 through a power management system, so as to implement functions such as charging management, discharging management, and power consumption management through the power management system. The power source may also include any components such as one or more DC or AC power sources, a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device 1 may also include various sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.

[0207] Furthermore, the electronic device 1 may further include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.

[0208] Optionally, the electronic device 1 may further include a user interface. The user interface may be a display, an input unit (such as a keyboard), and optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, which is used to display the information processed in the electronic device 1 and to display a visual user interface.

[0209] The program of the method for analyzing gestational diabetes data based on serum metabolic markers stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can implement:

[0210] Construct an original pregnancy data table set, where the original pregnancy data table set contains the serum metabolomes of different pregnant women at different pregnancy durations, and the original pregnancy data table set contains multiple pregnancy duration groups and multiple serum metabolome sets, and one pregnancy duration group corresponds to one serum metabolome set;

[0211] Extract the original pregnancy duration groups in each original pregnancy data table in the original pregnancy data table set to obtain an original pregnancy duration group set, and balance the pregnancy durations of the original pregnancy duration group set to obtain an effective duration group;

[0212] Extract multiple serum metabolome sets from the original pregnancy data set, where one original pregnancy data table corresponds to one serum metabolome set;

[0213] Perform metabolic marker analysis based on the multiple serum metabolome sets to obtain an effective metabolic marker set, where the effective metabolic marker set contains the names of serum metabolic markers and there are no duplicate names of serum metabolic markers;

[0214] Normalize the original pregnancy data set based on the effective duration set and the effective metabolic marker set to obtain a standard pregnancy data set;

[0215] Perform machine learning based on the standard pregnancy data set to obtain a predicted pregnancy data table, and complete the data analysis of gestational diabetes based on serum metabolic markers based on the predicted pregnancy data table.

[0216] Specifically, the specific implementation method of the above instructions by the processor 10 can refer to Figures 1 to 3 the description of the relevant steps in the corresponding embodiment, which will not be elaborated here.

[0217] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium can include: any entity or system that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory).

[0218] The present invention also provides a computer-readable storage medium, where the readable storage medium stores a computer program, and when the computer program is executed by a processor of an electronic device, it can implement:

[0219] Construct an original pregnancy data set, where the original pregnancy data set contains the serum metabolomes of different pregnant women at different pregnancy durations, and the original pregnancy data set contains multiple pregnancy duration groups and multiple serum metabolome sets, and one pregnancy duration group corresponds to one serum metabolome set;

[0220] Extract the original pregnancy duration groups in each original pregnancy data table in the original pregnancy data set to obtain an original pregnancy duration group set, and perform pregnancy duration balancing on the original pregnancy duration group set to obtain an effective duration set;

[0221] Extract multiple serum metabolome sets from the original pregnancy data set, where one original pregnancy data table corresponds to one serum metabolome set;

[0222] Based on the multiple serum metabolome sets, metabolic marker analysis is performed to obtain an effective metabolic marker set, wherein the effective metabolic marker set contains the names of serum metabolic markers and there are no duplicate names of serum metabolic markers;

[0223] Based on the effective duration set and the effective metabolic marker set, data normalization is performed on the original pregnancy data set to obtain a standard pregnancy data set;

[0224] Based on the standard pregnancy data set, machine learning is performed to obtain a predicted pregnancy data set, and based on the predicted pregnancy data set, data analysis of gestational diabetes mellitus based on serum metabolic markers is completed.

[0225] In several embodiments provided by the present invention, it should be understood that the disclosed devices, systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative, and there may be other division methods in actual implementation.

[0226] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0227] In addition, in each embodiment of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a hardware plus software functional module.

[0228] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0229] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for analyzing gestational diabetes data based on serum metabolic markers, characterized in that: The method comprises: Constructing an original pregnancy data table set, wherein the original pregnancy data table set includes serum metabolomes of different pregnant women at different pregnancy durations, and the original pregnancy data table set includes multiple pregnancy duration groups and multiple serum metabolome sets, and one pregnancy duration group corresponds to one serum metabolome set; Extracting the original pregnancy duration group in each original pregnancy data table from the original pregnancy data table set to obtain an original pregnancy duration group set, performing pregnancy duration balancing on the original pregnancy duration group set to obtain an effective duration group; Extracting multiple serum metabolome sets from the original pregnancy data table set, wherein one original pregnancy data table corresponds to one serum metabolome set; Performing a metabolite marker analysis according to the multiple serum metabolite group sets to obtain an effective metabolite marker group, wherein the effective metabolite marker group includes serum metabolite marker names, and there are no repeated serum metabolite marker names; Based on the effective duration group and the effective metabolic marker group, the original pregnancy data table set was standardized to obtain the standard pregnancy data table set; Based on the standard pregnancy data table set, machine learning is performed to obtain a predicted pregnancy data table, and based on the predicted pregnancy data table, gestational diabetes data analysis based on serum metabolic markers is completed.

2. The method for analyzing gestational diabetes data based on serum metabolic markers according to claim 1, characterized in that: The construction of the original pregnancy data table set includes: Acquiring a pregnancy database, wherein the pregnancy database includes serum metabolic information and pregnancy duration information of pregnant women during pregnancy; Retrieving the pregnant woman identity identifiers in the pregnancy database in sequence, and obtaining an original pregnancy data set corresponding to the pregnant woman identity identifier in the pregnancy database, wherein the original pregnancy data set includes: a pregnancy duration set and a corresponding serum metabolome set, and a pregnancy duration in the pregnancy duration set corresponds to a serum metabolome set in the serum metabolome set; Based on the original pregnancy data set, creating an original pregnancy data table for the pregnant woman identity identifier; When the number of original pregnancy data tables is equal to the preset number of standard tables, the step of retrieving the identity of the pregnant woman is stopped; The original pregnancy data tables are summarized to obtain an original pregnancy data table set.

3. The method for analyzing gestational diabetes data based on serum metabolic markers according to claim 2, characterized in that: The step of performing pregnancy duration balancing on the original pregnancy duration group set to obtain a valid duration group includes: Get the reference gestational duration; According to the reference pregnancy duration and the preset reference duration number, construct a plurality of original duration distribution intervals, wherein the interval length of each original duration distribution interval is the same, and the number of original duration distribution intervals is equal to the reference duration number plus 1; Using the original pregnancy duration group set, the multiple original duration distribution intervals are supplemented to obtain multiple visible duration distribution intervals; Performing interval adjustment on multiple visible duration distribution intervals to obtain multiple target duration distribution intervals; Calculate the median duration of each target duration distribution interval in the multiple target duration distribution intervals to obtain a median duration group, and record the median duration as a valid duration group, wherein the number of median durations in the median duration group is equal to the number of target duration distribution intervals, and the median duration is expressed as: Among them, t z represents the median duration, T max represents the maximum target duration of the target duration distribution interval, T min Indicates the minimum target duration in the target duration distribution interval.

4. The method for analyzing gestational diabetes data based on serum metabolic markers according to claim 3, characterized in that: The step of adjusting the multiple visible duration distribution intervals to obtain multiple target duration distribution intervals includes: The durations of multiple intervals of multiple visible duration distribution intervals are counted, and the interval discrete factor is calculated according to the multiple interval durations, wherein the interval discrete factor is expressed as: Among them, α represents the interval discrete factor, n represents the number of interval durations in multiple interval durations, represents the duration of the i-th interval, represents the duration of the jth interval; Determine whether the interval discrete factor is greater than a preset standard discrete factor; If the interval discrete factor is greater than the standard discrete factor, adjusting the multiple visible duration distribution intervals to obtain multiple adjusted duration distribution intervals; Using the multiple adjustment duration distribution intervals to update the multiple visible duration distribution intervals, and returning to the step of counting the multiple interval durations of the multiple visible duration distribution intervals; If the interval discrete factor is not greater than the standard discrete factor, the multiple visible duration distribution intervals are recorded as multiple target duration distribution intervals.

5. The method for analyzing gestational diabetes data based on serum metabolic markers according to claim 4, characterized in that: The adjusting the multiple visible duration distribution intervals to obtain multiple adjusted duration distribution intervals includes: Calculating the average duration of the multiple intervals, and sequentially extracting visible duration distribution intervals from the multiple visible duration distribution intervals; Obtaining the number of visible durations and the interval length in the visible duration distribution interval, and determining whether the visible duration distribution interval is a preset majority duration distribution interval, wherein the number of visible durations in the majority duration distribution interval is greater than the average duration number; If the visible duration distribution interval is not the majority duration distribution interval, the allocatable length is calculated based on the standard discrete factor, the number of visible durations and the interval length, where the allocatable length is expressed as: Among them, ΔL represents the allocatable length, α b represents the standard discrete factor, t s Indicates the viewing duration, L q Indicates the length of the interval, down(*) indicates the rounding down function; Identify the most recent duration distribution interval of the visible duration distribution interval, and determine whether the most recent duration distribution interval is the majority duration distribution interval; If the most recent duration distribution interval is not the majority duration distribution interval, based on the allocatable length, the visible duration distribution interval is used to supplement the most recent duration distribution interval to obtain a first allocated duration interval and a second allocated duration interval, wherein the first allocated duration interval is the visible duration distribution interval after allocation, and the second allocated duration interval is the most recent duration distribution interval after allocation; The first allocated duration interval and the second allocated duration interval are summarized to obtain a plurality of adjusted duration distribution intervals.

6. The method for analyzing gestational diabetes data based on serum metabolic markers according to claim 5, characterized in that: The step of performing metabolic marker analysis based on the multiple serum metabolome sets to obtain an effective metabolic marker set includes: sequentially obtaining the number of metabolite markers of each serum metabolite group in a plurality of serum metabolite group sets to obtain a metabolite marker number set; A minimum marker quantity is set, and multiple low-frequency marker quantities less than the minimum marker quantity are identified in the metabolic marker quantity set; Eliminating multiple low-frequency marker quantities from the metabolite marker quantity set to obtain a valid marker quantity set, and extracting multiple valid serum metabolite group sets corresponding to the valid marker quantity set from multiple serum metabolite group sets; Calculating the average marker quantity of the effective marker quantity set, and counting the metabolite frequency set of the multiple effective serum metabolite sets, wherein the metabolite frequency set includes a metabolite marker and the marker frequency of the metabolite marker appearing in the multiple effective serum metabolite sets; The metabolite frequency group sets are sorted to obtain a metabolite frequency group sequence set, wherein the metabolite frequency group with a higher marker frequency is sorted first; Based on the average number of markers, high-order extraction is performed in the metabolite frequency group sequence set to obtain an effective metabolite marker group, wherein the number of effective metabolite markers in the effective metabolite marker group is equal to the average number of markers.

7. The method for analyzing gestational diabetes data based on serum metabolic markers according to claim 6, characterized in that: Based on the effective duration group and the effective metabolic marker group, the original pregnancy data table set is standardized to obtain a standard pregnancy data table set, including: Obtaining the effective duration number of the effective duration group, and identifying the effective pregnancy data table sets corresponding to the multiple effective serum metabolic group sets in the original pregnancy data table set; The following operations are performed for each valid pregnancy data table in the valid pregnancy data table set: Extract the effective pregnancy duration group and the effective pregnancy metabolic group set from the effective pregnancy data table, obtain the effective pregnancy duration number of the effective pregnancy duration group, and calculate the effective difference between the effective pregnancy duration number and the effective duration number; Determine whether the effective pregnancy duration number and the effective difference number are respectively less than the effective duration number and the preset standard deviation number; If the effective pregnancy duration number and the effective difference number are not less than the effective duration number and the standard deviation number respectively, the pregnancy data of the effective pregnancy metabolic group set is standardized according to the effective duration group and the effective metabolic marker group to obtain the standard pregnancy duration group and the standard pregnancy metabolic group set; Using the standard pregnancy duration group and the standard pregnancy metabolic group set, the effective pregnancy data table is updated to obtain the standard pregnancy data table; The standard pregnancy data tables are summarized to obtain a set of standard pregnancy data tables.

8. The method for analyzing gestational diabetes data based on serum metabolic markers according to claim 7, characterized in that: The effective duration group and the effective metabolic marker group are used to standardize the pregnancy data of the effective pregnancy metabolic group set to obtain the standard pregnancy duration group and the standard pregnancy metabolic group set, including: Extracting effective durations in the effective duration groups in turn, and judging whether there is the same gestational duration in the effective gestational duration groups based on the effective durations; If the same pregnancy duration exists in the effective pregnancy duration group, the same pregnancy duration is recorded as the standard pregnancy duration, and the standard pregnancy metabolic group corresponding to the standard pregnancy duration is extracted from the effective pregnancy metabolic group set; If there is no identical pregnancy duration in the effective pregnancy duration group, the effective duration is recorded as the standard pregnancy duration, and the linear weight number of the effective pregnancy duration group is calculated according to the preset control factor and control value function, wherein the linear weight number is expressed as: Wherein, K represents the linear weight number, up(*) represents the upward rounding function, f(ab) represents the control value function, k represents the control factor, N1 represents the number of effective pregnancy durations in the effective pregnancy duration group, and N2 represents the number of effective durations in the effective duration group; Based on the linear weight number, identifying a linear pregnancy duration group of the effective duration in the effective pregnancy duration group, wherein the linear pregnancy duration group includes a plurality of effective pregnancy durations close to the effective duration, and the number of linear pregnancy durations in the linear pregnancy duration group is the linear weight number; The standard pregnancy metabolic group was calculated based on the linear pregnancy duration group, standard pregnancy duration group and effective metabolic marker group; The standard pregnancy duration and the standard pregnancy metabolic group are summarized respectively to obtain a standard pregnancy duration group and a standard pregnancy metabolic group set.

9. The method for analyzing gestational diabetes data based on serum metabolic markers according to claim 8, characterized in that: The standard pregnancy metabolic group is calculated according to the linear pregnancy duration group, the standard pregnancy duration group and the effective metabolic marker group, including: Identify the linear pregnancy metabolite group set corresponding to the linear pregnancy duration group, and integrate the linear pregnancy metabolite group set based on the effective metabolic marker group to obtain a linear metabolic content group set; The linear metabolic content groups are sequentially extracted from the linear metabolic content group set, and the linear metabolic factor set is calculated based on the linear gestational duration group and the linear metabolic content group, wherein the linear metabolic factor set is expressed as: G={…,G k ,…} Among them, G represents the linear metabolic factor set, G k represents the kth linear metabolic factor, represents the k+1th linear metabolic content in the linear metabolic content group, represents the kth linear metabolic content, represents the k+1th linear pregnancy duration in the linear pregnancy duration group, represents the length of the kth linear pregnancy period; Calculating the average metabolic factor of the linear metabolic factor set, identifying the most recent gestational duration closest to the standard gestational duration in the linear gestational duration group, and identifying the most recent metabolic content corresponding to the most recent gestational duration in the linear metabolic content group; The standard pregnancy metabolic content was calculated based on the average metabolic factor, the standard pregnancy duration, the most recent pregnancy duration, and the most recent metabolic content, where the standard pregnancy metabolic content was expressed as: Among them, W b Indicates the standard metabolic content during pregnancy. represents the average metabolic factor, T b represents the standard gestational period, T z Indicates the length of the most recent pregnancy, W z represents the recent metabolic content, and |*| represents the absolute value symbol; The standard pregnancy metabolic contents are summarized to obtain a standard pregnancy metabolic group.

10. A gestational diabetes data analysis system based on serum metabolic markers, characterized in that: The system comprises: A pregnancy table construction module, used to construct an original pregnancy data table set, wherein the original pregnancy data table set includes serum metabolomes of different pregnant women at different pregnancy durations, and the original pregnancy data table set includes multiple pregnancy duration groups and multiple serum metabolome sets, and one pregnancy duration group corresponds to one serum metabolome set; A pregnancy duration balancing module, used for extracting the original pregnancy duration groups in each original pregnancy data table in the original pregnancy data table set to obtain an original pregnancy duration group set, performing pregnancy duration balancing on the original pregnancy duration group set to obtain an effective duration group; A pregnancy metabolic analysis module, used to extract multiple serum metabolite groups from the original pregnancy data table set, wherein one original pregnancy data table corresponds to one serum metabolite group set, and perform metabolic marker analysis according to the multiple serum metabolite groups to obtain an effective metabolic marker group, wherein the effective metabolic marker group includes serum metabolic marker names, and there are no repeated serum metabolic marker names; The pregnancy data prediction module is used to perform data standardization on the original pregnancy data table set based on the effective duration group and the effective metabolic marker group to obtain a standard pregnancy data table set, and perform machine learning based on the standard pregnancy data table set to obtain a predicted pregnancy data table.