A biological environment monitoring data management system based on the Internet of Things

By collecting and cleaning biological environmental data through the Internet of Things, combining user needs and biological laws to judge the authenticity of the data, and screening out high-quality data for storage, the problem of screening out useless data in biological environmental monitoring data management is solved, and the accuracy and convenience of the data are improved.

CN119557564BActive Publication Date: 2025-10-03ZHAOQING TIANYING BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411728056.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-10-03
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

Existing biological environment monitoring data management technologies are unable to effectively filter out useless data, resulting in inconvenience in data use and affecting data accuracy and reliability.

Method used

Raw data is collected through the Internet of Things, and primary data is obtained after data cleaning. Useless data is filtered out according to user needs. The authenticity of the data is judged using biological environmental information and historical information database. Data is traced based on the correlation between monitoring objectives and research objectives, and high-level data is filtered out for storage.

Benefits of technology

It improves the accuracy and convenience of biological environment monitoring data, reduces the interference of false data, ensures the authenticity and environmental protection of data, and reduces resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119557564B_ABST
    Figure CN119557564B_ABST
Patent Text Reader

Abstract

This application discloses a bio-environment monitoring data management system based on the Internet of Things (IoT), relating to the technical field of monitoring data management. It includes: a data cleaning module that collects raw data through the IoT and cleans it to obtain primary data; a demand screening module that obtains user data requirements, deletes completely useless primary data, and obtains secondary data; a truth judgment module that extracts bio-environment information from the secondary data, determines whether the secondary data is true, and obtains a truth judgment result; a tracing judgment module that determines the importance of the secondary data and, based on the truth judgment result, determines whether to perform data tracing; a data tracing module that traces the secondary data, filters and deletes the data after tracing to obtain higher-level data; and a data storage module that stores the higher-level data to obtain a bio-environment monitoring database. This application improves the accuracy of IoT-based bio-environment monitoring data management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of monitoring data management, and in particular to a biological environment monitoring data management system based on the Internet of Things. Background Art

[0002] Bio-environmental monitoring data management refers to the use of modern information technology and database technologies to effectively manage the various types of data generated during bio-environmental monitoring. This data typically includes data on organisms' responses to environmental pollutants, changes in biocommunity structure, and the accumulation of pollutants within organisms. Effective data management ensures data accuracy, integrity, and accessibility, providing a scientific basis for environmental protection and decision-making. However, due to the variability of the bio-environment and the non-standard nature of monitoring methods, not all bio-environmental monitoring data is valid. Existing technologies, using only conventional data cleaning methods, struggle to filter out useless data, making its use inconvenient. Summary of the Invention

[0003] The purpose of the present invention is to provide a biological environment monitoring data management system based on the Internet of Things to solve the problems raised in the above background technology.

[0004] This application provides a biological environment monitoring data management system based on the Internet of Things, which adopts the following technical solutions:

[0005] The data cleaning module collects biological environment monitoring data through the Internet of Things and records it as raw data. The raw data is cleaned to obtain primary data;

[0006] The demand screening module obtains the user's data requirements, deletes completely useless primary data according to the user's data requirements, and obtains intermediate data;

[0007] A truth judgment module extracts biological environment information based on the intermediate data, judges whether the intermediate data is true based on the biological environment information, and obtains a truth judgment result;

[0008] The traceability judgment module obtains the importance of medium-sized data according to user needs, and determines whether to perform data traceability based on the actual judgment results to obtain the traceability judgment results;

[0009] The data tracing module performs data tracing on the medium-level data according to the tracing judgment results, and then filters and deletes the data to obtain the high-level data;

[0010] The data storage module stores the advanced data to obtain a biological environment monitoring database.

[0011] Preferably, the steps of extracting biological environment information based on the intermediate data, judging whether the intermediate data is true based on the biological environment information, and obtaining a true judgment result are specifically:

[0012] Extract biological environmental information from medium-sized data and determine whether the biological environmental information violates biological laws;

[0013] If the biological environment information does not violate biological laws, then obtain the historical information database corresponding to the biological environment information;

[0014] Determine whether the biological environment information appears in the historical information database. If the biological environment information appears in the historical information database, find the corresponding historical information and record it as reference information.

[0015] Obtaining the similarity of the control information, and evaluating the authenticity of the data based on the similarity of the control information;

[0016] If the biological environment information does not appear in the historical information database, the occurrence probability of the biological environment information is obtained and the occurrence probability is used as the data authenticity;

[0017] If the biological environment information violates the biological laws, the medium data is judged to be erroneous data and the authenticity is 0.

[0018] Preferably, the step of obtaining the similarity of the control information and evaluating the authenticity of the data based on the similarity of the control information is specifically as follows:

[0019] Compare the similarities between all the control information, calculate the average of the similarities, and obtain the standard similarity;

[0020] The average similarity between the biological environment information and all control information was compared and recorded as information similarity;

[0021] Obtain the total number of occurrences of the control information and record it as the total amount of information;

[0022] Obtain the monitoring location corresponding to the biological environment information, find the ratio of the number of control information at the monitoring location to the total amount of information, and record it as the location occurrence frequency;

[0023] The data authenticity is obtained by combining standard similarity, information similarity, total amount of information and frequency of occurrence.

[0024] Preferably, the step of obtaining the importance of medium data according to user needs, determining whether to perform data tracing in combination with the actual judgment result, and obtaining the tracing judgment result is specifically as follows:

[0025] Obtain user needs and extract the user's research objectives based on user needs;

[0026] Determine whether the biological environment information is the biological environment information of the research target; if not, extract the monitoring target corresponding to the biological environment information;

[0027] Obtain the correlation between the monitoring target and the research target, and record it as the target correlation;

[0028] If the target is biological environment information, the target relevance is the maximum value, and the information weight ratio of the target relevance and data authenticity is set respectively. The data traceability is calculated based on the information weight ratio.

[0029] Determine whether the data authenticity of the medium data is 0. If the data authenticity is 0, do not perform data tracing and delete the medium data.

[0030] If the data authenticity of the medium data is not 0, then a authenticity threshold is set to determine whether the data authenticity of the medium data reaches the authenticity threshold;

[0031] If the authenticity threshold is reached, the medium data will not be traced and will be recorded as high-level data;

[0032] If the authenticity threshold is not reached, a data traceability threshold is set. Medium data whose data traceability reaches the data traceability threshold will be traced, and medium data that does not reach the data traceability threshold will not be traced and will be deleted.

[0033] Preferably, the step of obtaining the correlation between the monitoring target and the research target and recording it as the target correlation is specifically as follows:

[0034] Determine whether the monitoring target and the research target are in a predator-prey relationship. If so, distinguish the predator target and the prey target from the monitoring target and the research target;

[0035] Obtain prey records of prey targets, count the frequency of prey targets in prey records and record it as target relevance;

[0036] If it is not a predator-prey relationship, then determine whether the monitoring target and the research target have a symbiotic relationship. If so, obtain the symbiotic frequency as the target association degree.

[0037] If it is not a symbiotic relationship, determine whether there is a predator-prey chain between the monitoring target and the research target;

[0038] If there is a predator-prey chain, obtain the difference between the trophic level of the monitoring target and the trophic level of the research target;

[0039] The niche overlap between the monitoring target and the research target was obtained, and the target correlation was obtained by combining the grade difference.

[0040] Preferably, the step of tracing back the medium data according to the tracing judgment result, and filtering and deleting the high-level data after tracing back is specifically as follows:

[0041] Tracing back to the data source of the intermediate data, filtering the intermediate data according to the data source to obtain the first data;

[0042] Tracing back the storage method of the intermediate data, filtering the first data according to the storage method to obtain the second data;

[0043] Trace the usage of medium data and filter the second data according to the usage to obtain high-level data.

[0044] Preferably, the step of tracing the data source of the intermediate data and filtering the intermediate data according to the data source to obtain the first data is specifically as follows:

[0045] Tracing the sources of medium-level data, including monitoring equipment, monitoring personnel, monitoring locations, and monitoring time points;

[0046] Extract the data accuracy of the medium data and determine whether the monitoring equipment has achieved the data accuracy. If not, delete the medium data.

[0047] If the data accuracy is achieved, the personnel monitoring error rate of the monitoring personnel is obtained;

[0048] Obtain historical monitoring data of the monitoring location, calculate the monitoring error rate of the historical monitoring data and record it as the location monitoring error rate;

[0049] Obtain the biological information type of the medium data, calculate the monitoring error rate of the biological information type at the monitoring time point and record it as the time monitoring error rate;

[0050] The monitoring weight ratios of the personnel monitoring error rate, location monitoring error rate, and time monitoring error rate are set respectively, and the information monitoring error rate is calculated based on the monitoring weight ratios;

[0051] According to the information monitoring error rate, medium data that reaches a preset error rate threshold is filtered and deleted to obtain first data.

[0052] Preferably, the storage mode of the tracing medium data and the step of filtering the first data to obtain the second data according to the storage mode are specifically as follows:

[0053] Trace the storage method of medium-sized data and calculate the data loss rate of storage method;

[0054] Calculate the storage duration of medium-sized data and combine it with the data loss rate to obtain data security;

[0055] The second data is obtained by filtering the first data according to the data security level.

[0056] Preferably, the step of tracing back the usage of the intermediate data and filtering the second data according to the usage to obtain the high-level data is specifically as follows:

[0057] Tracking the usage of medium data, including the using team, using project and frequency of use;

[0058] Obtain the highest authority of the using team and calculate the average project size of the using projects;

[0059] Obtain data monitored at the same time, by the same team, and in the same batch as the medium data as a complete set of monitoring data;

[0060] Extract the usage frequency of the entire set of monitoring data based on the usage frequency and record it as the usage frequency of the entire set;

[0061] The ratio of the usage frequency of the medium data to the usage frequency of the whole set is extracted and recorded as the data usage frequency;

[0062] The data usage value is obtained by combining the highest authority, average project size value, whole set usage frequency and data usage frequency;

[0063] The second data is filtered according to the data usage value to obtain the higher-level data.

[0064] In summary, this application includes at least one of the following beneficial technical effects:

[0065] 1. The data cleaning module collects raw data through the Internet of Things and performs data cleaning to obtain primary data. The demand screening module deletes completely useless primary data based on user data requirements to obtain secondary data. The authenticity judgment module determines whether the secondary data is authentic based on the biological environment information, obtaining a true judgment result. The traceability judgment module determines whether to conduct data tracing based on the importance of the secondary data and the true judgment result, obtaining a traceability judgment result. The data tracing module traces the source, storage method, and usage of the secondary data, and filters and deletes it to obtain high-level data. Finally, the data storage module stores the high-level data to obtain a biological environment monitoring database. This effectively reduces the interference of untrue and inaccurate data on users and improves the accuracy of biological monitoring environmental data.

[0066] 2. The authenticity of the intermediate data is determined by determining whether the biological environmental information extracted from the intermediate data violates biological laws and its similarity to the corresponding historical information. The authenticity of the intermediate data is then determined based on the correlation between the monitoring and research objectives to determine whether the intermediate data should be traced. Intermediate data that is not traceable is deleted or treated as high-level data and then traced. This ensures the accuracy and authenticity of the data, reduces resource waste in the tracing process, and improves the environmental friendliness of IoT-based biological environmental monitoring data management.

[0067] 3. Generate different target correlations based on the predator-prey relationship, symbiotic relationship, and predator-prey food chain relationship between the monitoring target and the research target, which better fits the actual situation of the monitoring target and the research target, and obtains more accurate results, which is conducive to traceability selection and subsequent data processing. This reduces the interference of invalid data, provides convenience for users to use biological environment monitoring data, and improves the convenience of biological environment monitoring data management based on the Internet of Things. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 This is a module connection diagram of an embodiment of a biological environment monitoring data management system based on the Internet of Things of the present invention. DETAILED DESCRIPTION

[0069] Below is a combination of the embodiments and Figure 1 The present invention will be described in further detail, but the embodiments of the present invention are not limited thereto.

[0070] The present invention discloses a biological environment monitoring data management system based on the Internet of Things, which specifically includes:

[0071] The data cleaning module collects biological environment monitoring data through the Internet of Things and records it as raw data. After data cleaning, the raw data is obtained to obtain primary data.

[0072] Data cleaning is a crucial step in data processing. Its purpose is to identify and correct errors, duplications, outliers, and inconsistencies in a dataset to improve its quality and reliability. Data cleaning involves removing obvious duplications, errors, and other data containing fundamental issues, and performing simple data processing.

[0073] The demand screening module obtains the user's data requirements, deletes completely useless primary data according to the user's data requirements, and obtains intermediate data.

[0074] For example, if a user needs to study air pollution, then heavy metals and pesticide residues in the soil, which are not relevant to air pollution monitoring, will be deleted. If a user only wants to study air pollution in area A, the air data in area B, which is farther away, will be useless for the research and will also need to be deleted.

[0075] The authenticity judgment module extracts biological environment information based on the intermediate data, judges whether the intermediate data is authentic based on the biological environment information, and obtains a authenticity judgment result.

[0076] The traceability judgment module obtains the importance of medium data according to user needs, and determines whether to perform data traceability based on the actual judgment results to obtain the traceability judgment results.

[0077] The data tracing module performs data tracing on the medium-level data according to the tracing judgment results, and then filters and deletes the data to obtain the high-level data.

[0078] The data storage module stores the advanced data to obtain a biological environment monitoring database.

[0079] In practical applications, the Internet of Things (IoT) refers to a network system that connects various physical devices, sensors, software, and other technologies via the internet, enabling them to communicate and exchange data. While the IoT makes monitoring bio-environmental data more convenient, it also makes it difficult to verify the authenticity of data. The use of false data can severely impact projects and results. Therefore, validating and processing data to extract more useful and accurate data can improve the accuracy of bio-environmental monitoring data, reduce the data screening process for users, and increase the convenience of using bio-environmental monitoring data.

[0080] The steps of extracting biological environment information based on the secondary data, judging whether the secondary data is true based on the biological environment information, and obtaining a true judgment result are specifically as follows:

[0081] Extract biological environmental information from medium-sized data and determine whether the biological environmental information violates biological laws.

[0082] For example, the extracted biological environmental information is that a bird stays on a tree. This phenomenon is normal and does not violate biological laws.

[0083] If the biological environment information does not violate biological laws, then a historical information database corresponding to the biological environment information is obtained.

[0084] For example, if the biological environment information is that plant No. 1 is found in an arid area, then the corresponding historical information is whether plant No. 1 is found in an arid area.

[0085] It is determined whether the biological environment information appears in the historical information database. If the biological environment information appears in the historical information database, the corresponding historical information is found and recorded as the comparison information.

[0086] The similarity of the control information is obtained, and the authenticity of the data is obtained based on the similarity evaluation of the control information.

[0087] If the biological environment information does not appear in the historical information database, the occurrence probability of the biological environment information is obtained and the occurrence probability is used as the data authenticity.

[0088] If biological environmental information does not appear in the historical information database, it means that this situation does not exist yet. However, this possibility cannot be ruled out due to reasons such as gene mutation. Therefore, the probability of occurrence is used as the authenticity of the data. These probabilities can be inferred based on current scientific knowledge and observations.

[0089] If the biological environment information violates the biological laws, the medium data is judged to be erroneous data and the authenticity is 0.

[0090] In practice, some biological environmental information that violates biological laws is assumed to be due to data errors caused by equipment or monitoring personnel errors, and thus has a veracity of 0. Some biological environmental information has no historical record of discovery, but it may occur due to factors such as genetic mutations and biological purification. Therefore, the data veracity is determined based on the probability of occurrence. For example, if organism A is grown in a high-temperature area for a long time, it may gradually evolve to be heat-resistant, achieving a biological phenomenon that can also occur in high-temperature areas. However, some biological environmental information has historical records, indicating that it is a normal biological phenomenon that can be observed. Therefore, the veracity is determined based on control information. For example, there is a probability that four-leaf clover will appear among three-leaf clover. Therefore, the discovery of four-leaf clover is also a normal biological phenomenon, and the data authenticity needs to be determined.

[0091] The steps of obtaining the similarity of the control information and evaluating the authenticity of the data based on the similarity of the control information are as follows:

[0092] Compare the similarities between all the control information, calculate the average of the similarities, and obtain the standard similarity.

[0093] Comparing the average similarity between all the control information can reflect whether the biological environment information is variable. For example, if the biological environment information shows that plant No. 1 was found at 25 degrees Celsius, and the control information is the temperature at which plant No. 1 was found, if the average similarity is low, it means that plant No. 1 is more adaptable to temperature and can survive in both low and high temperatures.

[0094] The average similarity between the biological environment information and all control information is compared and recorded as information similarity.

[0095] Comparing the average similarity between the biological environment information and the control information can reflect whether the extracted real-time biological environment information is similar to the historical information.

[0096] The total number of occurrences of the control information is obtained and recorded as the total amount of information.

[0097] Obtain the monitoring location corresponding to the biological environment information, find the ratio of the number of control information at the monitoring location to the total amount of information, and record it as the location occurrence frequency.

[0098] The data authenticity is obtained by combining standard similarity, information similarity, total amount of information and frequency of occurrence.

[0099] The weight ratios of standard similarity, information similarity, total information amount and frequency of occurrence are set respectively, and the data authenticity is calculated based on the weight ratios.

[0100] In practice, higher standard similarity indicates more variable biological environmental information, a greater likelihood of obtaining intermediate data, and a higher degree of authenticity for intermediate data. Higher information similarity also reflects a greater likelihood of intermediate data existing and a higher degree of authenticity. A greater amount of information and a greater frequency of occurrence both indicate a greater likelihood of the biological information occurring and, therefore, a higher degree of authenticity. For example, if sika deer have been photographed numerous times in Location A, the biological environmental information indicating their presence is highly likely to be true. However, if sika deer have not been photographed in Location B, the information indicating their presence is highly likely to be false. Similarly, penguins primarily inhabit cold regions. Therefore, if the biological environmental information indicates penguins are found in hot regions, the information similarity is low, indicating a low degree of authenticity.

[0101] The importance of medium-level data is determined based on user needs, and combined with the actual judgment results, it is determined whether to perform data tracing. The specific steps for obtaining the tracing judgment results are as follows:

[0102] Obtain user needs and extract user research objectives based on user needs.

[0103] Biological environment monitoring mainly monitors the living environment of organisms on Earth, as well as their distribution and activities. Because the biological environment is relatively large, users have their own research goals, such as sika deer, air pollution and other specific research goals.

[0104] Determine whether the biological environment information is the biological environment information of the research target. If it is not the biological environment information of the research target, extract the monitoring target corresponding to the biological environment information.

[0105] Sometimes, during research, additional data is needed to assist with the research process. Therefore, multi-target monitoring is performed, where targets outside the research objective are recorded as monitoring targets. For example, if the research target is tigers, but tigers eat prey such as rabbits, then rabbits and other prey also need to be monitored. These are monitoring targets.

[0106] Obtain the correlation between the monitoring target and the research target and record it as the target correlation.

[0107] If it is the biological environment information of the research target, the target relevance is the maximum value, and the information weight ratio of the target relevance and the data authenticity is set respectively, and the data traceability is calculated according to the information weight ratio.

[0108] If the biological environment information is the information of the research target itself, then this information is definitely needed, so the target relevance of this information is considered to be the maximum.

[0109] Determine whether the data authenticity of the medium data is 0. If the data authenticity is 0, do not perform data tracing and delete the medium data.

[0110] If the data authenticity is 0, it means that the data is wrong data or false data, and no further confirmation and verification is required, so it can be deleted directly.

[0111] If the data authenticity of the medium data is not 0, a authenticity threshold is set to determine whether the data authenticity of the medium data reaches the authenticity threshold.

[0112] If the authenticity threshold is reached, the medium data will not be traced and will be recorded as high-level data.

[0113] If the data authenticity reaches the authenticity threshold, it means that the data credibility is high, so there is no need for further verification and confirmation, and it can be directly recorded as high-quality data.

[0114] If the authenticity threshold is not reached, a data traceability threshold is set. Medium data whose data traceability reaches the data traceability threshold will be traced, and medium data that does not reach the data traceability threshold will not be traced and will be deleted.

[0115] In practice, if the authenticity threshold is not met, data traceability is used to determine whether traceability is necessary. Some data may have low authenticity and low relevance to the target, making them of little use in research and therefore worthy of discarding. Therefore, there's no need to waste time on traceability and data can be directly deleted. On the other hand, some data may have low authenticity but high relevance to the target, making them highly useful to users. Therefore, traceability verification is performed to confirm the data's authenticity and usability. If so, it's not deleted.

[0116] The steps to obtain the correlation between the monitoring target and the research target and record it as the target correlation are as follows:

[0117] Determine whether the monitoring target and the research target are in a predator-prey relationship. If so, distinguish the predator target and the prey target from the monitoring target and the research target.

[0118] If the monitoring target and the research target are in a predator-prey relationship, it is necessary to distinguish who preys on whom. For example, if the monitoring target and the research target are rabbits and tigers, the predator is the tiger, that is, the predator target, and the rabbit is the prey target.

[0119] Obtain the prey records of the prey target, count the frequency of the prey target in the prey records and record it as the target relevance.

[0120] In a predator-prey relationship, since prey targets typically have diverse prey, the more prey targets are preyed upon, the higher the target association. For example, a tiger preys on many species, such as deer, sheep, and wild boar. The probability of a tiger preying on a wild boar is 40%, the probability of a deer is 30%, and the probability of a rabbit is 5%. If the prey target is rabbit, the tiger does not primarily prey on rabbits, so the association with the tiger is not as high as with wild boars and deer.

[0121] If it is not a predator-prey relationship, then determine whether the monitoring target and the research target have a symbiotic relationship. If it is a symbiotic relationship, obtain the symbiotic frequency as the target association degree.

[0122] Some organisms form symbiotic relationships, which also influence research targets. For example, cattle egrets form symbiotic relationships with a variety of large herbivores, such as cattle, horses, and sheep. They hang on or near the backs of these animals to obtain food, such as insects. Therefore, these large herbivores can affect cattle egrets, and the probability of cattle egrets coexisting with these animals varies. Cattle egrets are more likely to coexist with cattle, so cattle are more relevant to the target.

[0123] If it is not a symbiotic relationship, determine whether there is a predator-prey chain between the monitoring target and the research target.

[0124] If there is a predator-prey chain, the difference between the trophic level of the monitoring target and the trophic level of the research target is obtained.

[0125] If there is a predator-prey chain between the targets, the targets and the organisms are also related. For example, if a cow eats grass and a tiger eats the cow, then the grass will also affect the tiger's survival, so the grass data will also play a role. The greater the difference in trophic level, the smaller the correlation.

[0126] The niche overlap between the monitoring target and the research target was obtained, and the target correlation was obtained by combining the grade difference.

[0127] Set the proportional coefficients for niche overlap and grade difference respectively, and calculate the target association based on the proportional coefficients. If there is no predator-prey chain, the niche overlap is used as the target association.

[0128] In practice, niche overlap refers to the phenomenon in which two or more species with similar ecological niches share or compete for common resources when living in the same space. Niche overlap occurs when two species utilize the same resource or share a resource element (such as food, nutrients, or space). In this case, a portion of the space is shared between the two niches. If two species have exactly the same niche, this is considered complete overlap. However, in most cases, niche overlap only occurs partially, meaning that some resources are shared while others are occupied by each species.

[0129] The degree of niche overlap is closely related to competition between organisms. The greater the degree of niche overlap, the more intense the competition between two species is likely to be, as they compete for the same resources. Competing organisms can also influence monitoring data, playing a significant role in research, leading to a high degree of target correlation.

[0130] According to the results of the retrospective judgment, the medium-level data is traced back, and the steps of filtering and deleting to obtain the high-level data are as follows:

[0131] Trace the data source of the intermediate data, and filter the intermediate data according to the data source to obtain the first data.

[0132] The source of data can reflect the authenticity of the data. Tracing the source of the intermediate data that needs to be traced can screen out inaccurate and untrue data, thereby improving the accuracy of the data.

[0133] The storage mode of the intermediate data is traced back, and the first data is filtered according to the storage mode to obtain the second data.

[0134] When the data storage environment is unsafe and unreliable, data will be illegally accessed or tampered with, thus affecting the authenticity of the data. If the data storage method is indeed unsafe, then for some medium-level data with low authenticity, tracing can be used to verify whether the data is safe and reliable.

[0135] Trace the usage of medium data and filter the second data according to the usage to obtain high-level data.

[0136] The usage of data can also reflect the authenticity of the data. The authenticity of the data can also be obtained by analyzing the usage of medium data, thereby screening the data.

[0137] In actual application, data tracing for some relatively important data with low authenticity can confirm the status of the verified data, thereby deleting some untrue and unreliable data, improving the accuracy of biological environment monitoring data, and providing convenience for users.

[0138] The steps of tracing the data source of the intermediate data and filtering the intermediate data according to the data source to obtain the first data are as follows:

[0139] Trace the data sources of medium-level data, including monitoring equipment, monitoring personnel, monitoring locations and monitoring time points.

[0140] The data accuracy of the medium data is extracted to determine whether the monitoring device has achieved the data accuracy. If the data accuracy has not been achieved, the medium data is deleted.

[0141] If the monitoring device does not achieve the required data accuracy, it indicates that the data is incorrect or has been tampered with, and the data is false. For example, the data monitored by the device is only accurate to two decimal places, but the actual data is accurate to three decimal places. This means that the data is not the data monitored by the monitoring device, or it has been tampered with later and is not the original data. Delete the data.

[0142] If the data accuracy is achieved, the personnel monitoring error rate of the monitoring personnel is obtained.

[0143] Different monitoring personnel have different monitoring levels and experiences, so the error rates of monitoring data are also different.

[0144] Obtain historical monitoring data of the monitoring location, calculate the monitoring error rate of the historical monitoring data and record it as the location monitoring error rate.

[0145] The accuracy of some biological environmental data can vary due to environmental changes, and the error rate can also vary when monitoring in different locations and environments. For example, the accuracy of monitoring equipment can be affected by the environment, resulting in different error rates.

[0146] Obtain the biological information type of the medium data, calculate the monitoring error rate of the biological information type at the monitoring time point, and record it as the time monitoring error rate.

[0147] The monitoring error rate of some biological environmental data is also affected by time. For example, biological environmental data such as images are prone to monitoring errors during nighttime monitoring due to lighting issues, which increases the monitoring error rate.

[0148] The monitoring weight ratios of the personnel monitoring error rate, location monitoring error rate and time monitoring error rate are set respectively, and the information monitoring error rate is calculated according to the monitoring weight ratios.

[0149] According to the information monitoring error rate, medium data that reaches a preset error rate threshold is filtered and deleted to obtain first data.

[0150] In actual applications, if some data with low authenticity is prone to monitoring errors during the monitoring process, the authenticity and reliability of the data will be even lower. Therefore, data with high monitoring error rates can be deleted to reduce the impact of inaccurate data.

[0151] Tracing back the storage method of the intermediate data, the steps of filtering the first data according to the storage method to obtain the second data are as follows:

[0152] Trace the storage method of medium-sized data and calculate the data loss rate of the storage method.

[0153] The storage duration of medium-sized data is calculated and combined with the data loss rate to obtain the data security.

[0154] Set the storage duration and data loss rate proportional coefficients respectively, and calculate the data security based on the proportional coefficients. The longer the storage duration, the greater the possibility of data loss, and therefore the lower the data security.

[0155] The second data is obtained by filtering the first data according to the data security level.

[0156] In practice, appropriate storage methods can ensure the integrity of bio-environmental monitoring data, meaning that data is protected from corruption or loss during storage. Improper storage methods, such as storage device failure or insufficient data backup, can lead to data loss or corruption, compromising data authenticity. Tracing the storage method of bio-environmental monitoring data can reflect its security, allowing for assessment of its reliability and authenticity and data screening. By setting a data security threshold, moderate data that does not meet the threshold is deleted to generate secondary data.

[0157] The steps for tracing the usage of medium data and filtering the second data to obtain high-level data based on the usage are as follows:

[0158] Track the usage of medium data, including the using team, using project and frequency of use.

[0159] Obtain the highest authority level of the using team and calculate the average project size of the using projects.

[0160] The degree of authority can be determined based on the beneficial effects of the user team and the number of people who follow it, while the project scale can be determined based on the number of project participants and the cost.

[0161] The data monitored at the same time, by the same team, and in the same batch as the medium data are obtained as a complete set of monitoring data.

[0162] The usage frequency of the entire set of monitoring data is extracted based on the usage frequency and recorded as the usage frequency of the entire set.

[0163] The ratio of the usage frequency of the medium data to the usage frequency of the whole set is extracted and recorded as the data usage frequency.

[0164] The data usage value is obtained by combining the highest authority level, average project size value, whole set usage frequency and data usage frequency.

[0165] The second data is filtered according to the data usage value to obtain the higher-level data.

[0166] Set the highest authority, average project size, overall usage frequency, and data usage frequency proportional coefficients, and calculate the data usage value based on the proportional coefficients. Set a data usage value threshold and delete the second data that does not reach the data usage value threshold to obtain high-quality data.

[0167] In practice, authoritative teams, due to the authority of the information they publish, will be more meticulous in verifying data. Therefore, the higher the authority of the using team, the lower the error rate, indicating that the data is more authentic and reliable. The larger the scale of the project, the greater the loss caused by using erroneous data. Therefore, data will be more rigorously reviewed to ensure its authenticity and accuracy. Therefore, the larger the scale, the higher the data accuracy. Furthermore, the more frequently a set of data from the same batch is used, the more widely accepted it is, and the more authentic and reliable it is. Furthermore, the more frequently a single data point within the set is used, the higher the trust users have in that data. For example, if temperature, humidity, and atmospheric data are monitored simultaneously, but most users do not use the temperature data, this suggests that the temperature data may be flawed and inaccurate, while humidity and atmospheric data are used more frequently and are more authentic and reliable.

[0168] The above are all preferred embodiments of the present application, and are not intended to limit the scope of protection of the present application. Therefore, any equivalent changes made based on the structure, shape, and principle of the present application should be included in the scope of protection of the present application.

Claims

1. A biological environment monitoring data management system based on the Internet of Things, characterized in that: include: The data cleaning module collects biological environment monitoring data through the Internet of Things and records it as raw data. The raw data is cleaned to obtain primary data; The demand screening module obtains the user's data requirements, deletes completely useless primary data according to the user's data requirements, and obtains intermediate data; A truth judgment module extracts biological environment information based on the intermediate data, judges whether the intermediate data is true based on the biological environment information, and obtains a truth judgment result; The traceability judgment module obtains the importance of medium-sized data according to user needs, and determines whether to perform data traceability based on the actual judgment results to obtain the traceability judgment results; The data tracing module performs data tracing on the medium-level data according to the tracing judgment results, and then filters and deletes the data to obtain the high-level data; The data storage module stores the advanced data to obtain a biological environment monitoring database.

2. The biological environment monitoring data management system based on the Internet of Things according to claim 1, characterized in that: The steps of extracting biological environment information based on the intermediate data, judging whether the intermediate data is true based on the biological environment information, and obtaining a true judgment result are specifically as follows: Extract biological environmental information from medium-sized data and determine whether the biological environmental information violates biological laws; If the biological environment information does not violate biological laws, then obtain the historical information database corresponding to the biological environment information; Determine whether the biological environment information appears in the historical information database. If the biological environment information appears in the historical information database, find the corresponding historical information and record it as reference information. Obtaining the similarity of the control information, and evaluating the authenticity of the data based on the similarity of the control information; If the biological environment information does not appear in the historical information database, the occurrence probability of the biological environment information is obtained and the occurrence probability is used as the data authenticity; If the biological environment information violates the biological laws, the medium data is judged to be erroneous data and the authenticity is 0.

3. The biological environment monitoring data management system based on the Internet of Things according to claim 2, characterized in that: The steps of obtaining the similarity of the control information and evaluating the authenticity of the data based on the similarity of the control information are specifically as follows: Compare the similarities between all the control information, calculate the average of the similarities, and obtain the standard similarity; The average similarity between the biological environment information and all control information was compared and recorded as information similarity; Obtain the total number of occurrences of the control information and record it as the total amount of information; Obtain the monitoring location corresponding to the biological environment information, find the ratio of the number of control information at the monitoring location to the total amount of information, and record it as the location occurrence frequency; The data authenticity is obtained by combining standard similarity, information similarity, total amount of information and frequency of occurrence.

4. The biological environment monitoring data management system based on the Internet of Things according to claim 3 is characterized in that: The steps of obtaining the importance of medium data according to user needs, determining whether to perform data tracing based on the actual judgment result, and obtaining the tracing judgment result are specifically as follows: Obtain user needs and extract the user's research objectives based on user needs; Determine whether the biological environment information is the biological environment information of the research target; if not, extract the monitoring target corresponding to the biological environment information; Obtain the correlation between the monitoring target and the research target, and record it as the target correlation; If the target is biological environment information, the target relevance is the maximum value, and the information weight ratio of the target relevance and data authenticity is set respectively. The data traceability is calculated based on the information weight ratio. Determine whether the data authenticity of the medium data is 0. If the data authenticity is 0, do not perform data tracing and delete the medium data. If the data authenticity of the medium data is not 0, then a authenticity threshold is set to determine whether the data authenticity of the medium data reaches the authenticity threshold; If the authenticity threshold is reached, the medium data will not be traced and will be recorded as high-level data; If the authenticity threshold is not reached, a data traceability threshold is set. Medium data whose data traceability reaches the data traceability threshold will be traced, and medium data that does not reach the data traceability threshold will not be traced and will be deleted.

5. The biological environment monitoring data management system based on the Internet of Things according to claim 4 is characterized in that: The steps of obtaining the correlation between the monitoring target and the research target and recording it as the target correlation are specifically as follows: Determine whether the monitoring target and the research target are in a predator-prey relationship. If so, distinguish the predator target and the prey target from the monitoring target and the research target; Obtain prey records of prey targets, count the frequency of prey targets in prey records and record it as target relevance; If it is not a predator-prey relationship, then determine whether the monitoring target and the research target have a symbiotic relationship. If so, obtain the symbiotic frequency as the target association degree. If it is not a symbiotic relationship, determine whether there is a predator-prey chain between the monitoring target and the research target; If there is a predator-prey chain, obtain the difference between the trophic level of the monitoring target and the trophic level of the research target; The niche overlap between the monitoring target and the research target was obtained, and the target correlation was obtained by combining the grade difference.

6. The biological environment monitoring data management system based on the Internet of Things according to claim 5, characterized in that: The steps of tracing back the medium data according to the tracing judgment result and filtering and deleting the high-level data after tracing back are specifically as follows: Tracing back to the data source of the intermediate data, filtering the intermediate data according to the data source to obtain the first data; Tracing back the storage method of the intermediate data, filtering the first data according to the storage method to obtain the second data; Trace the usage of medium data and filter the second data according to the usage to obtain high-level data.

7. The biological environment monitoring data management system based on the Internet of Things according to claim 6, characterized in that: The step of tracing the data source of the intermediate data and filtering the intermediate data according to the data source to obtain the first data is specifically as follows: Tracing the sources of medium-level data, including monitoring equipment, monitoring personnel, monitoring locations, and monitoring time points; Extract the data accuracy of the medium data and determine whether the monitoring equipment has achieved the data accuracy. If not, delete the medium data. If the data accuracy is achieved, the personnel monitoring error rate of the monitoring personnel is obtained; Obtain historical monitoring data of the monitoring location, calculate the monitoring error rate of the historical monitoring data and record it as the location monitoring error rate; Obtain the biological information type of the medium data, calculate the monitoring error rate of the biological information type at the monitoring time point and record it as the time monitoring error rate; The monitoring weight ratios of the personnel monitoring error rate, location monitoring error rate, and time monitoring error rate are set respectively, and the information monitoring error rate is calculated based on the monitoring weight ratios; According to the information monitoring error rate, medium data that reaches a preset error rate threshold is filtered and deleted to obtain first data.

8. The biological environment monitoring data management system based on the Internet of Things according to claim 7, characterized in that: The step of tracing back the storage mode of the medium data and filtering the first data to obtain the second data according to the storage mode is specifically as follows: Trace the storage method of medium-sized data and calculate the data loss rate of storage method; Calculate the storage duration of medium-sized data and combine it with the data loss rate to obtain data security; The second data is obtained by filtering the first data according to the data security level.

9. The biological environment monitoring data management system based on the Internet of Things according to claim 8, characterized in that: The step of tracing back the usage of the intermediate data and filtering the second data according to the usage to obtain the high-level data is specifically as follows: Tracking the usage of medium data, including the using team, using project and frequency of use; Obtain the highest authority of the using team and calculate the average project size of the using projects; Obtain data monitored at the same time, by the same team, and in the same batch as the medium data as a complete set of monitoring data; Extract the usage frequency of the entire set of monitoring data based on the usage frequency and record it as the usage frequency of the entire set; The ratio of the usage frequency of the medium data to the usage frequency of the whole set is extracted and recorded as the data usage frequency; The data usage value is obtained by combining the highest authority, average project size value, whole set usage frequency and data usage frequency; The second data is filtered according to the data usage value to obtain the higher-level data.

Citation Information

Patent Citations

  • Rapid validation method for internet of things big data

    CN106776103A

  • Customer classification method and system based on data mining

    CN107729377A