A cryosphere multi-source heterogeneous data semantic fusion method and system

By constructing environment and state vectors of multi-source heterogeneous data in the cryosphere, determining dynamic weights, identifying and correcting abnormal data, and using machine learning algorithms for fusion, the problem of the quality difference of multi-source heterogeneous data in the cryosphere affecting the fusion accuracy was solved, and accurate identification and long-term consistency of the cryosphere state were achieved.

CN120744836BActive Publication Date: 2025-12-23NORTHWEST INST OF ECO ENVIRONMENT & RESOURCES CAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511141589.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-12-23
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

The quality of multi-source heterogeneous cryosphere data from different sources varies greatly, and there is noise or error, which affects the accuracy of the fusion process and interferes with the understanding and judgment of the cryosphere state.

Method used

By collecting textual and environmental data from various monitoring points in the cryosphere, environmental and state vectors are constructed to determine the reliability and dynamic weight of the monitoring points. Data fusion is performed based on location distance and correlation to identify and correct abnormal data. Semantic fusion is then carried out using machine learning algorithms.

Benefits of technology

It enhances the semantic alignment capability of multimodal data, suppresses noise cascading amplification, accurately identifies functionally similar regions, improves the accuracy and reliability of data fusion, ensures the consistency of long-term sequences, and improves the accuracy of cryosphere state identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744836B_ABST
    Figure CN120744836B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of data processing, in particular to a cryosphere multi-source heterogeneous data semantic fusion method and system, which comprises the following steps: collecting various types of monitoring data of each monitoring point of the cryosphere, including text data and environmental data; constructing an environmental vector and a state vector of each to-be-tested word group; determining the same state possibility between any two monitoring points; acquiring the abnormality degree of any type of environmental data of any monitoring point, and determining the abnormal environmental data of each monitoring point collected each time; acquiring the semantic deviation of each to-be-tested word group of each monitoring point, and determining the abnormal word group of each monitoring point collected each time; acquiring the abnormal state vector in all state vectors of each monitoring point collected each time, and correcting the abnormal state vector; and using the corrected state vector of each monitoring point, combining a machine learning algorithm, and performing data fusion on the environmental data and the text data. The accuracy of cryosphere multi-source heterogeneous data semantic fusion is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, specifically to a method and system for semantic fusion of multi-source heterogeneous data in the cryosphere. Background Technology

[0002] The cryosphere refers to the regions of Earth covered by ice and snow, including polar ice caps, glaciers, permafrost, snow, sea ice, and ocean ice. These regions represent the solid form of water on Earth and play vital climatic and environmental functions. The cryosphere plays a crucial role in Earth's climate system, helping to regulate global climate.

[0003] Semantic fusion of multi-source heterogeneous data from the cryosphere can effectively combine data from different sources and formats, providing more comprehensive and accurate information for studying changes in the cryosphere. Furthermore, fusing multi-source data helps analyze the dynamic changes of the cryosphere from multiple dimensions, such as ice thickness, coverage, and physical properties, which is crucial for understanding the relationship between the cryosphere and the global climate system.

[0004] However, the quality of data from different sources varies greatly and may contain noise or errors, especially in long-term series data. These anomalous data can easily affect the accuracy of the final results during the fusion process, thereby interfering with the understanding and judgment of the state of the cryosphere. Summary of the Invention

[0005] To address the aforementioned technical problems, the purpose of this application is to provide a method and system for semantic fusion of multi-source heterogeneous data in the cryosphere. The specific technical solution adopted is as follows:

[0006] In a first aspect, embodiments of this application provide a semantic fusion method for multi-source heterogeneous data from the cryosphere, the method comprising the following steps:

[0007] Collect various monitoring data from various monitoring points in the cryosphere, including text data and environmental data;

[0008] Obtain each word group to be tested from the text data; combine all class environment data collected from each monitoring point each time to form an environment vector, and combine it with the word vector of each word group to be tested to obtain the state vector of each word group to be tested;

[0009] The correlation between the historical environmental vectors of each monitoring point and the environmental vectors of its neighboring monitoring points at the same time is analyzed to determine the reliability of the historical environmental vectors of each monitoring point; the dynamic weights of the environmental vectors of the current monitoring points are determined by combining the correlation between the current environmental vectors of each monitoring point and their historical environmental vectors, as well as the reliability; and the dynamic weights of each word group to be tested at the current monitoring points are obtained accordingly.

[0010] weighting state vectors of each to-be-tested phrase group by using all dynamic weights of current monitoring points; determining the same state likelihood between any two monitoring points based on correlation of weighted state vectors of the any two monitoring points and distance between the any two monitoring points;

[0011] classifying all monitoring points into regions by using the same state likelihoods corresponding to all monitoring points, and marking monitoring points in each region in each collection; determining difference of any type of environmental data of any monitoring point from remaining types of environmental data, analyzing deviation degree of the difference of the any monitoring point from the difference of remaining monitoring points in a region to which the any monitoring point belongs, and obtaining abnormal degree of the any type of environmental data of the any monitoring point in combination with a number of times that the any monitoring point is marked, to determine abnormal environmental data of each monitoring point in each collection;

[0012] analyzing correlation of weighted state vectors of each to-be-tested phrase group of each monitoring point and state vectors of remaining monitoring points in a region to which each monitoring point belongs, and abnormal degree of environmental data in the weighted state vectors of each to-be-tested phrase group, to obtain semantic deviation of each to-be-tested phrase group of each monitoring point, and determine abnormal phrase group of each monitoring point in each collection;

[0013] obtaining abnormal state vectors in all state vectors of each monitoring point in each collection in combination with the abnormal environmental data and the abnormal phrase group, and correcting the abnormal state vectors, to use corrected state vectors of each monitoring point and machine learning algorithm to perform data fusion on environmental data and text data.

[0014] In one embodiment, the determining of the credibility of the historical environmental vector of each monitoring point comprises:

[0015] calculating mean value of similarity of the historical environmental vector of each monitoring point and environmental vectors of all adjacent monitoring points at the same time, and determining range of similarity of the historical environmental vector of each monitoring point and environmental vectors of all adjacent monitoring points at the same time.

[0016] calculating ratio of the mean value and the range, and the credibility of the historical environmental vector of each monitoring point is positively correlated with the ratio.

[0017] In one embodiment, the determining of the dynamic weight of the environmental vector of each monitoring point comprises:

[0018] calculating similarity of the environmental vector of each monitoring point and the historical environmental vector of each monitoring point, denoted as first similarity, and the dynamic weight of the environmental vector of each monitoring point is a product of the first similarity and the credibility.

[0019] In one embodiment, the acquiring of the dynamic weight of each to-be-tested phrase of each monitoring point comprises:

[0020] The similarity between any to-be-tested phrase of each monitoring point and the phrase vector of all to-be-tested phrases of the historical environment vector at the same time is calculated, denoted as a second similarity, and the product of the maximum value in all the second similarities and the trust degree is taken as the dynamic weight of the any to-be-tested phrase of each monitoring point.

[0021] In one embodiment, the determination of the same-state possibility comprises:

[0022] The similarity between the weighted state vector of monitoring point i and the weighted state vector of monitoring point j is calculated, denoted as a third similarity, and the two state vectors corresponding to the maximum value in all the third similarities are taken as the matching state vectors of monitoring point i and monitoring point j.

[0023] The average value of the similarity between all the weighted state vectors of monitoring point i and the matching state vector of monitoring point i is calculated, denoted as a first average value, the metric distance between monitoring point i and monitoring point j is determined, and the ratio of the first average value and the normalized value of the metric distance is taken as the same-state possibility between monitoring point i and monitoring point j.

[0024] In one embodiment, the classification of all the monitoring points into regions and the marking of the monitoring points in each classified region at each collection comprise:

[0025] The two monitoring points with the normalized value of the same-state possibility greater than a first preset threshold are merged into the same region, and the monitoring points in each region are marked with the number of the region to which the monitoring points belong.

[0026] In one embodiment, the determination of the abnormal environment data of each monitoring point at each collection comprises:

[0027] Each type of environment data of each monitoring point is combined to form each time-series environment sequence, the difference between the zth time-series environment sequence and the (z+1)th time-series environment sequence of each monitoring point is calculated, denoted as a first difference, the average value of the first difference of all the remaining monitoring points in the region of each monitoring point is calculated, denoted as a second average value.

[0028] The difference between the first difference and the second average value is calculated, denoted as a second difference, the cumulative sum of the second difference between the zth time-series environment sequence and all the remaining time-series environment sequences of each monitoring point is calculated, and the ratio of the cumulative sum and the number of times each monitoring point is marked is taken as the abnormal degree of the corresponding environment data of the zth time-series environment sequence of each monitoring point.​​​​​

[0029] The environment data with a normalized value of the abnormality degree greater than a second preset threshold is taken as abnormal environment data.

[0030] In one embodiment, the determination of the abnormal phrase for each monitoring point in each collection includes:

[0031] The similarity between the weighted state vector of the xth to-be-tested phrase of the monitoring point i and all state vectors of the remaining monitoring points in the region to which the monitoring point i belongs is calculated, denoted as a fourth similarity, and the state vector with a similarity greater than a third preset threshold among all fourth similarities corresponding to the xth to-be-tested phrase of the monitoring point i is taken as a target vector of the xth to-be-tested phrase of the monitoring point i.

[0032] The average of the similarities between the weighted state vector of the xth to-be-tested phrase of the monitoring point i and all target vectors thereof is calculated, denoted as a third average, and the ratio between the average of the abnormality degrees of all environment data in the weighted state vector of the xth to-be-tested phrase of the monitoring point i and the third average is taken as the semantic deviation of the xth to-be-tested phrase of the monitoring point i.

[0033] The to-be-tested phrase with a normalized value of the semantic deviation greater than a fourth preset threshold is taken as an abnormal phrase.

[0034] In one embodiment, the data fusion of the environment data and the text data includes:

[0035] For the weighted state vectors of the current monitoring points, if there is abnormal environment data or an abnormal phrase, the corresponding state vector is marked as an abnormal state vector, and the state vectors of the monitoring points other than the abnormal state vectors are taken as normal state vectors.

[0036] The similarity between any abnormal state vector of the monitoring points and all normal state vectors of the remaining monitoring points in the region to which the monitoring points belong is calculated, denoted as a fifth similarity, and the normal state vector with a fifth similarity greater than a fifth preset threshold is taken as a correction reference vector of the any abnormal state vector.

[0037] The abnormal data position in the any abnormal state vector is obtained, the data average of all correction reference vectors of the any abnormal state vector that is the same as the abnormal data position is calculated, and the data average and the data at the abnormal data position are replaced to obtain a corrected state vector of the any abnormal state vector.

[0038] All normal state vectors and corrected abnormal state vectors of the monitoring points are taken as inputs of a random forest algorithm, and the state category of the cryosphere of each monitoring point is output.

[0039] In a second aspect, the embodiments of the present application also provide a cryosphere multi-source heterogeneous data semantic fusion system, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor implements the steps of the method according to any one of the preceding aspects when executing the computer program.

[0040] The present application has at least the following beneficial effects:

[0041] The present application enhances the semantic alignment capability of multi-modal data by mapping the text data and environmental data of the monitoring points into a unified state vector, captures the semantics of cryosphere professional terms through word vectors, and quantifies the physical state through environmental vectors, which are complementary to each other and improve the integrity of the feature expression of each monitoring point. Further, by determining the dynamic weight of the environmental vector of each monitoring point and the dynamic weight of each word group to be tested of each monitoring point, the dynamic distribution of data value is strengthened, and the cascade amplification of noise is inhibited. Based on the same state possibility corresponding to all monitoring points, all monitoring points are classified and divided into regions, which breaks through the limitation of geographical distance, accurately identifies the similar regions, and avoids distortion of the partition. Through abnormal detection on the environmental data and the word group to be tested in the state vector, the accurate positioning of the abnormal source is realized, and the accuracy and reliability of the subsequent data semantic fusion are improved. Further, through targeted correction of the abnormal state vector, the scientificity of data repair is improved, the consistency of long-term sequence is ensured, and the credibility of multi-source heterogeneous data is improved. Finally, the machine learning algorithm is used to realize semantic fusion, and the accuracy of cryosphere state recognition is improved. BRIEF DESCRIPTION OF DRAWINGS

[0042] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the drawings needed in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0043] Figure 1 A step flowchart of a cryosphere multi-source heterogeneous data semantic fusion method provided by an embodiment of the present application is shown in the figure.

[0044] Figure 2 An abnormal state vector correction flowchart is shown in the figure. DETAILED DESCRIPTION

[0045] In order to further illustrate the technical means and effects taken by the present application to achieve the predetermined object, the specific implementation, structure, features and effects of the cryosphere multi-source heterogeneous data semantic fusion method and system according to the present application are described in detail as follows in combination with the drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0047] The specific scheme of the cryosphere multi-source heterogeneous data semantic fusion method and system provided by the present application is described in detail below in combination with the drawings.

[0048] Please refer to Figure 1 which shows the step flowchart of the cryosphere multi-source heterogeneous data semantic fusion method provided by one embodiment of the present application, and the method comprises the following steps:

[0049] S1, collecting various types of monitoring data of each monitoring point of the cryosphere, including text data and environmental data.

[0050] This embodiment takes the cryosphere of the Qinghai-Tibet Plateau as an example to monitor the cryosphere and obtain its multi-source monitoring data, specifically: uniformly setting each monitoring point within the monitoring range of the cryosphere, installing a sensor at each monitoring point, and collecting various types of environmental data of each monitoring point. In this embodiment, the environmental data includes temperature, humidity, wind speed, snow depth, and snow water equivalent. The implementer can determine the types of monitored environmental data according to the actual situation, and this embodiment does not limit it.

[0051] In addition, by investigating the permafrost engineering disease records, glacial change analysis, and extreme weather description of each monitoring point, the text description information of each monitoring point is obtained as text data. The implementer can select the source of the text data. In this embodiment, a monitoring point is set every 200 kilometers, and the implementer can set it according to the actual situation, and this embodiment does not limit it.

[0052] S2, obtaining each to-be-tested word group in the text data; combining the environmental vector composed of all types of environmental data collected by each monitoring point each time with the word vector of each to-be-tested word group, to obtain the state vector of each to-be-tested word group.

[0053] The text data of each monitoring point is processed by using the jieba word segmentation tool to obtain each word group, and the word vectors of each word group are obtained by using the word2vec technology. In this embodiment, all the entries in the cryosphere dictionary are composed of the cryosphere vocabulary, the similarity between the word vectors of all word groups of each monitoring point and the word vectors of each entry in the cryosphere vocabulary is calculated, and the word groups with a similarity greater than a preset threshold are used as the to-be-tested word groups of each monitoring point.

[0054] It should be noted that the similarity in this embodiment is calculated by using the cosine similarity, and the implementer can select other feasible calculation methods of the similarity, such as the Pearson correlation coefficient, and the preset threshold in this embodiment is set to 0.5, which can be set by the implementer according to the actual situation, and this embodiment does not limit this.

[0055] For the monitoring data collected each time for each monitoring point, all the environmental data of each monitoring point are arranged in a fixed order to form an environmental vector of each monitoring point, that is, the environmental data at the same position in the environmental vectors of different monitoring points correspond to the same type. For example, any environmental vector β = [temperature, humidity, wind speed, …]. In addition, for the monitoring data collected each time for each monitoring point, the word vector of any to-be-tested word group is added to the last position in the environmental vector of the monitoring point to obtain a state vector of the any to-be-tested word group. For example, the state vector of the to-be-tested word group x is [temperature, humidity, wind speed, …, word vector of the to-be-tested word group x]. Therefore, the environmental vector of each monitoring point and the state vector of each to-be-tested word group of each monitoring point can be obtained.

[0056] S3, analyze the correlation between the historical environmental vector of each monitoring point and the environmental vector of the adjacent monitoring point at the same time, determine the credibility of the historical environmental vector of each monitoring point; combine the correlation between the current environmental vector of each monitoring point and its historical environmental vector, and the credibility, to determine the dynamic weight of the current environmental vector of each monitoring point; accordingly, the dynamic weight of each to-be-tested word group of the current each monitoring point is obtained.

[0057] Since there are differences in the dimensions and value ranges of the measured values and the word vectors, in order to avoid the text information being submerged after direct splicing, dynamic weights are set for the environmental vectors of each monitoring point and the word vectors of the to-be-tested word groups, which are specifically as follows:

[0058] The seasonal change trend of the cryosphere at the same monitoring point is relatively regular, the physical state of the cryosphere is relatively related under similar environmental conditions, and therefore, in this embodiment, the environmental vector β of the current monitoring point i is taken as an example for analysis, the environmental vectors collected by the monitoring point i before the current time and belonging to the same season are obtained, and the cosine similarity between the environmental vectors and the environmental vector β is calculated. The environmental vector corresponding to the maximum cosine similarity in the environmental vectors is recorded as the historical environmental vector of the monitoring point i. ; record the maximum cosine similarity as a first similarity The greater the first similarity is, the greater the credibility of the environment vector β is.

[0059] wherein, in order to avoid the historical environment vector having errors itself, if the cosine similarity of the environment vector of the monitoring point i with all the environment vectors of its adjacent monitoring points is greater, the credibility of the data of the monitoring point i is higher, therefore, the credibility of the historical environment vector of the monitoring point i is calculated, and the expression is: ; wherein, is the credibility of the historical environment vector of the monitoring point i, is the average value of the cosine similarity of the environment vector of the historical environment vector with the environment vectors of all the adjacent monitoring points of the monitoring point i at the same collection time, is the range of the cosine similarity of the environment vector of the historical environment vector with the environment vectors of all the adjacent monitoring points of the monitoring point i at the same collection time, the smaller the range is, the more similar the environment state of the monitoring point i and its adjacent monitoring points are, and sigmoid() is a normalization function.

[0060] It should be noted that the adjacent monitoring points in the embodiment are directly adjacent, that is, there is no monitoring point between the monitoring point i and its adjacent monitoring points.

[0061] The dynamic weight of the environment vector β of the current monitoring point i is .

[0062] In addition, when the environment states of the same monitoring point are similar, the corresponding text descriptions should also have certain similarity, therefore, all the test word groups collected at the corresponding moment of the historical environment vector of the monitoring point i are obtained, taking the test word group x of the current monitoring point i as an example, the cosine similarity of the test word group x with the word vectors of all the test word groups is calculated, recorded as a second similarity, and the maximum value of all the second similarities is recorded as The dynamic weight of the test word group x of the current monitoring point i is .

[0063] S4, all the dynamic weights of the current monitoring points are used to weight the state vectors of the test word groups; based on the correlation of the weighted state vectors of any two monitoring points and the position distance between the any two monitoring points, the possibility of the same state between the any two monitoring points is determined.

[0064] Further, based on the dynamic weight of the environment vector β of the current monitoring point i and the dynamic weight of the to-be-tested phrase group x, the state vector of the to-be-tested phrase group x of the current monitoring point i is weighted to obtain a coupling vector of the environment vector and the to-be-tested phrase group x of the current monitoring point i, denoted as the weighted state vector of the to-be-tested phrase group x of the current monitoring point i .

[0065] Since the cryospheric states of adjacent positions can be similar, the current monitoring point position at the current moment is merged based on the similarity of the weighted state vectors of the to-be-tested phrase groups of different monitoring points and the degree of proximity in geographical position, so as to realize the partition of the monitoring range. Therefore, the embodiment analyzes the possibility that the current monitoring point i and the monitoring point are in the same cryospheric state, denoted as the same state possibility , and specifically:

[0066] The similarity of the weighted state vectors of the monitoring point i and the weighted state vectors of the monitoring point is calculated, denoted as the third similarity. The two state vectors corresponding to the maximum value in all the third similarities are taken as the matching state vectors of the monitoring point i and the monitoring point , for example, the two state vectors corresponding to the maximum value in all the third similarities are and , respectively. , then is taken as the matching state vector of the monitoring point i, and is taken as the matching state vector of the monitoring point .

[0067] In the embodiment, the expression of the same state possibility between the current monitoring point i and the monitoring point is:

[0068] ; in the formula, is the average value of the similarities of all the weighted state vectors of the current monitoring point i and the matching state vectors of the monitoring point i, denoted as the first average value. The greater the first average value is, the more likely it is that the two monitoring points are in the similar cryospheric state. is the metric distance between the current monitoring point i and the monitoring point , that is, the measurement distance between the two monitoring points, and norm() is a normalization function.

[0069] S5, classifying all the monitoring points into regions by the same state likelihoods corresponding to all the monitoring points, and marking the monitoring points in each region divided for each collection; determining the difference between any one type of environmental data of any one monitoring point and the rest of the types of environmental data, analyzing the deviation degree of the difference of the any one monitoring point from the differences of the rest of the monitoring points in the region to which the any one monitoring point belongs, combining the number of times the any one monitoring point is marked, obtaining the abnormal degree of the any one type of environmental data of the any one monitoring point, and determining the abnormal environmental data of each monitoring point for each collection.

[0070] The same state likelihood between the current monitoring point i and the monitoring point is calculated as follows. The sigmoid normalization function is used for normalization to (0, 1), and if the normalized result is greater than a first preset threshold, the current monitoring point i and the monitoring point are divided into the same region.

[0071] It should be noted that the monitoring points of the determined region do not participate in the calculation of the division of other regions, and thus all the monitoring points of the cryosphere can be divided into regions, and if there is an isolated monitoring point, it is divided into the same region as the monitoring point corresponding to the maximum same state likelihood. In this embodiment, the first preset threshold is set to 0.7, and the implementer can determine it according to the actual situation, which is not limited in this embodiment.

[0072] Taking any one type of environmental data of monitoring point i as an example, all the collected data of the any one type of environmental data of the current and historical monitoring point i are arranged in time sequence to form a time sequence environmental sequence of the any one type of environmental data of the current monitoring point i, and if the region to which the monitoring point i belongs is unstable in the current season, and the difference between the time sequence environmental sequence of the any one type of environmental data and the time sequence environmental sequence of the rest of the types of environmental data is greater than the difference between the time sequence environmental sequence of the any one type of environmental data of the rest of the monitoring points in the region to which the monitoring point i belongs and the time sequence environmental sequence of the rest of the types of environmental data, it means that the any one type of environmental data of the current monitoring point i has a greater possibility of being abnormal.

[0073] Therefore, this embodiment first determines the historical attribution region of each monitoring point based on the historical stability of the cryosphere, specifically as follows.

[0074] The regions defined in two consecutive monitoring data collections are matched. For example, region j obtained in the k-th monitoring data collection is matched with the region with the most overlapping monitoring points and the closest center distance among all regions defined in the (k-1)-th monitoring data collection. After each match, each monitoring point in the region is marked. For example, if monitoring point i belongs to region j in the k-th monitoring data collection, then monitoring point i is marked as region j. If monitoring point i belongs to region a in the (k-1)-th monitoring data collection, then monitoring point i is marked as region a. If monitoring point i still belongs to region j, then monitoring point i is still marked as region j. Therefore, a monitoring point may always be marked by the same region or it may be marked by different regions. The number of times a monitoring point is marked is the number of regions marked in the current and historical processes.

[0075] Furthermore, the degree of anomaly in various types of environmental data at each monitoring point is calculated, specifically as follows:

[0076] Calculate the z-th time-series environmental sequence of each monitoring point and the z-th time-series environmental sequence of each monitoring point. The difference between each time-series environmental sequence is denoted as the first difference. The mean of the first difference of all other monitoring points in the area of ​​each monitoring point is calculated and denoted as the second mean.

[0077] The difference between the first difference and the second mean is calculated and denoted as the second difference. The sum of the second differences between the z-th time-series environmental sequence of each monitoring point and all other time-series environmental sequences is calculated. The ratio of the sum to the number of times each monitoring point is marked is used as the degree of abnormality of the environmental data corresponding to the z-th time-series environmental sequence of each monitoring point.

[0078] It should be noted that the difference represents the degree of difference between two variables and two time series. For two variables, the difference can be calculated using the absolute value of the difference, the square of the difference, the ratio, etc. For two time series, the difference can be calculated using Euclidean distance, DTW distance, etc. This embodiment does not impose any restrictions on this.

[0079] The expressions for the degree of anomaly in various types of environmental data at each monitoring point are as follows:

[0080] In the formula, This represents the degree of anomaly in the environmental data corresponding to the z-th time-series environmental sequence at monitoring point i during the k-th monitoring data collection. The z-th time-series environmental sequence of monitoring point i at the time of the k-th monitoring data collection and the z-th time-series environmental sequence of the monitoring point i. The DTW distance of each temporal environment sequence is denoted as the first difference. For the current k-th monitoring data collection, the z-th time-series environmental sequence of all other monitoring points within the area to which monitoring point i belongs is compared with the z-th time-series environmental sequence of the current k-th monitoring data collection. The mean DTW distance of each time-series environmental sequence is denoted as the second mean. This is denoted as the second difference. This represents the number of remaining time-series environmental sequences at monitoring point i, excluding the z-th time-series environmental sequence, during the current k-th monitoring data collection. This represents the number of times monitoring point i was marked during the current k-th monitoring data collection.

[0081] The degree of abnormality of various environmental data at each monitoring point is normalized to (0, 1) using the sigmoid function. Environmental data with a normalization result greater than a second preset threshold are considered abnormal environmental data. In this embodiment, the second preset threshold is set to 0.5. Implementers can set it according to the actual situation. This embodiment does not impose any restrictions on this.

[0082] S6. Analyze the correlation between the weighted state vector of each test word group at each monitoring point and the state vector of other monitoring points in the area to which each monitoring point belongs, as well as the degree of abnormality of environmental data in the weighted state vector of each test word group, to obtain the semantic deviation of each test word group at each monitoring point, and determine the abnormal word groups collected at each monitoring point each time.

[0083] Taking the word group x to be tested at the current monitoring point i as an example, the cosine similarity between the weighted state vector of the word group x to be tested at monitoring point i and all state vectors of all other monitoring points within the region to which monitoring point i belongs is calculated and denoted as the fourth similarity. State vectors with a fourth similarity greater than the third preset threshold corresponding to the word group x to be tested at monitoring point i are taken as the target vector of the word group x to be tested at monitoring point i. The similarity between the weighted state vector of the word group x to be tested and all its target vectors is analyzed. If the similarity is worse and the corresponding environmental vector is more likely to be abnormal, then the word group x to be tested is more likely to have semantic deviation. Therefore, the semantic deviation of each word group to be tested at each monitoring point is determined, specifically as follows:

[0084] In the formula, This represents the semantic bias of the word phrase x at monitoring point i during the current k-th monitoring data collection. This represents the mean of the anomaly severity of all environmental data types in the weighted state vector of the word phrase x to be tested at monitoring point i during the current k-th monitoring data collection. Let be the mean of the cosine similarity between the weighted state vector of the word group x to be tested at monitoring point i and all its target vectors at the current k-th monitoring data collection, denoted as the third mean.

[0085] The semantic deviation of each to-be-tested phrase of each monitoring point is normalized to the interval (0, 1) by using a sigmoid function, and a to-be-tested phrase with a normalized result greater than a fourth preset threshold is regarded as an abnormal phrase. In this embodiment, the fourth preset threshold is set to 0.5, and the implementer can set it according to the actual situation, which is not limited in this embodiment.

[0086] S7, in combination with the abnormal environment data and the abnormal phrase, an abnormal state vector in all state vectors of each monitoring point collected each time is obtained, and is corrected. The corrected state vectors of each monitoring point are combined with a machine learning algorithm to perform data fusion on the environment data and the text data.

[0087] For each region divided according to all monitoring points at present, if there is abnormal environment data or abnormal phrase in the weighted state vector of each monitoring point in the corresponding region, the state vector is marked as an abnormal state vector. The state vector in which there is no abnormal environment data and no abnormal phrase is marked as a normal state vector, that is, all state vectors of each monitoring point in each region are divided into abnormal state vectors and normal state vectors.

[0088] Taking any abnormal state vector R in region j as an example, the cosine similarity between the abnormal state vector R and all normal state vectors in region j is calculated, denoted as a fifth similarity. The normal state vector with a fifth similarity greater than a fifth preset threshold is regarded as a correction reference vector of the abnormal state vector R. In this embodiment, the fifth preset threshold is set to 0.7, and the implementer can set it according to the actual situation, which is not limited in this embodiment.

[0089] If the environment data at the pth position in the abnormal state vector R is abnormal environment data, the environment data at the pth position in the abnormal state vector R is corrected based on the environment data at the pth position in all correction reference vectors of the abnormal state vector R, and the correction is specifically:

[0090] In one embodiment, the mean value of the environment data at the pth position in all correction reference vectors of the abnormal state vector R is calculated as the environment data at the pth position in the abnormal state vector R, to obtain a corrected state vector of the abnormal state vector R.

[0091] In another embodiment, the cosine similarity between each correction reference vector of the abnormal state vector R and the abnormal state vector R is calculated, denoted as a sixth similarity. The sixth similarity is used as a weight to perform weighted averaging on the environment data at the pth position in all correction reference vectors of the abnormal state vector R, as the environment data at the pth position in the abnormal state vector R, to obtain a corrected state vector of the abnormal state vector R.

[0092] For the abnormal word group existing in the abnormal state vector, a modified reference vector corresponding to the fifth maximum value of the similarity of the abnormal state vector is obtained from all the modified reference vectors of the abnormal state vector, and the word vector of the to-be-tested word group in the modified reference vector is taken as the word vector at the corresponding position of the abnormal state vector, so as to complete the semantic modification of the abnormal word group. The abnormal state vectors of all the monitoring points are modified to obtain the modified state vectors of all the monitoring points. The abnormal state vector modification flowchart is shown in Figure 2

[0093] All the normal state vectors of the current monitoring points and the modified abnormal state vectors are taken as the input of the random forest algorithm, and the output is the cryosphere state category of each monitoring point. The cryosphere state category in the embodiment includes: frozen state, thawing state, and freeze-thaw alternating state. The implementer can set it according to the actual situation. The random forest algorithm is a known technology, and the implementer can select other feasible machine learning algorithms according to the actual situation, and the embodiment does not limit this.

[0094] Based on the same inventive concept as the above method, the embodiment of the present application also provides a cryosphere multi-source heterogeneous data semantic fusion system, which comprises a memory, a processor and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the method described in any one of the above methods are realized.

[0095] It should be noted that the above sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The above describes the specific embodiments of the present application. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are possible or may be advantageous.

[0096] Each embodiment in the present specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment mainly describes the differences from other embodiments.

[0097] The above is only the preferred embodiment of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the principles of the present application shall be included in the protection scope of the present application.​

Claims

1. A semantic fusion method for multi-source heterogeneous data from the cryosphere, characterized in that, The method includes the following steps: Collect various monitoring data from various monitoring points in the cryosphere, including text data and environmental data; Obtain each word group to be tested from the text data; combine all class environment data collected from each monitoring point each time to form an environment vector, and combine it with the word vector of each word group to be tested to obtain the state vector of each word group to be tested; Analyze the correlation between the historical environmental vectors of each monitoring point and the environmental vectors of its neighboring monitoring points at the same time to determine the credibility of the historical environmental vectors of each monitoring point; calculate the similarity between the current environmental vectors of each monitoring point and their historical environmental vectors, denoted as the first similarity, and determine the dynamic weight of the current environmental vectors of each monitoring point by multiplying the first similarity by the credibility; accordingly, obtain the dynamic weight of each word group to be tested at each current monitoring point. Using all the dynamic weights of each current monitoring point, the state vector of each word group to be tested is weighted; based on the correlation between the weighted state vectors of the word groups to be tested between any two current monitoring points, and the positional distance between the two monitoring points, the probability of the two monitoring points being in the same state is determined; Based on the probability of the same state corresponding to all monitoring points, all monitoring points are classified into different regions; the differences between different types of environmental data of each monitoring point are determined, the degree of deviation of each monitoring point from the differences of other monitoring points in its region is analyzed, and abnormal environmental data of each monitoring point are determined each time. Analyze the correlation between the weighted state vectors of the word groups to be tested at each monitoring point and the other monitoring points in its region to determine the abnormal word groups collected at each monitoring point each time. The state vectors corresponding to the abnormal environmental data and the abnormal word groups are corrected. Using the corrected state vectors of each monitoring point, combined with machine learning algorithms, the environmental data and text data are fused.

2. The semantic fusion method for multi-source heterogeneous data in the cryosphere as described in claim 1, characterized in that, The determination of the reliability of the historical environmental vectors for each monitoring point includes: Calculate the mean similarity between the historical environmental vector of each monitoring point and the environmental vector of all its neighboring monitoring points at the same time, and determine the range of the similarity between the historical environmental vector of each monitoring point and the environmental vector of all its neighboring monitoring points at the same time. The ratio of the mean to the range is calculated, and the reliability of the historical environmental vectors of each monitoring point is positively correlated with the ratio.

3. The semantic fusion method for multi-source heterogeneous data in the cryosphere as described in claim 1, characterized in that, The process of obtaining the dynamic weights of each word group to be tested at each monitoring point includes: Calculate the similarity between any word group to be tested at each monitoring point and the word vectors of all word groups to be tested at the same time as the historical environment vector, and record it as the second similarity. Multiply the maximum value of all second similarities by the confidence level, and use it as the dynamic weight of any word group to be tested at each monitoring point.

4. The semantic fusion method for multi-source heterogeneous data in the cryosphere as described in claim 1, characterized in that, The determination of the probability of the same state includes: Calculate the weighted state vectors of monitoring point i and the state vectors of monitoring point i. The weighted similarity of each state vector is denoted as the third similarity. The two state vectors corresponding to the maximum value of all the third similarities are taken as monitoring point i and monitoring point i. The two are mutually matching state vectors; Calculate the mean similarity between the weighted state vectors of monitoring point i and the matching state vectors of monitoring point i, denoted as the first mean, and determine the similarity between monitoring point i and monitoring point... The distance measurement is calculated by taking the ratio of the first mean to the normalized value of the distance measurement as the distance between monitoring point i and monitoring point i. The possibility of being in the same state between them.

5. The semantic fusion method for multi-source heterogeneous data in the cryosphere as described in claim 1, characterized in that, The classification and division of all monitoring points into various regions includes: Two monitoring points whose normalized values ​​for the probability of the same state are greater than a first preset threshold are merged into the same region, and the monitoring points in each region are marked with the region number to which they belong.

6. The semantic fusion method for multi-source heterogeneous data in the cryosphere as described in claim 5, characterized in that, The determination of abnormal environmental data collected at each monitoring point each time includes: The current and historical environmental data of each monitoring point are combined into time-series environmental sequences. The z-th time-series environmental sequence of each monitoring point is compared with the z-th time-series environmental sequence of the monitoring point. The difference between each time-series environmental sequence is denoted as the first difference. The mean of the first difference of all other monitoring points in the area of ​​each monitoring point is calculated and denoted as the second mean. Calculate the difference between the first difference and the second mean, and denote it as the second difference. Calculate the sum of the second differences between the z-th time-series environmental sequence of each monitoring point and all other time-series environmental sequences. Use the ratio of the sum to the number of times each monitoring point is marked as the degree of abnormality of the environmental data corresponding to the z-th time-series environmental sequence of each monitoring point. Environmental data whose normalized value of the degree of abnormality is greater than a second preset threshold are considered abnormal environmental data.

7. The semantic fusion method for multi-source heterogeneous data in the cryosphere as described in claim 6, characterized in that, The process of determining abnormal word groups at each monitoring point in each data collection includes: Calculate the similarity between the weighted state vector of the x-th word group to be tested at monitoring point i and all state vectors of all other monitoring points within the region to which monitoring point i belongs, and record it as the fourth similarity. Take the state vectors of all the fourth similarities corresponding to the x-th word group to be tested at monitoring point i that are greater than the third preset threshold as the target vector of the x-th word group to be tested at monitoring point i. The average similarity between the weighted state vector of the x-th test word group at monitoring point i and all its target vectors is calculated and denoted as the third average. The ratio of the average anomaly of all environmental data in the weighted state vector of the x-th test word group at monitoring point i to the third average is taken as the semantic bias of the x-th test word group at monitoring point i. The word groups whose normalized semantic deviation value is greater than the fourth preset threshold are considered abnormal word groups.

8. The semantic fusion method for multi-source heterogeneous data in the cryosphere as described in claim 1, characterized in that, The data fusion of environmental data and text data includes: For each weighted state vector of each monitoring point, if there is abnormal environmental data or abnormal word phrases, the corresponding state vector is marked as an abnormal state vector, and the state vectors of each monitoring point other than the abnormal state vectors are recorded as normal state vectors. Calculate the similarity between any abnormal state vector of each monitoring point and all normal state vectors of the other monitoring points in the area to which each monitoring point belongs, and record it as the fifth similarity. Use the normal state vectors with the fifth similarity greater than the fifth preset threshold as the correction reference vector of the any abnormal state vector. Obtain the abnormal data position in any abnormal state vector, calculate the average value of the data in all corrected reference vectors of any abnormal state vector that are the same as the abnormal data position, replace the average value with the data at the abnormal data position, and obtain the corrected state vector of any abnormal state vector. The normal state vectors and corrected abnormal state vectors of each monitoring point are used as input to the random forest algorithm, and the cryosphere state category of each monitoring point is output.

9. A semantic fusion system for multi-source heterogeneous data from the cryosphere, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Heterogeneous multi-source frozen circle big data exploration and analysis method and system based on space-time segmentation

    CN119202016A

  • Drainage basin ecological anomaly monitoring method based on clustering processing

    CN120449054A