Freezing circle multi-source heterogeneous data semantic fusion method and system

By constructing the environment and state vectors of multi-source heterogeneous data of the cryosphere, determining dynamic weights, detecting and correcting abnormal data, and combining machine learning algorithms, the problems of noise and error in the fusion of multi-source heterogeneous data of the cryosphere are solved, and high-accuracy and reliable data fusion is achieved, thereby improving the accuracy of cryosphere state identification.

CN120744836AActive Publication Date: 2025-10-03NORTHWEST INST OF ECO ENVIRONMENT & RESOURCES CAS
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511141589.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-10-03
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

The quality of multi-source heterogeneous cryosphere data from different sources varies greatly, and there are noise or errors, which affect the accuracy of the fusion process and interfere with the understanding and judgment of the state of the cryosphere.

Method used

By collecting text data and environmental data from various monitoring points in the cryosphere, constructing environmental vectors and state vectors, determining the dynamic weights of the monitoring points, dividing regions based on location and correlation, detecting abnormal data and making corrections, and combining machine learning algorithms for data fusion.

Benefits of technology

It enhances the semantic alignment capability of multimodal data, suppresses noise cascade amplification, accurately identifies functionally similar areas, improves the accuracy and reliability of data fusion, ensures the consistency of long-term sequences, and improves the accuracy of cryosphere state identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744836A_ABST
    Figure CN120744836A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a frozen circle multi-source heterogeneous data semantic fusion method and system.The method comprises the steps that various monitoring data, including text data and environment data, of all monitoring points of a frozen circle are collected; constructing an environment vector and a state vector of each word group to be tested; determining the same state possibility between any two monitoring points; obtaining the abnormal degree of any type of environmental data of any monitoring point, and determining the abnormal environmental data of each monitoring point collected each time; obtaining semantic deviation of each to-be-detected word group of each monitoring point, and determining abnormal word groups of each monitoring point collected each time; and acquiring an abnormal state vector in all state vectors of each monitoring point acquired each time, correcting the abnormal state vector, and performing data fusion on the environment data and the text data by utilizing the corrected state vector of each monitoring point and combining a machine learning algorithm. Therefore, the accuracy of semantic fusion of the multi-source heterogeneous data of the frozen circle is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method and system for semantic fusion of multi-source heterogeneous data in the cryosphere. Background Art

[0002] The cryosphere refers to Earth's ice-covered regions, including polar ice caps, glaciers, permafrost, snow, floating ice, and sea ice. These regions hold Earth's water in solid form and perform important climatic and environmental functions. The cryosphere plays a key role in Earth's climate system, helping to regulate global climate.

[0003] Semantic fusion of multi-source heterogeneous cryosphere data can effectively combine data from different sources and formats, providing more comprehensive and accurate information for studying cryosphere changes. Furthermore, fusing multi-source data facilitates multi-dimensional analysis of cryosphere dynamics, such as ice thickness, coverage, and physical properties. This is crucial for understanding the interrelationship between the cryosphere and the global climate system.

[0004] However, the data quality from different sources varies greatly and may contain noise or errors, especially in long time series data. These abnormal data can easily affect the accuracy of the final results during the fusion process, thereby interfering with the understanding and judgment of the state of the cryosphere. Summary of the Invention

[0005] To solve the above technical problems, the purpose of this application is to provide a method and system for semantic fusion of multi-source heterogeneous cryosphere data. The technical solutions adopted are as follows: In a first aspect, an embodiment of the present application provides a method for semantic fusion of multi-source heterogeneous cryosphere data, the method comprising the following steps: Collect various monitoring data from various monitoring points in the cryosphere, including text data and environmental data; Obtain each test phrase in the text data; combine all the class environment data collected at each monitoring point into an environment vector, and combine it with the word vector of each test phrase to obtain the state vector of each test phrase; Analyze the correlation between the historical environmental vector of each monitoring point and the environmental vectors of its adjacent monitoring points at the same time to determine the credibility of the historical environmental vector of each monitoring point; combine the correlation between the current environmental vector of each monitoring point and its historical environmental vector, as well as the credibility, to determine the dynamic weight of the current environmental vector of each monitoring point; accordingly, obtain the dynamic weight of each phrase to be tested at each current monitoring point; Using all the dynamic weights of the current monitoring points, the state vectors of the phrases to be tested are weighted; based on the correlation between the weighted state vectors of any two monitoring points and the position distance between the two monitoring points, the probability of the two monitoring points being in the same state is determined; Classify all monitoring points into regions based on the likelihood of the same state corresponding to all monitoring points, and mark the monitoring points in each region for each collection; determine the difference between any type of environmental data of any monitoring point and the remaining types of environmental data, analyze the degree of deviation between the difference of any monitoring point and the differences of the remaining monitoring points in the region to which the any monitoring point belongs, and obtain the degree of abnormality of any type of environmental data of any monitoring point in combination with the number of times the any monitoring point is marked, and determine the abnormal environmental data of each monitoring point each time; Analyze the correlation between the weighted state vector of each test phrase at each monitoring point and the state vectors of other monitoring points in the area to which each monitoring point belongs, as well as the degree of abnormality of the environmental data in the weighted state vector of each test phrase, to obtain the semantic deviation of each test phrase at each monitoring point, and determine the abnormal phrases at each monitoring point each time; Combining the abnormal environmental data with the abnormal phrases, the abnormal state vectors among all the state vectors collected at each monitoring point are obtained and corrected. The corrected state vectors of each monitoring point are used in combination with a machine learning algorithm to perform data fusion on the environmental data and the text data.

[0006] In one embodiment, determining the credibility of the historical environment vector of each monitoring point includes: Calculate the mean of the similarity between the historical environmental vector of each monitoring point and the environmental vectors of all adjacent monitoring points at the same time, and determine the range of the similarity between the historical environmental vector of each monitoring point and the environmental vectors of all adjacent monitoring points at the same time; The ratio of the mean to the range is calculated, and the credibility of the historical environment vector of each monitoring point is positively correlated with the ratio.

[0007] In one embodiment, determining the dynamic weight of the current environment vector of each monitoring point includes: The similarity between the current environment vector of each monitoring point and its historical environment vector is calculated and recorded as the first similarity. The dynamic weight of the current environment vector of each monitoring point is the product of the first similarity and the credibility.

[0008] In one embodiment, obtaining the dynamic weight of each phrase to be tested at each current monitoring point includes: Calculate the similarity between the word vectors of any word group to be tested at each monitoring point and all word groups to be tested at the same time as the historical environment vector, record it as the second similarity, and multiply the maximum value of all the second similarities by the credibility level as the dynamic weight of any word group to be tested at each monitoring point.

[0009] In one embodiment, determining the likelihood of the same state includes: Calculate the weighted state vectors of monitoring point i and the weighted state vectors of monitoring point i The weighted similarity of each state vector is recorded as the third similarity, and the two state vectors corresponding to the maximum value of all the third similarities are taken as the monitoring point i and the monitoring point are mutually matching state vectors; Calculate the mean of the similarity between all weighted state vectors of monitoring point i and the matching state vector of monitoring point i, record it as the first mean, and determine the similarity between monitoring point i and monitoring point i. The metric distance between monitoring point i and monitoring point i is the ratio of the first mean to the normalized value of the metric distance. The possibility of the same state between them.

[0010] In one embodiment, classifying all monitoring points into regions and marking the monitoring points in each region for each acquisition includes: Two monitoring points whose normalized values ​​of the same-state probability are greater than a first preset threshold are merged into the same area, and the monitoring points in each area are marked with the number of the area to which they belong.

[0011] In one embodiment, determining the abnormal environmental data collected at each monitoring point each time includes: The current and historical environmental data of each monitoring point are combined into various time series environmental sequences, and the zth time series environmental sequence and the zth time series environmental sequence of each monitoring point are calculated. The difference of the time series environment sequence is recorded as the first difference, and the mean of the first difference of all other monitoring points in the area of ​​each monitoring point is calculated and recorded as the second mean; Calculate the difference between the first difference and the second mean, record it as the second difference, calculate the cumulative sum of the second differences between the zth time series environment sequence of each monitoring point and all other time series environment sequences, and use the ratio of the cumulative sum to the number of times each monitoring point is marked as the abnormality level of the environmental data corresponding to the zth time series environment sequence of each monitoring point; Environmental data whose normalized value of the abnormality degree is greater than a second preset threshold is regarded as abnormal environmental data.

[0012] In one embodiment, determining abnormal phrases at each monitoring point for each acquisition includes: Calculate the similarity between the weighted state vector of the xth test phrase at monitoring point i and all state vectors of all other monitoring points in the region to which monitoring point i belongs, and record it as a fourth similarity. Select the state vectors corresponding to the xth test phrase at monitoring point i and having a fourth similarity greater than a third preset threshold as the target vector of the xth test phrase at monitoring point i. Calculate the mean of the similarities between the weighted state vector of the x-th test phrase at monitoring point i and all its target vectors, recorded as the third mean, and take the ratio of the mean of the abnormality of all class environment data in the weighted state vector of the x-th test phrase at monitoring point i to the third mean as the semantic deviation of the x-th test phrase at monitoring point i; The phrases to be tested whose normalized semantic deviation values ​​are greater than a fourth preset threshold are regarded as abnormal phrases.

[0013] In one embodiment, the data fusion of the environmental data and the text data includes: For each weighted state vector of each current monitoring point, if there is abnormal environmental data or abnormal phrase, the corresponding state vector is marked as an abnormal state vector, and the state vectors of each monitoring point except the abnormal state vector are recorded as normal state vectors; Calculate the similarity between any abnormal state vector of each monitoring point and all normal state vectors of other monitoring points in the area to which each monitoring point belongs, record it as a fifth similarity, and use the normal state vector whose fifth similarity is greater than a fifth preset threshold as a correction reference vector for any abnormal state vector; Obtaining an abnormal data position in any abnormal state vector, calculating a mean value of data at the same position as the abnormal data in all corrected reference vectors of the any abnormal state vector, replacing the data mean value with the data at the abnormal data position, and obtaining a corrected state vector of the any abnormal state vector; All normal state vectors and the corrected abnormal state vectors of each monitoring point are used as inputs of the random forest algorithm to output the cryosphere state category of each monitoring point.

[0014] In a second aspect, an embodiment of the present application also provides a cryosphere multi-source heterogeneous data semantic fusion system, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of any one of the above methods when executing the computer program.

[0015] This application has at least the following beneficial effects: This application enhances the semantic alignment capability of multimodal data by mapping the text data and environmental data of monitoring points into a unified state vector. The semantics of cryosphere professional terms are captured by word vectors, and the physical state is quantified by environmental vectors. The two complement each other and improve the integrity of the feature expression of each monitoring point. Furthermore, by determining the dynamic weight of the environmental vector of each current monitoring point and the dynamic weight of each test phrase of each current monitoring point, the dynamic distribution of data value is strengthened and the cascade amplification of noise is suppressed. Based on the possibility of the same state corresponding to all monitoring points, all monitoring points are classified into regions, breaking through the limitations of geographical distance, accurately identifying functionally similar areas, and avoiding zoning distortion. By performing anomaly detection on the environmental data and test phrases in the state vector, the root cause of the anomaly is accurately located, and the accuracy and reliability of subsequent data semantic fusion are improved. Furthermore, by performing targeted correction on the abnormal state vector, the scientific nature of data repair is improved, the consistency of long-term sequences is guaranteed, and the credibility of multi-source heterogeneous data is improved. Finally, semantic fusion is achieved using machine learning algorithms, which improves the accuracy of cryosphere state recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0017] Figure 1 A flowchart of the steps of a method for semantic fusion of multi-source heterogeneous cryosphere data provided in one embodiment of the present application; Figure 2 Correct the flow chart for abnormal state vector. DETAILED DESCRIPTION

[0018] To further illustrate the technical means and effectiveness of this application to achieve the intended invention objectives, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features, and effectiveness of a method and system for semantic fusion of multi-source heterogeneous cryosphere data proposed in this application. In the following description, different references to "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable manner.

[0019] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0020] The following describes in detail a method and system for semantic fusion of multi-source heterogeneous cryosphere data provided by the present application with reference to the accompanying drawings.

[0021] See also Figure 1 , which shows a flowchart of a method for semantic fusion of multi-source heterogeneous cryosphere data provided by an embodiment of the present application, the method comprising the following steps: S1, collects various monitoring data from various monitoring points in the cryosphere, including text data and environmental data.

[0022] This embodiment takes the Qinghai-Tibet Plateau cryosphere as an example to monitor the cryosphere and obtain its multi-source monitoring data. Specifically, monitoring points are evenly set within the monitoring range of the cryosphere, sensors are installed at each monitoring point, and various environmental data of each monitoring point are collected. In this embodiment, the environmental data includes temperature, humidity, wind speed, snow depth, and snow water equivalent. The implementer can determine the type of environmental data to be monitored based on actual conditions, and this embodiment does not impose any restrictions on this.

[0023] In addition, textual descriptions of each monitoring point are obtained by researching records of permafrost engineering diseases, analyzing glacier changes, and describing extreme weather events at each monitoring point, serving as text data. Implementers can select the source of this text data. In this embodiment, a monitoring point is set every 200 kilometers. Implementers can set this number based on their actual circumstances, and this embodiment does not impose any restrictions.

[0024] S2, obtaining each phrase to be tested in the text data; forming an environment vector from all the class environment data collected at each monitoring point each time, and combining it with the word vector of each phrase to be tested to obtain a state vector of each phrase to be tested.

[0025] The text data from each monitoring point was processed using the Jieba word segmentation tool to obtain each phrase, and the word vector for each phrase was obtained using word2vec technology. In this example, all entries in the Cryospheric Science Dictionary were combined into a cryosphere vocabulary. The similarity between the word vectors of all phrases at each monitoring point and the word vectors of each entry in the cryosphere vocabulary was calculated. Phrases with similarity greater than a preset threshold were selected as the test phrases for each monitoring point.

[0026] It should be noted that the similarities described in this embodiment are all calculated using cosine similarity. The implementer can choose other existing feasible similarity calculation methods, such as the Pearson correlation coefficient. The preset threshold in this embodiment is set to 0.5, and the implementer can set it according to actual conditions. This embodiment does not impose any restrictions on this.

[0027] For each monitoring data collected at each monitoring point, all the environmental data of each monitoring point are arranged in a fixed order to form the environmental vector of each monitoring point. That is, the environmental data at the same position in the environmental vectors of different monitoring points correspond to the same type. For example, any environmental vector β = [temperature, humidity, wind speed, ...]. In addition, for each monitoring data collected at each monitoring point, the word vector of any phrase to be tested is added to the last position in the environmental vector of the monitoring point to obtain the state vector of the phrase to be tested. For example, the state vector of the phrase to be tested x is [temperature, humidity, wind speed, ..., word vector of the phrase to be tested x]. Therefore, the environmental vector of each monitoring point and the state vector of each phrase to be tested at each monitoring point can be obtained.

[0028] S3, analyze the correlation between the historical environmental vector of each monitoring point and the environmental vector of its adjacent monitoring points at the same time, and determine the credibility of the historical environmental vector of each monitoring point; combine the correlation between the current environmental vector of each monitoring point and its historical environmental vector, as well as the credibility, to determine the dynamic weight of the current environmental vector of each monitoring point; accordingly, obtain the dynamic weight of each phrase to be tested at each current monitoring point.

[0029] Since the measured values ​​and word vectors have different dimensions and ranges, in order to avoid the text information being submerged after direct concatenation, dynamic weights are set for the environmental vectors of each monitoring point and the word vectors of the phrase to be measured. Specifically, The seasonal variation trend of the same monitoring point in the cryosphere is relatively regular. Under similar environmental conditions, the physical state of the cryosphere is relatively relevant. Therefore, this embodiment takes the environmental vector β of the current monitoring point i as an example for analysis, obtains the environmental vectors collected by the monitoring point i before the current one and in the same season as the current one, and calculates the cosine similarity between the environmental vectors and the environmental vector β. The environmental vector corresponding to the maximum cosine similarity among the environmental vectors is recorded as the historical environmental vector of the monitoring point i. ; Record the maximum value of the cosine similarity as the first similarity , the greater the first similarity, the greater the credibility of the environment vector β.

[0030] Among them, in order to avoid the historical environment vector If there is an error in itself, If the cosine similarity of the environmental vectors of all its adjacent monitoring points is large, the credibility of its data is high. Therefore, the historical environmental vector of monitoring point i is calculated. The credibility of is expressed as: ;in, is the historical environment vector of monitoring point i The credibility of Historical environment vector The mean of the cosine similarity of the environmental vectors of all adjacent monitoring points of monitoring point i with the same collection time, Historical environment vector The range of the cosine similarity of the environmental vectors of monitoring point i and all adjacent monitoring points with the same acquisition time. The smaller the range, the more similar the environmental states of monitoring point i and its adjacent monitoring points are. Sigmoid() is a normalization function.

[0031] It should be noted that the adjacent monitoring points in this embodiment are directly adjacent, that is, there is no monitoring point between the monitoring point i and its adjacent monitoring point.

[0032] Then the dynamic weight of the environment vector β of the current monitoring point i is for .

[0033] In addition, when the environmental status of the same monitoring point is similar, the corresponding text description should also have a certain degree of similarity. Therefore, the historical environmental vector of monitoring point i is obtained. For all the test phrases collected at the corresponding moment, take the test phrase x at the current monitoring point i as an example, calculate the cosine similarity between the test phrase x and the word vectors of all the test phrases, record it as the second similarity, and record the maximum value of all the second similarities as , then the dynamic weight of the phrase x to be tested at the current monitoring point i is for .

[0034] S4, using all the dynamic weights of the current monitoring points, weighting the state vectors of each phrase to be tested; based on the correlation of the weighted state vectors of any two current monitoring points and the position distance between the any two monitoring points; determining the possibility of the same state between the any two monitoring points.

[0035] Furthermore, based on the dynamic weight of the environment vector β of the current monitoring point i and the dynamic weight of the phrase x to be tested, the state vector of the phrase x to be tested at the current monitoring point i is weighted to obtain the coupling vector of the environment vector of the current monitoring point i and the phrase x to be tested, which is recorded as the weighted state vector of the phrase x to be tested at the current monitoring point i .

[0036] Since the cryosphere states of neighboring locations may be similar, the monitoring points at the current moment are merged based on the similarity of the weighted state vectors of the test phrases at different monitoring points and the degree of proximity of the geographical locations to achieve the partitioning of the monitoring range. Therefore, this embodiment analyzes the current monitoring point i and the monitoring point i. The probability of being in the same cryosphere state, denoted as the same state probability , specifically: Calculate the weighted state vectors of monitoring point i and the weighted state vectors of monitoring point i The weighted similarity of each state vector is recorded as the third similarity, and the two state vectors corresponding to the maximum value of all the third similarities are taken as the monitoring point i and the monitoring point are mutually matching state vectors, for example, the two state vectors corresponding to the maximum values ​​of all the third similarities are respectively With monitoring points of , then As the matching state vector of monitoring point i, As a monitoring point The matching state vector.

[0037] In this embodiment, the current monitoring point i and the monitoring point The possibility of the same state between The expression is: Where, is the mean of the similarity of the matching state vectors of all state vectors of the current monitoring point i after weighting, recorded as the first mean. The larger the first mean, the more likely it is to be in a similar cryosphere state. The current monitoring point i and the monitoring point The metric distance between two monitoring points is , that is, the measured distance between two monitoring points, and norm() is the normalization function.

[0038] S5. Classify all monitoring points into areas according to the same-state possibilities corresponding to all monitoring points, and mark the monitoring points in each area collected each time; determine the difference between any type of environmental data of any monitoring point and the remaining types of environmental data, analyze the degree of deviation between the difference of any monitoring point and the difference of the remaining monitoring points in the area to which any monitoring point belongs, and obtain the degree of abnormality of any type of environmental data of any monitoring point in combination with the number of times any monitoring point is marked, and determine the abnormal environmental data of each monitoring point collected each time.

[0039] Combine the current monitoring point i with the monitoring point The possibility of the same state between Use sigmoid normalization function to normalize to (0,1). If the normalization result is greater than the first preset threshold, the current monitoring point i is compared with the monitoring point Divided into the same area.

[0040] It should be noted that monitoring points in a defined region do not participate in the calculation of other region divisions. At this point, all monitoring points in the cryosphere can be divided into regions. If an isolated monitoring point exists, it is assigned to the same region as the monitoring point with the highest probability of the same state. In this embodiment, the first preset threshold is set to 0.7. Implementers can determine this threshold based on actual circumstances, and this embodiment does not impose any restrictions.

[0041] Taking any type of environmental data of monitoring point i as an example, all collected data of any type of environmental data of the current and historical monitoring point i are arranged in time series to form a time series environmental sequence of any type of environmental data of the current monitoring point i. If in the current season, the area to which monitoring point i belongs is unstable, and the difference between the time series environmental sequence of any type of environmental data and the time series environmental sequence of other types of environmental data is greater than the difference between the time series environmental sequence of any type of environmental data of other monitoring points in the area to which monitoring point i belongs and the time series environmental sequence of other types of environmental data, then it means that there is a greater possibility that any type of environmental data of the current monitoring point i is abnormal.

[0042] Therefore, this embodiment first determines the historical region of each monitoring point based on the historical stability of the cryosphere, specifically: Match the regions divided in two adjacent monitoring data collections. For example, match region j obtained during the k-th monitoring data collection with the region with the most overlapping monitoring points and the closest center distance to region j among all the regions obtained during the k-1-th monitoring data collection. After each match, mark each monitoring point in the region. For example, if monitoring point i belongs to region j during the k-th monitoring data collection, then monitoring point i is marked as region j. If monitoring point i belongs to region a during the k-1-th monitoring data collection, then monitoring point i is marked as region a. If monitoring point i still belongs to region j, then monitoring point i is still marked as region j. Therefore, for a monitoring point, it may always be marked by the same region or by different regions. The number of times a monitoring point is marked is the number of regions marked in the current and historical processes.

[0043] Furthermore, the abnormality degree of various environmental data of each current monitoring point is calculated, specifically: Calculate the zth time series environment sequence of each monitoring point and the zth The difference of the time series environment sequence is recorded as the first difference, and the mean of the first difference of all other monitoring points in the area of ​​each monitoring point is calculated and recorded as the second mean; Calculate the difference between the first difference and the second mean, record it as the second difference, calculate the cumulative sum of the second differences between the zth time series environment sequence of each monitoring point and all other time series environment sequences, and take the ratio of the cumulative sum to the number of times each monitoring point is marked as the abnormality degree of the environmental data corresponding to the zth time series environment sequence of each monitoring point.

[0044] It should be noted that the difference represents the degree of difference between two variables and two time series. For two variables, the difference can be calculated using the absolute value of the difference, the square of the difference, the ratio, etc.; for two time series, the difference can be calculated using the Euclidean distance, DTW distance, etc., which is not limited in this embodiment.

[0045] The expression of the abnormal degree of various environmental data at each monitoring point is: Where, is the abnormality degree of the environmental data corresponding to the z-th time series environmental sequence of monitoring point i during the current k-th monitoring data collection, is the zth temporal environment sequence of monitoring point i during the kth monitoring data collection and the The DTW distance of a time series environment sequence is recorded as the first difference, is the zth temporal environment sequence of all other monitoring points in the area to which monitoring point i belongs when the kth monitoring data is collected and the zth temporal environment sequence of all other monitoring points in the area to which monitoring point i belongs The mean of the DTW distances of the time series environment sequences is recorded as the second mean. Recorded as the second difference, is the number of remaining time series environment sequences except the zth time series environment sequence of monitoring point i during the current kth monitoring data collection, It is the number of times monitoring point i is marked during the current k-th monitoring data collection.

[0046] The abnormality degree of each type of environmental data at each current monitoring point is normalized to (0, 1) using the sigmoid function, and the environmental data whose normalized result is greater than the second preset threshold is regarded as abnormal environmental data. In this embodiment, the second preset threshold is set to 0.5. The implementer can set it according to the actual situation, and this embodiment does not impose any restrictions on this.

[0047] S6, analyzing the correlation between the weighted state vectors of each test phrase at each monitoring point and the state vectors of other monitoring points in the area to which each monitoring point belongs, as well as the abnormality of the environmental data in the weighted state vectors of each test phrase, to obtain the semantic deviation of each test phrase at each monitoring point, and determine the abnormal phrases collected at each monitoring point each time.

[0048] Taking the test phrase x of the current monitoring point i as an example, the cosine similarity between the weighted state vector of the test phrase x of the monitoring point i and all the state vectors of all other monitoring points in the area to which the monitoring point i belongs is calculated, which is recorded as the fourth similarity. The state vectors corresponding to all the fourth similarities of the test phrase x of the monitoring point i that are greater than the third preset threshold are used as the target vector of the test phrase x of the monitoring point i. The similarity between the weighted state vector of the test phrase x and all its target vectors is analyzed. If the similarity is worse and the corresponding environment vector is more likely to be abnormal, then the possibility of semantic deviation of the test phrase x is greater. Therefore, the semantic deviation of each test phrase at each monitoring point is determined as follows: Where, is the semantic deviation of the tested phrase x at monitoring point i during the current k-th monitoring data collection, is the mean value of the abnormality degree of all environmental data in the weighted state vector of the tested phrase x at the monitoring point i during the current k-th monitoring data collection. The third mean is the mean of the cosine similarities between the weighted state vector of the test phrase x at the monitoring point i and all its target vectors during the current k-th monitoring data collection.

[0049] The semantic deviation of each tested phrase at each current monitoring point is normalized to the range (0, 1) using a sigmoid function. Test phrases whose normalized results exceed a fourth preset threshold are considered abnormal phrases. In this embodiment, the fourth preset threshold is set to 0.5. The implementer can set this threshold based on actual circumstances, and this embodiment does not impose any restrictions on this.

[0050] S7, combining the abnormal environmental data with the abnormal phrase, obtaining the abnormal state vector from all state vectors collected at each monitoring point each time, and correcting it, using the corrected state vector of each monitoring point, combined with a machine learning algorithm, to perform data fusion on the environmental data and the text data.

[0051] For each area divided by all current monitoring points, for each weighted state vector of each monitoring point in the corresponding area, if there is abnormal environmental data or abnormal phrases in the state vector, the state vector is marked as an abnormal state vector, and the state vector without abnormal environmental data and abnormal phrases in the state vector is marked as a normal state vector, that is, all weighted state vectors of each monitoring point in each area are divided into abnormal state vectors and normal state vectors.

[0052] Taking any abnormal state vector R in region j as an example, the cosine similarity between the abnormal state vector R and all normal state vectors in region j is calculated, recorded as the fifth similarity. The normal state vectors whose fifth similarity is greater than the fifth preset threshold are used as correction reference vectors for the abnormal state vector R. In this embodiment, the fifth preset threshold is set to 0.7. Implementers can set it according to actual circumstances, and this embodiment does not impose any restrictions on this.

[0053] If the environmental data at the p-th position in the abnormal state vector R is abnormal environmental data, then based on the environmental data at the p-th position in all the corrected reference vectors of the abnormal state vector R, the environmental data at the p-th position in the abnormal state vector R is corrected, specifically: In one embodiment, the mean of the environmental data at the p-th position in all the corrected reference vectors of the abnormal state vector R is calculated as the environmental data at the p-th position in the abnormal state vector R, to obtain a corrected state vector of the abnormal state vector R; In another embodiment, the cosine similarity between each corrected reference vector of the abnormal state vector R and the abnormal state vector R is calculated, recorded as the sixth similarity, and the sixth similarity is used as a weight to perform weighted averaging on the environmental data of the p-th position in all corrected reference vectors of the abnormal state vector R as the environmental data of the p-th position in the abnormal state vector R, to obtain the corrected state vector of the abnormal state vector R.

[0054] For abnormal phrases in the abnormal state vector, the corrected reference vector corresponding to the fifth maximum similarity value of the abnormal state vector is obtained from all corrected reference vectors of the abnormal state vector, and the word vector of the phrase to be tested in the corrected reference vector is used as the word vector of the corresponding position of the abnormal state vector to complete the semantic correction of the abnormal phrase. The abnormal state vectors of all current monitoring points are corrected to obtain the corrected state vectors of all current monitoring points. The abnormal state vector correction flow chart is shown in FIG. Figure 2 shown.

[0055] All normal state vectors and the corrected abnormal state vectors for each monitoring point are used as input to the random forest algorithm, which outputs the cryosphere state category for each monitoring point. In this embodiment, the cryosphere state categories include frozen state, fused state, and alternating freeze-thaw state, which can be customized by the implementer based on actual circumstances. The random forest algorithm is a well-known technique, and implementers can choose other feasible existing machine learning algorithms based on actual circumstances; this embodiment does not limit this.

[0056] Based on the same inventive concept as the above method, an embodiment of the present application also provides a cryosphere multi-source heterogeneous data semantic fusion system, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above-mentioned cryosphere multi-source heterogeneous data semantic fusion methods are implemented.

[0057] It should be noted that the order in which the embodiments of the present application are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. Furthermore, the foregoing descriptions of specific embodiments of this specification are provided. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential sequence shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0058] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.

[0059] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A method for semantic fusion of multi-source heterogeneous cryosphere data, characterized by: The method comprises the following steps: Collect various monitoring data from various monitoring points in the cryosphere, including text data and environmental data; Obtain each test phrase in the text data; combine all the class environment data collected at each monitoring point into an environment vector, and combine it with the word vector of each test phrase to obtain the state vector of each test phrase; Analyze the correlation between the historical environmental vector of each monitoring point and the environmental vector of its adjacent monitoring points at the same time to determine the credibility of the historical environmental vector of each monitoring point; calculate the similarity between the current environmental vector of each monitoring point and its historical environmental vector, recorded as a first similarity, and multiply the first similarity by the credibility to determine the dynamic weight of the environmental vector of each current monitoring point; accordingly, obtain the dynamic weight of each to-be-tested phrase at each current monitoring point; Using all the dynamic weights of the current monitoring points, the state vectors of the phrases to be tested are weighted; based on the correlation between the weighted state vectors of any two monitoring points and the position distance between the two monitoring points, the probability of the two monitoring points being in the same state is determined; Classify all monitoring points into regions based on the likelihood of the same state corresponding to all monitoring points; determine the differences between different types of environmental data at each monitoring point, analyze the degree of deviation between the differences at each monitoring point and the rest of the monitoring points in its region, and determine abnormal environmental data at each monitoring point each time; Analyze the correlation between the weighted state vectors of the phrases to be tested of each monitoring point and the rest of the monitoring points in its area, and determine the abnormal phrases of each monitoring point collected each time; The state vectors corresponding to the abnormal environmental data and the abnormal phrases are corrected, and the corrected state vectors of each monitoring point are used in combination with a machine learning algorithm to perform data fusion on the environmental data and the text data.

2. The method for semantic fusion of multi-source heterogeneous cryosphere data according to claim 1, characterized in that: Determining the credibility of the historical environmental vector of each monitoring point includes: Calculate the mean of the similarity between the historical environmental vector of each monitoring point and the environmental vectors of all adjacent monitoring points at the same time, and determine the range of the similarity between the historical environmental vector of each monitoring point and the environmental vectors of all adjacent monitoring points at the same time; The ratio of the mean to the range is calculated, and the credibility of the historical environment vector of each monitoring point is positively correlated with the ratio.

3. The method for semantic fusion of multi-source heterogeneous cryosphere data according to claim 1, characterized in that: The step of obtaining the dynamic weight of each phrase to be tested at each current monitoring point includes: Calculate the similarity between the word vectors of any word group to be tested at each monitoring point and all word groups to be tested at the same time as the historical environment vector, record it as the second similarity, and multiply the maximum value of all the second similarities by the credibility level as the dynamic weight of any word group to be tested at each monitoring point.

4. The method for semantic fusion of multi-source heterogeneous cryosphere data according to claim 1, characterized in that: The determination of the same-state possibility includes: Calculate the weighted state vectors of monitoring point i and the weighted state vectors of monitoring point i The weighted similarity of each state vector is recorded as the third similarity, and the two state vectors corresponding to the maximum value of all the third similarities are taken as the monitoring point i and the monitoring point are mutually matching state vectors; Calculate the mean of the similarity between all weighted state vectors of monitoring point i and the matching state vector of monitoring point i, record it as the first mean, and determine the similarity between monitoring point i and monitoring point i. The metric distance between monitoring point i and monitoring point i is the ratio of the first mean to the normalized value of the metric distance. The possibility of the same state between them.

5. The method for semantic fusion of multi-source heterogeneous cryosphere data according to claim 1, characterized in that: All monitoring points are classified into various areas, including: Two monitoring points whose normalized values ​​of the same-state probability are greater than a first preset threshold are merged into the same area, and the monitoring points in each area are marked with the number of the area to which they belong.

6. The method for semantic fusion of multi-source heterogeneous cryosphere data according to claim 5, characterized in that: The determining of abnormal environmental data collected at each monitoring point each time includes: The current and historical environmental data of each monitoring point are combined into various time series environmental sequences, and the zth time series environmental sequence and the zth time series environmental sequence of each monitoring point are calculated. The difference of the time series environment sequence is recorded as the first difference, and the mean of the first difference of all other monitoring points in the area of ​​each monitoring point is calculated and recorded as the second mean; Calculate the difference between the first difference and the second mean, record it as the second difference, calculate the cumulative sum of the second differences between the zth time series environment sequence of each monitoring point and all other time series environment sequences, and use the ratio of the cumulative sum to the number of times each monitoring point is marked as the abnormality level of the environmental data corresponding to the zth time series environment sequence of each monitoring point; Environmental data whose normalized value of the abnormality degree is greater than a second preset threshold is regarded as abnormal environmental data.

7. The method for semantic fusion of multi-source heterogeneous cryosphere data according to claim 6, characterized in that: The step of determining abnormal phrases at each monitoring point during each collection includes: Calculate the similarity between the weighted state vector of the xth test phrase at monitoring point i and all state vectors of all other monitoring points in the region to which monitoring point i belongs, and record it as a fourth similarity. Select the state vectors corresponding to the xth test phrase at monitoring point i and having a fourth similarity greater than a third preset threshold as the target vector of the xth test phrase at monitoring point i. Calculate the mean of the similarities between the weighted state vector of the x-th test phrase at monitoring point i and all its target vectors, recorded as the third mean, and take the ratio of the mean of the abnormality of all class environment data in the weighted state vector of the x-th test phrase at monitoring point i to the third mean as the semantic deviation of the x-th test phrase at monitoring point i; The phrases to be tested whose normalized semantic deviation values ​​are greater than a fourth preset threshold are regarded as abnormal phrases.

8. The method for semantic fusion of multi-source heterogeneous cryosphere data according to claim 1, characterized in that: The data fusion of the environmental data and the text data includes: For each weighted state vector of each current monitoring point, if there is abnormal environmental data or abnormal phrase, the corresponding state vector is marked as an abnormal state vector, and the state vectors of each monitoring point except the abnormal state vector are recorded as normal state vectors; Calculate the similarity between any abnormal state vector of each monitoring point and all normal state vectors of other monitoring points in the area to which each monitoring point belongs, record it as a fifth similarity, and use the normal state vector whose fifth similarity is greater than a fifth preset threshold as a correction reference vector for any abnormal state vector; Obtaining an abnormal data position in any abnormal state vector, calculating a mean value of data at the same position as the abnormal data in all corrected reference vectors of the any abnormal state vector, replacing the data mean value with the data at the abnormal data position, and obtaining a corrected state vector of the any abnormal state vector; All normal state vectors and the corrected abnormal state vectors of each monitoring point are used as inputs of the random forest algorithm to output the cryosphere state category of each monitoring point.

9. A cryosphere multi-source heterogeneous data semantic fusion system, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Smart city optimization management method and system based on multi-source data fusion

    CN115861011A

  • Aquaculture environment monitoring method based on multi-source data fusion

    CN118171135A

  • Heterogeneous multi-source frozen circle big data exploration and analysis method and system based on space-time segmentation

    CN119202016A

  • Multi-modal data identification and analysis system based on deep learning

    CN120105259A

  • Data lake metadata management method based on semantic synthesis and text vectorization

    CN120181094A