A method and system for constructing a secure data warehouse supporting multi-dimensional combined queries
Patent Information
- Application Number
- CN202611097024.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-23
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2046-07-23
AI Technical Summary
[0003]然而,由于数据采集链路复杂、来源多样、日志数据表达方式不统一等原因,导致采集的消防数据中存在大量重复记录
本申请通过分析相似事件集合中单个维度消防数据被逐一移除前后的相似度变化量,构建维度扰动度并转化为维度贡献比,能够量化各维度消防数据对消防风险事件去重判别的重要程度,为包含关键判别信息的维度赋予越高权重,反之赋予越低权重,从而使具有区分性的维度在去重过程中发挥更大作用;
Smart Images

Figure CN122614967B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data warehouse construction technology, specifically to a secure data warehouse construction method and system that supports multi-dimensional combined queries. Background Technology
[0002] With the development of IoT and big data technologies, intelligent management using massive amounts of fire safety data has become a trend. By collecting fire safety data such as log data through IoT sensors, fire risk assessments can be conducted for enterprises and factories, and a data warehouse for emergency fire management can be built, providing analysis, query, management, and decision support for regulatory authorities and enterprises.
[0003] However, due to the complexity of data acquisition chains, diverse sources, and inconsistent log data representation methods, a large number of duplicate records exist in the collected fire protection data. For example, data retransmission caused by IoT devices due to network issues, or different descriptions of the same fire risk event from fire protection data collected from different ports, can all result in duplicate data for the same fire risk event. This duplicate data not only increases storage and computing overhead and reduces query efficiency, but also leads to deviations in fire risk assessment results, seriously affecting the efficiency and reliability of the data warehouse. Summary of the Invention
[0004] In light of the above, it is necessary to provide a method and system for constructing a secure data warehouse that supports multi-dimensional combined queries. Compared with traditional methods for constructing secure data warehouses that support multi-dimensional combined queries, this method significantly improves the reliability and query efficiency of the data warehouse by deduplicating fire risk events. In a first aspect, embodiments of this application provide a method for constructing a secure data warehouse that supports multi-dimensional combined queries, the method comprising the following steps: Acquire multi-dimensional fire safety data for each enterprise and factory under each fire risk event at each port, including the enterprise port and the regulatory port; Fire risk events are referred to as events. For each enterprise factory under each port, a set of similar events is obtained by filtering the semantic information of fire data across all dimensions of a single event and other events. By measuring the change in similarity between fire data of a single dimension in the set of similar events and fire data of the same dimension under the single event before and after removing fire data of a single dimension one by one, the dimensional perturbation degree of fire data of each dimension under the single event is obtained, and then the dimensional contribution ratio of fire data of each dimension under the single event is obtained. By using the semantic information of all words within all dimensions of fire protection data under a single event, the words are classified. By comparing the number of words in each category with the number of words under a single event, the lexical scarcity of each category is obtained, and then the distinguishability of each word in each category is obtained. Each word segment under a single event is encoded, and the weighted sequence string of the single event is obtained by combining the discriminative power and the dimensional contribution ratio; the event is deduplicated by the degree of difference between the weighted sequence strings of a single event and each event in its similar event set. A safety data warehouse is built based on the remaining fire protection data after deduplication.
[0005] In one embodiment, the filtering process for the set of similar events is as follows: The semantic feature vectors of all dimensions of fire protection data for each event are concatenated to obtain the semantic vector of each event; Events whose semantic vector similarity to a single event is greater than a preset threshold are denoted as similar events to the single event; all similar events of a single event are combined into a set of similar events. The preset threshold is the average similarity of the semantic vectors between a single event and all other events.
[0006] In one embodiment, the process of obtaining the dimensional perturbation degree is as follows: Extract fire data of any dimension under a single event, and form a data set of the same dimension from the set of similar events containing all fire data of the same dimension as the fire data of any dimension. Calculate the mean of the semantic feature vector similarity between any dimension of fire data and all fire data in the same dimension data set; remove fire data one by one from the same dimension data set without replacement in the order of timestamp from earliest to latest, until only one fire data remains in the same dimension data set; after each fire data is removed, recalculate the mean of the semantic feature vector similarity between any dimension of fire data and all remaining fire data in the same dimension data set; Calculate the difference between the mean values before and after each removal of fire data; use the average of the differences obtained before and after all removals of fire data as the dimensional perturbation of any dimension of fire data under a single event. In one embodiment, the dimensional contribution ratio is the percentage of the dimensional perturbation degree of each dimension of fire protection data under a single event in the dimensional perturbation degree of all dimensions of fire protection data.
[0007] In one embodiment, the word segmentation classification includes: using a clustering algorithm to cluster all words within all dimensions of fire data under a single event, wherein the distance metric is the difference between the semantic feature vectors of the words.
[0008] In one embodiment, the process of obtaining the vocabulary scarcity is as follows: Calculate the percentage of each type of word segmentation in all words under a single event; The scarcity of the vocabulary is negatively correlated with the proportion of the quantity.
[0009] In one embodiment, the process of obtaining the distinguishability is as follows: The vocabulary scarcity is amplified, and the amplification result is mapped to a positive number; The discrimination index is the rounded integer result of the positive number.
[0010] In one embodiment, the process of obtaining the weighted sequence string is as follows: The hash values of each word are weighted using the aforementioned distinguishability to obtain the sequence string of each word; The sequence strings of all segmented words within the fire protection data of each dimension are summed bit by bit to obtain the merged sequence string of the fire protection data of each dimension. The merged sequence strings of fire protection data from each dimension were normalized separately. The weighted sequence string of all dimensions of fire protection data under a single event is obtained by summing the combined sequence string bit by bit. The weight of each data in the combined sequence string of fire protection data of each dimension is the contribution ratio of that dimension.
[0011] In one embodiment, the event deduplication includes: Calculate the difference between a single event and each event in its set of similar events, and identify events whose difference with a single event is less than a preset difference threshold as duplicate events of the single event. Take a single event and all its repeating events to form a repeating event set, and calculate the arithmetic mean of the difference between each event in the repeating event set and all the other events; keep the event with the smallest arithmetic mean among the repeating event set and remove the rest of the events.
[0012] Secondly, embodiments of this application also provide a secure data warehouse construction system that supports multi-dimensional combined queries, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the secure data warehouse construction method that supports multi-dimensional combined queries described above.
[0013] This application has at least the following beneficial effects: This application analyzes the change in similarity before and after removing individual fire data in a set of similar events one by one, constructs the dimension perturbation degree and converts it into the dimension contribution ratio. This can quantify the importance of each dimension of fire data in the deduplication of fire risk events, assign higher weights to dimensions containing key discrimination information and lower weights to those containing less information, so that the discriminative dimensions can play a greater role in the deduplication process. Furthermore, based on the semantic information of word segmentation, clustering and comparing the proportion of each type of word segmentation, constructing vocabulary scarcity and converting it into distinguishability, can identify sensitive words under fire risk events. Higher weights are assigned to key descriptive words that are more likely to be distinctive, and lower weights are assigned to those that are less likely to be distinctive. This effectively solves the problem of a large number of professional terms obscuring key identification words and helps to improve the deduplication effect of fire risk events. Furthermore, by combining the dimensional contribution ratio and the discriminative power, the word segmentation encoding is weighted twice to obtain a weighted sequence string of fire risk events. The dimensional contribution ratio improves the discriminative power of fire data from the data dimension level, while the discriminative power improves the discriminative power of word segmentation from the vocabulary level. The synergistic effect of the two enhances the robustness of the generated weighted sequence string in describing differences, providing a more reliable feature representation for subsequent deduplication of fire risk events. Furthermore, by calculating the degree of difference in weighted sequence strings to determine duplicate fire risk events and deduplicating them, redundant data can be accurately identified and removed. This effectively eliminates duplicate records caused by network retransmission and differences in descriptions across multiple ports, while retaining the most representative events. After the deduplication operation, storage and computing overhead can be significantly reduced, query efficiency can be improved, and the interference of duplicate data on fire risk assessment results can be eliminated, ensuring the accuracy of the assessment. This lays the foundation for building a high-quality, highly reliable security data warehouse, thereby significantly improving the reliability and query efficiency of the data warehouse. Attached Figure Description
[0014] Figure 1 A flowchart illustrating the steps of a method for constructing a secure data warehouse that supports multi-dimensional combined queries, as provided in one embodiment of this application; Figure 2 A schematic diagram illustrating the process of obtaining the weighted sequence string; Figure 3 A flowchart illustrating the process of deduplicating fire risk events. Detailed Implementation
[0015] It should be noted that the terms "first" and "second" in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0016] The following, in conjunction with the accompanying drawings, details a specific scheme for a secure data warehouse construction method and system that supports multi-dimensional combined queries, as provided in this application.
[0017] Please see Figure 1 The diagram illustrates a flowchart of a method for constructing a secure data warehouse that supports multi-dimensional combined queries, according to an embodiment of this application. The method includes the following steps: Step 1: Obtain multi-dimensional fire protection data for each enterprise factory under each fire risk event at each port. The ports include the enterprise port and the regulatory port.
[0018] Acquire multi-dimensional fire protection data for each enterprise and factory under each fire risk event at each port, where the ports include enterprise ports and regulatory ports.
[0019] The multi-dimensional fire safety data for enterprise-level fire risk events includes: fire safety facility status data, fire risk investigation log data, fire safety rating data, and fire safety rectification log data. The multi-dimensional fire safety data for regulatory-level fire risk events includes: fire risk profile data, fire safety assessment data, and enforcement inspection record data for regulated enterprises and factories.
[0020] It should be noted that the multi-dimensional fire safety data from both the enterprise and regulatory sides revolve around each fire risk assessment, collectively forming the entire process of each fire risk assessment. For example, for any given fire risk assessment: the enterprise records whether the fire safety facilities are qualified, risk inspection logs, rating results, and rectification plan logs, etc.; the regulatory side records the enterprise's fire inspection records, assessment results, and subsequent enforcement records, etc., for any given fire risk assessment.
[0021] The fire protection data obtained from each port and each dimension are processed by word segmentation and stop word removal. For each fire risk event of each enterprise and factory under each port, the semantic feature vector of each dimension of fire protection data under each fire risk event and the semantic feature vector of each word segment within it are obtained.
[0022] In this embodiment, the jieba tool is used to perform word segmentation processing on fire protection data of various dimensions. The jieba tool is a well-known technology and will not be described in detail in this application. As other implementation methods, based on the ability to perform word segmentation processing on fire protection data of various dimensions, implementers may use other existing feasible technologies, such as the HanLP tool, etc. This application does not impose any special restrictions.
[0023] In this embodiment, the BERT (Bidirectional Encoder Representations from Transformers) model is used to obtain the semantic feature vectors of fire protection data in each dimension and the semantic feature vectors of each word segment within them. The BERT model is a well-known technology and will not be described in detail in this application. As other implementation methods, based on the ability to obtain the semantic feature vectors of fire protection data in each dimension and the semantic feature vectors of each word segment within them, implementers may use other existing feasible technologies. This application does not impose any special restrictions.
[0024] Deduplication of fire risk events obtained from different ports is performed separately. Taking the deduplication of fire risk events from the enterprise port as an example, subsequent analysis and processing are carried out.
[0025] Step 2: For each enterprise factory under each port, filter to obtain a set of similar events for a single fire risk event; obtain the dimensional perturbation degree and dimensional contribution ratio of fire data of each dimension under a single fire risk event; classify all words within all dimensions of fire data under a single fire risk event, obtain the vocabulary scarcity of each category, and obtain the distinguishability of each word in each category; remove duplicates from the fire risk events.
[0026] Step 2.1: For each enterprise factory under each port, a set of similar events for a single fire risk event is obtained by filtering the semantic information of all dimensions of fire data between a single fire risk event of a single enterprise factory and other fire risk events. By measuring the change in similarity between the single dimension of fire data in the set of similar events and the fire data of the same dimension under the single fire risk event before and after each removal of the single dimension of fire data, the dimensional perturbation degree of each dimension of fire data under the single fire risk event is obtained, and then the dimensional contribution ratio of each dimension of fire data under the single fire risk event is obtained.
[0027] A fire risk incident involves multiple business departments, each recording fire data from different dimensions, collectively forming a complete fire risk assessment and management process. For example, the department responsible for risk investigation records investigation logs, the department responsible for assessment records rating results, and the department responsible for rectification records rectification progress, etc.
[0028] Because different business departments describe the same fire risk event differently, the same fire risk event will show obvious different characteristics when deduplicating due to the differences in log descriptions, resulting in poor deduplication effect.
[0029] Based on the above analysis, for a single enterprise factory, a set of similar events for a single fire risk event is obtained by filtering the semantic information of fire data across all dimensions between the single fire risk event and other fire risk events of that single enterprise factory. The specific process is as follows: The semantic feature vectors of all dimensions of fire data for each fire risk event are concatenated to obtain the semantic vector of each fire risk event. Fire risk events whose semantic vector similarity to a single fire risk event is greater than a preset threshold are denoted as similar fire risk events of the single fire risk event; all similar fire risk events of a single fire risk event are combined into a similar event set.
[0030] The preset threshold is the average similarity of the semantic vectors between a single fire risk event in a single enterprise factory and all other fire risk events.
[0031] In this embodiment, the similarity of semantic vectors is cosine similarity. The calculation of cosine similarity is a well-known technique and will not be described in detail here. As other implementation methods, based on the ability to measure the similarity of semantic vectors, implementers may adopt other existing feasible techniques, such as the reciprocal of Euclidean distance, etc. This application does not impose any special restrictions.
[0032] It should be noted that the overall similarity is measured by using the semantic feature vectors of all dimensions of fire data under a single fire risk event. The reason is that all dimensions of fire data jointly describe a fire risk event, so the deduplication object should be the complete fire risk event, rather than the fire data of a single dimension.
[0033] Furthermore, for a single enterprise factory, the dimensional perturbation degree of each dimension of fire data under a single fire risk event is obtained by measuring the change in similarity between the fire data of a single dimension in a set of similar events for a single fire risk event and the fire data of the same dimension under the same fire risk event before and after each removal of fire data of a single dimension from the similar event set. The specific process is as follows: Extract fire data of any dimension under a single fire risk event, and form a data set of the same dimension from the set of similar events of the single fire risk event, which contains all fire data of the same dimension as the fire data of the stated dimension. Calculate the mean of the semantic feature vector similarity between any dimension of fire protection data and all fire protection data in the same dimension data set; Firefighting data is removed one by one from the same-dimensional data set without replacement, in the order of timestamps from earliest to latest, until only one firefighting data remains in the same-dimensional data set; After each fire data point is removed, the mean of the semantic feature vector similarity between the fire data in any dimension and all remaining fire data in the same dimension data set is recalculated. The difference between the mean semantic feature vector similarity before and after each removal of fire data is calculated to reflect the impact of each removed fire data on the original similarity structure. The larger the calculated difference, the more likely the removed fire data is to be key data, and the more significant the impact of its removal on the original similarity structure. Conversely, the smaller the calculated difference, the stronger the consistency of all fire data in the same dimension dataset, and the smaller the impact of removing a single fire data on the overall similarity judgment. The average of the differences obtained before and after removing all fire data is taken as the dimensional perturbation degree of any dimension of the fire data under a single fire risk event. In this embodiment, the similarity between semantic feature vectors is cosine similarity. As other implementation methods, based on the ability to measure the similarity of semantic feature vectors, implementers may adopt other existing feasible technologies, such as the reciprocal of Euclidean distance, etc. This application does not impose any special restrictions.
[0034] In this embodiment, the difference between the mean similarity values of semantic feature vectors is the absolute value of the difference. As another implementation method, based on the ability to measure the degree of difference between the mean similarity values of semantic feature vectors before and after each removal of fire data, the implementer may use other calculation methods, such as the square of the difference, the ratio, etc. This application does not impose any special restrictions.
[0035] It should be noted that: the larger the calculated dimensional perturbation, the more important the key features of the fire data in the corresponding dimension of the same dimension data set are semantically, and the stronger the ability of the fire data in this dimension to assess the repetition of the entire fire risk event; conversely, the smaller the calculated dimensional perturbation, the weaker the ability of the fire data in this dimension to assess the repetition of the entire fire risk event.
[0036] Furthermore, by analyzing the distribution of dimensional perturbations of all dimensions of fire data under a single fire risk event, the dimensional contribution ratio of each dimension of fire data under that single fire risk event is obtained, specifically: The proportion of the dimensional perturbation degree of each dimension of fire data under a single fire risk event to the dimensional perturbation degree of all dimensions of fire data is taken as the dimensional contribution ratio of each dimension of fire data under a single fire risk event.
[0037] It should be noted that: the greater the dimensional perturbation of each dimension of fire data under a single fire risk event, the higher the contribution of each dimension of fire data to the assessment of the single fire risk event, and the greater the dimensional contribution ratio of each dimension of fire data under a single fire risk event; conversely, the smaller the dimensional contribution ratio of each dimension of fire data under a single fire risk event.
[0038] It should be added that: if the amount of data in the same dimension of fire data under any fire risk event is less than or equal to 1, the dimension perturbation degree of fire data under any dimension under a single fire risk event is assigned to 0.
[0039] It should be added that: if the sum of the dimensional perturbations of all dimensions of fire data under a single fire risk event is 0, then the dimensional contribution ratio is equally distributed to all dimensions of fire data under a single fire risk event. That is, the dimensional contribution ratio of all dimensions of fire data under a single fire risk event is assigned to the reciprocal of the total number of dimensions of all fire data under a single fire risk event.
[0040] Step 2.2: Classify the words by using the semantic information of all words in all dimensions of fire data under a single fire risk event. By comparing the number of words in each category with the number of words under a single fire risk event, obtain the vocabulary scarcity of each category, and then obtain the distinguishability of each word in each category.
[0041] Data deduplication is a crucial step in data warehouse construction, and common methods include hash-based deduplication. The SimHash algorithm, for example, assigns weights to each word in a text for hash encoding, then weights, merges, and reduces the dimensionality of the hash values of each word. Finally, it calculates the similarity between two texts based on the dimensionality reduction results, thus achieving text deduplication. However, when deduplicating fire risk events, the content related to fire risk investigation logs, fire rectification logs, and fire assessment data for each fire risk event is relatively short and contains a large number of emergency fire-related technical terms. This causes distinctive words in the text describing fire risk events to be masked by other commonly used technical terms, resulting in a low weight for key distinguishing words when using the SimHash algorithm for deduplication, thus affecting the deduplication effect.
[0042] Based on the above analysis, for a single enterprise factory, the words are classified using the semantic information of all words within the fire data of all dimensions under a single fire risk event. The specific process is as follows: A clustering algorithm is used to cluster all word segments within all dimensions of fire data under a single fire risk event, where the distance metric is the difference between the semantic feature vectors of the word segments.
[0043] In this embodiment, the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm is used to cluster the word segments. Cosine distance is used to measure the difference between the semantic feature vectors of the word segments. The neighborhood radius in the clustering process is the median of the cosine distance between the semantic feature vectors of any two word segments within all dimensions of fire data under a single fire risk event. The minimum number of neighborhood points is 3. Based on the premise that the minimum number of neighborhood points is an integer within the range [2,5], the implementer can set the specific value of the minimum number of neighborhood points according to the actual situation. The DBSCAN algorithm and cosine distance are well-known technologies, and will not be described in detail in this application. As for other implementation methods, since the number of word segments and the number of word segmentation types are different in different fire risk events, it is necessary to use a clustering method that does not require specifying the number of clusters. Therefore, based on the premise of using a clustering method that does not require specifying the number of clusters to cluster the word segments, the implementer can choose other feasible technologies, such as density peak clustering algorithm, etc. This application does not impose any special restrictions.
[0044] It should be noted that the purpose of clustering word segments is that a fire risk event contains fire data recorded by different enterprise departments. For the same fire protection facility or the same hidden danger point, the string descriptions of different departments are different. Therefore, if word segmentation statistics are performed only based on the literal text, the sensitivity of some word segments may be overestimated. Therefore, we first cluster based on the semantic similarity of word segments, and then perform sensitivity statistical analysis on each category to improve the accuracy of the statistics.
[0045] Furthermore, since the word segments in each category are very similar in meaning, they may be describing the same topic. The statistical features of word segments in each category can be used to measure the distinguishing ability of word segments in each category when assessing the repetition of the entire fire risk event.
[0046] Based on the above analysis, for a single enterprise factory, the lexical scarcity of each category under a single fire risk event is obtained by comparing the number of word segments in each category with the total number of word segments under a single fire risk event. Specifically: Calculate the percentage of each type of word segmentation in all words under a single fire risk event; The scarcity of vocabulary for each category under a single fire risk event is negatively correlated with the proportion of the aforementioned quantities.
[0047] It should be noted that negative correlation means that the variables change in opposite directions; when one variable increases, the other decreases, and vice versa.
[0048] In this embodiment, the expression for the vocabulary scarcity of each type under a single fire risk event is as follows: In the formula, This indicates the vocabulary scarcity of the k-th category under a single fire risk event; ( ) denotes an exponential function with the natural constant as the base, used for quantification. and The negative correlation between them; This represents the proportion of the word segment in category k under a single fire risk event among all word segments under that single fire risk event. The natural constant is merely one embodiment of this application; implementers may replace it with other values greater than 1 according to actual circumstances, and this application does not impose any special restrictions.
[0049] It should be noted that the sensitivity of word segmentation in a particular category is measured by the proportion of each segmentation word in a single fire risk event. The smaller the calculated proportion, the lower the frequency of each segmentation word in a single fire risk event, indicating that the segmentation words in a particular category are more likely to be distinctive and sensitive words, and thus the greater the lexical scarcity of that category. Conversely, the larger the calculated proportion, the higher the frequency of each segmentation word in a single fire risk event, indicating that the segmentation words in a particular category have a lower role in distinguishing the repetition of fire risk events, and thus the smaller the lexical scarcity of that category.
[0050] Furthermore, for a single enterprise factory, the distinguishability of each word segment within each category under a single fire risk event is obtained by assessing the vocabulary scarcity of each category. The specific process is as follows: The scarcity of words in each category is amplified, and the amplified result is mapped to a positive number. The distinguishability of each word segment in each category is the rounded result of the positive number.
[0051] In this embodiment, the expression for the distinguishability of each word segment in each category is: In the formula, This represents the distinguishability of a single word in the k-th category under a single fire risk event; round() represents the rounding function, used to ensure that the distinguishability of the word is an integer; This represents a preset integer greater than 0, used to determine the degree to which vocabulary scarcity affects the distinguishability of word segmentation; This represents the lexical scarcity of the k-th class under a single fire risk event. The exponential function with base 2 is used to avoid situations where the lexical scarcity is too small, resulting in a discrimination factor of 0. 2 is merely one embodiment of this application; implementers can replace 2 with other values greater than 1 according to actual circumstances, and this application does not impose any special restrictions. This is an amplified result of vocabulary scarcity. This is the positive number obtained by mapping the amplified result.
[0052] In this embodiment, The value of 4, when there are too many word segments, can be used to prevent sensitive word segments from being overwhelmed by the large number of segments, thus increasing the efficiency. The value of .
[0053] Step 2.3: Encode each word segment under a single fire risk event, and combine the discriminative power and the dimensional contribution ratio to obtain the weighted sequence string of a single fire risk event.
[0054] Furthermore, for a single enterprise factory, each word segment under a single fire risk event is encoded, and combined with the discriminative power and the dimensional contribution ratio, a weighted sequence string of a single fire risk event is obtained. The specific process is as follows: The hash values of each word are weighted by the distinguishability of each word segment to obtain the sequence string of each word segment; The sequence strings of all segmented words within the fire protection data of each dimension are summed bit by bit to obtain the merged sequence string of the fire protection data of each dimension. The merged sequence strings of fire protection data from each dimension were normalized separately. The weighted sequence string of all dimensions of fire data under a single fire risk event is obtained by summing the combined sequence strings bit-wise and weighted by position. The weight of each data point in the combined sequence string of each dimension is its dimensional contribution ratio. A schematic diagram of the process for obtaining the weighted sequence string is shown below. Figure 2 As shown.
[0055] In this embodiment, the SimHash algorithm is used to encode each word segment. The SimHash algorithm is a well-known technology and will not be described in detail in this application.
[0056] In this embodiment, the specific process of weighting the hash values of each word by using the distinguishability of each word to obtain the sequence string of each word is as follows: For any word, traverse each bit of the hash value of the word. If any bit is 1, add the distinguishability of the word to the accumulator of the word. If any bit is 0, subtract the distinguishability of the word from the accumulator of the word.
[0057] In this embodiment, the maximum value normalization method is used to normalize the merged sequence string of fire protection data for each dimension. During the normalization process, the maximum value is the maximum value within the merged sequence string of fire protection data for each dimension. The maximum value normalization method is a well-known technology and will not be described in detail in this application.
[0058] Step 2.4: Deduplicat fire risk events by measuring the degree of difference between a single fire risk event and the weighted sequence strings of each fire risk event in its similar event set.
[0059] Furthermore, for a single enterprise factory, duplicate fire risk events are removed by assessing the degree of difference between a single fire risk event and the weighted sequence strings of other fire risk events within its similar event set. The specific process is as follows: Calculate the difference between a single fire risk event and the weighted sequence of each fire risk event in its similar event set. Fire risk events with a difference less than a preset difference threshold are considered as repeating events of the single fire risk event. A single fire risk event is grouped into a duplicate event set along with all its duplicate events. The arithmetic mean of the differences between each fire risk event in the duplicate event set and all other fire risk events is calculated. The fire risk event with the smallest arithmetic mean among all other fire risk events in the duplicate event set is retained, and the remaining fire risk events are removed, thus retaining the most representative fire risk event in the duplicate event set. A schematic diagram of the fire risk event deduplication process is shown below. Figure 3 As shown.
[0060] In this embodiment, the weighted sequence string of a single fire risk event is binarized by setting positions greater than 0 to 1 and others to 0, generating a binary weighted fingerprint string for the single fire risk event. The difference between the weighted sequence strings is specifically the Hamming distance between the binary weighted fingerprint strings. The calculation of the Hamming distance is a well-known technique. As other implementations, implementers can use other existing feasible techniques to measure the degree of difference between binary weighted fingerprint strings; this application does not impose any special restrictions. The preset difference threshold is 3, and the value of the preset difference threshold is derived from experimental data. The reason for binarizing the weighted sequence string is that, mathematically, the Hamming distance is only applicable to binary sequences of equal length.
[0061] Based on the deduplication method for fire risk events of individual enterprise factories under the enterprise side, deduplication is performed on the fire risk events of each enterprise factory under the enterprise side, and deduplication is also performed on the fire risk events of each enterprise factory under the regulatory side.
[0062] Step 3: Build a safety data warehouse based on the remaining fire protection data after deduplication.
[0063] During the data warehouse construction process, a database management system was selected for construction, and a star schema was adopted for the data model. An ETL tool was used to load the remaining fire safety data from both the enterprise and regulatory ends after deduplication of fire risk events into the database. The data in the database was then encrypted to ensure data security, thus completing the construction of the secure data warehouse. The construction of the data warehouse is a well-known technology in the field of big data, and the specific process will not be elaborated upon in this application.
[0064] In this embodiment, the database management system adopts MySQL. MySQL is a well-known technology and will not be described in detail in this application. As other implementation methods, under the premise that it can be used as the underlying database management system of the data warehouse, the implementer may adopt other existing feasible technologies, such as Oracle, etc. This application does not impose any special restrictions.
[0065] In this embodiment, the AES (Advanced Encryption Standard) algorithm is used for data encryption. The AES algorithm is a well-known technology and will not be described in detail in this application. As other implementation methods, implementers may use other existing feasible technologies on the basis of data encryption, and this application does not impose any special restrictions.
[0066] Based on the same inventive concept as the above method, this application embodiment also provides a secure data warehouse construction system that supports multi-dimensional combined queries, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above-described secure data warehouse construction methods that supports multi-dimensional combined queries.
[0067] In summary, this application analyzes the change in similarity before and after individual dimensions of fire data in a set of similar events are removed one by one, constructs the dimension perturbation degree and converts it into the dimension contribution ratio. This can quantify the importance of each dimension of fire data in the deduplication of fire risk events, assign higher weights to dimensions containing key discrimination information and lower weights to those containing less information, so that the dimensions with distinctiveness can play a greater role in the deduplication process. Furthermore, based on the semantic information of word segmentation, clustering and comparing the proportion of each type of word segmentation, constructing vocabulary scarcity and converting it into distinguishability, can identify sensitive words under fire risk events. Higher weights are assigned to key descriptive words that are more likely to be distinctive, and lower weights are assigned to those that are less likely to be distinctive. This effectively solves the problem of a large number of professional terms obscuring key identification words and helps to improve the deduplication effect of fire risk events. Furthermore, by combining the dimensional contribution ratio and the discriminative power, the word segmentation encoding is weighted twice to obtain a weighted sequence string of fire risk events. The dimensional contribution ratio improves the discriminative power of fire data from the data dimension level, while the discriminative power improves the discriminative power of word segmentation from the vocabulary level. The synergistic effect of the two enhances the robustness of the generated weighted sequence string in describing differences, providing a more reliable feature representation for subsequent deduplication of fire risk events. Furthermore, by calculating the degree of difference in weighted sequence strings to determine duplicate fire risk events and deduplicating them, redundant data can be accurately identified and removed. This effectively eliminates duplicate records caused by network retransmission and differences in descriptions across multiple ports, while retaining the most representative events. After the deduplication operation, storage and computing overhead can be significantly reduced, query efficiency can be improved, and the interference of duplicate data on fire risk assessment results can be eliminated, ensuring the accuracy of the assessment. This lays the foundation for building a high-quality, highly reliable security data warehouse, thereby significantly improving the reliability and query efficiency of the data warehouse.
[0068] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0069] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from its essential characteristics. Therefore, the embodiments described above should be considered exemplary and non-limiting in all respects.
Claims
1. A method for constructing a secure data warehouse that supports multi-dimensional combined queries, characterized in that, The method includes the following steps: Acquire multi-dimensional fire safety data for each enterprise and factory under each fire risk event at each port, including the enterprise port and the regulatory port; Fire risk events are referred to as events. For each enterprise factory under each port, a set of similar events is obtained by filtering the semantic information of fire data across all dimensions of a single event and other events. By measuring the change in similarity between fire data of a single dimension in the set of similar events and fire data of the same dimension under the single event before and after removing fire data of a single dimension one by one, the dimensional perturbation degree of fire data of each dimension under the single event is obtained, and then the dimensional contribution ratio of fire data of each dimension under the single event is obtained. By using the semantic information of all words within all dimensions of fire protection data under a single event, the words are classified. By comparing the number of words in each category with the number of words under a single event, the lexical scarcity of each category is obtained, and then the distinguishability of each word in each category is obtained. Each word segment under a single event is encoded, and the weighted sequence string of the single event is obtained by combining the discriminative power and the dimensional contribution ratio; the event is deduplicated by the degree of difference between the weighted sequence strings of a single event and each event in its similar event set. A safety data warehouse is built based on the remaining fire protection data after deduplication. The process of obtaining the dimensional perturbation degree is as follows: Extract fire data of any dimension under a single event, and form a data set of the same dimension from the set of similar events containing all fire data of the same dimension as the fire data of any dimension. Calculate the mean of the semantic feature vector similarity between any dimension of fire data and all fire data in the same dimension data set; remove fire data one by one from the same dimension data set without replacement in the order of timestamp from earliest to latest, until only one fire data remains in the same dimension data set; after each fire data is removed, recalculate the mean of the semantic feature vector similarity between any dimension of fire data and all remaining fire data in the same dimension data set; Calculate the difference between the mean values before and after each removal of fire data; use the average of the differences obtained before and after all removals of fire data as the dimensional perturbation of any dimension of fire data under a single event. The dimensional contribution ratio is the percentage of the dimensional perturbation of each dimension of fire protection data in a single event in the total dimensional perturbation of all dimensions of fire protection data.
2. The method for constructing a secure data warehouse supporting multi-dimensional combined queries as described in claim 1, characterized in that, The filtering process for the set of similar events is as follows: The semantic feature vectors of all dimensions of fire protection data for each event are concatenated to obtain the semantic vector of each event; Events whose semantic vectors are more than a preset threshold with a single event are denoted as similar events to the single event. Group all similar events of a single event into a set of similar events; The preset threshold is the average similarity of the semantic vectors between a single event and all other events.
3. The method for constructing a secure data warehouse supporting multi-dimensional combined queries as described in claim 1, characterized in that, The word segmentation classification includes: using a clustering algorithm to cluster all words within all dimensions of fire data under a single event, wherein the distance metric is the difference between the semantic feature vectors of the words.
4. The method for constructing a secure data warehouse supporting multi-dimensional combined queries as described in claim 1, characterized in that, The process of obtaining the vocabulary scarcity is as follows: Calculate the percentage of each type of word segmentation in all words under a single event; The scarcity of the vocabulary is negatively correlated with the proportion of the quantity.
5. The method for constructing a secure data warehouse supporting multi-dimensional combined queries as described in claim 1, characterized in that, The process of obtaining the discrimination index is as follows: The vocabulary scarcity is amplified, and the amplification result is mapped to a positive number; The discrimination index is the rounded integer result of the positive number.
6. The method for constructing a secure data warehouse supporting multi-dimensional combined queries as described in claim 1, characterized in that, The process of obtaining the weighted sequence string is as follows: The hash values of each word are weighted using the aforementioned distinguishability to obtain the sequence string of each word; The sequence strings of all segmented words within the fire protection data of each dimension are summed bit by bit to obtain the merged sequence string of the fire protection data of each dimension. The merged sequence strings of fire protection data from each dimension were normalized separately. The weighted sequence string of all dimensions of fire protection data under a single event is obtained by summing the combined sequence string bit by bit. The weight of each data in the combined sequence string of fire protection data of each dimension is the contribution ratio of that dimension.
7. The method for constructing a secure data warehouse supporting multi-dimensional combined queries as described in claim 1, characterized in that, The event deduplication includes: Calculate the difference between a single event and each event in its set of similar events, and identify events whose difference with a single event is less than a preset difference threshold as duplicate events of the single event. Take a single event and all its repeating events to form a repeating event set, and calculate the arithmetic mean of the difference between each event in the repeating event set and all the other events; keep the event with the smallest arithmetic mean among the repeating event set and remove the rest of the events.
8. A secure data warehouse construction system supporting multi-dimensional combined queries, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the secure data warehouse construction method that supports multi-dimensional combined queries as described in any one of claims 1-7.
Citation Information
Patent Citations
Task-independent interpreter learning method and system
CN121745164A
Computerized system for simulating the likelihood of technology change incidents
US20170243131A1