Danger source identification method based on cluster
By identifying hazardous sources in office areas based on clustering clustering, the problem of inaccurate identification in the prior art is solved, and the rapid and efficient identification of potential hazardous sources is achieved, and the risk of accidents is reduced.
Patent Information
- Application Number
- CN202411932342.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to effectively identify and control potential hazard sources in office areas, resulting in frequent accidents.
The cluster cluster-based hazard source identification method is adopted to identify potential hazard sources through data collection, preprocessing, feature selection and clustering analysis, and the accuracy of hazard sources and the probability of false alarms is improved through the analysis of cluster clusters and the judgment of early warning values.
It realizes rapid, efficient and accurate identification of hazardous sources in office areas, reducing the risk and losses of accidents.
Smart Images

Figure CN119939291A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of safety management and control, and specifically relates to a hazard source identification method based on clustering. Background Art
[0002] According to data released by the Ministry of Emergency Management in July 2024, there were 9,229 production safety accidents and 8,507 deaths in the first half of the year. Most of these accidents were caused by the failure to identify and control dangerous sources in time. At present, the mainstream form of office area layout in domestic enterprises is mostly to divide different functional areas in commercial buildings and concentrate a large number of people in the same area. As an important place for employees' daily work and communication, the safety of the office area is directly related to the life safety and physical health of employees. However, in a seemingly calm working environment, there are many potential sources of danger. These sources of danger include electrical sources of danger. Computers, printers, copiers, etc. may cause fires, electric shocks and other safety accidents due to improper operation or aging damage; environmental sources of danger, flammable and explosive items, chemicals, building materials, etc. may cause explosions, poisoning and diseases, seriously endangering the safety and health of employees; other types of dangerous sources, wire installation, high-altitude operations, etc., may also cause accidents, resulting in personnel safety and property losses. These sources of danger may be caused by improper use of equipment, unreasonable environmental layout, human factors and other reasons, posing a serious threat to the safety of office areas.
[0003] Many accidents are caused by the failure to promptly discover and control hazardous sources. Through hazard identification, we can grasp information such as the type, quantity, distribution and degree of hazard, and provide a basis for formulating targeted preventive measures. Therefore, the identification and control of hazardous sources is an important means to prevent accidents and reduce the extent of accident losses. Currently, the commonly used hazard identification methods mainly include safety inspection methods, pre-hazard analysis, and operating condition hazard assessment methods. These methods are suitable for different scenarios and needs, but in actual applications, the identification of hazardous sources is a continuous process that needs to be continuously updated and improved as the system or environment changes. Summary of the invention
[0004] 1. Technical issues to be resolved
[0005] The technical problem to be solved by the present invention is: how to provide a hazard source identification method based on clustering.
[0006] (II) Technical solution
[0007] To solve the above technical problems, the present invention provides a hazard source identification method based on clustering clusters. The method obtains a data set, classifies and establishes different clustering cluster data samples, collects hazard source data from the clustering cluster data set as target hazard source data, and judges whether the collected data or the data to be judged contains hazard information or contains hazard sources; all hazard source data whose distance to the target hazard source data is less than or equal to a preset neighborhood radius are screened out from the data set to be classified; if there is hazard source data that does not belong to any clustering cluster, the hazard source data that does not belong to any clustering cluster is regarded as abnormal hazard source data, and based on the relationship between the abnormal hazard source data and the corresponding warning value, it is judged whether the corresponding position of the hazard source data is a hazard source leakage point; thereby improving the accuracy of hazard sources and reducing the probability of false alarms.
[0008] The method specifically includes:
[0009] Step 1: Hazard source identification execution steps;
[0010] Step 2: Clustering algorithm classification step.
[0011] The steps for executing the hazard source identification in step 1 include:
[0012] Step 11: Data collection and processing: Collect relevant hazardous information data in the building to be tested and the area, including historical accident records, environmental parameters, and equipment status; then, pre-process the data, including data cleaning, missing value filling, standardization or normalization, to ensure data quality and consistency;
[0013] Step 12: Feature selection: Select features that have a key impact on hazard identification from the processed data; these behaviors should be able to reflect the essential attributes or behavior patterns of the hazard;
[0014] Step 13: Cluster analysis: Use clustering algorithms to perform cluster analysis on the selected features and divide the data into several clusters; the objects within each cluster have high similarities in features, while the objects between different clusters have obvious differences;
[0015] Step 14: Hazard source identification: Identify potential hazard sources by analyzing the characteristics of clusters; this involves statistical analysis of clusters, trend prediction, and comparison with known hazard sources; during the identification process, focus on clusters with similar characteristics to known hazard sources or abnormal behavior patterns;
[0016] Step 15: Evaluation and verification: Evaluate and verify the identified hazards to confirm their authenticity and danger;
[0017] Step 16: Develop response measures: Develop corresponding response measures and plans for the identified hazards, including hazard source monitoring, establishment of early warning systems, and improvement of emergency response mechanisms, to ensure that when hazards occur, they can be dealt with quickly and effectively.
[0018] Among them, in the step 1, regarding the method for identifying hazardous sources, a classified data set is established, and hazardous source data is selected as target hazardous source data from a given sample data source to determine whether the target hazardous source data is a hazardous core point; if the target hazardous source data is a hazardous core point, a cluster is established, and the target hazardous source data is added to the cluster, and all hazardous source data whose distance from the target hazardous source data is less than a preset field radius are screened out from the data set to be classified, and the screened hazardous source data are added to the corresponding cluster; if there is no hazardous source data, the hazardous source data that does not belong to any cluster is regarded as abnormal hazardous source data, and whether the corresponding position of the data is a hazardous source leakage point is determined based on the abnormal hazardous source data and the corresponding warning value; this idea is repeated to traverse all hazardous source data in the sample data set to be classified.
[0019] Among them, in the clustering algorithm classification step of step 2:
[0020] The clustering algorithm calculates the similarity or distance between data points; divides the data points into different clusters; there are many common similarity or distance metrics, and the step 2 uses the Euclidean distance as the metric; the goal of the clustering algorithm is to divide the data points in the data set into several clusters according to these metrics, so that the similarity of the data points within the cluster is maximized and the similarity of the data points between clusters is minimized; the core idea of the definition of clusters and the classification and identification of the data to be tested is to traverse the sample data, establish clusters centered on different core objects according to the feature values, and classify and output the data;
[0021] The original data set to be processed needs to be standardized and preprocessed using the following formula:
[0022]
[0023] Among them, x i is the internal object of the standardized dataset, x is the internal object of the original dataset, and x min is the minimum value in the data set, x max is the maximum value in the data set;
[0024] The core idea of the cluster center determination process is to minimize the sum of the Euclidean distances between each data point and its cluster point. The number of clusters, that is, the K value, can be determined by the elbow rule, and then the sum of the Euclidean distances is calculated by traversing the data. The calculation formula is as follows:
[0025]
[0026] Where x is the data point, c i is the i-th cluster center, d is the data dimension, x j and c ij are x and c respectively i The value in the jth dimension;
[0027] The clustering algorithm needs to be iteratively updated, recalculating each cluster center, and taking the mean of all data points in the cluster as the new cluster center for iterative update. The calculation formula is as follows:
[0028]
[0029] Among them, S i is the set of data points of the ith cluster, |S i | is the number of data points in the data set;
[0030] The classification standard is to calculate the distance between the data to be tested and the clusters of various types of dangerous data, and determine whether it meets the condition of being less than or equal to the radius of the cluster area:
[0031] dist(p,q)≤E ps ,q∈D
[0032] Where dist(p,q) represents the Euclidean distance between the target hazard source data p and other hazard source data q, E ps Represents the preset neighborhood radius, which is the length referenced when determining the density area range based on the current target hazard source data.
[0033] The clustering method identifies and classifies sample data. The execution process is as follows: Figure 1 shown.
[0034] Among them, the hazard source identification method establishes a model for the regional environment, conducts high-dimensional visualization exploration of the analyzed and processed hazard source data, and displays the visualization results, so that users can more intuitively understand the abnormal hazard source data.
[0035] The method is used to identify the sources of danger in network systems and office environments.
[0036] (III) Beneficial effects
[0037] Compared with the prior art, the present invention can establish data models according to different scenarios, process the data collected by Internet sensing devices through data identification and classification, and can quickly, efficiently and accurately identify the source of danger. The present invention has the following key innovations:
[0038] (1) Based on clustering, hazard source identification can be realized. The clustering results can be reasonably explained based on characteristic analysis, and the hazard source types and characteristics of each cluster can be clarified. Based on the clustering results, objects or events with similar hazard source characteristics can be identified as potential hazard sources;
[0039] (2) Clustering algorithms can be selected according to the situation. Clusters can be flexibly selected and established according to the scenario. Reasonable clustering effects can be obtained by reasonably setting the number of clusters K (elbow rule), the minimum point parameter minPts, the preset neighborhood radius Eps, etc.
[0040] (3) Data preprocessing and feature extraction: extracting effective features from complex data sets for cluster analysis. This includes operations such as data cleaning, standardization, and dimensionality reduction to ensure that the clustering algorithm can accurately identify objects or events with similar dangerous characteristics.
[0041] The present invention has the following beneficial effects:
[0042] (1) Automation and efficiency: The clustering method can classify data without manual intervention, greatly improving the efficiency of identifying hazard sources;
[0043] (2) Data-driven: grouping based on the similarities and differences of the data itself can objectively reflect the characteristics and distribution patterns of hazard sources;
[0044] (3) Strong flexibility: This method can process large-scale data sets and is not limited by data differences and distribution conditions. It can identify hazard source data sets in various complex situations.
[0045] (4) Interpretability: Clustering results are usually interpretable to a certain extent, which helps to understand the internal connections and potential risks between hazard sources;
[0046] (5) Decision support: Cluster analysis can be used to identify high-risk groups of hazardous sources, providing strong support for the formulation of targeted risk-based prevention and control strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 This is a flow chart of cluster establishment and data classification calculation in the technical solution of the present invention. DETAILED DESCRIPTION
[0048] In order to make the purpose, content, and advantages of the present invention more clear, the specific implementation methods of the present invention are further described in detail below in conjunction with the accompanying drawings and examples.
[0049] The present invention proposes to use the idea of clustering to identify dangerous source data, such as Figure 1As shown, it is possible to extract data features based on the collected scene data and existing experience data, identify various types of hazard source data, strengthen the identification and control of hazard sources in the scene, and better prevent safety accidents.
[0050] To solve the above technical problems, the present invention provides a hazard source identification method based on clustering clusters. The method obtains a data set, classifies and establishes different clustering cluster data samples, collects hazard source data from the clustering cluster data set as target hazard source data, and judges whether the collected data or the data to be judged contains hazard information or contains hazard sources; all hazard source data whose distance to the target hazard source data is less than or equal to a preset neighborhood radius are screened out from the data set to be classified; if there is hazard source data that does not belong to any clustering cluster, the hazard source data that does not belong to any clustering cluster is regarded as abnormal hazard source data, and based on the relationship between the abnormal hazard source data and the corresponding warning value, it is judged whether the corresponding position of the hazard source data is a hazard source leakage point; thereby improving the accuracy of hazard sources and reducing the probability of false alarms.
[0051] The method specifically includes:
[0052] Step 1: Hazard source identification execution steps;
[0053] Step 2: Clustering algorithm classification step.
[0054] The steps for executing the hazard source identification in step 1 include:
[0055] Step 11: Data collection and processing: Collect relevant hazardous information data in the building to be tested and the area, including historical accident records, environmental parameters, and equipment status; then, pre-process the data, including data cleaning, missing value filling, standardization or normalization, to ensure data quality and consistency;
[0056] Step 12: Feature selection: Select features that have a key impact on hazard identification from the processed data; these behaviors should be able to reflect the essential attributes or behavior patterns of the hazard;
[0057] Step 13: Cluster analysis: Use clustering algorithms to perform cluster analysis on the selected features and divide the data into several clusters; the objects within each cluster have high similarities in features, while the objects between different clusters have obvious differences;
[0058] Step 14: Hazard source identification: Identify potential hazard sources by analyzing the characteristics of clusters; this involves statistical analysis of clusters, trend prediction, and comparison with known hazard sources; during the identification process, focus on clusters with similar characteristics to known hazard sources or abnormal behavior patterns;
[0059] Step 15: Evaluation and verification: Evaluate and verify the identified hazards to confirm their authenticity and danger;
[0060] Step 16: Develop response measures: Develop corresponding response measures and plans for the identified hazards, including hazard source monitoring, establishment of early warning systems, and improvement of emergency response mechanisms, to ensure that when hazards occur, they can be dealt with quickly and effectively.
[0061] Among them, in the step 1, regarding the method for identifying hazardous sources, a classified data set is established, and hazardous source data is selected as target hazardous source data from a given sample data source to determine whether the target hazardous source data is a hazardous core point; if the target hazardous source data is a hazardous core point, a cluster is established, and the target hazardous source data is added to the cluster, and all hazardous source data whose distance from the target hazardous source data is less than a preset field radius are screened out from the data set to be classified, and the screened hazardous source data are added to the corresponding cluster; if there is no hazardous source data, the hazardous source data that does not belong to any cluster is regarded as abnormal hazardous source data, and whether the corresponding position of the data is a hazardous source leakage point is determined based on the abnormal hazardous source data and the corresponding warning value; this idea is repeated to traverse all hazardous source data in the sample data set to be classified.
[0062] Among them, in the clustering algorithm classification step of step 2:
[0063] The clustering algorithm calculates the similarity or distance between data points; divides the data points into different clusters; there are many common similarity or distance metrics, and the step 2 uses the Euclidean distance as the metric; the goal of the clustering algorithm is to divide the data points in the data set into several clusters according to these metrics, so that the similarity of the data points within the cluster is maximized and the similarity of the data points between clusters is minimized; the core idea of the definition of clusters and the classification and identification of the data to be tested is to traverse the sample data, establish clusters centered on different core objects according to the feature values, and classify and output the data;
[0064] The original data set to be processed needs to be standardized and preprocessed using the following formula:
[0065]
[0066] Among them, x i is the internal object of the standardized dataset, x is the internal object of the original dataset, and x min is the minimum value in the data set, x max is the maximum value in the data set;
[0067] The core idea of the cluster center determination process is to minimize the sum of the Euclidean distances between each data point and its cluster point. The number of clusters, that is, the K value, can be determined by the elbow rule, and then the sum of the Euclidean distances is calculated by traversing the data. The calculation formula is as follows:
[0068]
[0069] Where x is the data point, c i is the i-th cluster center, d is the data dimension, x j and c ij are x and c respectively i The value in the jth dimension;
[0070] The clustering algorithm needs to be iteratively updated, recalculating each cluster center, and taking the mean of all data points in the cluster as the new cluster center for iterative update. The calculation formula is as follows:
[0071]
[0072] Among them, S i is the set of data points of the ith cluster, |S i | is the number of data points in the data set;
[0073] The classification standard is to calculate the distance between the data to be tested and the clusters of various types of dangerous data, and determine whether it meets the condition of being less than or equal to the radius of the cluster area:
[0074] dist(p,q)≤E ps ,q∈D
[0075] Where dist(p,q) represents the Euclidean distance between the target hazard source data p and other hazard source data q, E ps Represents the preset neighborhood radius, which is the length referenced when determining the density area range based on the current target hazard source data.
[0076] The clustering method identifies and classifies sample data. The execution process is as follows: Figure 1 shown.
[0077] Among them, the hazard source identification method establishes a model for the regional environment, conducts high-dimensional visualization exploration of the analyzed and processed hazard source data, and displays the visualization results, so that users can more intuitively understand the abnormal hazard source data.
[0078] The method is used to identify the sources of danger in network systems and office environments.
[0079] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A cluster-based hazard source identification method, characterized in that: The method acquires a data set, classifies and establishes different cluster data samples, collects hazard source data from the cluster data set as target hazard source data, and determines whether the collected data or the data to be determined contains hazard information or contains a hazard source; and filters out all hazard source data whose distance to the target hazard source data is less than or equal to a preset neighborhood radius from the data set to be classified; If there is hazardous source data that does not belong to any cluster, the hazardous source data that does not belong to any cluster will be regarded as abnormal hazardous source data, and based on the relationship between the abnormal hazardous source data and the corresponding warning value, it is judged whether the corresponding position of the hazardous source data is a hazardous source leakage point; thereby improving the accuracy of hazardous sources and reducing the probability of false alarms.
2. The cluster-based hazard source identification method according to claim 1, characterized in that: The method specifically comprises: Step 1: Hazard source identification execution steps; Step 2: Clustering algorithm classification step.
3. The cluster-based hazard source identification method according to claim 2, characterized in that: The steps for executing the hazard source identification in step 1 include: Step 11: Data collection and processing: Collect relevant hazardous information data in the building to be tested and the area, including historical accident records, environmental parameters, and equipment status; then, pre-process the data, including data cleaning, missing value filling, standardization or normalization, to ensure data quality and consistency; Step 12: Feature selection: Select features that have a key impact on hazard identification from the processed data; these behaviors should be able to reflect the essential attributes or behavior patterns of the hazard; Step 13: Cluster analysis: Use clustering algorithms to perform cluster analysis on the selected features and divide the data into several clusters; the objects within each cluster have high similarities in features, while the objects between different clusters have obvious differences; Step 14: Hazard source identification: Identify potential hazard sources by analyzing the characteristics of clusters; this involves statistical analysis of clusters, trend prediction, and comparison with known hazard sources; during the identification process, focus on clusters with similar characteristics to known hazard sources or abnormal behavior patterns; Step 15: Evaluation and verification: Evaluate and verify the identified hazards to confirm their authenticity and danger; Step 16: Develop response measures: Develop corresponding response measures and plans for the identified hazards, including hazard source monitoring, establishment of early warning systems, and improvement of emergency response mechanisms, to ensure that when hazards occur, they can be dealt with quickly and effectively.
4. The cluster-based hazard source identification method according to claim 3, characterized in that: In the step 1, regarding the method for identifying hazardous sources, a classification data set is established, and hazardous source data is selected as target hazardous source data from a given sample data source to determine whether the target hazardous source data is a hazardous core point; if the target hazardous source data is a hazardous core point, a cluster is established, the target hazardous source data is added to the cluster, and all hazardous source data whose distance to the target hazardous source data is less than a preset field radius are screened out from the data set to be classified, and the screened hazardous source data are added to the corresponding cluster; if no hazardous source data exists, the hazardous source data that does not belong to any cluster is regarded as abnormal hazardous source data, and whether the corresponding position of the data is a hazardous source leakage point is determined based on the abnormal hazardous source data and the corresponding warning value; this idea is repeated to traverse all hazardous source data in the sample data set to be classified.
5. The method for identifying danger sources based on clustering as claimed in claim 4, characterized in that: In the clustering algorithm classification step of step 2: The clustering algorithm calculates the similarity or distance between data points; divides the data points into different clusters; the step 2 uses the Euclidean distance as the metric; the goal of the clustering algorithm is to divide the data points in the data set into several clusters according to these metrics, so that the similarity of the data points within the cluster is maximized and the similarity of the data points between clusters is minimized; the core idea of the definition of clusters and the classification and identification of the data to be tested is to traverse the sample data, establish clusters centered on different core objects according to the characteristic values, and classify and output the data; The original data set to be processed needs to be standardized and preprocessed using the following formula: Among them, x i is the internal object of the standardized dataset, x is the internal object of the original dataset, and x min is the minimum value in the data set, x max is the maximum value in the data set; The core idea of the cluster center determination process is to minimize the sum of the Euclidean distances between each data point and its cluster point. The number of clusters, that is, the K value, can be determined by the elbow rule, and then the sum of the Euclidean distances is calculated by traversing the data. The calculation formula is as follows: Where x is the data point, c i is the i-th cluster center, d is the data dimension, x j and c ij are x and c respectively i The value in the jth dimension; The clustering algorithm needs to be iteratively updated, recalculating each cluster center, and taking the mean of all data points in the cluster as the new cluster center for iterative update. The calculation formula is as follows: Among them, S i is the set of data points of the ith cluster, |S i | is the number of data points in the data set; The classification standard is to calculate the distance between the data to be tested and the clusters of various types of dangerous data, and determine whether it meets the condition of being less than or equal to the radius of the cluster area: dist(p,q)≤E ps ,q∈D Where dist(p,q) represents the Euclidean distance between the target hazard source data p and other hazard source data q, E ps Represents the preset neighborhood radius, which is the length referenced when determining the density area range based on the current target hazard source data.
6. The method for identifying danger sources based on clustering according to claim 5, characterized in that: The hazard source identification method establishes a model for the regional environment, performs high-dimensional visualization exploration on the analyzed and processed hazard source data, and displays the visualization results, so as to facilitate users to more intuitively recognize the abnormal hazard source data.
7. The method for identifying dangerous sources based on clustering according to claim 6, characterized in that: The method is used to identify the danger sources of network systems and office environments.
8. The cluster-based hazard source identification method according to claim 6, characterized in that: The method belongs to the technical field of safety management and control.
9. The cluster-based hazard source identification method according to claim 6, characterized in that: The method can establish data models according to different scenarios, process the data collected by Internet sensing devices through data identification and classification, and can identify dangerous sources quickly, efficiently and accurately.
10. The method for identifying dangerous sources based on clustering according to claim 6, characterized in that: The method realizes hazard source identification based on clustering clusters, can reasonably explain the clustering results according to characteristic analysis, clarify the hazard source type and characteristics of each cluster, and according to the clustering results, can identify objects or events with similar hazard source characteristics as potential hazard sources.