Mine user portrait and classification method based on security big data

By using a user profiling and classification method based on safety big data, the problem of insufficient data integration and analysis in mine management has been solved. A comprehensive and accurate user profile has been constructed, which has improved the ability to identify and prevent mine safety risks and enhanced the level of safe production.

CN121997186APending Publication Date: 2026-05-08YUNNAN KUNGANG ELECTRONICS INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YUNNAN KUNGANG ELECTRONICS INFORMATION TECH CO LTD
Filing Date
2025-11-28
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing mine management methods are inadequate in terms of data integration and analysis, making it difficult to effectively process and comprehensively analyze multi-dimensional information on miners, equipment, and the environment. This results in the inability to build comprehensive and accurate user profiles, leading to limited capabilities in identifying and controlling mine safety risks.

Method used

This paper adopts a user profiling and classification method based on safety big data. By collecting and preprocessing multi-source data from mines, a basic dataset of user profiles is constructed. A feature weight optimization method that integrates dynamic weighting and cluster analysis is used to generate user profile labels. Finally, a random forest classification model and DS evidence theory are used to calculate the safety risk level.

Benefits of technology

It has achieved a comprehensive and accurate profile of mining users, effectively identified and controlled safety risks, and improved the level of safe production in mines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997186A_ABST
    Figure CN121997186A_ABST
Patent Text Reader

Abstract

The invention relates to a mine user portrait and classification method based on safety big data, and belongs to the technical field of mine safety and data processing. The method comprises the steps of establishing a mine user portrait basic data set, constructing user portrait labels and classifying safety risks. According to the invention, through collection, integration and analysis of mine multi-dimensional data, a comprehensive and accurate mine user portrait is constructed, and scientific classification is carried out, so that powerful support is provided for mine safety management, safety risks are effectively identified, prevented and controlled, and the mine safety production level is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of mine safety and data processing technology, specifically relating to a method for profiling and classifying mine users based on safety big data. Background Technology

[0002] With the continuous expansion of mining scale and the advancement of intelligent mine construction, mine safety management faces numerous challenges. The mining environment is complex and ever-changing, mining operations are diverse, and equipment operating status changes in real time. How to comprehensively and accurately grasp the situation of all elements in the mine and achieve precise safety management has become an urgent problem to be solved in the mining industry.

[0003] Existing mine management methods have shortcomings in data integration and analysis, making it difficult to effectively process and comprehensively analyze multi-dimensional information related to miners, equipment, and the environment. The inability to construct comprehensive and accurate user profiles results in limited capabilities for identifying and controlling mine safety risks, failing to meet the demands of modern mine safety management. Therefore, a new method is urgently needed to utilize safety big data to construct and classify mine user profiles to improve mine safety management. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies and provide a method for profiling and classifying mine users based on security big data.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for profiling and classifying mine users based on safety big data includes the following steps: S1: Establish a basic dataset for mine user profiles: By collecting and preprocessing multi-source data from mines, effective feature data is obtained, thereby forming a basic dataset for mine user profiles; the multi-source data includes miner data, equipment data, environmental data, and text data; S2: Constructing a user profile tagging system: Based on the basic dataset of mine user profiles obtained from S1, a feature weight optimization method that combines dynamic weighting and cluster analysis is adopted to generate user profile tags covering miner behavior characteristics, equipment operating status, environmental safety indicators, and text-derived data. S3: Safety Risk Classification: User profile tags are quantified into feature vectors, input into a random forest classification model, and then the comprehensive risk value is calculated based on the analytic hierarchy process and the DS evidence theory to finally obtain the safety risk level of the mining user.

[0006] Furthermore, preferably, in S1, the preprocessing includes data cleaning, missing value completion, noise reduction, text data structuring, and data association; the data association is to associate the operation records in the miner data with the operation trajectory.

[0007] Furthermore, preferably, in S2, an improved weight calculation method is used to assign weights to the feature data, specifically: in, The adjusted weights are the weights of the j-th data item in data set i. This represents the number of records containing the j-th data; c is the weighting factor; N represents the total number of data records in the entire dataset; The result of the dynamically weighted adjustment of data frequency is as follows: In the formula, This refers to the frequency of characteristic data in critical security incidents; The frequency of the characteristic data in daily monitoring data; The frequency of the feature data in historical data; This refers to the weighting coefficient for critical security events. Weighting coefficients for daily monitoring data; Based on the clustering analysis results, a key feature data set S is identified from all features; for features belonging to set S, initial weights are assigned... Perform readjustment and calculate its final weight ω. ij′ ′: Where γ is the re-adjusted weighting coefficient for cluster analysis. The specific formula is: Where S is the set of key feature data. This represents the i-th feature data in the dataset.

[0008] Furthermore, preferably, in S2: The generation of miner behavior feature labels includes using the K-means algorithm to cluster mining operation types and / or using the PrefixSpan algorithm to mine dangerous behavior sequences of miners. The generation of equipment operating status labels includes using the isolated forest algorithm to detect abnormal equipment health status and / or using the Apriori algorithm to mine association rules for equipment failure risks; The generation of environmental safety indicator labels includes threshold determination based on real-time parameters and / or trend early warning using ARIMA time series models; The generation of text-derived tags includes calculating text keyword weights using an improved TF-IWF algorithm, using the LDA topic model for text security topic identification and feature enhancement, and finally generating tags based on the weighted ranking results.

[0009] Furthermore, preferably, the specific method of S3 is as follows: Quantize user profile tags into feature vectors; Using the random forest algorithm, an initial mapping relationship from label feature vectors to security risk levels is learned; The weights of four dimensions—miner behavior, equipment status, environmental indicators, and text information—were determined using the analytic hierarchy process (AHP). The four dimensions are treated as four independent sources of evidence. The risk level probability distributions output by each source of evidence are fused using the Dempster combination rule of the DS evidence theory. The comprehensive risk value is calculated based on the fused probability distribution, and the final security risk level is determined based on a preset threshold.

[0010] Furthermore, preferably, it also includes S4: Dynamic Update: Periodically update the user profile tags and random forest model to adapt to the dynamic changes in the mine safety status.

[0011] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the program to implement the steps of the above-described method for profiling and classifying mine users based on security big data.

[0012] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the above-described method for profiling and classifying mine users based on security big data.

[0013] Compared with the prior art, the beneficial effects of this invention are as follows: This invention provides a method for profiling and classifying mine users based on safety big data. By collecting, integrating and analyzing multi-dimensional mine data, a comprehensive and accurate mine user profile is constructed and scientifically classified, thereby providing strong support for mine safety management, effectively identifying and preventing safety risks, and improving the level of mine safety production. Attached Figure Description

[0014] Figure 1 This is a flowchart of the mine user profiling and classification method based on security big data according to the present invention. Figure 2 This is a schematic diagram of the electronic device structure of the present invention. Detailed Implementation

[0015] The present invention will now be described in further detail with reference to the embodiments.

[0016] Those skilled in the art will understand that the following embodiments are for illustrative purposes only and should not be construed as limiting the scope of the invention. Where specific techniques or conditions are not specified in the embodiments, they are performed in accordance with the techniques or conditions described in the literature in the field or according to the product instructions. Materials or equipment whose manufacturers are not specified are all conventional products that can be obtained by purchase.

[0017] A method for profiling and classifying mine users based on safety big data includes the following steps: S1: Establish a basic dataset for mine user profiles: By collecting and preprocessing multi-source data from mines, effective feature data is obtained, thereby forming a basic dataset for mine user profiles; the multi-source data includes miner data, equipment data, environmental data, and text data; S2: Constructing a user profile tagging system: Based on the basic dataset of mine user profiles obtained from S1, a feature weight optimization method that combines dynamic weighting and cluster analysis is adopted to generate user profile tags covering miner behavior characteristics, equipment operating status, environmental safety indicators, and text-derived data. S3: Safety Risk Classification: User profile tags are quantified into feature vectors, input into a random forest classification model, and then the comprehensive risk value is calculated based on the analytic hierarchy process and the DS evidence theory to finally obtain the safety risk level of the mining user.

[0018] In S1, the preprocessing includes data cleaning, missing value completion, noise reduction, text data structuring, and data association; the data association is to associate the operation records in the miner data with the operation trajectory.

[0019] In S2, an improved weight calculation method is used to assign weights to the feature data, specifically: in, The adjusted weights are the weights of the j-th data item in data set i. This represents the number of records containing the j-th data; c is the weighting factor; N represents the total number of data records in the entire dataset; The result of the dynamically weighted adjustment of data frequency is as follows: In the formula, This refers to the frequency of characteristic data in critical security incidents; The frequency of the characteristic data in daily monitoring data; The frequency of the feature data in historical data; This refers to the weighting coefficient for critical security events. Weighting coefficients for daily monitoring data; Based on the clustering analysis results, a key feature data set S is identified from all features; for features belonging to set S, initial weights are assigned... Perform readjustment and calculate its final weight ω. ij′ ′: Where γ is the re-adjusted weighting coefficient for cluster analysis. The specific formula is: Where S is the set of key feature data. This represents the i-th feature data in the dataset.

[0020] In S2: The generation of miner behavior feature labels includes using the K-means algorithm to cluster mining operation types and / or using the PrefixSpan algorithm to mine dangerous behavior sequences of miners. The generation of equipment operating status labels includes using the isolated forest algorithm to detect abnormal equipment health status and / or using the Apriori algorithm to mine association rules for equipment failure risks; The generation of environmental safety indicator labels includes threshold determination based on real-time parameters and / or trend early warning using ARIMA time series models; The generation of text-derived tags includes calculating text keyword weights using an improved TF-IWF algorithm, using the LDA topic model for text security topic identification and feature enhancement, and finally generating tags based on the weighted ranking results.

[0021] The specific method for S3 is as follows: Quantize user profile tags into feature vectors; Using the random forest algorithm, an initial mapping relationship from label feature vectors to security risk levels is learned; The weights of four dimensions—miner behavior, equipment status, environmental indicators, and text information—were determined using the analytic hierarchy process (AHP). The four dimensions are treated as four independent sources of evidence. The risk level probability distributions output by each source of evidence are fused using the Dempster combination rule of the DS evidence theory. The comprehensive risk value is calculated based on the fused probability distribution, and the final security risk level is determined based on a preset threshold.

[0022] It also includes S4: Dynamic Update: Periodically update user profile tags and random forest models to adapt to the dynamic changes in mine safety status.

[0023] Figure 2 This is a schematic diagram of the electronic device structure provided in an embodiment of the present invention, with reference to... Figure 2 The electronic device may include a processor 201, a communications interface 202, a memory 203, and a communication bus 204, wherein the processor 201, the communications interface 202, and the memory 203 communicate with each other via the communication bus 204. The processor 201 can call logical instructions in the memory 203 to execute the following methods: S1: Establish a basic dataset for mine user profiles: By collecting and preprocessing multi-source data from mines, effective feature data is obtained, thereby forming a basic dataset for mine user profiles; the multi-source data includes miner data, equipment data, environmental data, and text data; S2: Constructing a user profile tagging system: Based on the basic dataset of mine user profiles obtained from S1, a feature weight optimization method that combines dynamic weighting and cluster analysis is adopted to generate user profile tags covering miner behavior characteristics, equipment operating status, environmental safety indicators, and text-derived data. S3: Safety Risk Classification: User profile tags are quantified into feature vectors, input into a random forest classification model, and then the comprehensive risk value is calculated based on the analytic hierarchy process and the DS evidence theory to finally obtain the safety risk level of the mining user.

[0024] Furthermore, the logical instructions in the aforementioned memory 203 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0025] On the other hand, embodiments of the present invention also provide a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described methods for profiling and classifying mine users based on security big data, including, for example: S1: Establish a basic dataset for mine user profiles: By collecting and preprocessing multi-source data from mines, effective feature data is obtained, thereby forming a basic dataset for mine user profiles; the multi-source data includes miner data, equipment data, environmental data, and text data; S2: Constructing a user profile tagging system: Based on the basic dataset of mine user profiles obtained from S1, a feature weight optimization method that combines dynamic weighting and cluster analysis is adopted to generate user profile tags covering miner behavior characteristics, equipment operating status, environmental safety indicators, and text-derived data. S3: Safety Risk Classification: User profile tags are quantified into feature vectors, input into a random forest classification model, and then the comprehensive risk value is calculated based on the analytic hierarchy process and the DS evidence theory to finally obtain the safety risk level of the mining user.

[0026] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0027] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0028] This invention presents a method for profiling and classifying mine users based on safety big data. First, it establishes a basic dataset for mine user profiles by comprehensively collecting and preprocessing relevant mine data. Then, it uses machine learning, data mining, and sensor data analysis technologies to construct a user profile tag system covering multiple dimensions such as miners, equipment, and the environment. Finally, based on the constructed user profile tag system, it uses classification algorithms to classify mine users.

[0029] The specific method for establishing the basic dataset of mine user profiles is as follows: The collected mine source data undergoes data cleaning and preprocessing to remove noise and invalid data, filtering out valid feature data. Based on the operational needs of mine safety management, key features are extracted from the dataset, including miners' personal information, work behavior data, and health data; basic parameters, operating status data, and maintenance records of mining equipment; and geological data, air quality data, and temperature and humidity data of the mining environment. The specific implementation plan is as follows: For miners' static data, such as personal information, database management tools are used to process missing attributes and abnormal data. For example, default values ​​are used to fill in missing non-critical information, and abnormal identity information is verified and corrected. For miners' operational behavior data, data is collected through sensors and monitoring equipment deployed at the mine site. The unique identification device carried by the miners (such as a smart badge) is combined with the positioning system to identify the miner's identity for each piece of behavioral data. Abnormal behavior data (such as records of violations) is filtered and marked.

[0030] For mining equipment data, real-time operating status data is collected through the equipment's built-in sensors and monitoring devices. Equipment identification is performed using its serial number, and abnormal data generated during operation (such as equipment fault warning data) is filtered and recorded. Maintenance records are managed through a combination of manual entry and automatic system updates to ensure data accuracy and completeness.

[0031] For mine environmental data, geological data, air quality data, temperature and humidity data are collected by environmental monitoring sensors deployed in various areas of the mine. The sensor number is used to locate and identify the data source, and abnormal environmental data (such as air quality exceeding the standard) is marked and processed in a timely manner.

[0032] For mine text data, it is obtained through the mine safety management platform, including safety inspection reports, accident analysis logs, and miner feedback records. Abnormal text data (such as garbled characters) is marked and processed in a timely manner.

[0033] The method for constructing a user profile tag system covering multiple dimensions uses an improved weight calculation method to obtain the weight of each feature data, enhances the representativeness of feature data in user profiles, identifies feature data that is highly related to key factors of mine safety, and uses these feature data as important tags for constructing user profiles.

[0034] An improved dynamic weighting algorithm is employed to differentiate different types of data based on the importance of feature data to mine safety through dynamic weighting coefficients, thereby enhancing the importance of key data in user profile construction. The adjusted weighting calculation formula is as follows: in, The adjusted weights are the weights of the j-th data item in data set i. This represents the number of records containing the j-th data point; c is a weighting factor used for adjustment and smoothing. The value of is used to avoid the denominator being zero and to control the calculation of weights; N represents the total number of data records in the entire dataset; The result of the dynamically weighted adjustment of data frequency is as follows: In the formula, This refers to the frequency of characteristic data in critical security incidents; The frequency of the characteristic data in daily monitoring data; The frequency of the feature data in historical data; This refers to the weighting coefficient for critical security events. Weighting coefficients for daily monitoring data; and The adjustments are made dynamically based on the data type and importance.

[0035] The improved weight calculation method also incorporates a clustering analysis model to detect hidden feature patterns in the dataset. First, based on all the data collected from the mine, a "data-feature" matrix including various types of feature data is constructed. Then, through training the clustering analysis model, different cluster categories are generated, with each category representing a data feature pattern.

[0036] The training process of a clustering analysis model utilizes the similarity measure between data points to divide the data into clusters. During model training, the cluster centers are continuously adjusted until the data partitioning reaches an optimal state, forming stable clusters. The model's performance is evaluated by calculating the similarity and differences between data in different clusters. If the clustering results can clearly distinguish different data feature patterns, it indicates that the training effect of the clustering analysis model is satisfactory.

[0037] After determining the parameters of the clustering analysis model, representative data features are selected as the key feature data set based on the different cluster categories obtained. Then, the initial weights obtained based on the dynamic weighting algorithm are applied according to the key feature data set. After readjustment, the final feature data weights are: Where γ is the re-adjusted weighting coefficient for cluster analysis. The specific formula is: Where S is the set of key feature data. This represents the i-th feature data in the dataset; that is, when the data feature is a key feature data obtained from cluster analysis, it has higher importance. The weight coefficients are readjusted to obtain the final weight calculation result.

[0038] When classifying mine users using the classification algorithm, a classification model based on the hierarchical analysis fusion method is first constructed. The hierarchical analysis method is used in combination with fuzzy comprehensive evaluation theory to effectively integrate multi-dimensional data. The uncertainty of the data is eliminated by fusing multiple information, the weight of each dimension of data is clarified, and a classification model is established.

[0039] The analytic hierarchy process (AHP) compares each factor at each level pairwise and assigns relative importance through expert consultation, literature review, and on-site data analysis. It then constructs a judgment matrix to calculate the weights of each dimension of data, performs consistency checks, and forms the basic framework for classification and evaluation.

[0040] Fuzzy comprehensive evaluation theory is used to process the data of each level of indicators, determine the membership degree of each level of indicator data, and combine the weights obtained by the analytic hierarchy process (AHP) to calculate the comprehensive score of the mine users belonging to different categories. Based on the comprehensive score, the mine users are classified into five categories: low safety risk, relatively low safety risk, medium safety risk, relatively high safety risk, and high safety risk. A corresponding score range is set for each category, and the category to which the mine user belongs is determined based on the calculated comprehensive score.

[0041] The specific method for constructing the judgment matrix is ​​as follows: Construct a judgment matrix P to compare indicators at a certain level, as shown in the following formula: in, This represents the evaluation result of the relative importance ratio between indicator i and indicator j.

[0042] The indicator identification framework is as follows: Indicates the first-level indicator The weight, Indicates secondary indicators The weights, and and satisfy: ; ; 0< <1, 0< <1.

[0043] The formula for checking the consistency of a matrix is: in, Let CI represent the largest eigenvalue of the judgment matrix, and n represent the order of the judgment matrix. When CI = 0, i.e., ... When n = n, the matrix is ​​a consistent matrix.

[0044] Application Examples This example demonstrates how to transform multi-source heterogeneous mine data into a standardized dataset with usable labels using a "targeted data collection + precise cleaning + structured transformation" approach. The specific technical steps for each stage are as follows: Step 1: Multi-source data acquisition and preprocessing (providing standardized data for tag construction) The core of this step is to transform multi-source heterogeneous data from mines into a standardized dataset with labels through "targeted collection + precise cleaning + structured transformation". The technical details of each step are as follows: 1.1 Targeted Collection of Multi-Source Data (Clarifying Data Sources and Collection Rules) 1.2 Precise data cleaning and structuring (clarifying the processing methods for each type of data) 1.2.1 Structured data cleaning (miner static data, equipment fault records) (1) Tool selection: SparkSQL (version 3.3.0) was used to build a distributed cleaning task to process data with a volume of ≥1 million records.

[0045] (2) Handling missing values: Key fields (such as "qualification certificate validity period" and "equipment failure type") for miners: complete them by linking to historical databases (such as HRMS historical qualification table and CMMS failure dictionary); those that cannot be completed are marked as "pending verification" and updated after manual confirmation.

[0046] Non-critical fields (such as miners' "place of origin" and "mining secondary parameters") are filled with industry averages (e.g., if miners' "work experience" is missing, fill with the average work experience of the same job type, formula: work experience = current year - year of employment).

[0047] (3) Outlier handling: Set a rule base (based on the "Mine Safety Regulations" and "Equipment Technical Manual"): such as "miner's age <18 years old or >60 years old" and "equipment vibration frequency >50Hz (the upper limit of the rated value of a certain type of motor)" are judged as abnormal values.

[0048] Outlier handling logic: First, mark the outlier data, then trace back through the original acquisition logs (such as checking HRMS entry records and equipment sensor calibration records). If the error is confirmed, it is corrected; otherwise, it is removed.

[0049] 1.2.2 Dynamic Behavioral Data Alignment (Association of Miner Trajectory and Operation Records) (1) Trajectory data completion: For missing positioning periods (e.g., positioning device signal interruption ≤ 5 minutes): use linear interpolation to fill in the missing coordinates. The formula is: missing coordinate X = X before + (X after - X before) × (t missing - t before) / (t after - t before) (t is the timestamp, X is the horizontal axis coordinate, and the same applies to the Y axis).

[0050] Location missing for more than 5 minutes: Mark as “Trajectory missing”, associate with concurrent video surveillance, if the miner can be identified in the video, extract behavioral features (such as “holding tools” or “bending over to work”) through YOLOv5 (version 6.2) posture recognition, temporarily associate with “same area and same job group”, and bind to the specific miner ID after manual confirmation.

[0051] (2) Association between operation records and trajectory: Association logic: Using “miner’s unique ID + timestamp” as the key, match “operation equipment ID (e.g., E005)” with “trajectory coordinates (e.g., No. 3 mining area - East Lane, X=120, Y=80)” to ensure that “operation behavior” and “operation location” correspond one-to-one (time error ≤ 10 seconds).

[0052] 1.2.3 Noise Reduction of Time-Series Data (Equipment Operating Parameters, Environmental Data) Tool selection: Pandas (version 1.5.3) + SciPy (version 1.10.1) of Python were used to process time series data.

[0053] Noise reduction methods: The 3σ criterion is used to identify abnormal fluctuations: the mean μ and standard deviation σ of a certain parameter (such as gas concentration) are calculated, and values ​​that are “parameter value > μ + 3σ or < μ - 3σ” are marked as outliers.

[0054] Outlier handling: If the outlier lasts for ≤3 seconds (e.g., momentary sensor interference), replace it with the moving average of the previous 5 normal data; if it lasts for >3 seconds (e.g., actual gas exceedance), mark it as "to be verified" and trigger a real-time alarm to notify on-site personnel for verification.

[0055] 1.2.4 Structuring of Text Data (Safety Reports, Logs) (1)Word Segmentation and Stop Word Processing: Word Segmentation Tool: Jieba Word Segmentation (loading a mine-specific professional dictionary, including professional terms such as "gas leakage", "roof collapse", "bolt-shotcrete support", etc.), and the word segmentation granularity is set to "accurate mode".

[0056] Stop Word Removal: Based on a custom stop word list (including meaningless words such as "of", "in", "this time", etc.), the set filtering method is used to remove stop words.

[0057] (2)Text Vectorization: Tool: Word2Vec (based on Gensim 4.3.1), the training corpus is the text data of the mine in the past 3 years (about 50,000 articles), the vector dimension is set to 100, the window size is set to 5, and the minimum word frequency is set to 5 (only retaining words that appear ≥ 5 times), generating the vector representation of each word in the text.

[0058] Step 2: Construction of a Multidimensional Label System (clarifying the generation method and judgment criteria for each label) In this step, for the four core dimensions of "miners, equipment, environment, text", exclusive labels are generated through "algorithm modeling + rule definition", and each label clearly defines "input data, algorithm steps, output results, judgment criteria", as follows: 2.1 Miner Behavior Feature Labels (Core Input: Miner Dynamic Behavior Data + Static Data) 2.1.1 Basic Behavior Labels (Quantitative Labels, Directly Reflecting Behavior Risks) 2.1.2 Behavior Pattern Labels (Clustering / Sequential Mining Labels, Identifying Potential Risk Patterns) (1)"Operation Type" Clustering Label (K-means Algorithm) Step 1: Determine Clustering Features: Select 3 core features, namely "trajectory coincidence degree (overlap rate with the standard operation trajectory)", "operation compliance rate (number of compliant operations / total number of operations)", and "frequency of entering high-risk areas", and all are normalized to [0,1] through Min-Max normalization.

[0059] Step 2: Determine the Value of K: Use the elbow method to calculate the sum of squared errors within the cluster (SSE) when K = 2 - 5. When K = 3, the decline rate of SSE suddenly decreases (elbow point), so set K = 3.

[0060] Step 3: Execute K-means Clustering: Initialization: Use the "K-means++" method to initialize the cluster centers to avoid local optimality caused by random initialization.

[0061] Iteration: Calculate the Euclidean distance of each sample (miner) to the 3 centers and classify it into the closest category; update the cluster center to the feature mean of all samples in that category; repeat the iteration until the center change is <0.001 or the number of iterations is ≥100.

[0062] Step 4: Tag Definition: Category 1 (Trajectory overlap ≥ 0.8, Operational compliance rate ≥ 0.95): "Operation Type - Standardized" Category 2 (Trajectory overlap 0.6-0.8, Operational compliance rate 0.8-0.95): "Operation Type - General" Category 3 (Trajectory overlap < 0.6, Operational compliance rate < 0.8): "Operation Type - Risky" (2) “Dangerous Behavior Sequence” Labels (PrefixSpan Algorithm) Step 1: Behavior sequence construction: Sort the miner's hourly behavior by time to generate a sequence (e.g., "Inspect equipment → Enter low-risk area → Operate equipment → Enter high-risk area → Not wearing a safety helmet").

[0063] Step 2: Algorithm parameter settings: Minimum support = 5% (only retain behavioral sequences with an occurrence frequency ≥ 5% of the total number of sequences), Minimum confidence = 0.7 (the correlation strength between preceding and following behaviors in a sequence is ≥ 0.7).

[0064] Step 3: Sequence mining: Use PrefixSpan to mine frequent behavior sequences and filter out sequences containing "violations" (such as "failure to check equipment → entering a high-risk area" and "failure to wear a safety helmet → operating heavy equipment").

[0065] Step 4: Label determination: If a miner has exhibited the above dangerous sequence ≥ 2 times in the past 7 days, it is marked as "Dangerous Behavior Sequence - Existing"; otherwise, it is marked as "Dangerous Behavior Sequence - None".

[0066] 2.2 Equipment Operation Status Label (Core Input: Equipment Operation Parameters + Fault Records) 2.2.1 Health Status Labels (Anomaly Detection + Association Rule Mining) (1) “Device Status” label (Isolation Forest Algorithm) Step 1: Feature selection: Select four real-time parameters: “vibration frequency (Hz)”, “surface temperature (°C)”, “current (A)”, and “energy consumption (kW·h)”. Take the average value of the past hour as the sample feature and normalize it to [0,1].

[0067] Step 2: Model training: Isolation Forest is used, with sample size = normal operation data of a certain device for the past 30 days (approximately 8640 samples), number of trees = 100, and outlier ratio = 0.05 (setting the proportion of outliers in normal data to be ≤5%).

[0068] Step 3: Anomaly detection: Input real-time features into the model and calculate the anomaly score (0-1). A score > 0.8 is considered abnormal data. Calculate the percentage of abnormal data in the past hour (percentage of abnormal duration = number of abnormal data entries / total number of data entries).

[0069] Step 4: Tag determination: Abnormal duration percentage > 5% → “Device status - Abnormal”; otherwise → “Device status - Normal”.

[0070] (2) "Fault Risk" label (Apriori association rule algorithm) Step 1: Data preprocessing: Discretize the equipment operating parameters (e.g., temperature: ≤40℃ = "low", 40-60℃ = "medium", >60℃ = "high"), and encode the fault records by type (e.g., "motor fault" = 1, "bearing fault" = 2).

[0071] Step 2: Algorithm parameters: minimum support = 0.1 (retain itemsets with an occurrence frequency ≥ 10%), minimum confidence = 0.8 (rule confidence ≥ 80%).

[0072] Step 3: Association rule mining: Mining rules for "operating parameter itemset → fault type", such as "temperature = high, vibration frequency = high → motor fault" (confidence level = 0.85).

[0073] Step 4: Tag determination: If the real-time operating parameters match a certain fault rule (e.g., temperature = 65℃, vibration frequency = 55Hz), then mark it as "fault risk - high"; otherwise → "fault risk - low".

[0074] 2.2.2 Maintain requirement tags (time series calculation) Input data: Equipment historical maintenance records (interval time between the last 10 maintenance visits, such as average maintenance interval = 30 days), current running time (number of days since the last maintenance visit).

[0075] Calculation method: Remaining maintenance cycle = Average maintenance interval - Current running time.

[0076] Judgment criteria: Remaining maintenance cycle < 7 days → "Immediate maintenance required"; 7-15 days → "Planned maintenance required"; > 15 days → "Maintenance requirement - low".

[0077] 2.3 Environmental Safety Indicator Labels (Core Inputs: Real-time Environmental Parameters + Time-Series Predictions) 2.3.1 Real-time risk labeling (threshold determination, based on industry standards) 2.3.2 Trend Warning Labels (ARIMA Time Series Forecasting) Step 1: Data preparation: Take the time series data of environmental parameters (such as gas concentration) of a certain area for nearly 24 hours (1 second / data point, a total of 86,400 data points), and use "downsampling" to process it into 5 minutes / data point (a total of 288 data points) to avoid excessive data volume.

[0078] Step 2: Determine ARIMA model parameters: Stationarity test: If the data is non-stationary, perform first-order differencing (d=1) to make it stationary by using the ADF test (significance level=0.05).

[0079] Order selection: p=2 (ACF truncates to lag 2) is determined by the ACF (autocorrelation function) plot, and q=1 (PACF truncates to lag 1) is determined by the PACF (partial autocorrelation function) plot. The final model is ARIMA (p=2, d=1, q=1).

[0080] Step 3: Forecasting and Early Warning Prediction: Input the model to predict environmental parameters for the next 2 hours (1 prediction value every 5 minutes, for a total of 24 values).

[0081] Warning determination: If any predicted value exceeds the "high-risk threshold" (e.g., gas concentration > 0.8%), it is marked as "environmental risk - warning"; otherwise, it is marked as "environmental risk - no warning".

[0082] 2.4 Text-derived tags (Core input: Structured text data, improved TF-IWF+LDA) 2.4.1 Improved TF-IWF algorithm (optimized text feature weights) Text structure decomposition: A single text (such as a security inspection report) is split into three parts: "title, description, and body," and the term frequency (TF) of each word is calculated separately. title tf desc tf body ).

[0083] Dynamic position-weighted calculation of adjusted word frequencies: Weighting coefficients are set as follows: Title weight ω1 = 1.5 (core information), description weight ω2 = 1.2 (key summary), and body weight ω3 = 1.0 (detailed content).

[0084] Adjusted word frequency formula: tf' ij =ω1×tf title +ω2×tf desc +ω3×tf body (i is the text number, j is the word number).

[0085] IDF squaring (amplifying the weight of rare and risky keywords): IDF calculation formula: IDFj =log(N / (n j +c))(N=Total number of texts; n j = The number of texts containing word j; c = 0.1, to avoid a denominator of 0).

[0086] Improved weight formula: ω' ij =tf' ij ×(IDF j ) 2 (Squaring the weights of rare words, such as "gas leak", weakens the influence of high-frequency irrelevant words, such as "inspection").

[0087] 2.4.2 LDA Topic Model (Extracting Core Topics from Text) The number of topics K is determined: Perplexity calculation: For the LDA model with K=2-8, calculate the perplexity of the test set (accounting for 20% of the total text) using the formula: Perplexity=exp(-∑logP(w|θ,φ) / ∑Nw) (where w is the word, θ is the document-topic distribution, φ is the topic-word distribution, and Nw is the total number of words in the test set).

[0088] Elbow method judgment: When K=5, the decrease in perplexity slows down significantly (elbow point), therefore, the number of topics K is set to 5, as follows: Topic 1: Equipment Failure (Keywords: bearing, overheating, malfunction, repair, motor) Topic 2: Operational violations (keywords: not wearing a safety helmet, improper operation, failure to inspect, exceeding time limit) Topic 3: Environmental Anomalies (Keywords: gas, dust, exceeding standards, temperature and humidity, leakage) Topic 4: Roof Safety (Keywords: roof, support, collapse, crack, pressure) Topic 5: Training and Education (Keywords: training, assessment, qualification, safety knowledge) Additional adjustments to topic weighting: If word j belongs to the Top 30 feature words of a certain topic (e.g., "gas" belongs to topic 3), then its final weight is: ω'' ij =ω' ij +0.3 (γ=0.3, to enhance topic relevance).

[0089] 2.4.3 Text Tag Generation Step 1: For a single text, calculate the final weight ω'' ij Sort in descending order and extract the top 5 keywords with the highest weight (such as "gas leak", "illegal operation", "equipment overheating", "high-risk area", "uninspected").

[0090] Step 2: Convert the Top 5 keywords into tags in the format "text keyword-XX" (e.g., "text keyword-gas leak" or "text keyword-violation").

[0091] Step 3: Label-based security risk classification and prediction (clarifying label application logic and evaluation methods) This process involves four steps: "label quantification → feature fusion → model training → risk assessment," which transforms multidimensional labels into "risk levels" that can be directly used for security management. The technical details are as follows: 3.1 Label Quantization and Feature Fusion (converting labels into feature vectors that the model can input) 3.1.1 Label Quantization Rules (Unifying Feature Dimensions and Value Ranges) 3.1.2 Feature Vector Construction (Integrated and Quantized Labels by Dimension) Using "one hour of data from a single mining area (such as mining area No. 3)" as a single sample, construct a 12-dimensional feature vector, in the following format: Feature vector = [Miner risk propensity, dangerous behavior sequence, violation record, equipment status, failure risk, maintenance needs, gas risk, dust risk, temperature and humidity risk, environmental warning, average text keyword value, mining operation duration] Example: The quantized vector of 1 hour's data from a certain mining area is: [0.8,1,0.7,1,0.8,1,0.8,0.5,0.5,1,0.7,0.8] 3.2 Risk classification model training (random forest algorithm, to map labels to risk levels) 3.2.1 Training Data Preparation and Labeling (1) Data source: Historical data of the mine over the past two years, with a total of 10,000 samples collected (each sample contains “feature vector + actual safety event”).

[0092] (2) Labeling rules (based on actual security incidents): High risk (Level 5): Safety accidents (such as gas explosions or equipment injuries) occurred within the corresponding time period of the sample → marked with "5".

[0093] Higher risk (Level 4): Two or more high-risk warnings are triggered within the corresponding time period of the sample (such as excessive gas + equipment malfunction) → marked "4".

[0094] Medium risk (Level 3): One high-risk warning or two medium-risk warnings are triggered within the corresponding time period of the sample → marked with "3".

[0095] Lower risk (Level 2): ​​One medium warning is triggered within the corresponding time period of the sample → marked "2".

[0096] Low risk (Level 1): No warnings or abnormalities were detected within the corresponding time period of the sample → marked "1".

[0097] (3) Data partitioning: The data is divided into a training set (7000 data points), a validation set (2000 data points), and a test set (1000 data points) in a ratio of 7:2:1.

[0098] 3.2.2 Random Forest Model Parameter Optimization (1) Core parameter settings: Number of trees (n_estimators): When the validation set accuracy is highest through grid search (100-300, step size 50), n_estimators=200.

[0099] Tree depth (max_depth): Grid search (5-15, step size 1), max_depth=10 when the validation set accuracy is highest.

[0100] Minimum number of sample splits (min_samples_split): Set to 2 (default, to ensure the model does not overfit).

[0101] Minimum number of sample leaf nodes (min_samples_leaf): Set to 1 (default).

[0102] (2) Model training and validation: Training tool: Scikit-learn (version 1.2.2) for Python.

[0103] Validation method: 5-fold cross-validation (the training set is divided into 5 parts, and 4 parts are used for training and 1 part for validation in turn, and the average accuracy is taken).

[0104] Model performance: Test set accuracy ≥92%, recall rate (high-risk sample identification rate) ≥95%, meeting the needs of mine safety prediction.

[0105] 3.3 Safety Risk Assessment Calculation (AHP+DS Evidence Theory, Optimizing the Credibility of Risk Results) 3.3.1 AHP ​​determines the weight of each dimension's labels (addressing the issue of "differences in the importance of different labels") (1) Constructing a hierarchical structure: Target layer: Mine safety risk assessment.

[0106] Criteria Layer (4 dimensions): Miner behavior (A1), Equipment status (A2), Environmental indicators (A3), Text information (A4).

[0107] Solution level: Risk level (low / lower / medium / higher / higher).

[0108] (2) Constructing a judgment matrix (1-9 scale method, based on expert consultation + industry standards): Five mine safety experts (with ≥10 years of experience) were invited to conduct pairwise comparisons of the criteria layer, and the average value was used to construct a judgment matrix P. P=[ [1,1.2,0.8,1.5],#Importance ratio of A1 to A2, A3, and A4 [1 / 1.2,1,0.7,1.3],#The importance ratio of A2 to A1, A3, and A4 [1 / 0.8, 1 / 0.7, 1, 1.6], #Importance ratio of A3 to A1, A2, and A4 [1 / 1.5, 1 / 1.3, 1 / 1.6, 1]# The importance ratio of A4 to A1, A2, and A3 ] (3) Calculate weights and perform consistency checks: Weight calculation: The maximum eigenvalue of the judgment matrix is ​​obtained by eigenvalue method, λ_max=4.032. The corresponding eigenvector is normalized to become the criterion layer weight: W=[0.4,0.3,0.2,0.1] (i.e., miner behavior weight 40%, equipment status 30%, environment 20%, text 10%).

[0109] Consistency check: CI=(λ_max-n) / (n-1)=(4.032-4) / (4-1)=0.0107, random consistency index RI=0.9 (when n=4), consistency ratio CR=CI / RI=0.0119<0.1, the judgment matrix meets the consistency requirements, and the weights are effective.

[0110] 3.3.2 D-S Evidence Theory Fusion with Risk Probability (Solving the Problem of "Label Prediction Uncertainty") (1) Identify evidence sources: The label prediction results of the four criteria layers (miner, equipment, environment, and text) are used as four evidence sources; The Frame of Discernment is defined as a set of five security risk levels: Θ={R1,R2,R3,R4,R5} in: R1: Low risk R2: Lower risk R3: Medium Risk R4: Higher Risk R5: High Risk Suppose there are s = 4 independent sources of evidence, corresponding to: i=1: Miner behavior i=2: Device status i=3: Environmental indicators i=4: Text information Each evidence source i outputs a probability distribution about Θ: in, (risk level is) Rk |Evidence source i), and satisfies .

[0111] Each evidence source outputs "the probability that the sample belongs to one of the five risk levels" (e.g., the miner behavior evidence source output: P(low)=0.1, P(lower)=0.1, P(medium)=0.2, P(higher)=0.2, P(high)=0.4).

[0112] (2) Construction of Basic Probability Assignment (BPA): The predicted probability of each source of evidence is directly transformed into the Basic Probability Assignment (BPA) function: And it stipulates: m i (A) = 0 for all A ⊂ Θ and |A| = 1; m i (∅)=0; ; (3) Evidence composition (Dempster's combination rule): First, fuse two evidence sources (e.g., miner A1 and device A2), then fuse the result with environment A3, and finally fuse with text A4. Dempster's combination rule is used to progressively fuse multiple BPAs. Let... If the BPA is the result of fusing the first n pieces of evidence, then it is related to the (n+1)th piece of evidence. Synthesis results Defined as: Among them, the conflict coefficient K n for: Fusion process (sequential combination): The final result is the fused BPA: Define the quantification function of risk level : (4) Comprehensive risk value and level determination: Overall risk value E calculation: Based on the fused BPA m(R) k ), calculate the expected composite risk value: Final risk level assessment: Overall Risk Value E(X): The combined BPA value is weighted and summed according to "risk level - quantification value". The quantification level is: low = 0.2, lower = 0.4, medium = 0.6, higher = 0.8, and high = 1.0.

[0113] Risk level assessment: E∈[0,0.2): Low risk (Level 1) E∈[0.2,0.4): Lower risk (Level 2) E∈[0.4,0.6): Medium risk (Level 3) E∈[0.6,0.8): Higher risk (Level 4) E∈[0.8,1.0]: High risk (Level 5) Example: After fusion, BPA is m(low)=0.05, m(lower)=0.05, m(medium)=0.1, m(higher)=0.2, m(high)=0.6, E(X)=0.05×0.2+0.05×0.4+0.1×0.6+0.2×0.8+0.6×1.0=0.85 → judged as "high risk (level 5)".

[0114] Step 4: Dynamic update mechanism (to ensure the timeliness of labels and models) 4.1 Tags are dynamically updated (real-time / scheduled updates to reflect the latest status). Tag categories Update frequency Update logic Miner behavior tags Updated once per hour Collect the latest hourly trajectory and operation data, recalculate the "average daily working hours" and "percentage of stay in high-risk areas", and update the clustering and sequence labels. Device status label Updated every 10 minutes Collect the latest 10 minutes of operating parameters, re-detect abnormal values, calculate the remaining maintenance cycle, and update the health status and maintenance requirement tags. Environmental indicator labels Updated every minute Collect the latest environmental parameters for the past minute and update the real-time risk labels; rerun the ARIMA model every 30 minutes to update the trend warning labels. Text-derived tags Updated once daily at midnight Incorporate newly added security reports and logs from the previous day, re-execute the improved TF-IWF+LDA, and update text keyword tags. 4.2 Dynamic Model Updates (Regular Fine-tuning + Retraining to Ensure Prediction Accuracy) Weekly tweaks: Data input: Add 7 days of labeled samples (approximately 168 samples).

[0115] Fine-tuning logic: Fix the tree structure of the random forest and only update the decision rules of the leaf nodes; calculate the model error using new samples, and if the error is >5%, adjust parameters such as "tree depth" and "number of trees" (adjustment range ≤20%).

[0116] Monthly retraining: Data input: Add 30 days of labeled samples (approximately 720 samples) to replace the oldest 30-day samples in the training set (sliding window strategy to ensure data timeliness).

[0117] Retraining: Repeat the parameter optimization and 5-fold cross-validation in step 3.2 to generate a new model; validate it using a test set (100 newly collected samples). If the accuracy of the new model is ≥3% higher than that of the old model, then replace the old model.

[0118] Update verification: After each update, compare the "model predicted risk level" with the "actual security event", and calculate the accuracy and recall rate to ensure that the accuracy is ≥90% and the recall rate is ≥93%. Otherwise, backtrack to check for problems in the data cleaning or label building process.

[0119] This example addresses the ambiguity in existing technologies by clarifying the generation algorithms for each label, improving the text feature extraction method, and outlining the risk classification process, thereby achieving accurate assessment and prediction of mine safety risks.

[0120] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for profiling and classifying mine users based on safety big data, characterized in that, Includes the following steps: S1: Establish a basic dataset for mine user profiles: By collecting and preprocessing multi-source data from mines, effective feature data is obtained, thereby forming a basic dataset for mine user profiles; the multi-source data includes miner data, equipment data, environmental data, and text data; S2: Constructing a user profile tagging system: Based on the basic dataset of mine user profiles obtained from S1, a feature weight optimization method that combines dynamic weighting and cluster analysis is adopted to generate user profile tags covering miner behavior characteristics, equipment operating status, environmental safety indicators, and text-derived data. S3: Safety Risk Classification: User profile tags are quantified into feature vectors, input into a random forest classification model, and then the comprehensive risk value is calculated based on the analytic hierarchy process and the DS evidence theory to finally obtain the safety risk level of the mining user.

2. The method for profiling and classifying mine users based on safety big data according to claim 1, characterized in that, In S1, the preprocessing includes data cleaning, missing value completion, noise reduction, text data structuring, and data association; the data association is to associate the operation records in the miner data with the operation trajectory.

3. The method for profiling and classifying mine users based on safety big data according to claim 1, characterized in that, In S2, an improved weight calculation method is used to assign weights to the feature data, specifically: in, The adjusted weights are the weights of the j-th data item in data set i. This represents the number of records containing the j-th data; c is the weighting factor; N represents the total number of data records in the entire dataset; The result of the dynamically weighted adjustment of data frequency is as follows: In the formula, This refers to the frequency of characteristic data in critical security incidents; The frequency of the characteristic data in daily monitoring data; The frequency of the feature data in historical data; This refers to the weighting coefficient for critical security events. Weighting coefficients for daily monitoring data; Based on the clustering analysis results, a key feature data set S is identified from all features; for features belonging to set S, initial weights are assigned... Perform readjustment and calculate its final weight ω. ij′ ′: Where γ is the re-adjusted weighting coefficient for cluster analysis. The specific formula is: Where S is the set of key feature data. This represents the i-th feature data in the dataset.

4. The method for profiling and classifying mine users based on safety big data according to claim 1, characterized in that, In S2: The generation of miner behavior feature labels includes using the K-means algorithm to cluster mining operation types and / or using the PrefixSpan algorithm to mine dangerous behavior sequences of miners. The generation of equipment operating status labels includes using the isolated forest algorithm to detect abnormal equipment health status and / or using the Apriori algorithm to mine association rules for equipment failure risks; The generation of environmental safety indicator labels includes threshold determination based on real-time parameters and / or trend early warning using ARIMA time series models; The generation of text-derived tags includes calculating text keyword weights using an improved TF-IWF algorithm, using the LDA topic model for text security topic identification and feature enhancement, and finally generating tags based on the weighted ranking results.

5. The method for profiling and classifying mine users based on safety big data according to claim 1, characterized in that, The specific method for S3 is as follows: Quantize user profile tags into feature vectors; Using the random forest algorithm, an initial mapping relationship from label feature vectors to security risk levels is learned; The weights of four dimensions—miner behavior, equipment status, environmental indicators, and text information—were determined using the analytic hierarchy process (AHP). The four dimensions are treated as four independent sources of evidence. The risk level probability distributions output by each source of evidence are fused using the Dempster combination rule of the DS evidence theory. The comprehensive risk value is calculated based on the fused probability distribution, and the final security risk level is determined based on a preset threshold.

6. The method for profiling and classifying mine users based on safety big data according to claim 1, characterized in that, It also includes S4: Dynamic Update: Periodically update user profile tags and random forest models to adapt to the dynamic changes in mine safety status.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method for profiling and classifying mine users based on security big data as described in claim 1.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for profiling and classifying mine users based on security big data as described in claim 1.