Machine learning models for identifying and predicting health and safety risks in electronic communications

JP2024538508A5Pending Publication Date: 2025-08-13ORACLE INT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024515934
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-10-27
Filing Date
2022-09-15
Publication Date
2025-08-13

AI Technical Summary

Technical Problem

There is no conventional system for early detection of health and safety risks in construction and engineering projects based on electronic communications, which are often a precursor to larger incidents.

Method used

A computer-implemented method using machine learning classifiers to analyze email communications, tokenize and vectorize text, and classify potential safety risks by matching against prescribed vocabularies, generating probability risk values, and sending real-time alerts for safety risks.

Benefits of technology

Enables proactive intervention by identifying potential safety risks in near real-time, reducing the likelihood of accidents and violations through automated risk detection and notification systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Systems, methods and other embodiments related to a machine learning system for monitoring and detecting health and safety risks in electronic communications associated with a target field are described. In one embodiment, a method includes monitoring email communications over a network to identify emails associated with the target field. A machine learning classifier is initiated, the machine learning classifier configured to classify text from the email having a risk as including vocabulary that refers to a safety risk or a non-risk. The classifier generates a probability risk value that the email includes text that refers to a safety risk and labels the email as a safety risk or a non-risk based at least in part on the probability risk value. To provide an alert, an electronic notification is generated and transmitted to a remote device in response to the email being labeled as referring to a safety risk.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] background Incidents leading to health and safety risks are a common occurrence in most large projects, such as construction and engineering projects. These safety incidents cost owners, contractors, subcontractors, architects and consultants millions of dollars and affect the overall project. Early detection of potential problems may allow for proactive intervention, which may lead to avoiding accidents and safety violations at the work site.

[0002] For example, digital platforms are used to manage and communicate day-to-day electronic communications as projects progress. These electronic communications contain information that, if correctly deciphered, can provide early warning signs of potential issues that could lead to larger health and safety incidents. These early warning signs can be used to identify risks associated with each project and can act to provide early warning.

[0003] However, there are no prior art computer intelligent systems for identifying such early warning signs of risk for a project, nor are there prior art systems that can forecast or predict potential risks for a project based on electronic communications. Summary of the Invention

[0004] overview In one embodiment, a computer-implemented method is described that includes monitoring email communications over a network to identify emails. In response to receiving an email over the network, the method detects and identifies emails as related to a construction project. The method tokenizes text from the email into a plurality of words and vectorizes each of the plurality of words into a numeric vector that maps each word to a numeric value. A machine learning classifier configured to identify construction terminology and classify text having risk as referring to or discussing safety risks or not referring to safety risks, i.e., non-risk, is initiated, inputting the numeric vector generated from the email into the machine learning classifier. The machine learning classifier processes the numeric vector from the email by at least matching the numeric vector to a set of defined safety risk vocabulary and a set of defined non-risk vocabulary, and generates a probabilistic risk value by the machine learning classifier, and the email includes vocabulary that refers to or discusses safety risks. The email is labeled as a safety risk or non-risk based, at least in part, on a probability risk value that the email mentions a safety risk, and an electronic notification is generated and transmitted to a remote device in response to the email being labeled as mentioning a safety risk to provide an alert.

[0005] In another embodiment, the method further includes inputting construction terminology from a glossary or database of construction project terms into the machine learning classifier and training the machine learning classifier to identify safety risk text based, at least in part, on a first dataset of communications having known text related to or referencing safety risks and a second dataset of communications having known non-risk text that does not refer to health and safety risks.

[0006] In another embodiment, a computing system is described that includes at least one processor configured to execute instructions; at least one memory operatively connected to the at least one processor; a machine learning classifier configured to identify construction terminology and classify risky text as a safety risk or non-risk; and a non-transitory computer readable medium including stored computer executable instructions, the computer executable instructions when executed by the at least one processor to cause the computing device to monitor email communications over a network to identify sent emails; in response to receiving an email over the network, detect and identify the email as related to a construction project; and tokenize text from the email into a plurality of words. and inputting a plurality of words generated from the email into a machine learning classifier, the machine learning classifier being configured to evaluate the plurality of words from the email by at least matching the plurality of words to a set of defined safety risk vocabulary and a set of defined non-risk vocabulary, generating, by the machine learning classifier, a probability risk value that the email contains text that refers to a safety risk, labeling the email as a safety risk or non-risk based, at least in part, on the probability risk value that the email contains text that refers to a safety risk, and generating and sending an electronic notification to a remote device in response to the email being labeled as a safety risk to provide a near real-time alert in connection with receiving the email on the network.

[0007] In another embodiment, a computer-implemented method, computer system, or non-transitory computer-readable medium is described that includes or executes computer-executable instructions that, when executed by at least a processor of the computer, cause the computer to monitor email communications over a network to identify emails associated with a target field; initiate a machine learning classifier configured to classify text from the email having a risk as associated with a safety risk or non-risk; generate, by the machine learning classifier, a probability risk value that the email is associated with a safety risk; label, by the machine learning classifier, the email as a safety risk or non-risk based, at least in part, on the probability risk value indicating that the email is a safety risk; and generate and send an electronic notification to a remote device in response to the email being labeled as a safety risk to provide an alert.

[0008] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate various systems, methods and other embodiments of the disclosure. It will be appreciated that element boundaries (e.g., boxes, groups of boxes, or other shapes) depicted in the drawings represent one embodiment of the boundaries. In some embodiments, one element may be embodied as multiple elements, and multiple elements may be embodied as one element. In some embodiments, an element shown as an internal component of another element may be embodied as an external component, and vice versa. Additionally, elements may not be drawn to scale. [Brief description of the drawings]

[0009] [Figure 1] FIG. 1 illustrates one embodiment of a machine learning system relating to predicting risk from electronic communications. [Diagram 2]FIG. 1 illustrates one embodiment of a document vector matrix in Python for one communicating thread with four features. [Diagram 3] FIG. 1 illustrates one embodiment of a graph showing probability of safety risk cutoff values ​​versus a graph showing sensitivity (tpr) and isomerism (tnr) for selecting an initial threshold value. [Figure 4] FIG. 1 illustrates an embodiment of a method related to detecting potential risks from electronic communications in a construction project. [Diagram 5] FIG. 1 illustrates another embodiment of a method relating to detecting potential safety risks from electronic communications in a target field. [Figure 6] FIG. 1 illustrates an embodiment of a computing system configured with the disclosed example systems and / or methods. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0010] Detailed Description Systems and methods for implementing an artificial intelligence (AI)-based monitoring and prediction / detection system are described herein. In one embodiment, a machine learning model is configured to monitor network communications and predict health and safety risks in exchanged electronic communications / communications related to a target field or project. In one embodiment, the target field is a construction and engineering project. For example, the present systems and methods identify potential health and safety risks using application-specific artificial intelligence configured for semantic natural language understanding, specifically generated, trained, and optimized for health and safety risk detection in text from construction and engineering project communications.

[0011] As used herein, the term "health and safety risk" is also referred to as "safety risk" or "risk" for brevity. "Safety risk" refers to the determination or classification that a communication or text / language relates to, refers to, or discusses a potential health and safety risk, for example, where a health and safety risk has occurred or may occur in an environment (e.g., a construction project or site).

[0012] The term "non-risk" refers to a determination or classification that the communication or text / language is not considered (or is not related to or discussed as) a potential health and safety risk.

[0013] In one embodiment, the system monitors (e.g., in near real-time) network and electronic communications in an ongoing project or collection of projects in an organization's portfolio. Information in the electronic communications is decoded by a machine learning model to identify and detect language indicative of a safety risk. The machine learning model makes a prediction of whether the language will be at a threshold level of risk based on at least a set of trained data. If a risk is predicted for a communication, the system automatically generates an alert in near real-time and labels the associated communication thread as a safety risk (e.g., including vocabulary / text that discusses a safety risk). This may include labeling an email thread if there is a safety risk associated with the last communication sent in a thread.

[0014] In one embodiment, the system may combine the identified communications with contextual project metadata to associate the predicted risks with project processes. The identified communications and information about the identified projects may then be sent and / or displayed in a graphical user interface and / or sent in electronic messages to a remote computer so that the user has access to the information in near real-time. In another embodiment, the system provides a feedback process where the user can change the labels associated with each communication thread if the system's prediction is incorrect as determined based on the user's experience and intuition. The changed labels are then fed back into the system model as new training data to improve prediction accuracy over time.

[0015] 1, one embodiment of a safety risk detection system 100 configured to monitor network communications and predict safety risks in electronic communications is shown. Initially, the system 100 includes training a machine learning model (described below) with a known dataset of project communications that includes known safety risk language and known non-risk language. The training configures the machine learning model to identify and predict safety risks associated with a particular project based on the monitored electronic communications.

[0016] In one embodiment, after the model is deployed and operates to monitor communications, any identified health and safety risks are categorized based on the likelihood that the identified risks will result in a health or safety incident / event. The identified risks and associated communications may be provided and displayed for validation to allow for correction of any erroneous predictions. Any corrected predictions are then fed back into the machine learning model to learn from the corrected predictions. This allows the system to evolve over time to identify safety risks based on monitored communications, and the system can associate safety risks to specific construction projects. A more detailed description is provided below.

[0017] Training Phase 1, components of the initial training phase are indicated by dashed line 105. In one embodiment, training data 110 is input to a machine learning model that includes multiple independently operating base machine learning classifiers / algorithms. Each classifier produces an output that classifies the communication being evaluated, and all outputs are combined to produce an ensemble majority voting classifier 130.

[0018] In FIG. 1, the risk detection system 100 includes an odd number (three) of machine learning models / classifiers / algorithms 115, 120, and 125. In other embodiments, a different number of base classifiers may be used. In one embodiment, each base classifier is selected based on operating from a different theoretical background from the other classifiers to avoid bias and redundancy. For example, the three classifiers shown include: (1) a logistic regression classifier with L1 regularization 115, which is a parametric classifier; (2) a gradient boosting classifier XGBoost 120, which uses a gradient boosting framework; and (3) a random forest classifier 125, which is an ensemble learning method that operates by building multiple decision trees and implements machine learning algorithms under a bootstrap aggregation framework.

[0019] Training data 110 is input to each of the machine learning classifiers 115, 120, and 125. For example, the training data may include a known dataset of construction project communications that includes known safety risk text / language and known non-risk language. Known safety risk text / language includes text, language, and / or phrases that are known to refer to or discuss health and safety issues and / or events.

[0020] Classification problem structure A labeled dataset of over 40,000 health and safety risk related text datasets and about 6,000 non-risk related text datasets was prepared and used to test and train the machine learning models / classifiers 115, 120, and 125. Labeled refers to communications that are classified and known to have text that is either a health and safety risk or a non-safety risk. The training dataset was split into a 90% "training" and a 10% "testing" dataset. In one embodiment, the ratio of risk samples to non-risk samples was the same in both the test and training datasets. Of course, different ratios may be used and different amounts of samples may be used. The risk related samples included text data from communications from different construction projects and text data from health and safety risk related injury reports obtained from the Occupational Safety and Health Administration (OSHA) website. The non-risk samples also included text data from communications from different construction projects but did not include any health and safety risks associated with them.

[0021] In one embodiment, the communication text from each record was cleaned by removing stop words, punctuation, numbers, and HTML tags, and all words were stemmed to roots with all lowercase letters. Each communication was vectorized using a model that represents each document as a vector. For example, each communication was vectorized into a vector size of a selected number of features (e.g., several thousand features) using Gensim's Doc2Vec in Python into a document vector matrix where each row represents a unique communication thread and each column represents a feature in vector space. A simplified example of a document vector matrix is ​​shown in FIG. 2.

[0022] Each feature was normalized to a mean of 0 and a standard deviation of 1. The dataset was split into 90% for training and 10% for testing. After model regularization, the machine learning model made predictions on a testing dataset of 10,000 previously unseen records to determine its accuracy in prediction.

[0023] A number of pairs of observations (x i ,y i ) i = 1,...,n, where

[0024]

number

[0025] and y∈Y={health and safety risk communications, non-risk communications} are observed. X is the predictor space (or attributes) and Y is the response space (or classes).

[0026] In this case, the number of attributes is a feature of the vector obtained upon vectorization of each communication thread text. In one embodiment, the pre-trained vectorization model uses Gensim's Doc2Vec library for document / text vectorization, topic modeling, word embedding, and similarity. Text2vec may also be used. The first step is to vectorize the text using vocabulary-based vectorization, where unique terms are collected from a group of input documents (e.g., a group of email communications and threads), and each term is marked with a unique ID. For example, this may be performed using the create_vocabulary() function, which identifies and collects unique terms and collects statistics for the terms.

[0027] The risk detection system 100 then generates a vocabulary-based document-term matrix (DTM) using a pre-trained vectorization model Doc2Vec. This process converts each communication thread (e.g., an email thread or a single email communication) into a numerical representation of the communication in a vector space, also called a text embedding.

[0028] This text embedding process converts text into a numerical representation (embedding) of the semantic meaning of the text. If two words or documents have similar embeddings, they are semantically similar. Thus, using the numerical representation, the risk detection system 100 can capture the word's context, semantic and syntactic similarity, association with other words, etc. in the document.

[0029] In one embodiment, the entire dataset is converted into an [M×N] matrix (see Table 1), where M is the number of communication threads and N is the total number of features in the vector space. Each communication thread is represented by an N-dimensional vector. As seen in Table 1, each row represents a unique communication thread (Doc1, Doc2, Doc3...), and each column represents a unique term / attribute feature or term (T1, T2, T3...) present in the document set. The feature sets shown in Table 1 for the eight features are for display purposes only. The values ​​shown in each column are the frequency of the term occurring in the set of documents in the vector space (e.g., use T1 occurs twice in Doc1). Each row is a vector corresponding to the associated Doc representing the frequency of each term. Some datasets may have thousands of terms, which can be a large processing task for training a machine learning model. In one embodiment, the vectors may be filtered based on the features.

[0030] Table 1 shows an exemplary document term matrix.

[0031] [Table 1]

[0032] In another embodiment, the text embedding process converts the text into a numerical representation (embedding) into a document vector matrix rather than a document term matrix. Referring to FIG. 2, an example of a document vector matrix for one email communication thread with four features is shown. In FIG. 2, the matrix is ​​a matrix of document vectors or embeddings, with each row representing a vector representation of a unique communication thread in vector space. For example, the features are some characteristics of the document other than the terms and their associated numerical representations as resulting from the vectorization process. As mentioned above, the numerical representations are the semantic meaning of the text. As mentioned above, a thread may have hundreds or thousands of features (e.g., 7000 features). In one embodiment, the document vector matrix is ​​obtained from the corpus of communication or text data records used in the sample dataset, for example, using the Doc2vec library in Python.

[0033] A communication ID 205 is assigned to each particular communication thread (data 210). Four exemplary features are listed as feature 1, feature 2, feature 3, and feature 4, where each feature in the vector space is a feature from the communication generated by a pre-trained vectorization model. The generic terms "feature 1," "feature 2," etc. are used for simplicity and for descriptive purposes only. The labels for each of these features may be generated by the model and do not have any physical meaning in this description. The labels may instead be represented as other types of strings based on how the model is configured to generate such labels.

[0034] The goal is to estimate the relationship between X and Y, and thus use these observations to predict X from Y. The relationships are then used to define classification rules. h j (X) = argmaxP(y|X,θ j ),j=1,...,3 (Equation 1) It is shown as:

[0035] where P(.,.) is the probability distribution of the observed pairs, θ is a parameter vector for each base classifier, and j is the number of base classifiers. Because the implementation of risk detection system 100 has three base classifiers 115, 120, and 125, there are three classification rules, one classification rule for each base classifier. Thus, j=3.

[0036] In Figure 2, numbers such as -0.0155624, -0.0561929, etc. are shown under the columns for Features 1-4. These numbers represent example values ​​of each feature in a document vector (a vectorized representation of each document).

[0037] Data Preparation In one embodiment, the labeled dataset is generated from a sample group of communications having known safety risk text (e.g., communications identified with known text or phrases relating to or discussing health and safety issues) and a sample group of communications having known non-risk text (e.g., communications identified as not having known text or phrases relating to or discussing health and safety issues). For example, the labeled dataset may include about 1000 unique records of communication threads generated from about a 50%-50% ratio of known safety risk communications (having known safety risk text) and non-risk communications not relating to or discussing health or safety issues. Of course, different amounts of data records may be used in the dataset.

[0038] In addition to having known safety risk text and known non-risk text, the communications from the dataset may include known construction and / or engineering vocabulary and terminology. For example, construction terminology can be collected and input from an existing glossary or database of construction project terminology. This allows the machine learning model to learn and identify whether a received email communication is related to a construction project or not. This feature may be useful when the system operates on a generic email system that includes non-construction communications that should be removed to avoid unnecessary classification and use of computing resources (e.g., avoid using machine classifiers, avoid processor time, memory, etc.).

[0039] In another embodiment, the system may be trained to identify and target different types of communications instead of construction and engineering. Communications from the dataset may include known vocabulary and terminology from different target fields, such as aviation, marine, transportation, or other selected target fields. This allows the machine learning model to learn and identify whether a received email communication is related to or not relevant to the target field. In one embodiment, once it is determined that the email / communication is not relevant to the target field, the email / communication may be removed or not provided for further analysis.

[0040] The text of the correspondence from each transcript is cleaned by removal of stop words, punctuation, numbers, and HTML tags. Remaining words are stemmed to their roots in all lowercase.

[0041] The vocabulary was generated from a sample set of communication threads (e.g., 500+ emails and / or threads) that defined known safety risk vocabulary and known non-risk vocabulary. The sample set included a subset of communications that were known to discuss or mention safety risk issues and therefore included known safety risk vocabulary. Another subset of communication threads was known not to discuss or mention safety risk issues and therefore had known non-risk vocabulary.

[0042] Examples of vocabulary for known health and safety risks may include "accident," "fatal," "injury," "harmful," "danger," "damage," "hazard," "unsafe," "negligence," etc. Examples of vocabulary for known non-health and safety risks may include "contract," "document," "supplies," "employee," "road," "building," etc. In one embodiment, the sample set included the same or approximately the same ratio of safety risk communication samples and non-risk samples, although other ratios (e.g., 60% risk to 40% non-risk) can be used. The ratio is not relevant as long as it is used to train the machine learning model to accurately identify and classify between safety risk and non-risk communications to at least a specified threshold level. In this example, the sample set of communications was also based on a target field that was construction and engineering. Thus, the sample set of communications included a combination of construction / engineering vocabulary and safety risk / non-risk vocabulary. Each of the sample communication threads was then vectorized to generate a document vector matrix (see Figure 2) using the Doc2Vec library (or other vectorization function), where each row represents a unique communication thread and each column represents a feature in vector space.

[0043] In one embodiment, each feature was normalized to a mean of 0 and a standard deviation of 1. The dataset was divided into 90% for training and 10% for testing. The ratio of safety risk to non-risk communication was the same (or nearly the same) in both the testing and training datasets. The training dataset was provided as input to the machine learning model mentioned below. After model regularization, predictions were made on the testing dataset and a dataset of previously unseen records (e.g., 10,000 previously unseen records with unknown text). Any records predicted as non-risk by all three models were added back to the original labeled dataset records as a joint training dataset to increase the size of the labeled training and testing datasets for building the model.

[0044] With continued reference to FIG. 1 , the following includes a description of the machine learning algorithms for each of the machine learning models: logistic regression 115, XGBoost 120, and random forest 125.

[0045] Machine learning models / algorithms 1. Logistic regression model with L1 regularization 115 The input for the logistic regression model 115 is the scaled document-term matrix generated by the risk detection system 100 as described above [Table 1 and Figure 2]. A penalized logistic regression model 115 with L1 regularization was constructed, which penalizes the logistic model for having too many variables. The coefficients of some less contributing variables are forced to exactly zero in the Lasso regression. In one embodiment, only the most important variables are kept in the final model.

[0046] In logistic regression, the C parameter describes the inverse of the regularization strength. In the present model, the C parameter was found to be optimal at a value of 50. 10-fold cross-validation was performed. The maximum number of iterations it takes for the solver to converge was set to 1000 and the tolerance for optimization was taken as 1e-4.

[0047] The output of the logistic regression model 115 is a probability risk value of the communication thread for being a health and safety risk. The initial threshold for the probability risk value of the communication thread for being a safety risk was taken to be any value higher than 0.999. The initial threshold was chosen to be a cutoff value for the probability that the sensitivity, specificity, and accuracy are very close to each other using a grid search (see FIG. 3, which shows a graph of the probability of the safety risk cutoff value against the sensitivity (TPR-True Positive Rate) and specificity (TNR-True Negative Rate)). The threshold was then slightly modified according to how the model performed over unseen datasets (unknown / unclassified datasets). The threshold may be lowered to increase the number of communications classified as safety risks. However, this may decrease the accuracy of the model by potentially increasing the number of communications that are incorrectly classified as safety risks.

[0048] Table 2 shows the scale of evaluation in the test data by the logistic regression model.

[0049] [Table 2]

[0050] The ROC-AUC score is a measure of evaluation to assess the performance of a classification model. AUC is the "Area Under the ROC Curve". AUC measures the entire two-dimensional area under the entire ROC curve from (0,0) to (1,1). The ROC curve stands for "Receiver Operating Characteristic" curve and is a graph that shows the performance of a classification model at all classification thresholds. The ROC curve plots two parameters, namely True Positive Rate (TPR) and False Positive Rate (FPR), and the curve plots TPR vs. FPR at different classification thresholds.

[0051] 2. Gradient Boosting Algorithm Ensemble Algorithm Using XGBoost The input for the XGBoost model 120 is the scaled document-term matrix generated by the system 100 as described above (e.g., FIG. 2, Table 1). XGBoost is an optimized distributed gradient boosting library designed to be highly efficient, flexible and portable. It implements machine learning algorithms under a gradient boosting framework. In the XGBoost model 120, this is implemented by the gradient boosting trees algorithm. The input matrix is ​​the same document-term matrix (document-feature matrix) as mentioned above.

[0052] In tree-based ensemble methods such as XGBoost or Random Forest, each feature is evaluated as a potential split feature, which makes them robust to unimportant / irrelevant variables, since such variables that cannot distinguish between events / non-events will not be selected as split variables and thus will also be very low in the variable importance graph.

[0053] The "AUC" under the "ROC curve" was used as the evaluation measure. The number of trees was selected as 50 and the learning rate was selected as 0.2 for hyperparameter tuning for regularization after grid search over a range of values. The threshold for the probability risk value of a communication thread to be a safety risk was taken to be any value higher than 0.999. The threshold was selected based on the model performance over unseen datasets in the same way as was done for the logistic regression model. The output of the XGBoost model 120 is the probability risk value of a communication thread to be a safety risk.

[0054] Table 3 shows the evaluation measures for XGBoost on the test data.

[0055] [Table 3]

[0056] 3. Random Forest Classifier125 The input for the random forest classifier model 125 is a scaled document-term matrix (document-frequency matrix) as described above. In one embodiment, the random forest classifier 125 is constructed with four variables available for splitting at each tree node, selected through a grid search over a range of values.

[0057] The "AUC" under the "ROC curve" was taken as the evaluation measure and the number of trees was taken as 500. The number of features to consider when searching for the best split was taken to be 50 for hyperparameter tuning for regularization after grid search over a range of values. The threshold for the probability risk value of a communication thread for being a safety risk was taken to be any value higher than 0.999. The threshold was selected based on the model performance over unseen datasets in the same way as was done for the logistic regression model. The output of the random forest classifier model 125 is the probability risk value of a communication thread for being a safety risk.

[0058] Table 4 shows the evaluation measures for the Random Forest classifier on the test data.

[0059] [Table 4]

[0060] Ensemble Majority Voting Classification Continuing to refer to FIG. 1, each of the three base classifiers 115, 120, and 125 is an expert in a different region of the predictor space because each classifier processes the attribute space under a different rationale. The risk detection system 100 combines the outputs of the three classifiers 115, 120, and 125 to produce an ensemble majority voting classifier 130 that outperforms any of the individual classifiers and their rules. Thus, with an odd number of three classifiers, the majority vote / prediction for the final result requires at least two classifiers to vote / predict the same result (e.g., predicting either "safety risk" or "non-risk" for communication).

[0061] In operation, an electronic communication (e.g., an email thread) is evaluated by each model 115, 120, 125, and the output of each model is the probability of the email thread being a health and safety risk. Based on the probability compared to a threshold, a label is assigned for each email thread. For example, if the probability of the email thread being a safety risk as predicted by the model is greater than the threshold of probability of being a risk for that particular model, the label is "1". If the probability of the email thread being a safety risk as predicted by the model is less than the threshold of probability of being a risk for that particular model, the label is "0". As a result, if the email thread is of a non-risk nature (not a health and safety risk), the label is "0".

[0062] Of course, other labels may be used to indicate safety risk or no risk. In one embodiment, by using "1" and "0" as labels, the labels can be used as votes, and these votes may then be combined from multiple machine learning classifiers to generate the majority voting scheme of the ensemble model 130, as described below.

[0063] In one embodiment, the risk detection system 100 uses the following formula to combine the outputs from the three base classifiers into the ensemble model 130:

[0064] C(X)=h1(X)+h2(X)+h3(X) (Formula 2) where C(X) is the sum of the weighted outputs of the three individual classifiers, and h1(X), h2(X), and h3(X) are the outputs of the Random Forest 125, XGBoost Gradient Boosting 120, and Logistic Regression Classifier 115, respectively. Here, C, h1, h2, and h3 are all functions of X that represent features or attributes identified from the electronic communication being evaluated. In another embodiment, one or more of the classifier outputs can be given a weighting value in Equation 2, such as 2*H1(X).

[0065] In one embodiment, the system 100 classifies an electronic communication as a safety risk if C(X)≧2. If C(X)<2, the communication is classified as a non-risk (probably not a health and safety risk). Thus, if any two of the three base classifiers classify the communication as a health and safety risk, the ensemble model predicts that the communication is a safety risk. In another embodiment, as an extension to this binary classification, each classifier may classify each risk email into a high, medium, and low level of risk depending on their strength / severity of the risk.

[0066] Table 5 shows the evaluation measures for the ensemble model 130.

[0067] [Table 5]

[0068] Model Deployment - Operation / Execution Phase Continuing to refer to FIG. 1, in one embodiment, once the ensemble model 130 is constructed and trained, the ensemble model 130 is deployed for operation (block 135). During operation, communications are monitored and evaluated in real-time or near real-time for safety risk and non-risk content (block 140, also FIG. 4). The ensemble model 130 generates a risk prediction for each communication based on the text as described above and generates an associated label as safety risk or non-risk. Components of the deployed model are indicated by dashed lines 145.

[0069] When the ensemble model 130 determines and predicts that a communication is a safety risk, an electronic notification (block 160) is generated via the graphical user interface (block 140). The deployment and operation of the ensemble model is described with reference to FIG.

[0070] 1 may be configured to provide a number of additional features that are generated and displayed in a graphical user interface. These features may include dashboards 165, alerts and issued tracking 179, recommendations 175, and / or aggregations 180.

[0071] For example, a dashboard (block 165) may be generated to graphically display one or more types of results and / or summary data from the ensemble model 130. For example, a summary may include a number of safety risk emails exchanged in a project visible to a particular organization / individual with the project within a particular time interval. Other types of summary reports / information regarding analyzed communications, statistics, and / or data analysis reports may be included in the dashboard 165 as graphical information. The display may also include whether each of the emails or communications displayed in the dashboard 165 has safety risk content.

[0072] Alerts and Issue Tracking (Block 170): In one embodiment, the system 100 highlights topics and keywords that potentially point to why an email or communication has potential safety risk content as identified by the machine learning model. Alerts and issue tracking 170 may be combined with recommendations 175.

[0073] Recommendation (block 175): In one embodiment, the system 100 may categorize a project as high, medium or low risk category depending on the number of safety risk emails exchanged in the project visible to a particular organization / individual with the project within a particular time interval. This allows the stakeholders to take appropriate measures as quickly as possible.

[0074] Aggregation (block 180): In one embodiment, the system 100 may determine the percentage of safety risk emails among all emails exchanged in a project that are visible to a particular organization / individual with the project within a particular time interval.

[0075] 4, one embodiment of a method 400 illustrating the operation of ensemble model 130 during deployment and execution is shown. As previously described, ensemble model 130 is configured to monitor electronic communications and detect safety risks from electronic communications related to a construction or engineering project. In one embodiment, ensemble model 130 is configured as part of a selected computing platform and / or email network that receives the monitored electronic communications.

[0076] Generally, after the machine learning classifiers are constructed from the training data set (as described with respect to Figs. 1-3), new incoming email communications are automatically passed through each implemented classifier. In the system of Fig. 1, three classifiers 115, 120 and 125 are included. After analyzing the communications, each classifier classifies / labels each incoming communication as either a safety risk or a non-risk. In another embodiment, each classifier may classify identified risks as low, medium or high with a level of risk severity / intensity that is a safety risk. The ensemble model 130 may continuously learn from user feedback that helps validate the results, and these results are then fed back into the system for retraining. A more detailed description is provided below.

[0077] 4, at block 410, method 400 begins and, once operational at a targeted computing platform, network communications are monitored to identify electronic communications received by the computing platform. For example, emails or other electronic communications are identified by the associated email system in which the system operates.

[0078] At 420, the system detects and identifies the email and its associated construction project. For example, the system 100 may have a list of identified projects, and the system identifies which project the email belongs to. As previously described, the machine learning models 115, 120, and 125 are trained to identify construction and engineering vocabulary and terminology. This type of identification may help eliminate emails or email threads that are not related to construction projects.

[0079] As another example, an organization may have one or more ongoing construction projects, each with a defined name and / or other metadata stored in the system that identifies each project. The system may parse and scan the text from the received email to identify any known words or phrases that match existing project IDs and metadata. If found, the received email is related to an existing project. Other methods of identification may include having the project ID in the email.

[0080] Each incoming email communication is further passed through a number of functions to programmatically clean the communication. For example, in block 430, each email may be cleaned by removal of all non-Latin alphabet characters, html tags, punctuation, numbers, and stop words. The email text may be tokenized by identifying the communication text and breaking it down into words, punctuation marks, numbers, other objects in the text, etc. If the email contains at least one word that is more than four letters, each word in the email is stemmed to a root.

[0081] At 440, the tokenized text words from the email are vectorized and feature scaled. In one embodiment, vectorizing the text involves converting each word into a number, which is a numeric vector. Vectorization maps words or phrases from a vocabulary to a corresponding vector of real numbers, which may be used for word prediction, word similarity, and / or semantic discovery. The vector of numbers (i.e., features) are features and may be scaled by normalizing the range of the features in the data.

[0082] At 450, after cleaning and feature scaling the communication text, the ensemble machine learning classifier of FIG. 1 is initiated to identify construction terms and classify the safety risk of the communication. The communication text is passed through each of the three machine learning classifiers 115, 120 and 125 of system 100 (FIG. 1). Each classifier makes an individual prediction of whether the email text is a safety risk or non-risk based on the learned training data.

[0083] In one embodiment at block 460, the numeric vectors generated at block 440 are mapped to numeric vectors associated with a defined data set of known safety risk vocabulary and known non-risk vocabulary (e.g., from a previously generated document-term matrix (or document-vector matrix)). In other words, the machine learning classifier processes the numeric vectors generated from the email communication by at least matching and comparing the numeric vectors to known numeric vectors generated from a set of defined safety risk vocabulary and a set of defined non-risk vocabulary.

[0084] At block 470, each of the three classifiers 115, 120 and 125 independently evaluates the communication and generates a prediction of a probability risk value for the communication evaluated as previously described above. If the probability risk value exceeds a defined threshold set for the associated classifier, then that classifier labels the communication as referring to or discussing a safety risk (e.g., a label value of "1") or as non-risk (e.g., a label value of "0" does not discuss a safety risk). In general, the output label is viewed as a "vote" since the output is either "1" (safety risk "yes") or "0" (safety risk "no"). The multiple "votes" generated by the multiple classifiers are then combined for a majority decision.

[0085] At block 480, the three labels / votes output by the three classifiers are then combined using a majority voting scheme (e.g., Equation 2) as part of the ensemble model 130. Based on the combined label, the email communication is given a final label by the system as a safety risk or non-risk based on the majority vote of the individual votes of the three classifiers. In another embodiment, a different number of classifiers may be used and / or the selected classifier may have its output votes given a weighted value in Equation 2 to avoid ties in the votes.

[0086] The ensemble classifier includes an odd number of independent machine learning classifiers (three classifiers in the above embodiment). Each of the independent machine learning classifiers generates an output that classifies the email communication as a safety risk or non-risk. The outputs from each of the independent machine learning classifiers are all combined, at least in part, based on a majority voting scheme to generate a final label for the email as being a safety risk or non-risk.

[0087] At block 490, the system is configured to generate an electronic notification in response to the final label indicating that the communication is a safety risk. In one embodiment, the electronic notification includes data identifying the communication, the associated construction project, and an alert message regarding the potential safety risk. The electronic notification may also include additional data, such as the email sender and recipient. The electronic notification may highlight or visually distinguish text from the email communication associated with the safety risk vocabulary as identified by the machine learning classifier. The electronic notification is then sent to a remote device and / or displayed on a graphical user interface, so that a user can receive the notification and have access to the communication in near real-time, so that action can be taken to resolve the issue in the communication.

[0088] In another embodiment, the system sends an electronic notification to a designated remote device (e.g., via an address, mobile phone number, or other device ID) that includes at least an identifier of the email and a label indicating that the email contains text that mentions or discusses a safety risk or non-risk. In response to receiving the electronic notification, the remote device provides a user interface that displays data from the electronic notification and allows input to verify the label and change the label if the user believes the label is incorrect. This may include viewing any identified suspicious text from the email communication to allow the user to determine whether the text is a safety risk or non-risk based on the user's judgment. The user interface allows the label to be selected and changed. In response to the label being changed via the user interface, the system may then send the changed label and the corresponding email as feedback to the machine learning classifier to retrain the machine learning classifier. The verification mechanism is further described in the following section.

[0089] Validation and Continuous Learning Referring again to FIG. 1, in one embodiment, the communication text for one or more predictions made by the ensemble model 130 may be made available to a user of the system 100. This provides a validation mechanism so that the user can apply human decision-making to validate the predictions and associated labels (safety risk or non-risk). As part of the validation mechanism, the system 100 provides a feedback user interface 150 that allows the user to input corrections to re-tag or otherwise re-label selected communications as safety risk or non-risk if the user disagrees with the predictions and labels made by the ensemble model 130.

[0090] A continuous learning process is performed to retrain the ensemble model 130 with new feedback data that changes the previous labels. The ensemble model 130 receives as input the label changes and other feedback data to be combined with and retrained by the existing training dataset of classified data (block 155). This feedback data 155 is used to retrain the ensemble model 130 with the previous and newly labeled communication bodies. If the retrained model outperforms the existing model, the retrained ensemble model 130 replaces the existing model. This may be based on performing multiple comparison tests to determine the accuracy of the model in prediction. Using this feedback mechanism, the risk detection system 100 learns to classify communications more accurately over a period of time.

[0091] 5, another embodiment of a method 500 is shown that illustrates the operation of the ensemble model 130 during deployment and execution. As previously described, the ensemble model 130 (from FIG. 1) is configured to monitor electronic communications and detect security risks therefrom. In the method 500, the ensemble model 130 is trained and configured to identify communications associated with a target field.

[0092] The target field may be a selected field or activity, such as a construction project, an aviation, a shipping activity, a warehouse project, a product manufacturing project, or any other selected field or activity. In one embodiment, the ensemble model 130 is configured as part of a selected computing platform and / or email network that receives monitored electronic communications for the target field. Thus, the ensemble model 130 from FIG. 1 has been previously trained with a vocabulary associated with the target field (as described with respect to FIGS. 1-3).

[0093] Generally, after the machine learning classifiers are trained, new incoming email communications that are monitored are automatically passed through each classifier implemented. In the system of Figure 1, three classifiers 115, 120 and 125 are included. After analyzing the communications, each classifier classifies / labels each incoming communication as either a safety risk or a non-risk.

[0094] 5, at block 510, method 500 begins and, functioning on a targeted computing platform, network communications are monitored to identify electronic communications received by the computing platform associated with the target field. For example, emails or other electronic communications are detected by the associated email system on which the system operates.

[0095] In one embodiment, each incoming email communication is passed through a number of functions to programmatically clean the communication. For example, each email may be cleaned by removal of all non-Latin alphabet characters, html tags, punctuation, numbers, and stop words. The email text is tokenized by identifying the communication text and breaking it down into words, punctuation marks, numbers, other objects in the text, etc. If the email contains at least one word that is more than four characters, each word in the email may be stemmed to their root. Based on the remaining terms in the communication, the system identifies whether the vocabulary matches the known vocabulary of the target field as previously described. This type of identification may help eliminate emails or email threads that are not related to the target field.

[0096] In block 520, the ensemble machine learning classifiers of Figure 1 begin to classify text from risky email communications as related to a safety risk or non-risk. The communication text is passed through each of the three machine learning classifiers 115, 120, and 125 of system 100 (Figure 1). Each classifier makes an individual prediction of whether the email text mentions or discusses a safety risk or is non-risk based on learned training data.

[0097] At block 530, each of the three classifiers 115, 120 and 125 independently evaluates the communication and generates a prediction of a probability risk value for the communication evaluated as previously described above. If the probability risk value exceeds a defined threshold set for the associated classifier, then at block 540, the classifier labels the communication as a safety risk (e.g., a label value of "1") or non-risk (e.g., a label value of "0"). In general, the output label is viewed as a "vote" because the output is either "1" (safety risk "yes") or "0" (safety risk "no"), based at least in part on the probability risk value indicating that the email is a safety risk. The multiple "votes" generated by the multiple classifiers are then combined for a majority decision.

[0098] At block 540, the labels / votes output by the machine learning classifiers are then combined using a majority voting scheme (e.g., Equation 2) as part of the ensemble model 130. Based on the combined labels, the email communication is given a final label by the system as a safety risk or non-risk based on the majority vote of the individual votes of the three classifiers. In another embodiment, a different number of classifiers may be used and / or the selected classifier may have its output votes given a weighted value in Equation 2 to avoid ties in the votes.

[0099] At block 550, the system generates an electronic notification in response to the final label indicating that the communication is a safety risk. In one embodiment, the electronic notification includes data identifying the communication, the associated project (if determined), and an alert message regarding the potential safety risk. The electronic notification may include additional data, such as the email sender and recipient. The electronic notification may highlight or visually distinguish text from the email communication that is associated with the safety risk vocabulary as identified by the machine learning classifier. The electronic notification is then sent to a remote device and / or displayed on a graphical user interface, so that a user can receive the notification and have access to the communication in near real-time, so that action can be taken to resolve the issue in the communication.

[0100] In another embodiment, the system sends an electronic notification to a designated remote device (e.g., via an address, mobile phone number, or other device ID) that includes at least an identifier of the email and a label indicating the email as a safety risk or non-risk. In another embodiment, method 500 may perform one or more of the functions or sub-functions as described with reference to FIG. 4 and method 400.

[0101] In response to a final label indicating that the communication is a non-safety risk, no electronic notification may be generated and the system continues to process the next communication. The labeled non-risk communication may be sent to a remote device-associated verification, such as the verification mechanism described in FIG. 1, where someone may review the communication and verify that the label is correct.

[0102] Through the present system and method, email communications may be classified as safety risk or non-risk in real-time or near real-time. Such communications classified as safety risk may indicate early indications of potential issues that may lead to larger health and safety incidents that are most likely to have adverse or catastrophic impacts on the project / asset under construction. Thus, the present system enables early action to be taken to effectively mitigate this safety risk whenever it is proactively identified by the present system.

[0103] No act or function described or claimed herein is performed by the human mind, and any interpretation that any act or function can be performed in the human mind is inconsistent and contrary to this disclosure.

[0104] Cloud or Enterprise Implementation In one embodiment, the safety risk detection system 100 is a computing / data processing system including an application or a collection of distributed applications for an enterprise organization. The application and the safety risk detection system 100 may be configured to operate with or be implemented as a cloud-based networking system, a software as a service (SaaS) architecture, or other types of networked computing solutions. In one embodiment, the risk detection system is a centralized server-side application that provides at least the functions disclosed herein and is accessed by many users via computing devices / terminals that communicate with the risk detection system 100 (which functions as a server) over a computer network.

[0105] In one embodiment, one or more of the components described herein are configured as program modules stored on a non-transitory computer-readable medium, the program modules comprising stored instructions that, when executed by at least a processor, cause a computing device to perform corresponding functions as described herein.

[0106] Computing Device Embodiments In one embodiment, FIG. 6 illustrates a computing system 600 configured and / or programmed as a special-purpose computing device comprising one or more components of the safety risk prediction / detection systems and methods described herein and / or equivalents.

[0107] The exemplary computing system 600 may be a computer 605 including a hardware processor 610, a memory 615, and input / output ports 620 operatively connected by a bus 625. In one example, the computer 605 is configured with the safety risk prediction / detection system 100 as shown and described with reference to Figures 1-4. In different examples, the safety risk prediction / detection system 100 may be implemented in hardware, a non-transitory computer-readable medium with stored instructions, firmware, and / or combinations thereof.

[0108] In one embodiment, the risk prediction / detection system 100 and / or the computer 605 are means (e.g., structure, hardware, non-transitory computer-readable medium, firmware) for performing the described operations. In some embodiments, the computing device may be a server operating in a cloud computing system, a server configured in a Software as a Service (SaaS) architecture, a smartphone, a laptop, a tablet computing device, etc.

[0109] The risk prediction / detection system 100 may be implemented as stored computer-executable instructions that are temporarily stored in memory 615 and then provided to the computer 605 as data 640 that are executed by the processor 610.

[0110] Although an exemplary configuration of computer 605 has been generally described, processor 610 may be a wide variety of processors, including dual microprocessors and other multi-processor architectures. Memory 615 may include volatile and / or non-volatile memory. Non-volatile memory may include, for example, ROM, PROM, EPROM, EEPROM, etc. Volatile memory may include, for example, RAM, SRAM, DRAM, etc.

[0111] The storage disk 635 may be operatively connected to the computer 605, for example, via an input / output (I / O) interface (e.g., card, device) 645 and an input / output port 1020. The disk 635 may be, for example, a magnetic disk drive, a solid state disk drive, a floppy disk drive, a tape drive, a zip drive, a flash memory card, a memory stick, etc. Furthermore, the disk 635 may be a CD-ROM drive, a CD-R drive, a CD-RW drive, a DVD ROM, etc. The memory 615 may store, for example, processes 650 and / or data 640. The disk 635 and / or the memory 615 may store an operating system that controls and allocates resources of the computer 605.

[0112] The computer 605 may interact with input / output (I / O) devices through an I / O interface 645 and input / output ports 620. Communications between the processor 610 and the I / O interfaces and ports 620 are managed by an input / output controller 647. The input / output ports 620 may include, for example, serial ports, parallel ports, and USB ports.

[0113] The computer 605 can operate in a networked environment and thus can be connected to a network device 655 via the I / O interface 645 and / or the I / O port 620. Through the network device 655, the computer 605 can interact with a network 660. Through the network 660, the computer 605 can be logically connected to a remote computer 665. The networks with which the computer 605 can interact include, but are not limited to, a LAN, a WAN, and other networks.

[0114] The computer 605 can send and receive information and signals from one or more output or input devices through the I / O ports 620. The output devices include one or more displays 670, printers 672 (such as inkjet, laser, or 3D printers), and audio output devices 674 (such as speakers or headphones). The input devices include one or more text input devices 680 (such as a keyboard), cursor control devices 682 (such as a mouse, touchpad, or touch screen), audio input devices 684 (such as a microphone), video input devices 686 (video and still cameras), or other input devices such as a scanner 688. The input / output devices may further include disks 635, network devices 655, and the like. In some cases, the computer 605 can be controlled by information or signals generated or provided by input or output devices, such as the text input devices 680, cursor control devices 682, audio input devices 684, disks 635, and network devices 655.

[0115] Definitions and Further Embodiments In another embodiment, the described methods and / or their equivalents may be implemented by computer-executable instructions in the form of an executable application (either a standalone application or part of a larger system). Thus, in one embodiment, a non-transitory computer-readable / storage medium is configured with stored computer-executable instructions of an algorithm / executable application that, when executed by a machine, causes the machine (and / or associated components) to perform the method. Exemplary machines include, but are not limited to, processors, computers, servers operating in a cloud computing system, servers configured in a Software as a Service (SaaS) architecture, smartphones, etc. In one embodiment, a computing device is implemented by one or more executable algorithms configured to perform any of the disclosed methods.

[0116] In one or more embodiments, the disclosed methods or their equivalents are performed by either computer hardware configured to perform the methods, or by computer instructions embodied in modules stored on a non-transitory computer-readable medium, where the instructions are configured as an executable algorithm configured to perform the methods when executed by at least a processor of a computing device.

[0117] For ease of explanation, the methods illustrated in the figures are shown and described as a series of algorithmic blocks, but it should be appreciated that the methods are not limited by the order of the blocks. Some blocks may occur in a different order and / or concurrently with other blocks than those shown and described. Furthermore, fewer than all of the blocks shown may be used to implement an example method. The blocks may be combined or separated into multiple operations / components. Furthermore, additional and / or alternative methods may use additional operations not shown in the blocks.

[0118] The following contains definitions of selected terms used herein. The definitions include various examples and / or forms of components that fall within the scope of the term and may be used for implementation. The examples are not intended to be limiting. Both singular and plural terms may be within the definitions.

[0119] References to "one embodiment," "embodiment," "one example," "example," etc. indicate that the embodiment or example so described may include a particular feature, structure, characteristic, attribute, element, or limitation, but not necessarily all embodiments or examples include that particular feature, structure, characteristic, attribute, element, or limitation. Further, repeated use of the phrase "in one embodiment" does not necessarily refer to the same embodiment, but may.

[0120] As used herein, a "data structure" is an organization of data in a computing system stored in a memory, storage device, or other computerized system. A data structure may be, for example, any one of a data field, a data file, a data array, a data record, a database, a data table, a graph, a tree, a linked list, etc. A data structure may be formed from and include many other data structures (e.g., a database includes many data records). Other examples of data structures are possible according to other embodiments.

[0121] As used herein, a "computer-readable medium" or a "computer storage medium" refers to a non-transitory medium that stores instructions and / or data that, when executed, are configured to perform one or more of the disclosed functions. Data may function as instructions in some embodiments. Computer-readable media may take forms including, but not limited to, non-volatile media and volatile media. Non-volatile media may include, for example, optical disks, magnetic disks, and the like. Volatile media may include, for example, semiconductor memory, dynamic memory, and the like. Common forms of computer-readable media may include, but are not limited to, floppy disks, flexible disks, hard disks, magnetic tapes, other magnetic media, application specific integrated circuits (ASICs), programmable logic devices, compact disks (CDs), other optical media, random access memory (RAM), read only memory (ROM), memory chips or cards, memory sticks, solid state storage devices (SSDs), flash drives, and other media with which a computer, processor, or other electronic device can function. Each type of media may include stored instructions of an algorithm that, when selected for implementation in one embodiment, is configured to perform one or more of the functions disclosed and / or claimed.

[0122] As used herein, "logic" refers to a component implemented with computer or electrical hardware, non-transitory media with stored instructions of an executable application or program module, and / or combinations thereof, to perform any of the functions or operations as disclosed herein and / or to perform functions or operations from other logic, methods and / or systems as disclosed herein. Equivalent logic may include firmware, a microprocessor programmed with an algorithm, discrete logic (e.g., ASIC), at least one circuit, analog circuit, digital circuit, programmed logic device, memory device containing algorithmic instructions, etc., any of which may be configured to perform one or more of the disclosed functions. In one embodiment, logic may include one or more gates, combinations of gates, or other circuit components configured to perform one or more of the disclosed functions. Where multiple logics are described, it may be possible to incorporate multiple logics into one logic. Similarly, where one logic is described, it may be possible to distribute the one logic among multiple logics. In one embodiment, one or more of these logics are corresponding structures associated with performing the disclosed and / or claimed functions. The selection of which type of logic to perform may be based on desired system requirements or specifications. For example, if higher speed is a consideration, hardware is selected to perform the function. If lower cost is a consideration, stored instructions / executable applications are selected to perform the function.

[0123] An "operable connection" or a connection by which entities are "operably connected" is one in which signals, physical communications, and / or logical communications may be transmitted and / or received. An operable connection may include a physical interface, an electrical interface, and / or a data interface. An operable connection may include various combinations of interfaces and / or connections sufficient to permit operable control. For example, two entities can be operably connected to communicate signals to each other directly or through one or more intermediate entities (e.g., a processor, an operating system, logic, non-transitory computer-readable media). Logical and / or physical communication channels can be used to create an operable connection.

[0124] As used herein, a "user" includes, but is not limited to, one or more people, computers, or other devices, or any combination thereof.

[0125] Although the disclosed embodiments have been shown and described in considerable detail, it is not intended to restrict or in any way limit the scope of the appended claims to such details. Of course, it is not possible to describe every conceivable combination of components or methods to describe various aspects of the subject matter. Thus, the disclosure is not limited to the specific details or illustrative examples shown and described. Thus, the disclosure is intended to embrace changes, modifications and variations that fall within the scope of the appended claims.

[0126] To the extent the term "includes" or "including" is used in the detailed description or claims, it is intended to be inclusive in the same manner as the term "comprising" is interpreted when that term is used as a transitional phrase in a claim.

[0127] To the extent that the term "or" is used in the detailed description or claims (e.g., A or B), it is intended to mean "A or B or both." If the applicant intends to indicate "only A or B but not both," the phrase "only A or B but not both" is used. Thus, use of the term "or" herein is inclusive and not exclusive.

Claims

1. 1. A computer-implemented method executed by at least one computing device, comprising: monitoring email communications on the network to identify emails; responsive to receiving the email over the network, detecting and identifying the email as related to a construction project; tokenizing text from the email into words; vectorizing each of the plurality of words into a numeric vector that maps each word to a numeric value; initiating a machine learning classifier configured to identify construction terminology and to classify risky text as referring to a safety risk or a non-risk, and inputting the numeric vector generated from the email into the machine learning classifier; the machine learning classifier processes the numeric vectors from the email by at least matching the numeric vectors to a set of defined safety risk vocabulary and a set of defined non-risk vocabulary; The method comprises: generating, by the machine learning classifier, a probability risk value that the email contains vocabulary that refers to a safety risk; labeling the email as a safety risk or non-risk based, at least in part, on the probability risk value that the email mentions a safety risk; The computer-implemented method further includes generating and sending an electronic notification to a remote device in response to the email being labeled as referring to the safety risk to provide an alert.

2. the machine learning classifier comprises an ensemble classifier comprising a plurality of independent machine learning classifiers; each of the independent machine learning classifiers is configured to identify construction terminology; The method comprises: generating an output by each of the independent machine learning classifiers classifying the email as referring to a safety risk or non-risk; 10. The method of claim 1, further comprising: combining the output from each of the independent machine learning classifiers based at least in part on a majority vote to generate the label for the email as referring to the safety risk or the non-risk.

3. further comprising initiating a second machine learning classifier and a third machine learning classifier, both configured to identify construction terminology and predictively classify text as being a safety risk or non-risk; Each machine learning classifier is implemented in a different theoretical context to avoid bias and redundancy during classification; The method comprises: generating, by each of the machine learning classifiers, an individual prediction indicating whether the email mentions a safety risk or a non-risk to generate at least three individual predictions; The method of claim 1 or 2, further comprising labeling the email as referring to the safety risk or the non-risk based on a majority vote of the at least three individual predictions.

4. generating the electronic notification to include an identifier of the email and the label indicating the email as a safety risk or non-risk; providing a user interface for validating the label and allowing input to modify the label; 3. The method of claim 1 or 2, further comprising: in response to the label being changed via the user interface, feeding back the changed label and corresponding email to the machine learning classifier to retrain the machine learning classifier.

5. inputting the construction terminology into the machine learning classifier from a glossary or database of construction project terms; 3. The method of claim 1 or 2, further comprising: training the machine learning classifier to identify safety risk text based, at least in part, on a first dataset of communications having known text relating to or mentioning safety risks and a second dataset of communications having known non-risk text that does not mention health and safety risks.

6. 3. The method of claim 1 or 2, further comprising detecting and identifying the email as related to the construction project by evaluating the text from the email in relation to at least a trained dataset of construction terminology implemented by the machine learning classifier.

7. at least one processor configured to execute instructions; at least one memory operatively connected to said at least one processor; a machine learning classifier configured to identify construction terminology and classify risky text as a safety risk or non-risk; a non-transitory computer-readable medium containing computer-executable instructions stored thereon, the computer-executable instructions, when executed by the at least one processor, causing the computing system to: monitoring email communications on the network to identify emails sent; responsive to receiving the email over the network, detecting and identifying the email as related to a construction project; tokenizing text from the email into words; inputting the plurality of words generated from the email into the machine learning classifier; the machine learning classifier is configured to evaluate the plurality of words from the email by at least matching the plurality of words to a set of defined safety risk vocabulary and a set of defined non-risk vocabulary; generating, by the machine learning classifier, a probability risk value that the email contains text that mentions a safety risk; labeling the email as a safety risk or non-risk based, at least in part, on the probability risk value that the email contains text that mentions the safety risk; and generating and transmitting an electronic notification to a remote device in response to the email being labeled as a safety risk to provide a near real-time alert in connection with receiving the email on the network.

8. the machine learning classifier comprises an ensemble classifier comprising a plurality of independent machine learning classifiers; each of the independent machine learning classifiers is configured to identify construction terminology; each of the independent machine learning classifiers is configured to generate an output classifying the email as being a safety risk or a non-risk; 8. The computing system of claim 7, wherein the ensemble classifier is configured to combine the output from each of the independent machine learning classifiers based at least in part on a majority vote to generate the label for the email as being a safety risk or non-risk.

9. The machine learning classifier includes at least a first machine learning classifier, a second machine learning classifier, and a third machine learning classifier; each of the machine learning classifiers is configured to identify construction terminology and predictively classify text as a safety risk or non-risk; Each machine learning classifier is implemented in a different theoretical context to avoid bias and redundancy between classifications; each of the machine learning classifiers is configured to generate an individual prediction of whether the email is a safety risk or a non-risk to generate at least three individual predictions; 9. The computing system of claim 7 or 8, wherein the computing system is configured to label the email as a safety risk or non-risk based on a majority vote of the at least three individual predictions.

10. When executed by the at least one processor, the method causes the at least one processor to: sending the electronic notification to the remote device including an identifier of the email and the label indicating the email as a safety risk or non-risk; providing a user interface that allows input to verify and modify the label; 9. The computing system of claim 7 or 8, further comprising instructions for causing: in response to the label being changed via the user interface, feeding back the changed label and corresponding email to the machine learning classifier to retrain the machine learning classifier.

11. When executed by the at least one processor, the method causes the at least one processor to:

9. The computing system of claim 7 or 8, further comprising instructions to train the machine learning classifier to (i) identify safety risk text from a first dataset of communications having known text that references health and safety issues, and (ii) identify non-risk text from a second dataset of communications having known non-risk text that does not reference health and safety issues.

12. A program for causing a computer to execute the method described in claim 1 or 2.