A construction method of an occupational health analysis model based on big data
By constructing a multimodal machine learning model and feature transformation matrix fusion health coefficient, the problem of insufficient data processing capabilities in the existing technology is solved, efficient and accurate occupational health analysis is achieved, and the efficiency and adaptability of health management is improved.
Patent Information
- Application Number
- CN202411293372.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-09-14
AI Technical Summary
The existing occupational health analysis model based on big data has problems such as insufficient data processing capabilities, single analysis model, lack of dynamic adjustment and self-learning ability, resulting in insufficient efficiency and accuracy of occupational health management.
By obtaining occupational health data for preprocessing, extracting health features and labeling feature labels, building a multimodal machine learning model, combining feature transformation matrix and health coefficient for data fusion, building an occupational health analysis model based on big data, using feature transformation matrix and health coefficient for feature fusion, and building occupational health scores and evaluations.
It improves the efficiency and accuracy of occupational health data analysis, enhances the adaptability of the model, can achieve accurate and personalized health management, and provides better health management support for enterprises and employees.
Smart Images

Figure CN119864171B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of big data processing and artificial intelligence, and particularly to a method for constructing an occupational health analysis model based on big data. Background Art
[0002] Occupational health refers to ensuring the physical and mental health of workers and preventing the occurrence of occupational diseases through preventive measures and health management in the workplace. Analysis models based on big data mainly collect, integrate, process, and analyze large-scale data sets to provide a scientific basis for decision-making, risk assessment, trend prediction, etc.
[0003] In the context of modern industrialization and informatization, occupational health issues have received wide attention. Traditional occupational health management methods often rely on regular physical examinations and questionnaires, and the efficiency and real-time nature of this method are limited, and the roles of mental health and environmental factors are ignored. The rise of big data technology has provided new ideas and methods for occupational health management. Although some occupational health analysis models based on big data have been invented, there are still the following problems: insufficient data processing capabilities, single analysis models, and lack of dynamic adjustment and self-learning capabilities. Therefore, designing an occupational health analysis model based on big data to overcome the deficiencies of existing analysis models and achieve comprehensive and dynamic analysis of occupational health, and providing accurate and personalized health management services for enterprises and employees is of great significance. Summary of the Invention
[0004] The object of the present invention is to provide a method for constructing an occupational health analysis model based on big data.
[0005] To achieve the above object, the present invention is implemented according to the following technical solutions:
[0006] The present invention includes the following steps:
[0007] Obtain occupational health data and preprocess the occupational health data; the occupational health data includes historical data and environmental data;
[0008] Classify the historical data, extract features from the historical data to obtain health features, label feature tags for the health features according to the values of the health features, and construct a feature tag library according to the feature tags;
[0009] Construct a multi-modal machine learning model, and input the health features into the multi-modal machine learning model to obtain a health risk probability;
[0010] Interact the health characteristics with the environmental database composed of the environmental data according to the health risk probability to obtain interaction data, construct a feature transformation matrix based on the interaction data and the health characteristics, and input the interaction data and the mental health characteristics into a health function to obtain a health coefficient;
[0011] Fuse the feature transformation matrix and the health coefficient to obtain a fused feature, construct a big data-based occupational health analysis model based on the fused feature and feature labels, and input the occupational health data to be analyzed into the occupational health analysis model to obtain an occupational health score and a health evaluation.
[0012] Further, the method for extracting health characteristics from the historical data includes:
[0013] Perform hot encoding on non-numerical data, and use the random forest algorithm to divide the data. The data types include physical health data and mental health data;
[0014] Use principal component analysis to extract physical health characteristics from physical health data; the physical health characteristics are comprehensive characteristics related to physical health;
[0015] Use factor analysis to extract mental health characteristics from mental health data; the mental health characteristics are latent factors reflecting the meaning of mental health scores;
[0016] Divide multiple feature labels according to the values of the health characteristics and label the corresponding health characteristics with feature labels; the feature labels are health evaluations of the health characteristics;
[0017] Combine the feature labels to form a feature label library. The feature label library consists of a physical health label group, a mental health label group, and an environmental label group. Each label group contains multiple feature labels with different feature indicators, and the same feature indicator contains multiple feature labels with different values;
[0018] Combine the physical health characteristics and mental health characteristics extracted from all historical data to form a health characteristic library, divide the health characteristic library into a probability model set and a health characteristic set, and divide the probability model set into a probability training set and a probability test set.
[0019] Further, the method for obtaining the health risk probability includes:
[0020] Define that the first probability training set and the first probability test set only contain physical health characteristics, the second probability training set and the second probability test set only contain mental health characteristics, and the third probability training set and the third probability test set contain both physical health characteristics and mental health characteristics;
[0021] A multi-modal machine learning model is constructed by using gradient boosting trees, multi-layer perceptrons, and convolutional neural networks combined with long short-term memory networks as base models respectively;
[0022] The first probability training set is input into the gradient boosting tree, the second probability training set is input into the multi-layer perceptron, and the third probability training set is input into the convolutional neural network combined with the long short-term memory network for independent training. K-fold cross-validation is used to optimize the parameters of the base models. After training is completed, the first probability test set, the second probability test set, and the third probability test set are used to evaluate the performance of the base models respectively. The weighted average method is used as the integration strategy to combine the prediction results of the three base models to calculate the final prediction probability, and the multi-modal machine learning model is output;
[0023] The health feature set is input into the multi-modal machine learning model for prediction to obtain the health risk probability.
[0024] Furthermore, the method for constructing the feature transformation matrix according to the interaction data and the health features includes:
[0025] When the health risk probability is less than the preset probability threshold, it is determined to be occupational health;
[0026] When the health risk probability is greater than the preset probability threshold, it is determined that there is an occupational health risk. The interaction data is obtained by interacting the environmental database composed of the health features and the environmental data according to the personal information. The interaction data is preprocessed, and the random forest is used to extract features from the interaction data to obtain environmental features. Multiple feature labels are divided according to the values of each environmental feature, and the feature labels are marked for them according to the values of the environmental features; the interaction data includes the working environment temperature, environmental quality, working hours, salary level, occupational hazard probability, and occupational injury rate;
[0027] The physical health features, mental health features, and environmental features are combined into a feature sample, and the original training matrix is defined It is represented that the sample data matrix X contains n samples and d features. By adding a regularization sparse regression framework, adding a non-negative matrix factorization framework, preserving the local structure, and constraining the similarity matrix, an objective function is constructed, and the objective function is optimized to obtain the feature transformation matrix W. The expression of the objective function is:
[0028]
[0029] s.t.Σ j S ij =1, 0≤S ij ≤1, for all i, P, Q, T≥0, A T A=I, B T B=I
[0030] where is the feature transformation matrix, m = c is the number of clusters, is the cluster center matrix, is the coefficient matrix, η ≥ 0 is the regularization parameter, ‖W‖ 2,1 is to use l 2,1 norm to constrain the matrix W, the first and second terms are the regularized sparse regression framework; XP is the basis matrix, k is the cardinality, is the coefficient matrix, the third term is the non - negative matrix factorization framework; α ≥ 0 is the regularization coefficient, is the auxiliary matrix, the role of B is to ensure the orthogonality of T; L is the graph Laplacian matrix of the original data matrix, R = l k×k -l k , γ ≥ 0, λ ≥ 0 and ε ≥ 0 are regularization parameters, the fifth and sixth terms are the locally - preserving structure, and the seventh term is the constraint term to reduce the correlation of the coefficient matrix B; S is the similarity matrix, O is the initial similarity matrix, and β is the learning weight parameter of the model;
[0031] The optimization of the objective function involves a total of six variables (W, A, B, P, Q, S), resulting in the non - convexity of the objective function. Fixing the variables (A, B, P, Q, S) makes the objective function convex. Setting the derivative of the variable W to 0 and solving for the optimal solution of the problem, the update formula for the variable feature transformation matrix W is obtained:
[0032] W←(XX T +γXLX T +ηD)XBA T
[0033] where is a diagonal matrix, and the matrix element d ij =(1 / 2‖W i ‖2), and the derivative of ‖W‖ 2,1 with respect to the variable W is 2DW.
[0034] Furthermore, the method of obtaining the health coefficient by inputting the interaction data and the mental health characteristics into the health function includes:
[0035] Align the environmental characteristics and the mental health characteristics in the feature dimension, and input the feature data into the health function to obtain the health coefficient. The expression of the health function is:
[0036]
[0037] where δ i is the health coefficient of the i - th group of data, w1 and w2 are the weight coefficients, and S i is the similarity between the i - th group of environmental feature vectors and the mental health feature vectors, is the mean of the similarity of n groups of feature vectors, where n is the number of groups of feature vectors, and σ s is the standard deviation of the similarity of n groups of feature vectors.
[0038] Furthermore, the method for fusing the feature transformation matrix and the health coefficient to obtain the fused feature includes:
[0039] Adding the health coefficient δ i to the i-th row of the feature transformation matrix W according to the corresponding data group to form a fusion matrix m represents the number of data groups, (d + 1) represents the dimension of the row vector, and the input vector is set as the row vector y i , the kernel function is K(·), the feature vector set is Y, and the output vector is u i , and the optimal hyperplane expression for obtaining the data is:
[0040] u = τ i u i K(y i , Y) + θ
[0041] where u is the obtained optimal hyperplane, τ i is the weighted coefficient of the i-th row vector, and θ is the bias term;
[0042] Based on the obtained optimal hyperplane, the fusion matrix data is divided into h groups, and different groups correspond to a new fused feature. Setting the actual data value as u, the training output as U, and the generalization parameter as χ, a health data fusion model of SVM is established, and the expression is:
[0043]
[0044] In the formula, J is the health data fusion model, l i and l j are the optimal solution marking forms of data i and j, and are the optimal solution mean marking forms;
[0045] By determining the optimal solution of the feature data, a decision function of the feature data is established to complete the model solution and realize data fusion, and the expression is:
[0046]
[0047] where f(x) is the constructed decision function, is the set of means of h groups of data, and l h is the optimal solution marking form of the data;
[0048] The fusion matrix Y is input according to the row vector y i and decomposed into h groups of data, and m groups of fused features with a dimension of h are obtained through the health data fusion model.
[0049] Further, a method for constructing an occupational health analysis model based on big data according to the fusion features and historical tags includes:
[0050] Combine the fusion features and feature tags into a feature set, divide the feature set into an attention training set and an attention test set, and construct an occupational health analysis model. The specific structure includes an input layer, a feature preprocessing layer, a multi-head attention fusion layer, an adaptive feature fusion layer, and an output layer;
[0051] Input the attention training set into the occupational health analysis model. The data in the attention training set undergoes a linear transformation in the feature preprocessing layer to obtain a feature representation with unified dimensions and forms. After feature extraction of the linearly transformed data, it is input into the multi-head attention fusion layer. The APReLU activation function is used to adjust the weights of each head to generate weighted features. The weighted features and the original features are input into the adaptive feature fusion layer. The ASFF is used to integrate the features of different layers and input them into the output layer. Two fully connected layers are set in the output layer. One fully connected layer generates an occupational health score, and the other fully connected layer is followed by a softmax activation function to predict the feature tags, outputting the health score and health evaluation;
[0052] Use the mean squared error loss and the SGD optimizer for model evaluation and weight optimization to obtain the attention mechanism function, and use the attention test set to evaluate the attention mechanism function to output the attention mechanism model;
[0053] Input the occupational health data to be analyzed into the occupational health analysis model to obtain the occupational health score and health evaluation.
[0054] The beneficial effects of the present invention are:
[0055] The present invention is a method for constructing an occupational health analysis model based on big data. Compared with the prior art, the present invention has the following technical effects:
[0056] Through steps such as feature extraction, labeling feature tags, constructing a feature transformation matrix, obtaining a health coefficient, feature fusion, and model construction, the present invention can improve the data preprocessing ability and enhance the model adaptability in occupational health data analysis, thereby improving the efficiency and accuracy of occupational health data analysis. Optimizing the occupational health data analysis technology can greatly save resources and improve work efficiency. It can realize the analysis of occupational health data, provide strong support for occupational health management, and is of great significance for occupational health data analysis. It can meet the terminal occupational health data analysis needs of different big data-based occupational health analysis systems and different users' big data-based occupational health analysis systems, and has a certain universality. Description of the Drawings
[0057] Figure 1It is a flowchart of the steps for a method of constructing an occupational health analysis model based on big data according to the present invention. Detailed implementation manners
[0058] The present invention will be further described below through specific embodiments. The illustrative embodiments and descriptions of the present invention are used to explain the present invention, but do not limit the present invention.
[0059] A method for constructing an occupational health analysis model based on big data according to the present invention includes the following steps:
[0060] As Figure 1 shown, in this embodiment, it includes the following steps:
[0061] Obtain occupational health data and preprocess the occupational health data; the occupational health data includes historical data and environmental data;
[0062] Classify the historical data, extract features from the historical data to obtain health features, label feature tags for the health features according to the values of the health features, and construct a feature tag library according to the feature tags;
[0063] Construct a multimodal machine learning model, and input the health features into the multimodal machine learning model to obtain a health risk probability;
[0064] Interact the environmental database composed of the health features and the environmental data according to the health risk probability to obtain interaction data, construct a feature transformation matrix according to the interaction data and the health features, and input the interaction data and the mental health features into a health function to obtain a health coefficient;
[0065] Fuse the feature transformation matrix and the health coefficient to obtain a fused feature, construct an occupational health analysis model based on big data according to the fused feature and the feature tags, and input the occupational health data to be analyzed into the occupational health analysis model to obtain an occupational health score and a health evaluation.
[0066] In this embodiment, the method for extracting health features from the historical data includes:
[0067] Perform one-hot encoding on non-numerical data, and use the random forest algorithm to divide the data. The data types include physical health data and mental health data;
[0068] Use principal component analysis to extract physical health features from the physical health data; the physical health features are comprehensive features related to physical health;
[0069] Feature extraction is performed on mental health data using factor analysis to obtain mental health features; the mental health features are latent factors reflecting the meaning of mental health scores.
[0070] Multiple feature labels are divided according to the values of the health features, and corresponding feature labels are marked for the health features; the feature labels are health evaluations of the health features.
[0071] The feature labels are combined to form a feature label library, which consists of a physical health label group, a mental health label group, and an environmental label group. Each label group contains multiple feature labels with different feature indicators, and the same feature indicator contains multiple feature labels with different values.
[0072] The physical and mental health features extracted from all historical data are combined to form a health feature library. The health feature library is divided into two parts: a probability model set and a health feature set. The probability model set is divided into a probability training set and a probability test set.
[0073] In the actual evaluation, taking the occupational health data of the employees of Company X as an example, a set of historical data is randomly given for classification to obtain health data. After feature extraction of the health data, a set of health features is obtained: Gender / Male, Weight / 75 kg, Age / 25 years, (Physical health features: Pulse rate 72 beats per minute, Blood pressure 120 / 80 mmHg, Vision - Normal, Hearing - Normal, Facial features - Normal, Electrocardiogram - Sinus arrhythmia, Slight abnormality, Require regular review, B-ultrasound - Normal, Blood - Normal blood routine, Uric acid 450 μmol / L is slightly high, Internal medicine - Normal, Surgery - Normal, Nervous system - Normal), (Mental health features: Work pressure - Moderate pressure, within the controllable range, Job satisfaction - Overall satisfied with the job, Emotional stability - Slight fluctuations in emotions due to work pressure or personal affairs, Coping ability - Strong). Labeling operations are performed on the health data.
[0074] In this embodiment, the method for obtaining the health risk probability includes:
[0075] It is defined that the first probability training set and the first probability test set only contain physical health features, the second probability training set and the second probability test set only contain mental health features, and the third probability training set and the third probability test set contain both physical health features and mental health features.
[0076] Gradient boosting trees, multi-layer perceptrons, and convolutional neural networks combined with long short-term memory networks are used as base models respectively to construct a multi-modal machine learning model.
[0077] Input the first probability training set into the gradient boosting tree, input the second probability training set into the multi-layer perceptron, and input the third probability training set into the convolutional neural network combined with the long short-term memory network for independent training. Use k-fold cross-validation to optimize the parameters of the base model. After training, use the first probability test set, the second probability test set, and the third probability test set to evaluate the performance of the base model respectively. Use the weighted average method as the integration strategy to combine the prediction results of the three base models to calculate the final prediction probability, and output the multi-modal machine learning model;
[0078] Input the health feature set into the multi-modal machine learning model for prediction to obtain the health risk probability;
[0079] In the actual evaluation, score this group of random health features on a 10-point scale: gender / male, weight / 75 kg, age / 25 years, (physical health features: pulse rate 10 points, blood pressure 10 points, vision 10 points, hearing 10 points, facial features 10 points, electrocardiogram 8 points, B-ultrasound 10 points, blood 8 points, internal medicine 10 points, surgery 10 points, nervous system 10 points), (mental health features: work pressure 7 points, job satisfaction 8 points, emotional stability 8 points, coping ability 9 points);
[0080] Input this group of data into the trained multi-modal machine learning model. The health risk probabilities given by the three base models are 5%, 10%, and 8% respectively. Use the weighted average method (the weight coefficients of the output results of the three base models are 0.3, 0.2, and 0.5 respectively) as the integration strategy to combine the prediction results of the three base models to calculate the final predicted health risk probability of 7.5%.
[0081] In this embodiment, the method for constructing the feature transformation matrix according to the interaction data and the health features includes:
[0082] When the health risk probability is less than the preset probability threshold, it is determined as occupational health;
[0083] When the health risk probability is greater than the preset probability threshold, it is determined that there is an occupational health risk. According to the personal information, interact the environmental database composed of the health features and the environmental data to obtain interaction data, preprocess the interaction data, use the random forest to extract features from the interaction data to obtain environmental features, divide multiple feature labels according to the value of each environmental feature, and label the feature labels according to the value of the environmental feature; the interaction data includes working environment temperature, environmental quality, working hours, salary level, occupational hazard probability, and occupational injury rate;
[0084] Combine the physical health features, mental health features, and environmental features into a feature sample, and define the original training matrix It is indicated that the sample data matrix X contains n samples and d features. By adding a regularized sparse regression framework, adding a non-negative matrix factorization framework, preserving the local structure, and constraining the similarity matrix, an objective function is constructed. The optimal solution of the objective function is obtained to get the feature transformation matrix W. The expression of the objective function is:
[0085]
[0086] s.t. ∑ j S ij =1, 0 ≤ S ij ≤ 1, for all i, P, Q, T ≥ 0, A T A=I, B T B=I
[0087] where is the feature transformation matrix, m = c is the number of clusters, is the cluster center matrix, is the coefficient matrix, η ≥ 0 is the regularization parameter, ‖W‖ 2,1 is the matrix W constrained by the l 2,1 norm. The first and second terms are the regularized sparse regression framework; XP is the basis matrix, k is the cardinality, is the coefficient matrix, and the third term is the non-negative matrix factorization framework; α ≥ 0 is the regularization coefficient, is the auxiliary matrix, and the role of B is to ensure the orthogonality of T; L is the graph Laplacian matrix of the original data matrix, R = l k×k -l k , γ ≥ 0, λ ≥ 0, and ε ≥ 0 are regularization parameters. The fifth and sixth terms are the local structure preservation, and the seventh term is the constraint term to reduce the correlation of the coefficient matrix B; S is the similarity matrix, O is the initial similarity matrix, and β is the learning weight parameter of the model;
[0088] The optimization of the objective function involves a total of six variables (W, A, B, P, Q, S), resulting in the non-convexity of the objective function. Fixing the variables (A, B, P, Q, S) makes the objective function convex. Setting the derivative of the variable W to 0 and solving the optimal solution of the problem, the update formula of the variable feature transformation matrix W is obtained:
[0089] W ← (XX T + γXLX T + ηD)XBA T
[0090] where is a diagonal matrix, and the matrix element d ij =(1 / 2‖W i ‖2), ‖W‖ 2,1 The derivative of the variable W is 2DW;
[0091] In the actual assessment, the health risk probability of this group of data is 7.5%, which is greater than the preset probability threshold of 5%. It is determined that there is an occupational health risk. Interactive data is matched in the large database according to personal information: (working environment temperature -22°C, environmental quality - good, working hours - 60 h / week, slightly higher than the industry average working hours, salary level - 10% lower than the industry average level, occupational hazard probability - medium, occupational injury rate - slightly higher than the industry average level, increased cardiovascular burden caused by long-term high-intensity work);
[0092] Similarly, on a 10-point scale: (working environment temperature 10 points, environmental quality 10 points, working hours 4 points, salary level 5 points, occupational hazard probability 7 points, occupational injury rate 6 points). Thus, a complete set of characteristic samples including physical health indicators, mental health indicators, and environmental indicators is obtained. Processing multiple groups of data samples constitutes a characteristic sample set. The row vector corresponding to this group of data in the training matrix is [10, 10, 10, 10, 10, 8, 10, 8, 10, 10, 10, 7, 8, 8, 9, 10, 10, 4, 5, 7, 6]. After feature transformation, the row vector corresponding to it in the feature transformation matrix W is [100, 100, 100, 100, 100, 64, 100, 64, 100, 100, 100, 49, 64, 64, 81, 100, 10, 0, 16, 25, 49, 36].
[0093] In this embodiment, the method for obtaining the health coefficient by inputting the interactive data and the mental health characteristics into the health function includes:
[0094] Align the environmental characteristics and mental health characteristics in the feature dimension, and input the feature data into the health function to obtain the health coefficient. The expression of the health function is:
[0095]
[0096] where δ i is the health coefficient of the i-th group of data, w1 and w2 are weight coefficients, S i is the similarity between the i-th group of environmental feature vectors and mental health feature vectors, is the mean of the similarities of n groups of feature vectors, n is the number of feature vector groups, and σ s is the standard deviation of the similarities of n groups of feature vectors;
[0097] In the actual assessment, according to the environmental feature dimension of 6 and the mental health feature dimension of 4, the mental health feature dimension is extended to 6, and the feature data is input into the health function to calculate that the health coefficient of this group of data is 1.0469.
[0098] In this embodiment, the method for fusing the feature transformation matrix and the health coefficient to obtain a fused feature includes:
[0099] Adding the health coefficient η i To the i-th row of the feature transformation matrix W according to the corresponding data group to form a fusion matrix m represents the number of data groups, (d + 1) represents the dimension of the row vector, and the input vector is set as the row vector y i , the kernel function is K(·), the feature vector set is Y, and the output vector is u i , and the optimal hyperplane expression for obtaining the data is:
[0100] u = τ i u i K(y i , Y) + θ
[0101] Where u is the obtained optimal hyperplane, τ i Is the weighted coefficient of the i-th row vector, and θ is the bias term;
[0102] Based on the obtained optimal hyperplane, the fusion matrix data is divided into h groups, and different groups correspond to a new fused feature. Setting the actual data value as u, the training output as U, and the generalization parameter as χ, a health data fusion model of SVM is established, and the expression is:
[0103]
[0104] In the formula, J is the health data fusion model, l i And l j Are the optimal solution marking forms of data i and j, And Are the optimal solution mean marking forms;
[0105] By determining the optimal solution of the feature data, a decision function of the feature data is established to complete model solving and realize data fusion. The expression is:
[0106]
[0107] Where f(x) is the constructed decision function, Is the set of h-group data means, and l h Is the optimal solution marking form of the data;
[0108] The fusion matrix Y is input according to the row vector y i , decomposed into h groups of data, and m groups of fused features with a dimension of h are obtained through the health data fusion model;
[0109] In the actual evaluation, the fusion matrix is divided into 12 groups of data for data fusion, and the row vectors of the fusion matrix corresponding to the above data are subjected to feature fusion to obtain the fusion feature [100, 100, 64, 100, 100, 100, 100, 49, 64, 100, 16, 12].
[0110] In this embodiment, a big data-based occupational health analysis model is constructed according to the fusion feature and the historical label, including:
[0111] The fusion feature and the feature label are combined into a feature set, the feature set is divided into an attention training set and an attention test set, and an occupational health analysis model is constructed. The specific structure includes an input layer, a feature preprocessing layer, a multi-head attention fusion layer, an adaptive feature fusion layer, and an output layer;
[0112] The attention training set is input into the occupational health analysis model. The data of the attention training set undergoes a linear transformation in the feature preprocessing layer to obtain a feature representation with a unified dimension and form. After the feature extraction of the linearly transformed data, it is input into the multi-head attention fusion layer. The APReLU activation function is used to adjust the weights of each head to generate weighted features. The weighted features and the original features are input into the adaptive feature fusion layer. The ASFF is used to integrate the features of different layers and input them into the output layer. Two fully connected layers are set in the output layer. One fully connected layer generates an occupational health score, and the other fully connected layer is followed by a softmax activation function to predict the feature label, and the health score and health evaluation are output;
[0113] The mean square error loss and the SGD optimizer are used for model evaluation and weight optimization to obtain the attention mechanism function, and the attention test set is used to evaluate the attention mechanism function to output the attention mechanism model;
[0114] The occupational health data to be analyzed is input into the occupational health analysis model to obtain the occupational health score and health evaluation;
[0115] In actual evaluation, a set of health data to be analyzed is randomly generated: gender / male, weight / 80 kg, age / 45 years old, (physical health characteristics: pulse rate 85 beats per minute, slightly fast, blood pressure 145 / 95 mmHg, stage 1 hypertension, vision - normal, hearing - normal, facial features - normal, electrocardiogram - sinus arrhythmia, left ventricular high voltage, B - ultrasound - normal, blood - normal blood routine, uric acid 500 μmol / L is on the high side, internal medicine - mild fatty liver, surgery - normal, nervous system - symptoms of neurasthenia, easy fatigue, inattentiveness, sleep disorder), (mental health characteristics: work pressure - high pressure, difficult to control, job satisfaction - overall dissatisfaction with the job, emotional stability - large emotional fluctuations, coping ability - average), (working environment temperature - 20°C, environmental quality - average, working hours - 50 h / week, slightly higher than the industry average working hours, salary level - 20% higher than the industry average level, probability of occupational hazard - high, occupational injury rate - slightly higher than the industry average level, cervical problems and muscle strain caused by long - term sitting position), input into the occupational health analysis model to obtain a health score of: 57 points, and the overall health status is at a medium - low level. According to the predicted characteristic labels, the health evaluation is: there are problems such as hypertension, high uric acid and fatty liver in physical health, there are further mental health problems, there are additional work pressures and expectations, and attention needs to be paid to occupational health and safety.
[0116] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for constructing an occupational health analysis model based on big data, characterized in that, It includes the following steps: S1. Obtain occupational health data and preprocess the occupational health data; the occupational health data includes historical data and environmental data; S2. Classify the historical data, extract features from the historical data to obtain health features, label feature tags for the health features according to the values of the health features, and construct a feature tag library according to the feature tags; the health features include physical health features and mental health features; S3. Construct a multimodal machine learning model, and input the health features into the multimodal machine learning model to obtain a health risk probability; S4. Interact the environmental database composed of the health features and the environmental data according to the health risk probability to obtain interaction data, construct a feature transformation matrix according to the interaction data and the health features, and input the interaction data and the mental health features into a health function to obtain a health coefficient; S5. Fuse the feature transformation matrix and the health coefficient to obtain a fusion feature, construct a big data-based occupational health analysis model according to the fusion feature and the feature tag, and input the occupational health data to be analyzed into the occupational health analysis model to obtain an occupational health score and a health evaluation; The method for constructing a feature transformation matrix according to the interaction data and the health features includes: When the health risk probability is less than a preset probability threshold, it is determined to be occupational health; When the health risk probability is greater than the preset probability threshold, it is determined that there is an occupational health risk. Interact the environmental database composed of the health features and the environmental data according to personal information to obtain interaction data, preprocess the interaction data, extract environmental features from the interaction data using a random forest, divide multiple feature tags according to the values of each environmental feature, and label feature tags for them according to the values of the environmental features; the interaction data includes working environment temperature, environmental quality, working hours, salary level, occupational hazard probability, and occupational injury rate; The physical health characteristics, mental health characteristics, and environmental characteristics are combined to form a feature sample, and the original training matrix is defined , represents the sample data matrix contains samples and features. By adding a regularized sparse regression framework, adding a non-negative matrix factorization framework, preserving the local structure, and constraining the similarity matrix, an objective function is constructed, and the feature transformation matrix is obtained by optimizing the objective function. The expression of the objective function is as follows: =I Among them is the feature transformation matrix, = c is the number of clusters, is the cluster center matrix, is the coefficient matrix, 0 is the regularization parameter, is to utilize the norm constraint matrix , the first and second terms are the regularized sparse regression framework; is the basis matrix, , is the cardinality, is the coefficient matrix, and the third term is the non - negative matrix factorization framework; is the regularization coefficient, is the auxiliary matrix, The role of is to ensure orthogonality; is the graph Laplacian matrix of the original data matrix, , , and are regularization parameters, the fifth and sixth terms are the locally - preserving structure, and the seventh term is the constraint term to reduce the correlation of the coefficient matrix ; is the similarity matrix, is the initial similarity matrix, is the learning weight parameter of the model; The objective function involves a total of six variables optimization, resulting in the objective function being non-convex. Fixing the variable makes the objective function convex. Setting the derivative of the variable to 0 and finding the optimal solution to the problem gives the variable feature transformation matrix update formula: Among them is a diagonal matrix, and the matrix elements , For the variable The derivative is ; The method for constructing a big data-based occupational health analysis model according to the fusion feature and the historical label includes: Form a feature set with the fusion feature and the feature tag, divide the feature set into an attention training set and an attention test set, and construct an occupational health analysis model. The specific structure includes an input layer, a feature preprocessing layer, a multi-head attention fusion layer, an adaptive feature fusion layer, and an output layer; Input the attention training set into the occupational health analysis model. The data of the attention training set undergoes a linear transformation in the feature preprocessing layer to obtain a feature representation with a unified dimension and form. After extracting features from the linearly transformed data, input it into the multi-head attention fusion layer. Use the APReLU activation function to adjust the weights of each head to generate weighted features. Input the weighted features and the original features into the adaptive feature fusion layer. Use ASFF to integrate the features of different layers and input them into the output layer. Set two fully connected layers in the output layer. One fully connected layer generates an occupational health score, and the other fully connected layer is followed by a softmax activation function to predict the feature tag, and output the health score and the health evaluation; The mean squared error loss and SGD optimizer are used for model evaluation and weight optimization to obtain the attention mechanism function, and the attention mechanism function is evaluated using the attention test set to output the attention mechanism model; The occupational health data to be analyzed is input into the occupational health analysis model to obtain the occupational health score and health evaluation.
2. The construction method of an occupational health analysis model based on big data according to claim 1, characterized in that A method for extracting features from the historical data to obtain health features includes: Performing hot encoding on non-numerical data and using the random forest algorithm to divide the data, where the data types include physical health data and mental health data; Performing feature extraction on the physical health data using principal component analysis to obtain physical health features; the physical health features are comprehensive features related to physical health; Performing feature extraction on the mental health data using factor analysis to obtain mental health features; the mental health features are latent factors reflecting the meaning of mental health scores; Dividing multiple feature labels according to the values of the health features and labeling the corresponding health features with the feature labels; the feature labels are health evaluations of the health features; Combining the feature labels to form a feature label library, which consists of a physical health label group, a mental health label group, and an environmental label group. Each label group contains multiple feature labels with different feature indicators, and the same feature indicator contains multiple feature labels with different values; The physical health features and mental health features extracted from all historical data are combined to form a health feature library, which is divided into two parts: a probability model set and a health feature set. The probability model set is divided into a probability training set and a probability test set.
3. The construction method of an occupational health analysis model based on big data according to claim 1, characterized in that, A method for obtaining the health risk probability includes: Defining that the first probability training set and the first probability test set only contain physical health features, the second probability training set and the second probability test set only contain mental health features, and the third probability training set and the third probability test set contain both physical health features and mental health features; Using gradient boosting trees, multi-layer perceptrons, and convolutional neural networks combined with long short-term memory networks as base models to construct a multi-modal machine learning model; Inputting the first probability training set into the gradient boosting tree, the second probability training set into the multi-layer perceptron, and the third probability training set into the convolutional neural network combined with the long short-term memory network for independent training. Using k-fold cross-validation to optimize the parameters of the base models. After training, using the first probability test set, the second probability test set, and the third probability test set to evaluate the performance of the base models respectively. Using the weighted average method as the integration strategy to combine the prediction results of the three base models to calculate the final prediction probability and output the multi-modal machine learning model; Inputting the health feature set into the multi-modal machine learning model for prediction to obtain the health risk probability.
4. The construction method of an occupational health analysis model based on big data according to claim 1, characterized in that, A method for obtaining the health coefficient by inputting the interaction data and the mental health features into the health function includes: Aligning the environmental features and the mental health features in the feature dimension, and inputting the feature data into the health function to obtain the health coefficient. The expression of the health function is: where is the health coefficient of the th group of data, and are the weight coefficients, is the similarity between the th group of environmental feature vectors and mental health feature vectors, is the mean value of the similarity of the groups of feature vectors, is the standard deviation of the similarity of the 5. The construction method of an occupational health analysis model based on big data according to claim 1, characterized in that A method for fusing the feature transformation matrix and the health coefficient to obtain the fused features includes: Add the health coefficient to the corresponding data group and add it to the feature transformation matrix of to form a fusion matrix within the row , where represents the number of data groups, represents the dimension of the row vector, and the input vector is set as a row vector , the kernel function is , the feature vector set is , and the optimal hyperplane expression for obtaining the data is: Among them is the optimal hyperplane obtained is the weighted coefficient of the th row vector is the bias term; Based on the obtained optimal hyperplane, the fusion matrix data is divided into groups, and each group corresponds to a new fusion feature. Set the actual value of the data as , the training output as , and the generalization parameter as . Establish a health data fusion model of SVM, and the expression is: wherein is a health data fusion model, and are the optimal solution marking forms of data and , and and are the optimal solution mean marking forms; By determining the optimal solution of the feature data, establishing the decision function of the feature data, completing the model solution, and realizing data fusion. The expression is: wherein is the constructed decision function, is the set of group data means, is the data optimal solution marking form; Fusion matrix By row vector Input, decomposed into Groups of data, obtained through the health data fusion model Groups with dimensions of Fusion features
Citation Information
Patent Citations
Major disease risk evaluating method for electric power career population and system
CN110189829A
Method for establishing an occupational health data analysis model
CN113284620A