Semantic boundary modeling-based repeated hidden danger automatic identification method
By constructing a standard library of repeated hidden dangers and semantic boundary modeling, and using prototype networks and joint loss function optimization models, the coarse grain size and sparse sample number of repeated hidden dangers in the existing technology are solved, and high-precision repetitive hidden danger identification and unknown category refusal recognition are achieved, which improves the pertinence of enterprise hidden danger management.
Patent Information
- Application Number
- CN202510479656.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-25
AI Technical Summary
Most of the existing repetitive hidden danger classification methods are based on coarse grain size, making it difficult to identify repetitive problems in specific situations, and a large amount of training data is required to achieve good results, and lacks the ability to identify unknown categories.
By constructing a standard library of repeated hidden dangers, using the prototype network for semantic boundary modeling, and using the Episode training method of support sets and query sets, a joint loss function including classification loss, multi-relational marginal loss and rejection loss is designed, and the sample distribution in the embedded space is optimized to realize the rejection of unknown categories.
It improves the granularity and matching accuracy of hidden danger identification, enhances the ability to identify unknown categories, avoids mistakenly classified into known categories, and improves the pertinence and effectiveness of enterprise hidden danger management.
Smart Images

Figure CN120373299A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text classification, and particularly to an automatic recognition method for repeated hidden dangers based on semantic boundary modeling. Background Art
[0002] With the continuous improvement of the informatization level of industrial safety production, enterprises are accelerating their development towards digitization and automation in hidden danger management. Especially in high-risk industries such as the chemical industry, the frequent occurrence of repetitive safety hazards has become one of the important inducements for major accidents. Being able to quickly and accurately identify repeated hidden dangers has become the core technical requirement faced by enterprises in safety management.
[0003] Currently, most enterprises combine natural language technology to classify and manage repeated hidden dangers. For example, Patent Publication No.: CN113139057A discloses a domain adaptation method and system for short text classification of chemical safety hidden dangers. This application obtains a number of short texts to be classified in the field of chemical safety hidden danger investigation; extracts vectors for each short text to be classified to obtain the initial text vector corresponding to each short text to be classified; inputs the initial text vectors corresponding to all short texts to be classified into the trained short text classification model, and outputs the short text classification result. This application uses GRU+HAN to learn the information fusion representation of words, phrases, and sentences at different levels of short texts in a specific field, making up for the domain information deviation problem of short texts in general corpora, and showing better classification effects in the classification task of chemical safety hidden danger investigation.
[0004] Patent Publication No.: CN113535906B discloses a method for text classification of hidden danger events in the power field and related devices. This application constructs a risk hidden danger library including labeled samples and an unlabeled sample library including samples to be classified; trains a preset text classification network with the preprocessed labeled samples to obtain a text classification model; classifies the preprocessed samples to be classified through the text classification model, and obtains the confidence level according to the classification category probability; adds the first preset number of samples to be classified with the highest confidence level to the risk hidden danger library, and returns the remaining samples to be classified to the unlabeled sample library; classifies the labeled samples in the updated risk hidden danger library through the text classification model and obtains the confidence level; adds the labeled samples with the lowest confidence level in the updated risk hidden danger library with the second preset data volume to the risk hidden danger library recycle bin, improving the existing text of risk hidden danger events in the power field and solving the technical problems of low efficiency and long time consumption in the manual review method.
[0005] Although the existing hidden danger classification methods have achieved automatic classification of hidden dangers to a certain extent, most of them are based on coarse-grained classification and are difficult to identify repetitive problems in specific situations. Moreover, due to the scarcity of the number of samples in each category in the repetitive hidden danger recognition task, it is difficult to support the effective training of traditional models, resulting in a decline in classification performance. At the same time, such methods usually require all samples to be classified into known categories, lacking a recognition and exclusion mechanism for non-repetitive hidden danger samples that do not belong to any category, which is prone to misjudgment. The above problems limit the applicability of the existing automatic classification methods in the repetitive hidden danger recognition scenario. Summary of the Invention
[0006] 1. Technical problems to be solved by the invention
[0007] First, most of the existing repetitive hidden danger classification methods are based on coarse-grained, only making large-category divisions, and it is difficult to identify repetitive problems in specific situations. The present invention proposes fine-grained repetitive hidden danger recognition. By abstracting and classifying specific hidden dangers that repeatedly appear in history, a standard library of repetitive hidden dangers is constructed. The entries in the standard library are used as fine-grained categories, and newly reported hidden danger records will be classified according to this standard library to determine whether they are repetitive hidden dangers, which helps to improve the pertinence of enterprise hidden danger management.
[0008] Second, the existing repetitive hidden danger classification methods automatically classify newly emerging samples into existing categories and require a large amount of training data to achieve good results. The present invention proposes repetitive hidden danger recognition based on semantic boundary modeling. By using a prototype network to construct prototypes and generate semantic boundaries under few-shot conditions. According to the relative positions of new samples and the semantic boundaries of each category, it is judged whether they belong to a certain repetitive hidden danger category. When a sample is far from the semantic boundaries of all categories, the model can automatically reject classification and has the ability to recognize unknown categories. Furthermore, a joint loss function including classification loss, multi-relation margin loss, and rejection loss is designed to optimize the sample distribution in the embedding space, making the semantic boundaries clearer and enabling the model to have stronger category discrimination ability.
[0009] 2. Technical solutions
[0010] To achieve the above object, the technical solution provided by the present invention is as follows:
[0011] An automatic repetitive hidden danger recognition method based on semantic boundary modeling of the present invention includes the following steps:
[0012] Step 1: Conduct fine-grained abstraction and classification on specific hidden dangers that repeatedly appear in historical hidden danger data, use each specific hidden danger as an independent category to generate standard library entries, and construct a standard library of repetitive hidden dangers;
[0013] Step 2: Semantically match and label the historical hidden danger texts with the standard library entries. For the successfully matched samples, label the corresponding category tags, and for the unmatched samples, label them as non-repeated hidden dangers. Retain the description text and the corresponding category tag information of each hidden danger to generate a training sample set;
[0014] Step 3: Adopt the episode training method consisting of a support set and a query set. Extract features through a text encoder to generate category prototypes, calculate the semantic boundary radius of each category, and optimize the model using a joint loss function that includes classification loss, multi-relation margin loss, and open space repulsion loss to complete model training;
[0015] Step 4: Extract features from the newly reported hidden danger texts, calculate their distances from the category prototypes of each category, and determine whether they belong to repeated hidden dangers based on the semantic boundary radius, and output whether it is a repeated hidden danger and the entry information in the corresponding repeated hidden danger standard library.
[0016] Furthermore, the construction of the repeated hidden danger standard library in Step 1 specifically includes:
[0017] According to industry specifications, initially divide the historical hidden danger data into major categories of hidden dangers, and then gradually refine them into minor categories of hidden dangers;
[0018] On the basis of the existing division of minor categories of hidden dangers, introduce the risk level standard, and re-divide the minor category data in terms of risk dimensions to form sub-categories of hidden dangers with clear risk attributes;
[0019] Abstract and summarize the hidden dangers with similar semantics or similar manifestation forms to form specific standard library entries;
[0020] Each entry serves as a fine-grained repeated hidden danger category, constituting the final repeated hidden danger standard library.
[0021] Furthermore, the risk level standards in Step 1 include major hidden dangers, relatively large hidden dangers, and general hidden dangers;
[0022] Each standard library entry contains fields such as index number, hidden danger description, category information, risk analysis, and legend.
[0023] Furthermore, the annotation process in Step 2 adopts a mechanism of independent annotation by two people and cross-validation, and uses Cohen's Kappa coefficient for annotation consistency analysis; the category label value of non-repeated hidden danger samples is assigned as 0, and non-repeated hidden danger samples do not participate in the modeling as actual categories, but are only used for the training of the open space repulsion loss function.
[0024] Furthermore, in the model training process of step 3, first, the texts in the sample set are preprocessed by standardization, and text augmentation operations such as synonym replacement, random insertion, and deletion are performed on the duplicate risk samples; then, M class labels are randomly selected from the class labels of the duplicate risks, and N samples are randomly selected from the samples corresponding to each class label to construct a support set; the query set consists of duplicate risk samples and non-duplicate risk samples, and the selection ratio of duplicate risk samples to non-duplicate risk samples is 1:1.
[0025] Furthermore, in step 3, the category prototypes are generated by extracting features through a text encoder. The specific process of calculating the semantic boundary radius of each category is as follows:
[0026] (1) Use the pre-trained text encoder BERT model to extract features from the risk texts in the support set and the query set. For each risk text, after being processed by the encoder, it is mapped to a feature vector of a fixed dimension;
[0027] (2) Take the mean of the feature vectors of the N duplicate risk samples of each category in the support set after feature extraction to obtain the category prototype of each category;
[0028] (3) Use the cosine similarity to calculate the average similarity between the prototype vector of each category and the N sample feature vectors in the support set of this category, and the result is used as the semantic density of this category;
[0029] (4) Obtain the set of Euclidean distances between each category prototype vector and the feature vectors of its support set samples; then, combine the corresponding semantic density to scale and adjust this distance set to obtain the semantic boundary radius of this category.
[0030] Furthermore, the specific process of the model optimization in step 3 is as follows:
[0031] (1) For each query sample, first calculate the distance between it and all category prototype vectors, and subtract the semantic boundary radius of the corresponding category as the matching score; after this score result is normalized by Softmax, the classification probabilities of this query sample for each category and the corresponding classification losses are obtained;
[0032] (2) Introduce the inter-class separation loss and the intra-class compactness loss to obtain the multi-relationship margin loss; this loss term reflects the relative distribution density of the samples within the category boundary by calculating the Euclidean distance between the support samples and the category prototype and using the semantic boundary radius of this category as the ratio standard;
[0033] (3) Introduce the open space rejection loss, which is used to penalize non-duplicate potential hazard samples close to any class prototype in the embedding space. For all non-duplicate potential hazard samples in the query set, calculate the Euclidean distance between their embedding vectors and each class prototype, and combine the semantic boundary radius and the rejection intensity threshold to determine whether the sample is too close to a certain class prototype; if a non-duplicate sample is less than the sum of the class semantic boundary radius and the rejection intensity from a certain class prototype, the corresponding rejection loss is generated.
[0034] (4) Optimize the model by weighted summation of the classification loss, multi-relationship margin loss, and open loss to form the total loss.
[0035] Furthermore, in step 3, use the trained text encoder to extract the embeddings of all support set samples, calculate the class prototype and semantic boundary radius of each class, and form a triple (i, c i , r i ) with the class label, class prototype, and semantic boundary radius for structured storage and use when loading during inference.
[0036] Furthermore, in step 4, new reported potential hazards can be entered through a mobile terminal and support on-site photo upload. All input information is encapsulated in a unified format and submitted to the system backend.
[0037] Furthermore, in step 4, show whether it is a duplicate potential hazard and the corresponding entry information in the duplicate potential hazard standard library on the mobile terminal, and the new potential hazard record is automatically synchronized to the background management system.
[0038] 3. Beneficial effects
[0039] Adopting the technical solution provided by the present invention, compared with the existing well-known technologies, it has the following remarkable effects:
[0040] (1) A method for automatic identification of duplicate potential hazards based on semantic boundary modeling in the present invention. By constructing a duplicate potential hazard standard library, the specific potential hazards that have repeatedly occurred in history are finely grained and classified, breaking through the limitation of the existing potential hazard classification method based only on coarse-grained division. New potential hazard records can be matched and classified according to the standard library, so as to achieve accurate identification of duplicate potential hazards. Compared with traditional methods, the present invention significantly improves the granularity and matching accuracy of potential hazard identification, and enhances the pertinence and effectiveness of potential hazard governance work.
[0041] (2) An automatic recognition method for repeated hidden dangers based on semantic boundary modeling of the present invention uses a prototype network to construct prototype representations of various categories, introduces a semantic boundary modeling mechanism, and realizes effective rejection of unknown categories through the relative position relationship between samples and the boundary, avoiding misclassifying non-repeated hidden dangers into known categories. The jointly designed classification loss, multi-relationship margin loss, and rejection loss further optimize the sample distribution in the embedding space, making the semantic boundary clearer and the model having stronger category discrimination ability. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a flowchart of automatic hidden danger recognition based on a standard library of repeated hidden dangers proposed by the present invention;
[0043] Figure 2 is a legend of an entry in the standard library of repeated hidden dangers;
[0044] Figure 3 is a structural diagram of a repeated hidden danger recognition model based on semantic boundary modeling proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] To further understand the content of the present invention, the present invention will be described in detail with reference to the accompanying drawings and embodiments.
[0046] As Figure 1 shown, the steps of the automatic recognition method for repeated hidden dangers based on semantic boundary modeling in this embodiment are as follows:
[0047] Step 1: Based on the fine-grained abstraction and classification of specific repeated hidden dangers in historical hidden danger data, construct a standard library of repeated hidden dangers;
[0048] First, according to the classification framework in the "Guidelines for the Investigation and Governance of Safety Risks and Hidden Dangers of Hazardous Chemical Enterprises", combined with the enterprise's own production scenarios and management practices, the historical hidden danger data is initially divided into 8 major hidden danger categories. Each major category is further divided into 5-10 minor hidden danger categories to form a more business-targeted hierarchical structure.
[0049] Secondly, on the basis of the existing minor category division, introduce risk level standards, specifically including: major hidden dangers, relatively large hidden dangers, and general hidden dangers; re-divide the minor category data in terms of risk dimensions to form hidden danger sub-categories with clear risk attributes.
[0050] Finally, abstract and summarize the hidden dangers with similar semantics or highly similar manifestation forms in the results of the three rounds of division to generate representative and general standard library entries. Each entry serves as a fine-grained repeated hidden danger category, constituting the final standard library of repeated hidden dangers.
[0051] In this embodiment, the standard library contains a total of 50 typical repeated hidden dangers. Each entry in the standard library corresponds to a representative description of a repeated hidden danger, which serves as the basis for subsequent judgment. The content of the standard library includes, but is not limited to, fields such as index number, hidden danger description, category information, risk analysis, and legend.
[0052] For example:
[0053] Index number: 18
[0054] Hidden danger description: The power cord and signal line are damaged or not protected.
[0055] Category information: Electrical and instrumentation
[0056] Risk analysis: Electric shock to personnel, signal loss.
[0057] Legend: See Figure 2 .
[0058] When new hidden danger information is generated, if its hidden danger description matches any entry in the library, it is determined as a repeated hidden danger; if there is no matching entry, it is regarded as a non-repeated hidden danger.
[0059] Step 2: Based on the previously constructed standard library of repeated hidden dangers, organize an expert team to carry out the standardization annotation work on historical hidden danger data. The annotation process is based on the semantic matching between the historical hidden danger text and the standard library. For each historical hidden danger description, first judge whether it forms a matching relationship with any entry in the standard library at the semantic level. If the match is successful, assign the category label corresponding to the matching entry to the sample; if there is no match, assign the category label value 0, indicating that the sample is a non-repeated hidden danger.
[0060] In the subsequent category prototype calculation and classification modeling process, samples with a label of 0 do not participate in the modeling as actual categories; such samples are only used for the training of the rejection loss function to enhance the model's ability to identify non-repeated hidden dangers.
[0061] The annotation process follows a unified annotation specification and hidden danger assignment standard, and multiple rounds of manual review and cross-validation are carried out to improve the accuracy and consistency of data annotation. In this embodiment, two people independently annotate and conduct cross-evaluation on the same data sample. To evaluate the annotation consistency, Cohen's Kappa coefficient is used in this embodiment for annotation consistency analysis. The consistency coefficient of the final annotation result is 0.83, indicating a high level of consistency.
[0062] After the annotation is completed, the system only retains the description text of each hidden danger and the corresponding category label information from the original data, and the rest of the fields do not participate in the subsequent training process. Finally, a sample set with a unified structure is formed, and each sample consists of a "text-category label" pair for subsequent model training.
[0063] Step 3: Using the sample set constructed in Step 2, train the model in the form of an Episode composed of a support set and a query set. As the basic unit of few-shot training, each Episode constructs a set of support set and query set internally, and completes the whole process from feature extraction, prototype calculation to boundary modeling and loss optimization. The entire training phase consists of 50 Epochs, and each Epoch contains 200 Episodes. The model continuously improves its discrimination ability for repeated and non-repeated potential hazards through cross-task contrast optimization. The training process of the model in a single Episode is as Figure 3 shown.
[0064] The specific process of model training is introduced below:
[0065] Step 3.1: First, perform standardized preprocessing on the text in the sample set to clean special characters, invalid characters, redundant spaces, etc. After the data cleaning is completed, for repeated potential hazard samples, perform semantic-based text augmentation operations, including synonym replacement, random insertion and deletion of text. Each repeated potential hazard sample is expanded to generate 3 augmented samples.
[0066] For example: Initial text: "The operator did not operate according to the safety regulations."
[0067] Data cleaning: "The operator did not operate according to the safety regulations"
[0068] Data augmentation: "The staff did not follow the safety operating procedures", "The operator did not operate according to the work safety", "The operator did not operate according to the safety".
[0069] Step 3.2: Randomly select 5 class labels from the class labels of repeated potential hazards, and randomly select 5 samples from the samples corresponding to each class label to construct a support set. Subsequently, construct a query set, which consists of the following two parts:
[0070] (1) Repeated potential hazard samples: From the current 5 selected class labels, randomly select 5 samples that are not included in the support set respectively, for a total of 25 samples;
[0071] (2) Non-repeated potential hazard samples: Randomly select 25 samples from all non-repeated potential hazard samples with a class label of 0.
[0072] Finally, each query set contains 50 samples.
[0073] Step 3.3: Use the pre-trained text encoder BERT model (bert-base-chinese) to extract features from the hidden danger texts in the support set and the query set. For each hidden danger text, after being processed by the encoder, it is mapped into a feature vector with a fixed dimension. The encoder parameters are optimized and updated through the backpropagation mechanism during the training process.
[0074] Specifically in this embodiment, after the text is tokenized and encoded, it is input into the BERT encoder, and the output of its hidden layer is taken as the semantic feature representation of the text. The output is a feature vector with a dimension of 768, and the feature vector f(x) is expressed as:
[0075] f(x) = BERT φ (x)
[0076] where φ represents the trainable parameters of the BERT model. All the generated feature vectors are in the same embedding space, and this space serves as the basic feature space for the model to uniformly perform representation learning during the calculation of class prototypes, sample matching, and loss optimization.
[0077] Step 3.4: Take the mean of the feature vectors of the 5 repeated hidden danger samples of each category in the support set after feature extraction to obtain the category prototype of each category.
[0078] For example:
[0079] Repeated hidden danger standard library entry: Power cord and signal line damaged, no protection measures taken. Label: 18.
[0080] The hidden danger support set samples and their feature representations are as follows:
[0081] "The socket wiring of the fixed fire area is exposed" [0.04564485 -0.26038316 -0.8279663...]
[0082] "The wiring of the hydrochloric acid frame pump is exposed without explosion-proof measures" [0.00665599 0.01773121 -0.16657946...]
[0083] "The motor cable of hydrogen peroxide phase I is exposed" [0.29158145 -0.24638432 -0.71879584...]
[0084] "Part of the power cord of the liquid chlorine loading platform is exposed" [0.0854884 0.19063099 -0.54460084...]
[0085] "The lighting switch in the freezing substation is damaged, causing the power line to be exposed" [0.27513453 -0.5460279 -0.36012578...]
[0086] The corresponding class prototype is represented as follows: [0.14090104 -0.16888663 -0.52361357...]
[0087] Step 3.5: Use cosine similarity to calculate the average similarity between the prototype vector of each class and the 5 sample feature vectors in the support set of that class, and the result is used as the semantic density of that class.
[0088] For example: In actual calculation, taking the hidden danger class with label 18 as an example, its semantic density result is 0.9305.
[0089] Step 3.6: The calculation process of the semantic boundary radius is as follows: Obtain the set of Euclidean distances between the prototype vector of each class and the sample feature vectors in its support set; then combine the corresponding semantic density to scale and adjust this distance set to obtain the semantic boundary radius of that class:
[0090]
[0091] Among them, Quantile p represents the quantile function, which is used to calculate the quantile value of the coverage ratio p. ò is the numerical stability term. ρ i is the semantic density of class i. K represents the number of samples in each class. c i is the class prototype of class i.
[0092] For example: In the hidden danger class with label 18, the set of Euclidean distances between 5 samples and the class prototype is {6.4398, 5.4568, 6.2008, 5.0368, 5.8815}, taking the 80% quantile, the semantic density calculated in the previous step is 0.9305, and the stability term coefficient is 10 -6 , then calculate the semantic boundary radius:
[0093]
[0094] Step 3.7: For each query sample, first calculate the distance between it and all class prototype vectors, and subtract the semantic boundary radius of the corresponding class as the matching score. After normalizing this score result by Softmax, the classification probability of this query sample for each class is obtained:
[0095]
[0096] Among them, τ is the temperature coefficient, which controls the smoothness of the distribution, and is set to 0.6 here. x q represents the query set sample.
[0097] The corresponding classification loss is:
[0098]
[0099] Step 3.8: To improve the category discrimination ability of the model, the present invention introduces the inter-class separation loss and the intra-class compactness loss as auxiliary optimization objectives. The inter-class separation loss is used to constrain sufficient spacing between different category prototypes to avoid feature overlap between categories, thereby reducing confusion misjudgment. The intra-class compactness loss aims to make the support samples of the same category as close as possible to their category prototypes in the embedding space. This loss term calculates the Euclidean distance between the support samples and the category prototypes, and uses the semantic boundary radius of the category as the ratio standard, thereby reflecting the relative distribution density of the samples within the category boundary. The smaller this ratio is, the more the samples are concentrated around the prototype, and the more compact the category is internally. To simultaneously model these two key relationships of inter-class separation and intra-class compactness, the overall multi-relationship margin loss is defined as follows:
[0100]
[0101] where μ represents the threshold of the minimum inter-class spacing, which is set to 0.8.
[0102] Step 3.9: To improve the rejection ability of the model for non-repeating hidden danger samples, the open space rejection loss is introduced. This loss term is used to penalize non-repeating hidden danger samples that are close to any category prototype in the embedding space. For all non-repeating hidden danger samples in the query set, calculate the Euclidean distance between their embedding vectors and each category prototype, and combine the semantic boundary radius and the rejection intensity threshold to determine whether the sample is too close to a certain category prototype. If a non-repeating sample is less than the sum of the semantic boundary radius of the category and the rejection intensity from a certain category prototype, the corresponding rejection loss is calculated by the following formula:
[0103]
[0104] where δ is a hyperparameter that controls the rejection intensity and is set to 0.5. |Q open | represents the number of non-repeating hidden danger samples in the query set.
[0105] Step 3.10: The classification loss, the multi-relationship margin loss, and the open loss are weighted and summed to form the total loss. The final loss function is:
[0106] L = L cls + λ1L margin + λ2L open
[0107] Different loss terms are balanced by the weight coefficients λ1 and λ2, where λ1 is set to 0.4 and λ2 is set to 0.5.
[0108] The model is trained using the Adam optimizer, and the initial learning rate is set to 10 -4 。
[0109] Step 3.11. Use the trained text encoder to extract embeddings for all support set samples, calculate the class prototype and semantic boundary radius for each class, and form a triple (i, c i , r i ) for structured storage and use during inference.
[0110] Step 4. When enterprise employees conduct hidden danger investigation and discover new hidden dangers, they enter the new hidden danger information, including fields such as hidden danger content and location, through the handheld terminal filling page, and support on-site photo upload. All input information is encapsulated in a unified format and submitted to the system backend.
[0111] For example: An employee of a company discovers a hidden danger during the hidden danger investigation, opens the handheld device, and enters the following information:
[0112] Hidden danger content: "The outlet valve of the phosphoric acid transfer pump leaks phosphoric acid", location: "Hydrogen peroxide plant area". Then click submit to upload this hidden danger.
[0113] Step 5. The new hidden danger description is automatically input into the model for the process of identifying duplicate hidden dangers, specifically as follows:
[0114] Step 5.1. After the model receives the hidden danger text to be identified, it first performs data cleaning on the input text, including cleaning special characters, invalid characters, invalid numbers, and extra spaces.
[0115] Step 5.2. After the cleaning is completed, call the trained text encoder to extract features from the new hidden danger text and generate a feature vector with a fixed dimension.
[0116] Step 5.3. The system loads all the class prototypes and corresponding semantic boundary radii saved in Step 3.11, calculates the Euclidean distance between the feature vector of the current sample and each class prototype, and makes a matching judgment based on the semantic boundary radius.
[0117] Step 5.4. If the sample feature vector falls within the semantic boundary radius of any class, it is determined that the sample is a duplicate hidden danger belonging to that class. Among all the classes that meet the conditions, select the class with the smallest distance as the prediction result; if the sample does not fall within the semantic boundary radius of any class, it is determined as a non-duplicate hidden danger.
[0118] Step 5.5. The system organizes the model output into a structured unified format, including a duplicate identifier and a prediction label.
[0119] For example: Input the hidden danger text: "The outlet valve of the phosphoric acid transfer pump leaks phosphoric acid"
[0120] The result obtained after the module is identified is: repeated hidden danger, the matching category is 21, and the corresponding entry in the repeated hidden danger standard library is "polluting the environment, endangering human health, and wasting resources".
[0121] Step 6: Display the prediction information on the mobile handheld terminal, including whether there is a repeated hidden danger and the entry information in the corresponding repeated hidden danger standard library. At the same time, automatically synchronize the new hidden danger in the hidden danger management background system.
[0122] The above schematically describes the present invention and its implementation manners. This description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Therefore, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments to this technical solution without creative efforts without departing from the purpose of the present invention, they should all fall within the protection scope of the present invention.
Claims
1. An automatic recognition method for repeated hidden dangers based on semantic boundary modeling, characterized in that, It includes the following steps: Step 1: Conduct fine-grained abstraction and classification on the specific hidden dangers that repeatedly appear in the historical hidden danger data. Take each specific hidden danger as an independent category to generate standard library entries, and construct a standard library for repeated hidden dangers; Step 2: Conduct semantic matching annotation on the historical hidden danger text and the standard library entries. Mark the successfully matched samples with the corresponding category labels, mark the unmatched samples as non-repeated hidden dangers, and retain the description text and corresponding category label information of each hidden danger to generate a training sample set; Step 3: Adopt the episode training method composed of a support set and a query set. Extract features through a text encoder to generate category prototypes, calculate the semantic boundary radius of each category, and optimize the model using a joint loss function that includes classification loss, multi-relation margin loss, and open space repulsion loss to complete model training; Step 4: Extract features from the newly reported hidden danger text, calculate its distance from the category prototypes of each category, and judge whether it belongs to a repeated hidden danger according to the semantic boundary radius, and output whether it is a repeated hidden danger and the entry information in the corresponding standard library for repeated hidden dangers.
2. The automatic recognition method for repeated hidden dangers based on semantic boundary modeling according to claim 1, characterized in that: The construction of the standard library for repeated hidden dangers described in Step 1 specifically includes: According to industry specifications, initially divide the historical hidden danger data into major hidden danger categories, and then gradually refine them to form minor hidden danger categories; On the basis of the existing division of minor hidden danger categories, introduce a risk level standard, and re-divide the minor category data in terms of risk dimension to form hidden danger sub-categories with clear risk attributes; Abstract and summarize hidden dangers with similar semantics or similar manifestation forms to form specific standard library entries; Each entry serves as a fine-grained repeated hidden danger category to form the final standard library for repeated hidden dangers.
3. The automatic recognition method for repeated potential hazards based on semantic boundary modeling according to claim 2, wherein: The risk level standards described in Step 1 include major hidden dangers, relatively large hidden dangers, and general hidden dangers; Each standard library entry includes fields such as index number, hidden danger description, category information, risk analysis, and legend.
4. The automatic recognition method for repeated potential hazards based on semantic boundary modeling according to claim 1, characterized in that: In the annotation process of Step 2, a two-person independent annotation and cross-validation mechanism is adopted, and the Cohen's Kappa coefficient is used for annotation consistency analysis; the category label value of non-repeated hidden danger samples is assigned as 0, and non-repeated hidden danger samples do not participate in modeling as actual categories, but are only used for the training of the open space repulsion loss function.
5. An automatic recognition method for repeated hidden dangers based on semantic boundary modeling according to any one of claims 1-4, characterized in that: In the model training process of Step 3, first, preprocess the text in the sample set, and perform text enhancement operations such as synonym replacement, random insertion, and deletion on the repeated hidden danger samples; then randomly select M category labels from the category labels of the repeated hidden dangers, and randomly select N samples from the samples corresponding to each category label to construct a support set; the query set is composed of repeated hidden danger samples and non-repeated hidden danger samples, and the selection ratio of repeated hidden danger samples and non-repeated hidden danger samples is 1:
1.
6. The automatic recognition method for repeated hazards based on semantic boundary modeling according to claim 5, characterized in that: The specific process of calculating the semantic boundary radius of each category by extracting features through a text encoder in Step 3 is as follows: (1) Use the pre-trained text encoder BERT model to extract features from the hidden danger texts in the support set and the query set. For each hidden danger text, after being processed by the encoder, it is mapped into a feature vector of a fixed dimension; (2) Take the mean value of the feature vectors of the N repeated hidden danger samples of each category in the support set after feature extraction to obtain the category prototype of each category; (3) Calculate the average similarity between the prototype vector of each category and the N sample feature vectors in the support set of that category using cosine similarity, and the result is used as the semantic density of that category. (4) Obtain the set of Euclidean distances between the prototype vector of each category and the sample feature vectors in its support set; then, combine the corresponding semantic density to scale and adjust this distance set to obtain the semantic boundary radius of that category.
7. The automatic recognition method for repeated hidden dangers based on semantic boundary modeling according to claim 6, characterized in that: The specific process of step 3 for model optimization is as follows: (1) For each query sample, first calculate the distance between it and the prototype vectors of all categories, and subtract the semantic boundary radius of the corresponding category as the matching score; after normalizing this score result by Softmax, obtain the classification probability of this query sample for each category and the corresponding classification loss. (2) Introduce the inter-class separation loss and the intra-class compactness loss to obtain the multi-relationship margin loss; this loss term reflects the relative distribution density of samples within the category boundary by calculating the Euclidean distance between the support samples and the category prototype and using the semantic boundary radius of that category as the ratio standard. (3) Introduce the open space repulsion loss, which is used to penalize non-duplicate potential hazard samples close to the prototype of any category in the embedding space. For all non-duplicate potential hazard samples in the query set, calculate the Euclidean distance between their embedding vectors and the prototype of each category, and combine the semantic boundary radius and the repulsion intensity threshold to determine whether the sample is too close to the prototype of a certain category. If the distance between a non-duplicate sample and the prototype of a certain category is less than the sum of the semantic boundary radius of that category and the repulsion intensity, a corresponding repulsion loss is generated. (4) Optimize the model by weighted summing the classification loss, the multi-relationship margin loss, and the open loss to form the total loss.
8. The automatic recognition method for repeated potential hazards based on semantic boundary modeling according to claim 7, characterized in that: Step 3 uses the trained text encoder to extract embeddings for all support set samples, calculates the class prototype and semantic boundary radius for each class, and forms a triple (i, c i , r i ) with the class label, class prototype, and semantic boundary radius for structured storage and use during inference.
9. The automatic recognition method for repeated hidden dangers based on semantic boundary modeling according to claim 8, characterized in that: In step 4, the newly reported potential hazards can be entered through the mobile terminal, and on-site photo uploading is supported. All input information is encapsulated in a unified format and submitted to the system backend.
10. The automatic recognition method for repeated potential hazards based on semantic boundary modeling according to claim 9, characterized in that: In step 4, it is shown on the mobile terminal whether there are duplicate potential hazards and the entry information in the corresponding duplicate potential hazard standard library, and the new potential hazard records are automatically synchronized to the background management system.
Citation Information
Patent Citations
Domain-adaptive chemical potential safety hazard short text classification method and system
CN113139057A
A Text Classification Method and Related Device for Potential Safety Hazards in the Power Sector
CN113535906B