Automatic container terminal abnormity and risk analysis method based on BERT-BiLSTM-CRF
Through the model architecture based on BERT-BiLSTM-CRF, the problem of insufficient accuracy of abnormal identification and risk management in automated container terminals is solved, efficient identification and management of abnormalities and risks is achieved, and the safety and stability of container terminals are improved.
Patent Information
- Application Number
- CN202510291697.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-07-11
AI Technical Summary
In the abnormal identification and risk management of automated container terminals, the prior art has problems of insufficient accuracy and real-time response capabilities, making it difficult to effectively deal with complex dynamic operation scenarios.
The model architecture based on BERT-BiLSTM-CRF is adopted, combining deep feature extraction, sequence dependency processing and conditional random field layer to achieve efficient analysis of multiple data sources and identify abnormalities and risks of container terminals.
It improves the accuracy and timeliness of abnormal and risk identification, enhances the ability to understand complex text semantics and sequence relationships, and improves the safety management and risk control capabilities of container terminals.
Smart Images

Figure CN120297722A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of anomaly management and risk analysis. Specifically, it particularly relates to a method for analyzing anomaly problems and potential risks in an automated container terminal based on BERT-BiLSTM-CRF. Background Art
[0002] With the booming development of global trade, container terminals, as key nodes for cargo import and export, their operation efficiency and safety management have become increasingly important. In recent years, the degree of automation in container terminals has been continuously improving. The automation of equipment and business processes not only enhances the operation efficiency but also significantly reduces the probability of human operation errors. However, with the expansion of the scale of automated equipment, the operating environment of automated container terminals has become increasingly complex, and various potential anomalies often occur during operation. The anomalies that occur in the ACT operating environment may pose a threat to the safety of equipment and systems, thereby affecting the operation progress, production efficiency, and operation quality of the terminal. Therefore, it is particularly necessary to construct a container terminal anomaly and risk identification model and method based on BERT-BiLSTM-CRF to promote the improvement of overall operation efficiency and risk prevention and control capabilities.
[0003] In recent years, the progress of artificial intelligence technology and the improvement of big data processing capabilities have opened up new possibilities for anomaly and risk management in container terminals. Especially in the field of natural language processing, by deeply analyzing text data, potential anomalies and risks can be automatically identified. However, the existing technologies are still insufficient in terms of dynamic response and adaptability to operation scenarios, which limits their effectiveness in practical applications. Therefore, designing an effective model architecture to improve the accuracy and timeliness of anomaly and risk identification has important technical application value. The analysis method based on the BERT-BiLSTM-CRF model proposed by the present invention aims to comprehensively improve the response ability of container terminals to anomaly problems and potential risks through multi-level feature extraction and deep learning mechanisms, and contribute to the promotion of intelligent logistics. Summary of the Invention
[0004] In view of the above technical problems in the existing technology, which are insufficient in the accuracy of anomaly identification and real-time response ability, a method for analyzing anomaly problems and potential risks in an automated container terminal based on BERT-BiLSTM-CRF is provided. The present invention combines the deep feature extraction ability of the BERT pre-trained model, the sequence dependence processing advantage of the bidirectional long short-term memory network, and the sequence labeling function of the conditional random field layer to achieve efficient analysis of multiple data sources, thereby quickly identifying anomalies and risks in port operations.
[0005] The technical means adopted by the present invention are as follows:
[0006] An automated container terminal anomaly and risk analysis method based on BERT - BiLSTM - CRF, comprising the following steps:
[0007] Step 1: Determine entity extraction requirements and relationship modeling;
[0008] Step 2: Data preprocessing and dataset construction; Collect research materials and actual operation data from the container terminal to generate an accident corpus with specific logical semantics;
[0009] Step 3: Anomaly problem and risk hazard identification; Use the BERT - BiLSTM - CRF framework to process unstructured data and achieve automated extraction of accident - related entities and their relationships;
[0010] Step 4: Anomaly and risk data storage and application analysis; Store the extracted knowledge elements in the Neo4j graph database to ensure systematic storage and management of knowledge; Through the use of Cypher query language, achieve multi - dimensional visual analysis for the exploration and management of the accident knowledge graph.
[0011] Furthermore, the process of determining entity extraction requirements and relationship modeling in Step 1 includes the following steps:
[0012] Step 11: According to the basic requirements of the container terminal, identify and refine operation anomaly events and related risks;
[0013] Step 12: Conduct a comprehensive analysis of existing accident data, logs, and monitoring records to identify relevant key entities and potential relationships, and form a preliminary knowledge framework;
[0014] Step 13: Construct a multi - level knowledge ontology to refine the definitions of entities and relationships structurally and semantically, ensuring that the diversity and complexity of anomaly events can be accurately reflected in the model.
[0015] Furthermore, the process of data preprocessing and dataset construction in Step 2 includes the following steps:
[0016] Step 21: Obtain research materials and actual operation data of the container terminal;
[0017] Step 22: Introduce a multi - level data processing method to extract useful information from the original data; Conduct in - depth analysis of the text through semantic analysis to identify key topics and entities; At the same time, segment long sentences and complex structures through long - sentence segmentation to convert them into short sentences; In addition, conduct preliminary cleaning and unified formatting processing of the data through format alignment to ensure consistency and readability;
[0018] Step 23: Generate a corpus set of abnormal problems and potential risks in the container terminal with specific logical semantics; the corpus set of abnormal problems and potential risks in the container terminal includes: various context information related to abnormal events, ensuring that the knowledge in the corpus can be fully utilized during model training;
[0019] Step 24: Divide the corpus set of abnormal problems and potential risks in the container terminal into a training data set, a test data set, and a validation data set at a ratio of 7:2:1;
[0020] Step 25: Process the processed data for knowledge processing and organize it into triples to complete the construction of the abnormal problem and potential risk data set in the container terminal.
[0021] Further, Step 25 organizes entities and relationships into triples, including the following steps:
[0022] Step 251: Conduct a preliminary check on the entity and relationship data extracted by the model; the preliminary check includes: spelling error correction and consistency verification;
[0023] Step 252: For entities and relationships that have passed manual verification and automated rule screening, where h represents the main entity, t represents the tail entity, and r represents the relationship between the two;
[0024] Step 253: Through the ACT domain-specific business scenarios and operation processes, ensure that the definition and applicable scenarios of each relationship are supported by clear context.
[0025] Further, Step 3 includes the following steps:
[0026] Step 31: Use a deep learning model based on the Transformer with bidirectional encoder representations to extract features from the cleaned and standardized text data;
[0027] Step 32: Transfer the feature vectors extracted from the bidirectional encoder to a bidirectional long short-term memory network;
[0028] Step 33: Introduce a conditional random field in the output layer of the BiLSTM and handle the optimal distribution of labels in sequence annotation by combining the properties of the conditional random field;
[0029] Step 34: Adopt a transfer learning method and combine domain-specific annotation data to fine-tune the overall model; through pre-training BERT and tuning on specific data sets, improve the recognition accuracy and model robustness in the container terminal environment.
[0030] Further, Step 31 includes the following steps:
[0031] Step S311: Word Segmentation and Tokenization; Extract the input statement from the original text data, perform word segmentation on it, and add specific special tokens;
[0032] Step S312: Word Embedding Generation; Use the word embedding layer of BERT to convert the segmented sequence into word vectors. The word embedding layer combines token embeddings, segment embeddings, and position embeddings;
[0033] Step S313: Multi-Layer Transformer Encoding; Through the multi-layer Transformer encoder of BERT, use the self-attention mechanism to enable the model to focus on the relationships between words;
[0034] Step S314: Feature Representation Optimization; Generate word vectors containing context information by processing the hidden state outputs of each word.
[0035] Further, the step 32 includes the following steps:
[0036] Step S321: Bidirectional Encoding; Encode the sequence in two directions in the BiLSTM, namely the forward LSTM and the backward LSTM, to ensure that each hidden state contains complete context information;
[0037] Step S322: Memory Gate Mechanism; Adjust the storage and update of information through mechanisms such as forget gates, input gates, and output gates, capture long-distance dependencies, and improve the ability to recognize sequence patterns;
[0038] Step S323: Hidden State Calculation; Calculate the hidden state at each time point based on the memory states of the previous and next time steps.
[0039] Further, the step 33 includes the following steps:
[0040] Step S331: Conditional Probability Modeling: In the CRF layer, perform conditional probability modeling on the sequence output generated by the BiLSTM;
[0041] Step S332: Viterbi Algorithm Decoding: Use the Viterbi algorithm to find the highest conditional probability path for each label sequence, and optimize through the transitions between label sequences;
[0042] Step S333: Sequence Label Optimization: Optimize the processing of the hidden state sequence y at each time step to improve the overall accuracy of recognition and achieve precise recognition of abnormal events and potential risks.
[0043] Further, the step 4 includes the following steps:
[0044] Step S41: Knowledge Storage; Store the extracted knowledge elements, including entities and their relationships, in the Neo4j graph database;
[0045] Step S42: Multi-dimensional analysis; Use the Cypher query language to perform multi-dimensional analysis on the stored graph data;
[0046] Step S43 Coupling mechanism evaluation; By deeply analyzing the entities and their relationships in the visualization results, detect the coupling mechanism between the abnormal problems and potential risk hazards in the container terminal, and evaluate the propagation paths and influence of potential risks.
[0047] Compared with the prior art, the present invention has the following advantages:
[0048] The ACT abnormal problem and risk hazard analysis model based on BERT-BiLSTM-CRF provided by the present invention can analyze and manage the abnormal events and risks of ACT more efficiently compared with the prior art. This model uses the BERT-BiLSTM-CRF architecture to enhance the in-depth understanding ability of complex text semantics and sequence relationships, so as to better parse the diverse information in port operations. The architecture design of this model is simple and clear, which is convenient for integration with the existing container terminal management system, thereby effectively improving the stability and security of daily operations, reducing unnecessary economic losses and potential safety hazards. With the continuous update and iteration of data, the model ensures the high efficiency and accuracy in abnormal detection and risk analysis. In summary, the present invention has significant application value and practical benefits in the field of safety management and risk control of container terminals. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 It is the overall framework diagram of the ACT abnormal problem and risk hazard analysis model based on BERT-BiLSTM-CRF of the present invention;
[0051] Figure 2 It is the framework diagram of data preprocessing and corpus construction of ACT abnormal problems and risk hazards of the present invention;
[0052] Figure 3 It is the framework diagram of abnormal and risk hazard extraction based on BERT-BiLSTM-CRF of the present invention;
[0053] Figure 4 It is the visualization result of the ACT abnormal problem and risk hazard analysis model based on BERT-BiLSTM-CRF of the present invention. Detailed implementation manners
[0054] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0055] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0056] As Figures 1-4 shown, the present invention provides an automated container terminal anomaly and risk analysis method based on BERT-BiLSTM-CRF, including the following steps:
[0057] Step 1: Determine entity extraction requirements and relationship modeling. In this application, the process of determining entity extraction requirements and relationship modeling in Step 1 includes the following steps:
[0058] Step 11: Identify and refine operation anomaly events and related risks according to the basic requirements of the container terminal;
[0059] Step 12: Conduct a comprehensive analysis of existing accident data, logs and monitoring records, identify relevant key entities and potential relationships, and form a preliminary knowledge framework.
[0060] Extract data from the port operation system, including risk accident sheets, daily inspection and detection records, historical monitoring video logs, and related accident records; conduct a comprehensive analysis of the data, including the types of abnormal events (such as equipment failures, scheduling anomalies, extreme weather, etc.), descriptions of risk hazards (such as specific hazard phenomena and possible consequences), and the production process stages where they occur (such as loading and unloading operations, transportation, yard storage, etc.). Through text mining combined with expert experience, annotate and refine the key entities in the abnormal events (i.e., representing the categories, affected areas, hazard contents and locations of abnormal events, etc.) and potential causal relationships or associations (such as abnormal events triggering risk hazards, the impact relationship of specific risks on process nodes, the mutual association between abnormal events or the conduction relationship to specific risk chains, etc.), thereby forming a preliminary knowledge framework.
[0061] Step 13: Construct a multi-level knowledge ontology, refine the definitions of entities and relationships in terms of structure and semantics to ensure that the diversity and complexity of abnormal events can be accurately reflected in the model.
[0062] Construct a multi-level knowledge ontology, which includes three levels: (1) Abnormal problem entity layer: Refine the classification of abnormal problems, such as crane equipment failures, container damage, truck scheduling anomalies, etc., and clarify the attribute characteristics of each type of abnormality; (2) Risk hazard entity layer: Identify possible hazard characteristics for each type of abnormal problem, such as reduced operation efficiency, yard congestion, equipment downtime, etc.; (3) Operation process entity layer: Associate the process nodes with the occurrence locations or stages of abnormal problems and hazards, and combine the definition of the operation process, such as automated equipment operations in the shore operation stage, container flipping and repositioning in the yard loading and unloading environment, etc., to describe the relationship between how abnormal events trigger or affect process nodes. The structure of the multi-level knowledge ontology is defined through horizontal and vertical semantic relationships to ensure that the diversity (including various relationships such as triggering, conduction, mutual exclusion, etc.) and complexity (reflecting the characteristics of multiple causes / multiple consequences and cross-impacts of multiple nodes in accidents) of abnormal events and their impact chains can be comprehensively captured, and its accuracy is verified through actual scenario data annotation.
[0063] Furthermore, Step 2: Data preprocessing and dataset construction; collect research materials and actual operation data from container terminals to generate an accident corpus with specific logical semantics.
[0064] The process of Step 2 data preprocessing and dataset construction includes the following steps:
[0065] Step 21: Obtain research materials and actual operation data of container terminals;
[0066] Obtain research materials and actual operation data of container terminals; among them, research materials include literature, technical reports, patents, industry standard documents, and port accident case analyses related to container terminals, which are collected by consulting public materials and patent databases; actual operation data includes daily port operation data, monitoring log records, operation process-related documents, and accident report summaries, etc., which are obtained by extracting from the port management system and monitoring platform;
[0067] Step 22: Introduce a multi-level data processing method to extract useful information from the original data; conduct in-depth analysis of the text through semantic analysis to identify key themes and entities; at the same time, split long sentences and complex structures through long sentence segmentation to convert them into short sentences; in addition, conduct preliminary cleaning and unified formatting of the data through format alignment to ensure consistency and readability;
[0068] Introduce a multi-level data processing method to the obtained multi-source data to extract useful information related to the research topic, that is, the content closely related to the normal operation of ACT extracted from the obtained original data, including but not limited to risk hazards that may lead to process anomalies, descriptions of abnormal events, triggering conditions of risk events, and their possible influence ranges, etc. The specific process of the multi-level data processing method includes:
[0069] 1) Conduct in-depth semantic analysis of the text data through semantic analysis technology to identify key themes and entities related to abnormal operations and potential risk hazards in automated container terminals (ACT);
[0070] 2) For long sentences and complex structures in the text, convert them into short sentences through clause segmentation to simplify the difficulty of the model in understanding complex information;
[0071] 3) Conduct preliminary data cleaning and unified formatting, including operations such as deleting meaningless symbols, filling in missing information, standardizing time and unit formats, etc., to ensure data consistency and readability.
[0072] Step 23: Generate a corpus of abnormal problems and risk hazards in container terminals with specific logical semantics; the corpus of abnormal problems and risk hazards in container terminals includes: various context information related to abnormal events, ensuring that the knowledge in the corpus can be fully utilized during model training;
[0073] Generate a corpus of ACT abnormal problems and risk hazards with specific logical semantics based on the processed data; the corpus of ACT abnormal problems and risk hazards contains abnormal events and their context information, that is, the location, time, process nodes involved, related equipment, and the causal relationship that may affect the transfer, etc., of the abnormal or risk occurrence, ensuring that the knowledge information embedded in the corpus data can be fully utilized during model training.
[0074] Step 24: Divide the container terminal anomaly problem and risk hidden danger corpus into a training data set, a test data set, and a validation data set at a ratio of 7:2:1;
[0075] Step 25: Perform knowledge processing on the processed data and organize it into triples to complete the construction of the container terminal anomaly problem and risk hidden danger data set.
[0076] Analyze and model the key points (such as events, impacts, hidden dangers, etc.) and association relationships further extracted from the ACT anomaly problem and risk hidden danger corpus, and organize them into a knowledge expression in the form of triples. Structurally construct triples (h, r, t), where h represents the main entity, such as equipment or operation node, t represents the tail entity related to the main entity, such as a fault or risk factor, and r is the semantic relationship between the two, such as "trigger" or "intensify", etc., so as to complete the structured form expression of the ACT anomaly problem and risk hidden danger.
[0077] Step 25 organizes entities and relationships into triples, including the following steps:
[0078] Step 251: Conduct a preliminary check on the entity and relationship data extracted by the model; the preliminary check includes: spelling error correction and consistency verification;
[0079] Step 252: For the entities and relationships screened by manual verification and automated rules, where h represents the main entity, t represents the tail entity, and r represents the relationship between the two;
[0080] Step 253: Through the ACT domain-specific business scenarios and operation processes, ensure that the definition and applicable scenarios of each relationship are clearly supported by the context.
[0081] Through the ACT domain-specific business scenarios and operation processes, including but not limited to: shore operation (such as ship loader and unloader operation), horizontal transportation (such as container tractor transportation), and yard operation (such as stacker operation or yard storage management), ensure that the definition of the relationship in each triple is clear and the semantics conform to the business scenario and operation process.
[0082] Furthermore, Step 3: Anomaly problem and risk hidden danger identification; Use the BERT-BiLSTM-CRF framework to process unstructured data and realize the automated extraction of accident-related entities and their relationships; Step 3 includes the following steps:
[0083] Step 31: Use a deep learning model based on the Transformer with bidirectional encoder representations to extract features from the cleaned and standardized text data; Step 31 includes the following steps:
[0084] Step S311: Word Segmentation and Tokenization; Extract the input statement from the original text data, perform word segmentation on it, and add specific special tokens;
[0085] Step S312: Word Embedding Generation; Use the word embedding layer of BERT to convert the segmented sequence into word vectors. The word embedding layer combines token embedding, segment embedding, and position embedding;
[0086] Step S313: Multi-layer Transformer Encoding; Through the multi-layer Transformer encoder of BERT, using the self-attention mechanism, enable the model to focus on the relationships between words;
[0087] Step S314: Feature Representation Optimization; Generate word vectors containing context information by processing the hidden state output of each word.
[0088] Step 32: Pass the feature vectors extracted from the bidirectional encoder to the bidirectional long short-term memory network; Step 32 includes the following steps:
[0089] Step S321: Bidirectional Encoding; Encode the sequence in two directions in the BiLSTM, namely the forward LSTM and the backward LSTM, to ensure that each hidden state contains complete context information;
[0090] Step S322: Memory Gate Mechanism; Adjust the storage and update of information through mechanisms such as forget gate, input gate, and output gate, capture long-distance dependencies, and improve the ability to recognize sequence patterns;
[0091] Step S323: Hidden State Calculation; Calculate the hidden state at each time point according to the memory states of the previous and next time steps.
[0092] Step 33: Introduce a conditional random field in the output layer of the BiLSTM, and combine the properties of the conditional random field to process the optimal distribution of labels in sequence labeling; Step 33 includes the following steps:
[0093] Step S331: Conditional Probability Modeling: In the CRF layer, perform conditional probability modeling on the sequence output generated by the BiLSTM;
[0094] Step S332: Viterbi Algorithm Decoding: Use the Viterbi algorithm to find the highest conditional probability path for each label sequence, and optimize through the transitions between label sequences;
[0095] Step S333: Sequence Label Optimization: Optimize the processing of the hidden state sequence y at each time step to improve the overall accuracy of recognition and achieve precise recognition of abnormal events and potential risks.
[0096] Step 34: Use the transfer learning method and combine domain-specific labeled data to fine-tune the overall model; improve the recognition accuracy and model robustness in the container terminal environment through pre-training BERT and tuning on a specific dataset.
[0097] Further, step 4: Abnormal and risk data storage and application analysis; store the extracted knowledge elements in the Neo4j graph database to ensure systematic storage and management of knowledge; through the use of the Cypher query language, achieve multi-dimensional visual analysis for the exploration and management of the accident knowledge graph.
[0098] Step 4 includes the following steps:
[0099] Step S41: Knowledge storage; store the extracted knowledge elements, including entities and their relationships, in the Neo4j graph database;
[0100] Step S42: Multi-dimensional analysis; perform multi-dimensional analysis on the stored graph data using the Cypher query language;
[0101] Step S43 Coupling mechanism evaluation; through in-depth analysis of entities and their relationships in the visualization results, detect the coupling mechanism between abnormal problems and risk hazards in the container terminal, and evaluate the propagation path and influence of potential risks.
[0102] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0103] In the above embodiments of the present invention, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0104] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0105] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0106] In addition, in each embodiment of the present invention, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0107] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.
[0108] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of each embodiment of the present invention.
Claims
1. An automated container terminal anomaly and risk analysis method based on BERT-BiLSTM-CRF, characterized in that Including the following steps: Step 1: Determine entity extraction requirements and relationship modeling; Step 2: Data preprocessing and dataset construction; Collect research materials and actual operation data from container terminals to generate an accident corpus with specific logical semantics; Step 3: Identify abnormal problems and potential risks; Use the BERT-BiLSTM-CRF framework to process unstructured data and achieve automated extraction of accident-related entities and their relationships; Step 4: Store and apply analysis of abnormal and risk data; Store the extracted knowledge elements in the Neo4j graph database to ensure systematic storage and management of knowledge; Through the use of the Cypher query language, achieve multi-dimensional visual analysis for the exploration and management of the accident knowledge graph.
2. The automated container terminal anomaly and risk analysis method based on BERT-BiLSTM-CRF according to claim 1, characterized in that, The process of determining entity extraction requirements and relationship modeling in Step 1 includes the following steps: Step 11: Identify and refine abnormal operation events and related risks according to the basic requirements of container terminals; Step 12: Conduct a comprehensive analysis of existing accident data, logs, and monitoring records to identify relevant key entities and potential relationships, and form a preliminary knowledge framework; Step 13: Construct a multi-level knowledge ontology to refine the definitions of entities and relationships structurally and semantically to ensure that the diversity and complexity of abnormal events can be accurately reflected in the model.
3. The automated container terminal anomaly and risk analysis method based on BERT-BiLSTM-CRF according to claim 1, characterized in that, The process of data preprocessing and dataset construction in Step 2 includes the following steps: Step 21: Obtain research materials and actual operation data of container terminals; Step 22: Introduce a multi-level data processing method to extract useful information from the original data; Conduct in-depth analysis of the text through semantic analysis to identify key topics and entities; At the same time, segment long sentences and complex structures through long sentence segmentation to convert them into short sentences; In addition, conduct preliminary cleaning and unified formatting processing of the data through format alignment to ensure consistency and readability; Step 23: Generate a corpus of abnormal problems and potential risks in container terminals with specific logical semantics; The corpus of abnormal problems and potential risks in container terminals includes: various context information related to abnormal events to ensure that the knowledge in the corpus can be fully utilized during model training; Step 24: Divide the corpus of abnormal problems and potential risks in container terminals into a training dataset, a test dataset, and a validation dataset in a ratio of 7:2:1; Step 25: Process the processed data for knowledge and organize it into triples to complete the construction of the dataset of abnormal problems and potential risks in container terminals.
4. According to an automated container terminal abnormal and risk analysis method based on BERT-BiLSTM-CRF as described in claim 3, Step 25 of organizing entities and relationships into triples includes the following steps: Step 251: Conduct a preliminary check on the entity and relationship data extracted by the model; The preliminary check includes: spelling error correction and consistency verification; Step 252: For entities and relationships that have passed manual verification and automated rule screening, where h represents the main entity, t represents the tail entity, and r represents the relationship between the two; Step 253: Ensure that the definition and applicable scenarios of each relationship are supported by clear context through ACT domain-specific business scenarios and operation processes.
5. The automated container terminal anomaly and risk analysis method based on BERT-BiLSTM-CRF according to claim 1, characterized in that, The said Step 3 includes the following steps: Step 31: Use a deep learning model based on a Transformer with bidirectional encoder representations to extract features from the cleaned and standardized text data; Step 32: Pass the feature vectors extracted from the bidirectional encoder to a bidirectional long short-term memory network; Step 33: Introduce a conditional random field in the output layer of the BiLSTM, and process the optimal distribution of labels in sequence annotation in combination with the properties of the conditional random field; Step 34: Adopt a transfer learning method and fine-tune the overall model in combination with domain-specific annotation data; improve the recognition accuracy and model robustness in the container terminal environment through pre-training BERT and tuning on specific datasets.
6. The automated container terminal anomaly and risk analysis method based on BERT - BiLSTM - CRF according to claim 5, characterized in that, The said Step 31 includes the following steps: Step S311: Word segmentation and tokenization; Extract input statements from the original text data, perform word segmentation on them, and add specific special tokens; Step S312: Word embedding generation; Use the word embedding layer of BERT to convert the segmented sequence into word vectors. The word embedding layer combines token embeddings, segment embeddings, and position embeddings; Step S313: Multi-layer Transformer encoding; Through the multi-layer Transformer encoder of BERT, use the self-attention mechanism to enable the model to focus on the relationships between each word; Step S314: Feature representation optimization; Generate word vectors containing context information by processing the hidden state output of each word.
7. An automated container terminal anomaly and risk analysis method based on BERT-BiLSTM-CRF according to claim 5, characterized in that, The said Step 32 includes the following steps: Step S321: Bidirectional encoding; Encode the sequence in two directions in the BiLSTM, namely the forward LSTM and the backward LSTM, to ensure that each hidden state contains complete context information; Step S322: Memory gate mechanism; Adjust the storage and update of information through mechanisms such as forget gates, input gates, and output gates, capture long-distance dependencies, and improve the ability to recognize sequence patterns; Step S323: Hidden state calculation; Calculate the hidden state at each time point according to the memory states of the previous and subsequent time steps.
8. An automated container terminal anomaly and risk analysis method based on BERT-BiLSTM-CRF according to claim 5, characterized in that The said Step 33 includes the following steps: Step S331: Conditional probability modeling: In the CRF layer, perform conditional probability modeling on the sequence output generated by the BiLSTM; Step S332: Viterbi algorithm decoding: Use the Viterbi algorithm to find the highest conditional probability path for each label sequence, and optimize through the transition between label sequences; Step S333: Sequence label optimization: Optimize the processing of the hidden state sequence y at each time step to improve the overall accuracy of recognition and achieve precise recognition of abnormal events and potential risks.
9. The automated container terminal anomaly and risk analysis method based on BERT-BiLSTM-CRF according to claim 1, characterized in that The said Step 4 includes the following steps: Step S41: Knowledge storage; Store the extracted knowledge elements, including entities and their relationships, in the Neo4j graph database; Step S42: Multi-dimensional analysis; Perform multi-dimensional analysis on the stored graph data using the Cypher query language; Step S43: Coupling mechanism evaluation; By deeply analyzing the entities and their relationships in the visualization results, detect the coupling mechanism between the abnormal problems and potential risks in the container terminal, and evaluate the propagation paths and influence of potential risks.