Staged large model construction method for multi-modal information of petrochemical energy marine transportation industry
Through multimodal data integration and preprocessing, phased training and knowledge enhancement methods, the problems of complex data formats and model lag in petrochemical energy sea transportation have been solved, and the real-time, compliance and security of the model have been improved, meeting the efficient transportation needs of petrochemical energy sea transportation.
Patent Information
- Application Number
- CN202510606639.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-09-26
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence, natural language processing, and maritime shipping risk management, and specifically to a method for constructing a phased large-scale model for multimodal information in the petrochemical energy shipping industry. Background Art
[0002] Petrochemical energy shipping refers to the process of transporting petrochemical energy by ship at sea. Petrochemical energy mainly includes liquid or gaseous energy extracted from fossil fuels, such as oil, natural gas, and coal.
[0003] Currently, when transporting petrochemical energy and hazardous chemicals by sea, efficient scheduling must be taken into account while also strictly adhering to safety regulations and operating standards. However, the maritime transport process often involves the following challenges:
[0004] 1. Complex data formats: Surveillance images, ship communication recordings, shipping logs, and sensor time series data all come from different channels, making it difficult for traditional single-modality algorithms to integrate them.
[0005] 2. High risk and strict compliance: Oil and chemical products and processes require no errors during loading and unloading, navigation, and port docking. Furthermore, the frequent updates of relevant regulations and standards can easily lead to a lag in model knowledge.
[0006] 3. Significant real-time requirements: Unexpected weather, sea conditions, and international market fluctuations can instantly impact shipping decisions, requiring the model to iterate immediately and output actionable recommendations.
[0007] 4. Large-scale training resource consumption: Massive amounts of business data continue to emerge. Without phased training and flexible resource management, the model iteration cycle is too long and forgetfulness is prone to occur.
[0008] Based on the above, a phased large-scale model construction method for multimodal information in the petrochemical energy shipping industry is invented to strengthen professional knowledge in the field of petrochemical shipping, and supplemented by compliance and safety control mechanisms to help ensure the reliability and efficiency of maritime transportation activities. Summary of the Invention
[0009] To solve the above technical problems, according to one aspect of the present invention, the present invention provides the following technical solutions:
[0010] A phased large-scale model construction method for multimodal information in the petrochemical energy shipping industry includes the following specific steps:
[0011] S1, multimodal data acquisition and preprocessing: First, the navigation data is collected. After collection, the data will be quality-filtered and labeled. After the tags are added, they will be encrypted during data transmission and storage. At the same time, the data can be enhanced.
[0012] S2, phased training and knowledge reinforcement: first, unsupervised cross-modal pre-training, followed by industry fine-tuning and regulatory integration, and finally, instructional application and online learning;
[0013] S3, model construction and deployment optimization: first carry out multimodal fusion architecture, then carry out layered modular design, and then carry out model slimming and compliance management as well as knowledge distillation and terminal deployment.
[0014] As a preferred solution of the method for constructing a multi-modal large-scale model for the petrochemical energy shipping industry described in the present invention, the specific steps of S1 are as follows:
[0015] S11, heterogeneous data integration: first collect data and pre-process the data;
[0016] S12, quality filtering and labeling: First, perform quality filtering on the data, and then label the data;
[0017] S13, Privacy and Security Strategy: During data transmission and storage, use TLS or IPSec protocols to anonymize or homomorphically encrypt sensitive content to ensure compliance with regulations and industry compliance requirements;
[0018] S14, data enhancement: first enhance the text data, then enhance the image / video data, and finally enhance the audio data.
[0019] As a preferred solution of the method for constructing a multi-modal large model for the petrochemical energy shipping industry according to the present invention, the specific steps of S11 are as follows:
[0020] S111: First, acquire text, image / video, audio, and time series measurement data in real time from shipping information systems, field sensors, and external organizations;
[0021] S112: Unify the format and perform basic cleaning on all data.
[0022] As a preferred solution of the method for constructing a multi-modal large model for the petrochemical energy shipping industry according to the present invention, the specific steps of S12 are as follows:
[0023] S121: After converting non-text modalities into parseable text using OCR and ASR technologies, noise or duplicate data is removed by combining word frequency distribution and similarity measurement methods;
[0024] S122: Label key fields such as hazardous chemical categories, port regulations, and ship types to facilitate subsequent model retrieval and semantic understanding.
[0025] As a preferred solution of the method for constructing a multi-modal large model for the petrochemical energy shipping industry according to the present invention, the specific steps of S14 are as follows:
[0026] S141: performing synonym replacement and rewriting on text data;
[0027] S142: cropping and noise injection of images / videos;
[0028] S143: Perform pitch perturbation or spectrum analysis on the audio data.
[0029] As a preferred solution of the method for constructing a multi-modal large model for the petrochemical energy shipping industry described in the present invention, the specific steps of S2 are as follows:
[0030] S21, unsupervised cross-modal pre-training: first conduct large-scale self-supervised learning, followed by elastic computing power and parallelization;
[0031] S22, industry fine-tuning and regulatory integration: supervised fine-tuning is performed first, followed by reinforcement learning and human review feedback, and finally knowledge base integration;
[0032] S23, Instructional Application and Online Learning: First, multiple rounds of interactive instructions are performed, followed by online iterations.
[0033] As a preferred solution of the method for constructing a multi-modal large model for the petrochemical energy shipping industry according to the present invention, the specific steps of S21 are as follows:
[0034] S211, Large-Scale Self-Supervised Learning: Leveraging massive amounts of unlabeled multimodal data to train basic models to obtain universal cross-modal representations;
[0035] S212, Elastic Computing and Parallelization: Dynamically allocate GPU / TPU nodes in a containerized environment, and reduce resource usage and accelerate training convergence through mixed precision or pipeline parallelism.
[0036] As a preferred solution of the method for constructing a multi-modal large model for the petrochemical energy shipping industry according to the present invention, the specific steps of S22 are as follows:
[0037] S221, supervised fine-tuning: Introducing petrochemical and hazardous chemical transportation cases and maritime regulations with labeled data, and then conducting deep learning on the model to enable it to grasp safety regulations and emergency response points;
[0038] S222, Reinforcement Learning and Human Feedback: Using expert scoring or on-site operation results as feedback, the model continuously refines its decisions on high-risk scenarios such as hazard identification and loading and unloading process control.
[0039] S223, knowledge base linkage: For newly released industry documents or hazardous materials classification information, the model can be aligned and updated at any time to ensure that the output keeps pace with the times.
[0040] As a preferred solution of the method for constructing a multi-modal large model for the petrochemical energy shipping industry according to the present invention, the specific steps of S23 are as follows:
[0041] S231, Multi-round Interactive Instructions: Scenarios such as route scheduling, capacity dispatch, and weather warnings are designed as dialogue instruction sets, and continuous dialogue tasks are used to further hone the model's reasoning and execution capabilities.
[0042] S232, Online Iteration: After the system is put into use, the model can perform incremental learning based on real-time feedback data and quickly update weights using container clusters or distributed architectures to ensure the immediacy of scheduling and risk prediction.
[0043] As a preferred solution of the method for constructing a multi-modal large model for the petrochemical energy shipping industry described in the present invention, the specific steps of S3 are as follows:
[0044] S31, Multimodal Fusion Architecture: Setting up cross-modal attention or feature fusion modules in the model backbone to achieve unified encoding and collaborative reasoning between text, images, audio, and sensor sequences;
[0045] S32, layered modular design: Specific functions are implemented as pluggable sub-modules, supporting flexible expansion while ensuring the stability of the core semantic layer;
[0046] S33, Model Slimming and Compliance Management: Quantization and pruning technologies are used on the trained master model to reduce memory and latency. A hierarchical access and log audit mechanism is also configured to trace dangerous operations and key inferences, complying with industry regulatory and security requirements.
[0047] S34, knowledge distillation and terminal deployment: The professional capabilities of the large model are transferred to the small model through distillation, so that the latter can be deployed on terminal devices with limited bandwidth or computing power to meet the real-time decision-making needs of the work site.
[0048] Compared with existing technologies:
[0049] 1. Through multimodal data integration and preprocessing, not only can the problem of diverse data source channels and complex forms be effectively solved, providing a unified and standardized data foundation for subsequent model processing, but it can also achieve unified encoding and collaborative reasoning between text, images / audio and sensor sequences, further integrating multimodal data, making full use of information from different modal data, and improving the model's ability to process complex data.
[0050] 2. Through industry fine-tuning, regulatory implantation, reinforcement learning, and human review feedback, the model can be updated with newly released industry documents or hazardous materials classification information at any time, ensuring that the output keeps pace with the times. This ensures that the model can strictly comply with safety regulations and operating standards, adapt to the frequent updates of regulatory standards, and reduce the risks caused by lagging model knowledge.
[0051] 3. Through model slimming and compliance management, it is possible to trace dangerous operations and key inferences based on hierarchical access and log audit mechanisms, complying with industry regulatory and safety requirements. This management mechanism ensures the compliance and safety of the model in the high-risk petrochemical energy shipping sector.
[0052] 4. Through elastic computing power and parallelization, the efficiency of model training is improved, allowing the model to be updated and optimized more quickly to respond to situations requiring immediate decision-making, such as sudden weather events, sea conditions, and changes in international markets. In addition, through online iteration, the model can promptly output feasible suggestions based on real-time changes, meeting the real-time needs of petrochemical energy shipping.
[0053] 5. Through phased training and knowledge reinforcement, it can dynamically allocate resources in a containerized environment and effectively utilize resources, avoiding the long model iteration cycle and forgetting problems caused by the lack of phased training and elastic resource management.
[0054] 6. Through model slimming and knowledge distillation, the system can quantize and prune the trained main model to reduce memory and latency. It can also transfer the professional capabilities of the large model to the small model through distillation, allowing the small model to be deployed on edge devices with limited bandwidth or computing power. While ensuring model performance, it reduces the demand for large-scale training resources and improves resource utilization efficiency. DETAILED DESCRIPTION
[0055] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention are described in further detail below.
[0056] The present invention provides a phased large-scale model construction method for multimodal information in the petrochemical energy shipping industry, comprising the following specific steps:
[0057] S1, multimodal data acquisition and preprocessing: First, the navigation data is collected. After collection, the data will be quality-filtered and labeled. After the tags are added, they will be encrypted during data transmission and storage. At the same time, the data can be enhanced.
[0058] The specific steps of S1 are as follows:
[0059] S11, heterogeneous data integration: first collect data and pre-process the data;
[0060] The specific steps of S11 are as follows:
[0061] S111: First, obtain text (such as logbooks, documents, etc.), image / video (such as monitoring or inspection images), audio (such as wireless communication), and time series measurement data (such as temperature, liquid level, speed, etc.) in real time from shipping information systems, on-site sensors, and external organizations;
[0062] S112: Unify the format and perform basic cleaning of all data;
[0063] S12, quality filtering and labeling: First, perform quality filtering on the data, and then label the data;
[0064] The specific steps of S12 are as follows:
[0065] S121: After converting non-text modalities into parseable text using OCR and ASR technologies, noise or duplicate data is removed by combining word frequency distribution and similarity measurement methods;
[0066] S122: Label key fields such as hazardous chemical categories, port regulations, and ship types to facilitate subsequent model retrieval and semantic understanding;
[0067] S13, Privacy and Security Strategy: During data transmission and storage, use TLS or IPSec protocols to anonymize or homomorphically encrypt sensitive content to ensure compliance with regulations and industry compliance requirements;
[0068] S14, data enhancement: first enhance the text data, then enhance the image / video data, and finally enhance the audio data;
[0069] The specific steps of S14 are as follows:
[0070] S141: performing synonym replacement and rewriting on text data;
[0071] S142: cropping and noise injection of images / videos;
[0072] S143: performing pitch perturbation or spectrum analysis on the audio data;
[0073] S2, phased training and knowledge reinforcement: first, unsupervised cross-modal pre-training, followed by industry fine-tuning and regulatory integration, and finally, instructional application and online learning;
[0074] The specific steps of S2 are as follows:
[0075] S21, unsupervised cross-modal pre-training: first conduct large-scale self-supervised learning, followed by elastic computing power and parallelization;
[0076] The specific steps of S21 are as follows:
[0077] S211, Large-Scale Self-Supervised Learning: Leveraging massive amounts of unlabeled multimodal data to train basic models to obtain universal cross-modal representations;
[0078] S212, Elastic Computing and Parallelization: Dynamically allocate GPU / TPU nodes in a containerized environment, and reduce resource usage and accelerate training convergence through mixed precision or pipeline parallelization;
[0079] S22, industry fine-tuning and regulatory integration: supervised fine-tuning is performed first, followed by reinforcement learning and human review feedback, and finally knowledge base integration;
[0080] The specific steps of S22 are as follows:
[0081] S221, supervised fine-tuning: Introducing petrochemical and hazardous chemical transportation cases and maritime regulations with labeled data, and then conducting deep learning on the model to enable it to grasp safety regulations and emergency response points;
[0082] S222, Reinforcement Learning and Human Feedback: Using expert scoring or on-site operation results as feedback, the model continuously refines its decisions on high-risk scenarios such as hazard identification and loading and unloading process control.
[0083] S223, Knowledge Base Linkage: For newly released industry documents or hazardous materials classification information, the model is aligned and updated at any time to ensure that the output keeps pace with the times;
[0084] S23, Instructional Application and Online Learning: Multiple rounds of interactive instruction followed by online iteration;
[0085] The specific steps of S23 are as follows:
[0086] S231, Multi-round Interactive Instructions: Scenarios such as route scheduling, capacity dispatch, and weather warnings are designed as dialogue instruction sets, and continuous dialogue tasks are used to further hone the model's reasoning and execution capabilities.
[0087] S232, Online Iteration: After the system is put into use, the model can conduct incremental learning based on real-time data and quickly update weights using container clusters or distributed architectures to ensure the immediacy of scheduling and risk prediction.
[0088] S3, model construction and deployment optimization: First, develop a multimodal fusion architecture, then perform layered modular design, and then conduct model slimming and compliance management, knowledge distillation, and terminal deployment.
[0089] The specific steps of S3 are as follows:
[0090] S31, Multimodal Fusion Architecture: Setting up cross-modal attention or feature fusion modules in the model backbone to achieve unified encoding and collaborative reasoning between text, images, audio, and sensor sequences;
[0091] S32, layered modular design: Specific functions (such as hazardous chemical alarms and port queue predictions) are implemented as pluggable sub-modules, supporting flexible expansion while ensuring the stability of the core semantic layer;
[0092] S33, Model Slimming and Compliance Management: Quantization and pruning technologies are used on the trained master model to reduce memory and latency. A hierarchical access and log audit mechanism is also configured to trace dangerous operations and key inferences, complying with industry regulatory and security requirements.
[0093] S34, knowledge distillation and terminal deployment: The professional capabilities of the large model are transferred to the small model through distillation, so that the latter can be deployed on terminal devices with limited bandwidth or computing power to meet the real-time decision-making needs of the work site.
[0094] Based on the above, the present invention includes but is not limited to the following implementation cases:
[0095] Multimodal data acquisition and preprocessing example:
[0096] Scenario 1: Surveillance video and sensor fusion:
[0097] Extract cargo status or deck environment information (OCR identification of safety warnings) from port surveillance videos, combine it with sensor readings (storage temperature, tank level, etc.), and automatically verify and remove abnormal data points;
[0098] Scenario 2: Voice transcription and text association:
[0099] After crew conversations or emergency call recordings are transcribed into text by ASR, key risk terms (such as "leakage" and "fire") are annotated, the time sequence is mapped with the shipping log, and possible duplicate or noisy segments are removed.
[0100] Phased training and knowledge enhancement:
[0101] Stage 1: Unsupervised cross-modal pre-training:
[0102] Put pre-processed multimodal data into a container cluster for parallel training to evaluate the model's cross-modal alignment (such as text-image retrieval and audio-text matching) and general feature learning effects;
[0103] Phase 2: Industry fine-tuning and regulatory implementation:
[0104] Load industry regulations and typical hazardous chemicals cases for supervised training, focusing on the accuracy and timeliness of high-risk decisions (such as loading and unloading safety restrictions);
[0105] Through continuous fine-tuning using reinforcement learning, the model's ability to predict dangerous scenarios has gradually increased, and the incidence of illegal operations has been significantly reduced in simulation evaluations.
[0106] Phase 3: Instructional Application and Online Learning:
[0107] Test multi-round dialogue commands in the system, such as "What additional standards need to be followed when loading certain types of chemicals?" to evaluate the model's understanding of the commands and consistency of the continuous dialogue;
[0108] When there are sudden changes in sensors or market conditions, the model automatically samples new data, performs incremental learning online, and deploys hot updates, continuously improving the accuracy of capacity scheduling recommendations.
[0109] Typical application scenarios:
[0110] Risk warning and emergency guidance: Once the ship's sensors detect abnormal liquid levels or temperatures, the model immediately links with the maritime regulations knowledge base to generate corresponding emergency strategy prompts;
[0111] Route planning and dynamic scheduling: Based on weather forecasts, maritime traffic conditions, and real-time ship messages, the model automatically recommends the optimal sailing route or berthing port, reducing delays and fuel waste;
[0112] Comprehensive regulatory compliance assistance: Operators can consult the model on compliance issues such as hazardous chemical transportation restrictions and customs clearance procedures. The model, combined with the latest knowledge base, generates compliance operation guidelines and protective recommendations, eliminating tedious search work.
[0113] Although the present invention has been described above with reference to embodiments, various modifications may be made thereto and equivalent components may be substituted without departing from the scope of the present invention. In particular, as long as there are no structural conflicts, the various features of the embodiments disclosed herein may be combined with each other in any manner, and the omission of an exhaustive description of such combinations in this specification is solely for the sake of space and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A phased large-scale model construction method for multimodal information in the petrochemical energy shipping industry, characterized by: The specific steps are as follows: S1, multimodal data acquisition and preprocessing: First, the navigation data is collected. After collection, the data will be quality-filtered and labeled. After the tags are added, they will be encrypted during data transmission and storage. At the same time, the data can be enhanced. S2, phased training and knowledge reinforcement: first, unsupervised cross-modal pre-training, followed by industry fine-tuning and regulatory integration, and finally, instructional application and online learning; S3, model construction and deployment optimization: first carry out multimodal fusion architecture, then carry out layered modular design, and then carry out model slimming and compliance management as well as knowledge distillation and terminal deployment.
2. The method for constructing a multi-modal large-scale model for the petrochemical energy shipping industry according to claim 1 is characterized in that: The specific steps of S1 are as follows: S11, heterogeneous data integration: first collect data and pre-process the data; S12, quality filtering and labeling: First, perform quality filtering on the data, and then label the data; S13, Privacy and Security Strategy: During data transmission and storage, use TLS or IPSec protocols to anonymize or homomorphically encrypt sensitive content to ensure compliance with regulations and industry compliance requirements; S14, data enhancement: first enhance the text data, then enhance the image / video data, and finally enhance the audio data.
3. The method for constructing a multi-modal large-scale model for the petrochemical energy shipping industry according to claim 2 is characterized in that: The specific steps of S11 are as follows: S111: First, acquire text, image / video, audio, and time series measurement data in real time from shipping information systems, field sensors, and external organizations; S112: Unify the format and perform basic cleaning on all data.
4. The method for constructing a multi-modal large-scale model for the petrochemical energy shipping industry according to claim 2 is characterized in that: The specific steps of S12 are as follows: S121: After converting non-text modalities into parseable text using OCR and ASR technologies, noise or duplicate data is removed by combining word frequency distribution and similarity measurement methods; S122: Label key fields such as hazardous chemical categories, port regulations, and ship types to facilitate subsequent model retrieval and semantic understanding.
5. The method for constructing a multi-modal large-scale model for the petrochemical energy shipping industry according to claim 2 is characterized in that: The specific steps of S14 are as follows: S141: performing synonym replacement and rewriting on text data; S142: cropping and noise injection of images / videos; S143: Perform pitch perturbation or spectrum analysis on the audio data.
6. The method for constructing a multi-modal large-scale model for the petrochemical energy shipping industry according to claim 1 is characterized in that: The specific steps of S2 are as follows: S21, unsupervised cross-modal pre-training: first conduct large-scale self-supervised learning, followed by elastic computing power and parallelization; S22, industry fine-tuning and regulatory integration: supervised fine-tuning is performed first, followed by reinforcement learning and human review feedback, and finally knowledge base integration; S23, Instructional Application and Online Learning: First, multiple rounds of interactive instructions are performed, followed by online iterations.
7. The method for constructing a multi-modal large-scale model for the petrochemical energy shipping industry according to claim 6 is characterized in that: The specific steps of S21 are as follows: S211, Large-Scale Self-Supervised Learning: Leveraging massive amounts of unlabeled multimodal data to train basic models to obtain universal cross-modal representations; S212, Elastic Computing and Parallelization: Dynamically allocate GPU / TPU nodes in a containerized environment, and reduce resource usage and accelerate training convergence through mixed precision or pipeline parallelism.
8. The method for constructing a multi-modal large-scale model for the petrochemical energy shipping industry according to claim 6 is characterized in that: The specific steps of S22 are as follows: S221, supervised fine-tuning: Introducing petrochemical and hazardous chemical transportation cases and maritime regulations with labeled data, and then conducting deep learning on the model to enable it to grasp safety regulations and emergency response points; S222, Reinforcement Learning and Human Feedback: Using expert scoring or on-site operation results as feedback, the model continuously refines its decisions on high-risk scenarios such as hazard identification and loading and unloading process control. S223, knowledge base linkage: For newly released industry documents or hazardous materials classification information, the model can be aligned and updated at any time to ensure that the output keeps pace with the times.
9. The method for constructing a multi-modal large-scale model for the petrochemical energy shipping industry according to claim 1, characterized in that: The specific steps of S23 are as follows: S231, Multi-round Interactive Instructions: Scenarios such as route scheduling, capacity dispatch, and weather warnings are designed as dialogue instruction sets, and continuous dialogue tasks are used to further hone the model's reasoning and execution capabilities. S232, Online Iteration: After the system is put into use, the model can perform incremental learning based on real-time feedback data and quickly update weights using container clusters or distributed architectures to ensure the immediacy of scheduling and risk prediction.
10. The method for constructing a multi-modal large-scale model for the petrochemical energy shipping industry according to claim 1, characterized in that: The specific steps of S3 are as follows: S31, Multimodal Fusion Architecture: Setting up cross-modal attention or feature fusion modules in the model backbone to achieve unified encoding and collaborative reasoning between text, images, audio, and sensor sequences; S32, layered modular design: Specific functions are implemented as pluggable sub-modules, supporting flexible expansion while ensuring the stability of the core semantic layer; S33, Model Slimming and Compliance Management: Quantization and pruning technologies are used on the trained master model to reduce memory and latency. A hierarchical access and log audit mechanism is also configured to trace dangerous operations and key inferences, complying with industry regulatory and security requirements. S34, knowledge distillation and terminal deployment: The professional capabilities of the large model are transferred to the small model through distillation, so that the latter can be deployed on terminal devices with limited bandwidth or computing power to meet the real-time decision-making needs of the work site.