Offshore wind power and ocean engineering large model application method, device, equipment and medium

By constructing a domain-specific large model that integrates professional domain pre-training, supervised fine-tuning, and reinforcement learning algorithms based on reward models, the problem of insufficient specialization and generalization ability of large models in the fields of marine engineering and offshore wind power is solved, and efficient and accurate task processing is achieved.

CN122114190BActive Publication Date: 2026-07-31ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2026-04-28
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing large-scale pre-trained language models lack high-quality domain data, have limited understanding of professional terms, and lack generalization ability for interdisciplinary tasks in the application of marine engineering and offshore wind power, making it difficult to meet the high requirements of professionalism and accuracy.

Method used

By acquiring task requirement data in the fields of offshore wind power and marine engineering, and integrating professional domain pre-training, supervised fine-tuning, and reinforcement learning algorithm training based on reward models, a domain-specific large model is constructed to solve the problems of lack of domain expertise and insufficient task adaptability of general models.

Benefits of technology

It improves the professionalism and accuracy of domain task processing, reduces task processing costs, and enhances the generalization ability of the model, providing efficient and reliable technical support for intelligent practices in the fields of offshore wind power and marine engineering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114190B_ABST
    Figure CN122114190B_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, equipment, and medium for applying a large-scale model in offshore wind power and marine engineering. Relating to the field of artificial intelligence technology, the method includes: acquiring task requirement data in the field of offshore wind power and marine engineering; inputting the task requirement data into a large-scale model in the field of offshore wind power and marine engineering for processing, and obtaining professional processing results corresponding to the task requirement data; wherein, the large-scale model in the field of offshore wind power and marine engineering is a domain-specific large-scale model trained by integrating professional domain pre-training, supervised fine-tuning, and reinforcement learning algorithms based on reward models. This method can improve the professionalism and accuracy of domain task processing, reduce task processing costs, and enhance the model's generalization ability, providing efficient and reliable technical support for intelligent practices in the field of offshore wind power and marine engineering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, equipment and medium for the application of large-scale models of offshore wind power and marine engineering. Background Technology

[0002] With the rapid development of artificial intelligence technology, large-scale pre-trained language models have demonstrated outstanding performance in natural language processing, computer vision, and multimodal tasks. However, the application of general-purpose large models in specific vertical fields (such as marine engineering and offshore wind power) still faces many challenges. The main problems include: lack of high-quality domain data, limited understanding of technical terms, and insufficient generalization ability for interdisciplinary tasks.

[0003] Marine engineering and offshore wind power, as important directions for the development of marine economy and new energy, involve a complex multidisciplinary knowledge system encompassing engineering mechanics, marine meteorology, structural design, and equipment operation and maintenance. The scenarios are also highly variable (such as deep-sea engineering environments and wind power operation and maintenance under extreme weather conditions), placing higher demands on intelligent systems in terms of professional understanding and cross-domain reasoning. Currently, the practical application of large-scale models in this field is still in the exploratory stage and cannot meet the high accuracy and professionalism requirements of tasks such as marine platform design, mooring system analysis, wind turbine fault diagnosis, and engineering scheme optimization.

[0004] Therefore, there is an urgent need for a training method that can effectively improve the performance of large models in the fields of marine engineering and offshore wind power, so as to enhance their ability to apply knowledge, reason about complex problems and generate standardized content in professional scenarios, thereby promoting the deep application of artificial intelligence in this field. Summary of the Invention

[0005] This invention provides a method, apparatus, equipment, and medium for applying large-scale models of offshore wind power and marine engineering, in order to improve the professionalism and accuracy of task processing in the field, reduce task processing costs, and enhance the generalization ability of the model.

[0006] According to one aspect of the present invention, a method for applying large-scale models of offshore wind power and marine engineering is provided, comprising:

[0007] Acquire task requirement data in the field of offshore wind power and marine engineering;

[0008] The task requirement data is input into a large model in the field of offshore wind power and marine engineering for processing to obtain professional processing results corresponding to the task requirement data.

[0009] The large model for offshore wind power and marine engineering is a domain-specific large model trained by integrating professional domain pre-training, supervised fine-tuning, and reinforcement learning algorithms based on reward models.

[0010] According to another aspect of the present invention, a large-scale application device for offshore wind power and marine engineering is provided, comprising:

[0011] The requirement data acquisition module is used to acquire task requirement data in the field of offshore wind power and marine engineering.

[0012] The large model processing module is used to input the task requirement data into a large model in the field of offshore wind power and marine engineering for processing, and to obtain professional processing results corresponding to the task requirement data.

[0013] The large model for offshore wind power and marine engineering is a domain-specific large model trained by integrating professional domain pre-training, supervised fine-tuning, and reinforcement learning algorithms based on reward models.

[0014] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0015] At least one processor;

[0016] and memory that is communicatively connected to at least one processor;

[0017] The memory stores a computer program that can be executed by at least one processor, which is then executed by the at least one processor to enable the at least one processor to execute the offshore wind power and marine engineering large model application method according to any embodiment of the present invention.

[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement the offshore wind power and marine engineering large model application method of any embodiment of the present invention.

[0019] The technical solution of this invention acquires task requirement data in the field of offshore wind power and marine engineering, inputs this data into a large-scale model for processing, and obtains professional processing results corresponding to the task requirement data. This large-scale model is a domain-specific model trained by integrating professional domain pre-training, supervised fine-tuning, and a reward-based reinforcement learning algorithm. This technical solution constructs a domain-specific large-scale model through multi-stage targeted training. First, professional domain pre-training injects core domain knowledge into the basic model, addressing the lack of domain-specific accumulation in general models. Then, supervised fine-tuning adapts the model to actual task scenarios in the domain, compensating for the insufficient task adaptability of general models. Finally, a reward-based reinforcement learning algorithm optimizes the model's generation capabilities, addressing the problem that supervised learning models only learn fixed mapping relationships and cannot handle diverse and reasonable responses and dynamic context changes. Ultimately, this improves the professionalism and accuracy of domain task processing, reduces task processing costs, and enhances model generalization ability, providing efficient and reliable technical support for intelligent practices in the field of offshore wind power and marine engineering.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 A flowchart illustrating a method for applying a large-scale model of offshore wind power and marine engineering, as provided in an embodiment of the present invention;

[0023] Figure 2 A flowchart illustrating the training process of a large-scale model in the field of offshore wind power and marine engineering, provided as an embodiment of the present invention;

[0024] Figure 3 This is a flowchart of the original text data cleaning process provided in an embodiment of the present invention;

[0025] Figure 4 This is a flowchart of the GRPO reinforcement learning algorithm provided in an embodiment of the present invention;

[0026] Figure 5 A schematic diagram illustrating the classification of benchmark test sets provided in embodiments of the present invention;

[0027] Figure 6 This is a schematic diagram of the structure of a large-scale offshore wind power and marine engineering application device provided in an embodiment of the present invention;

[0028] Figure 7 A schematic diagram of the electronic device used to implement the large-scale model application method for offshore wind power and marine engineering in this embodiment of the invention. Detailed Implementation

[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0031] Figure 1 This is a flowchart illustrating a method for applying a large-scale model of offshore wind power and marine engineering, provided by an embodiment of the present invention. This embodiment is applicable to situations where various tasks in the field of offshore wind power and marine engineering are handled using a dedicated large-scale model. This method can be executed by an application device for the large-scale model of offshore wind power and marine engineering, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method specifically includes the following steps:

[0032] S110. Obtain task requirement data in the field of offshore wind power and marine engineering.

[0033] Among them, the task requirement data can be relevant data on various tasks to be processed in the field of offshore wind power and marine engineering, such as text-based questions asked by users through human-computer interaction.

[0034] S120. Input the task requirement data into the large model of offshore wind power and marine engineering for processing to obtain the professional processing results corresponding to the task requirement data.

[0035] Among them, the professional processing result can be understood as the processing result output by the large model in response to the input task requirement data, which meets the professional requirements of the domain; the large model of offshore wind power and marine engineering is a domain-specific large model trained by integrating professional domain pre-training, supervised fine-tuning and reinforcement learning algorithm based on reward model.

[0036] Domain-specific pre-training can be training a basic model with knowledge specific to the offshore wind power and marine engineering fields. Supervised fine-tuning can be targeted training of the model based on labeled domain data. Reinforcement learning algorithms based on reward models refer to learning algorithms that optimize the model's strategy by rewarding the quality evaluation results of the content generated by the model. Domain-specific large models can be large models that are adapted to the professional needs of the offshore wind power and marine engineering fields and have the ability to reason about knowledge and process tasks in that field.

[0037] Specifically, the acquired task requirement data can be input into a large-scale model specifically trained for the offshore wind power and marine engineering field. The model performs professional analysis and processing on the data, and finally outputs professional processing results corresponding to the task requirement data. This large-scale model integrates professional domain pre-training, supervised fine-tuning, and reinforcement learning algorithms based on reward models for training. It can accurately adapt to the professional needs of the offshore wind power and marine engineering field, and significantly improve the professionalism and accuracy of domain task processing.

[0038] The technical solution of this invention acquires task requirement data in the field of offshore wind power and marine engineering, inputs this data into a large-scale model for processing, and obtains professional processing results corresponding to the task requirement data. This large-scale model is a domain-specific model trained by integrating professional domain pre-training, supervised fine-tuning, and a reward-based reinforcement learning algorithm. This technical solution constructs a domain-specific large-scale model through multi-stage targeted training. First, professional domain pre-training injects core domain knowledge into the basic model, addressing the lack of domain-specific accumulation in general models. Then, supervised fine-tuning adapts the model to actual task scenarios in the domain, compensating for the insufficient task adaptability of general models. Finally, a reward-based reinforcement learning algorithm optimizes the model's generation capabilities, addressing the problem that supervised learning models only learn fixed mapping relationships and cannot handle diverse and reasonable responses and dynamic context changes. Ultimately, this improves the professionalism and accuracy of domain task processing, reduces task processing costs, and enhances model generalization capabilities, providing efficient and reliable technical support for intelligent practices in the field of offshore wind power and marine engineering.

[0039] Figure 2 This is a flowchart illustrating the training process of a large-scale model in the field of offshore wind power and marine engineering, provided as an embodiment of the present invention. This embodiment further optimizes the training process of the large-scale model in the field of offshore wind power and marine engineering described in the previous embodiments. For specific implementation details, please refer to the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here. Figure 2 As shown, the method specifically includes the following steps:

[0040] S210. Collect raw text data in the field of offshore wind power and marine engineering, and perform multi-layer cleaning and filtering processing on the raw text data to obtain the target dataset.

[0041] The raw text data can be unprocessed text information collected from relevant channels in the field of offshore wind power and marine engineering, such as 150,000 documents in the field of offshore wind power crawled from the Internet. It should be noted that the acquisition methods of the above-mentioned raw text data all comply with relevant legal regulations.

[0042] In this embodiment of the invention, raw text data in the fields of offshore wind power and marine engineering can be collected, and the raw text data can be cleaned and filtered in multiple layers to obtain a high-quality target dataset. The target dataset can be used for training various models in the future so that a large model that can handle the task requirements data in the fields of offshore wind power and marine engineering can be obtained when the training is completed, namely, a large model in the field of offshore wind power and marine engineering.

[0043] The target dataset includes a pre-training dataset, a domain dialogue dataset, a reward model dataset, and a benchmark dataset. The pre-training dataset refers to the dataset used for pre-training the basic language model in the domain. The domain dialogue dataset refers to the dataset containing question-and-answer dialogue content in the field of offshore wind power and marine engineering. The reward model dataset is used to construct multi-level answer quality labeling rules and train the basic reward model. The benchmark dataset is used to evaluate the performance of the initial large-scale model after training.

[0044] In some possible implementations, the collection of raw text data in the field of offshore wind power and marine engineering, and the multi-layer cleaning and filtering processing of the raw text data to obtain the target dataset, includes: performing basic format normalization and text feature quantization processing on the collected raw text data to obtain preprocessed text data; sequentially performing rule filtering and model quality scoring processing on the preprocessed text data to remove low-quality data and obtain filtered text data; performing algorithm deduplication and quality inspection processing on the filtered text data to obtain qualified text data; dividing the qualified text data according to data purpose to obtain the training dataset, the pre-training dataset, the domain dialogue dataset, the reward model dataset, and the benchmark dataset, and merging them to obtain the target dataset.

[0045] Among them, basic format normalization refers to the standardization of format and encoding of raw text data in the field of offshore wind power and marine engineering. Text feature quantization can be the calculation of features such as perplexity, number of tokens, and stop word ratio of raw text data. Preprocessed text data is the text data obtained after basic format normalization and text feature quantization.

[0046] In this embodiment, rule filtering can be a screening process of preprocessed text data based on preset labels, sorting rules, filtering rules, etc., and model quality scoring refers to using a model to score the quality of preprocessed text data in order to remove low-quality data; the screened text data is the text data obtained after rule filtering and model quality scoring.

[0047] Algorithm deduplication can be performed by using keyword matching and MinHash algorithms to remove duplicate data from the selected text data. Quality inspection refers to the quality verification of the deduplicated selected text data. Qualified text data is text data that meets the requirements for model training and evaluation after algorithm deduplication and quality inspection.

[0048] Specifically, basic format normalization and text feature quantification processing can be performed on the collected raw text data in the field of offshore wind power and marine engineering. Basic format normalization includes deleting email addresses, links, and IP information from the text, completing the encoding conversion from traditional Chinese to simplified Chinese and from Unicode to ASCII, and unifying the data format.

[0049] Furthermore, text feature quantification can involve calculating metrics such as perplexity, token count, and the ratio of stop words to alphanumeric characters. These two steps yield preprocessed text data. Next, rule-based filtering and model quality scoring are performed on the preprocessed text data. First, initial screening is conducted based on labels, ranking, and preset filtering rules. Then, a second screening is performed using the model's text quality scoring to remove low-quality data, resulting in filtered text data. Next, algorithmic deduplication and quality control are applied to the filtered text data. Keyword matching combined with the MinHash algorithm is used for text deduplication, and finally, quality verification is performed using data distribution statistics to obtain qualified text data.

[0050] This multi-layered screening approach can effectively improve data quality, providing clean and suitable high-quality data for large-scale models in the field of offshore wind power and marine engineering.

[0051] Finally, the qualified text data can be divided according to its intended use, resulting in a pre-training dataset, a domain-specific dialogue dataset, a reward model dataset, and a benchmark dataset. Different types of datasets are used to train the corresponding models. It should also be noted that the aforementioned text quality scoring model can be a lightweight model for general natural language processing.

[0052] For example, Figure 3 A flowchart for cleaning raw text data provided in an embodiment of the present invention, such as... Figure 3 As shown, it includes:

[0053] S2101: Perform basic cleaning, such as deleting emails, links, and IPs; converting between traditional and simplified Chinese characters; converting encodings (Unicode to ASCII); and unifying data formats.

[0054] S2102: Calculate text metrics such as perplexity, number of tokens (including maximum / average), and stop word / alphanumeric ratio to provide a quantitative basis for subsequent screening.

[0055] S2103: Filtering task, using labels, sorting, and rules to filter data and initially remove obviously unqualified content.

[0056] S2104: Quality scoring task, which uses model filtering to further filter out low-quality data.

[0057] S2105: Text deduplication task, using keyword matching and MinHash algorithm to remove duplicates and avoid interference from repetitive data.

[0058] S2106: Quality inspection task, manual sampling + data distribution statistics to ensure data quality.

[0059] S2107: After layers of processing, the "target dataset" is obtained, which can be used for model training.

[0060] S220. Based on the pre-trained dataset, perform professional domain pre-training on the basic language model to obtain a domain pre-trained model.

[0061] Among them, the basic language model can be a mature general-purpose language model, for example, DeepSeek-R1-0528-Qwen3-8B can be selected as the basic language model; the domain pre-trained model can be a model that has been pre-trained in the domain professionally and has the ability to understand basic domain knowledge, that is, a model that has the ability to understand offshore wind power and marine engineering knowledge.

[0062] In some possible implementations, the step of performing domain-specific pre-training on the basic language model based on the pre-training dataset to obtain a domain-specific pre-trained model includes: loading the basic language model and initializing the model parameters, word segmenter, and training framework of the basic language model; constructing the pre-training dataset as standardized training samples and dividing the training batches according to text features; setting a pre-training objective that integrates autoregressive language modeling and domain-specific question answering, and performing pre-training on the initialized basic language model using a training strategy that combines mixed precision training and gradient accumulation; dynamically adjusting and optimizing the pre-training process through training monitoring metrics; deriving the first model weights after training convergence; and obtaining the domain-specific pre-trained model based on the first model weights.

[0063] Among them, model parameters, word segmenter and training framework can be understood as the basic configuration required for training the basic language model; standardized training samples refer to samples that meet the requirements of model training input after format processing; text features can be the language type, length, semantic complexity and other features of text data; training batches refer to dividing the training samples into different batches according to preset rules, that is, multiple groups of samples for model training can be obtained by dividing and matching.

[0064] Autoregressive language modeling refers to a modeling approach that teaches the model to predict subsequent text based on context. Domain instruction question answering refers to a task type in offshore wind power and marine engineering that involves answering questions based on instructions. Mixed precision training can be a method of training the model using numerical values ​​of different precisions. Gradient accumulation training strategy refers to a strategy of accumulating gradients from multiple training iterations before updating the model parameters.

[0065] Training monitoring metrics can be indicators such as perplexity and loss value used to monitor the model training process. Optimization strategies refer to various training parameters and methods adjusted to improve the model training effect. The first model weights refer to the model weight parameters exported after the domain pre-trained model has been trained.

[0066] In this embodiment of the invention, a basic language model can be loaded, and its model parameters, word segmenter, and training framework can be initialized. After initialization, the pre-training dataset can be constructed as standardized training samples, and training batches can be divided according to text features. Furthermore, a pre-training objective that integrates autoregressive language modeling and domain instruction question answering can be set. Then, a training strategy combining mixed precision training and gradient accumulation can be used to pre-train the initialized basic language model to improve the model's language generation capability and domain instruction understanding capability.

[0067] Among them, the hybrid precision training and gradient accumulation strategy can reduce computational resource consumption and improve training efficiency while ensuring training effectiveness.

[0068] During the training process, the training strategy can be dynamically adjusted through monitoring indicators such as perplexity and loss value until the model converges. Then, the first model weights are derived, and finally, a domain pre-trained model with basic professional knowledge of offshore wind power and marine engineering is obtained.

[0069] For example, the training of a basic language model can be achieved through the following process:

[0070] S2201: Loading the base language model and building the training framework. DeepSeek-R1-0528-Qwen3-8B is selected as the pre-training base model. Its pre-training parameters, word segmenter, and model structure configuration file are loaded. A training framework adapted to this model is built, supporting multi-GPU parallel training and breakpoint resume training mechanisms, completing the model initialization and environment preparation work before pre-training.

[0071] S2202: Constructing and preparing the pre-training dataset. Using collected and cleaned text data in the field of offshore wind power and marine engineering, the data is divided into multiple training subsets based on corpus length, language type, and content structure features. Input samples with a uniform format are generated using strategies such as padding, truncation, and fragment splicing. Training batches are determined according to the divided training subsets to ensure that the model can learn the contextual logical relationships across sentences and documents, adapting to subsequent pre-training needs.

[0072] S2203: Set pre-training objectives and training strategies. Configure the training task as an autoregressive language modeling task, enabling the model to learn the ability to predict the next token based on context without providing labeled answers. Simultaneously, incorporate a small number of domain-specific micro-samples to guide the model to adapt to specific command response patterns in the offshore wind power and marine engineering domains, achieving the pre-training objective of integrating autoregressive language modeling and domain-specific command question answering. Set training hyperparameters including learning rate, batch size, number of training epochs, optimizer type, gradient accumulation strategy, and mixed precision configuration. Employ a mixed precision training and gradient accumulation strategy to ensure training stability and resource efficiency.

[0073] S2204: Perform pre-training and save model weights. During training, monitor the perplexity and loss value trends in real time, dynamically adjust the learning rate to improve model convergence speed, and dynamically adjust the optimization strategy for the pre-training process through training monitoring metrics. After training convergence, export the final model weights, configuration file, and tokenizer in HuggingFace format for subsequent supervised fine-tuning and reinforcement learning optimization stages. Based on the exported model weights, obtain the domain-specific pre-trained model.

[0074] S230. Based on the domain dialogue dataset, perform supervised fine-tuning on the domain pre-trained model to obtain a supervised fine-tuned model.

[0075] Among them, supervised fine-tuning model is a domain pre-trained model that has been adapted to domain dialogue task processing after supervised fine-tuning.

[0076] In some possible implementations, the step of performing supervised fine-tuning on the domain pre-trained model based on the domain dialogue dataset to obtain a supervised fine-tuned model includes: loading the first model weights of the domain pre-trained model and constructing a supervised fine-tuning training framework adapted to domain instruction-response data; inputting the domain dialogue dataset into the supervised fine-tuning training framework in batches, using dialogue history and user questions as model inputs and standard reference answers as model outputs; performing supervised fine-tuning on the domain pre-trained model using parameter-efficient fine-tuning techniques combined with loss function optimization and training stability assurance mechanisms; deriving the second model weights after training and obtaining the supervised fine-tuned model based on the second model weights.

[0077] Among them, the supervised fine-tuning training framework is a training and operation environment built for the command question-and-answer scenario in the field of offshore wind power and marine engineering. The domain dialogue dataset is professional question-and-answer data that has been cleaned and labeled. The standard reference answer is the standard answer text with professional norms in this field. The parameter efficient fine-tuning technology is a method to achieve efficient fine-tuning without changing all the parameters of the model. The loss function optimization and training stability guarantee mechanism are used to improve the convergence speed of the model and the stability of the training process. The second model weight is the model weight exported after the supervised fine-tuning stage is completed.

[0078] Specifically, the first model weights corresponding to the domain pre-trained model can be loaded, a supervised fine-tuning training framework adapted to domain instruction-response type data can be built, the domain dialogue dataset can be input into the training framework in batches, dialogue history and user questions can be used as model inputs, and professional standard reference answers can be used as output targets. The model can be supervisedly fine-tuned by using efficient parameter fine-tuning technology and loss function optimization and stability guarantee mechanism, so that the model can accurately understand and respond to various task instructions in the field of offshore wind power and marine engineering. After training, the second model weights are exported, thus obtaining a supervised fine-tuned model with stable domain question-answering capabilities.

[0079] For example, a supervised fine-tuning model can be obtained through the following process:

[0080] S2301: Model initialization and hyperparameter settings.

[0081] First, the pre-trained model from the previous stage is used as the initial weights, retaining its general language modeling capabilities. Based on this, domain task knowledge is injected through supervised signaling to achieve professional transfer. In the fine-tuning stage, the parameter-efficient fine-tuning method LoRA (Low-Rank Adaptation) is employed to reduce training resource consumption and improve adaptability. Key hyperparameter settings are as follows: learning rate is set to 6e-5, scheduling strategy is cosine annealing, number of training epochs is 5, single-device training batch size is set to 8, and gradient accumulation steps are 32, to achieve equivalent large-batch training and stabilize model updates.

[0082] S2302: Supervised training.

[0083] During the training phase, labeled dialogue samples related to marine engineering and offshore wind power are input into the model in batches. The model generates predicted outputs based on the context and question content. By comparing these predicted outputs with manually labeled standard answers, the cross-entropy loss is calculated, and backpropagation is performed to iteratively update the model parameters. The training objective of this phase is to enhance the model's ability to interpret command intent and accurately generate domain knowledge.

[0084] S2303: Fine-tuning effect evaluation and iterative optimization.

[0085] After each training round, the model's performance is evaluated using a reserved validation set. Evaluation metrics are set according to the task type; for example, accuracy and Top-k hit rate are used for question-answering, while semantic similarity metrics such as ROUGE, BLEU, and BERTScore are used for generation. For some key tasks, manual evaluation by domain experts can also assist in validation. Based on the evaluation results, the model's performance in different sub-tasks is analyzed. Combining the loss trend and generalization ability, the learning rate or number of training rounds is dynamically adjusted to achieve optimal convergence.

[0086] S240. Construct multi-level answer quality labeling rules based on the reward model dataset, and perform supervised fine-tuning on the basic reward model based on the multi-level answer quality labeling rules to obtain the reward fine-tuning model.

[0087] Among them, the multi-level answer quality labeling rule refers to the labeling criteria for dividing answers into different quality levels according to domain professional standards, and the basic reward model can be an initial model used to evaluate the quality of domain answers.

[0088] In some possible implementations, the step of constructing multi-level answer quality labeling rules based on the reward model dataset and performing supervised fine-tuning on the basic reward model based on the multi-level answer quality labeling rules to obtain a reward fine-tuning model includes: performing quality level labeling on answers in the reward model dataset based on the domain professional evaluation dimension to construct the multi-level answer quality labeling rules; loading the basic reward model, initializing the model parameters and word segmenter of the basic reward model, and adjusting the model output layer to adapt to the multi-level answer quality labeling rules; inputting the labeled reward model dataset into the adjusted basic reward model, and performing supervised fine-tuning with the loss function as the optimization objective; monitoring training metrics and dynamically adjusting the training strategy through the validation set, deriving the third model weights after training convergence, and obtaining the reward fine-tuning model based on the third model weights.

[0089] Among them, the domain professional evaluation dimension can be the professionalism, accuracy, logic, etc. of the evaluation answer quality in the field of offshore wind power and marine engineering; the quality level labeling refers to the quality labeling of the answer according to the preset level standard; the model output layer refers to the network layer in the model used to output the prediction results; the validation set can be a dataset used to monitor the model training process and evaluate the model training effect; and the third model weight refers to the model weight parameters derived after the reward fine-tuning model training converges.

[0090] In this embodiment of the invention, the answers in the reward model dataset can be labeled with quality levels based on domain-specific evaluation dimensions, thereby constructing multi-level answer quality labeling rules and providing a quality evaluation basis for reward model training.

[0091] Furthermore, the basic reward model is loaded, and its model parameters and word segmenter are initialized. The model output layer is adjusted to adapt to multi-level answer quality labeling rules, enabling the model to output evaluation results that conform to the quality level classification. Then, the labeled reward model dataset is input into the adjusted basic reward model, and supervised fine-tuning is performed with the loss function as the optimization objective, allowing the model to learn accurate domain answer quality evaluation capabilities.

[0092] Furthermore, training metrics can be monitored and training strategies dynamically adjusted using a validation set. Once training converges, the weights of the third model can be derived to obtain the reward-fine-tuned model. This model can accurately evaluate the quality of domain-specific answers and provide reward signals for reinforcement learning.

[0093] S250. Using the reward fine-tuning model as a quality assessment model, the generalized reward strategy optimization algorithm is used to perform reinforcement learning optimization training on the supervised fine-tuning model to obtain the initial large model.

[0094] Among them, the Generalized Reward Policy Optimization (GRPO) algorithm is a reinforcement learning algorithm that optimizes the model policy based on reward signals. The initial large-scale model can be an initial large-scale model that has been trained by reinforcement learning after supervised fine-tuning but has not undergone final performance evaluation. Multi-dimensional performance evaluation refers to a comprehensive evaluation of the performance of the initial large-scale model from multiple professional dimensions.

[0095] Specifically, the reward-based fine-tuning model can be used as a quality assessment model. A generalized reward strategy optimization algorithm can be used to perform reinforcement learning optimization training on the supervised fine-tuning model to further improve the professionalism and stability of the generated content and obtain the initial large-scale model.

[0096] In some possible implementations, the step of using the reward-fine-tuned model as a quality assessment model and performing reinforcement learning optimization training on the supervised fine-tuned model using a generalized reward policy optimization algorithm to obtain an initial large-scale model includes: loading the second model weights of the supervised fine-tuned model, using the second model weights as the initial policy network for reinforcement learning, and constructing a reinforcement learning dialogue prompt dataset; loading the third model weights of the reward-fine-tuned model, constructing a custom reward function compatible with multi-dimensional quality scoring, and using it to perform real-time quality assessment on the content generated by the policy network; configuring the training hyperparameters of the generalized reward policy optimization algorithm, enabling the initial policy network to generate candidate responses based on the dialogue prompt dataset, performing quality scoring on the candidate responses through the custom reward function and outputting reward values; calculating the advantage value by combining the advantage estimation method, constructing a loss function containing a truncation mechanism and a policy entropy adjustment term, and performing parameter updates and iterative training on the initial policy network based on the reward value and the advantage value; introducing an early stopping mechanism to monitor the training effect, dynamically saving the fourth model weights with the best training performance, and obtaining the initial large-scale model based on the fourth model weights.

[0097] The initial policy network can be a model network used as the initial policy in reinforcement learning. The reinforcement learning dialogue prompt dataset refers to a dataset that contains only user instructions or dialogue history and is used for reinforcement learning training. The custom reward function can be a reward function constructed according to domain requirements that can evaluate the quality of the content generated by the model from multiple dimensions.

[0098] Training hyperparameters refer to various parameters used to control the reinforcement learning training process; candidate responses refer to multiple responses to be evaluated generated by the initial policy network based on the dialogue prompt dataset; and reward value refers to the score output by the custom reward function after scoring the quality of the candidate responses.

[0099] Advantage estimation methods refer to methods used to calculate the advantage value of model behavior. The advantage value can be understood as the degree of improvement of the model's current behavior compared to the average behavior. The truncation mechanism refers to the mechanism used to limit the variation of model policy parameters. The policy entropy adjustment term refers to the adjustment term used to encourage the model policy to maintain a certain degree of randomness.

[0100] The loss function refers to the function used to calculate the error of the model strategy optimization. Parameter update and iterative training refer to the process of updating the model parameters and training repeatedly based on the calculation results of the loss function. The early stopping mechanism can be a mechanism to stop training when the model training effect no longer improves significantly. The fourth model weight refers to the model weight parameters that perform best during the initial training of the large model.

[0101] Specifically, the second model weights of the supervised fine-tuning model can be loaded and used as the initial policy network for reinforcement learning. At the same time, a reinforcement learning dialogue prompt dataset can be constructed to provide a policy foundation and data support for reinforcement learning training.

[0102] Furthermore, the third model weights of the reward fine-tuning model are loaded to construct a custom reward function compatible with multi-dimensional quality scoring. This function can be used to perform real-time quality evaluation on the content generated by the policy network. Then, the training hyperparameters of the generalized reward policy optimization algorithm are configured so that the initial policy network generates candidate responses based on the dialogue prompt dataset. The custom reward function is then used to perform quality scoring on the candidate responses and output the reward value.

[0103] Then, the advantage value can be calculated by combining the advantage estimation method, and a loss function with truncation mechanism and policy entropy adjustment term can be constructed. Based on the reward value and advantage value, the parameters of the initial policy network are updated and iteratively trained. The truncation mechanism ensures the stability of training, and the policy entropy adjustment term avoids the model from over-converging. Finally, an early stopping mechanism is introduced to monitor the training effect and dynamically save the weights of the fourth model with the best training performance. Based on the weights of the fourth model, the initial large model is obtained. The early stopping mechanism can effectively avoid model overfitting and ensure the generalization ability of the initial large model.

[0104] For example, Figure 4 The flowchart of the GRPO reinforcement learning algorithm provided in this embodiment of the invention shows how the model can be further optimized using the GRPO reinforcement learning algorithm:

[0105] This step aims to address the problem that supervised learning models only learn fixed mapping relationships and cannot cope with diverse and reasonable answers and dynamic context changes. The adopted Generalized Reward Policy Optimization (GRPO) algorithm integrates intra-group relative comparison mechanisms and policy entropy adjustment terms, enabling stable and effective reinforcement optimization in open-ended generation scenarios without reference answers. Specific optimization steps include:

[0106] S2501: Initialize the base model and reinforcement learning dataset.

[0107] A supervised, fine-tuned language model is used as the initial policy network to retain its language understanding and command adaptation capabilities in offshore wind power and marine engineering tasks. An interactive dataset required for reinforcement learning is constructed, selecting samples covering different task types and language scenarios from benchmark sets, including typical commands such as professional question answering, operation and maintenance suggestions, and equipment reasoning, as the input set for the state. In reinforcement learning, the state represents the input scenario currently faced by the model. Diverse state samples ensure the diversity and representativeness of samples during training, enhancing the policy generalization ability.

[0108] S2502: Configure the reward model and build the reward function.

[0109] The reward fine-tuning model (QRM-Llama3.1-8B-v2) fine-tuned using the multi-quality-level answer scoring system constructed in the previous steps is used to score the responses generated by the supervised fine-tuning model during reinforcement training.

[0110] The scoring dimensions include logical rigor, professional accuracy, and completeness of language expression, and the final output is a comprehensive reward value r. t The reward value is the overall quality score corresponding to a single generated response, and it is also the core signal guiding model optimization in reinforcement learning. To improve the stability of the score and its sensitivity to policy changes, the reward function adopts a batch invocation method and supports high-dimensional semantic alignment calculation to ensure the fairness and anti-interference ability of the evaluation.

[0111] S2503: Strategy Sampling and Trajectory Collection.

[0112] For each interaction sample, its corresponding state is randomly sampled based on the current strategy, which is represented by the current network to be optimized and its corresponding network parameters. The sampling temperature is set to Temperature=0.8, a parameter used to control the randomness of the generated content, thereby balancing the diversity and reasonableness of the generated content.

[0113] Each input state sample generates 5-10 candidate responses, forming multiple trajectories. Trajectories are used to record continuous sequences of states, actions, and rewards in reinforcement learning, with actions representing candidate responses generated by the model in the corresponding states.

[0114] Generalized advantage estimation (GAE) is introduced to estimate the advantage value at each time step. GAE is a method for accurately calculating the advantage value, which represents the degree of improvement of the current action relative to the average policy and is also the core incentive basis for policy optimization. The core formula is:

[0115]

[0116] The temporal difference error is calculated by combining the single-step reward, the discount factor, and the state value. γ is the discount factor, and λ is the GAE parameter. The discount factor is used to adjust the weight of future rewards, and the GAE parameter is used to balance the bias and variance in the estimation process. The advantage function reflects the relative improvement of the current behavior relative to the average policy and is a key incentive signal for reinforcement learning optimization.

[0117] S2504: Policy Optimization and Parameter Update. Construct the loss function corresponding to the generalized reward policy optimization algorithm, update the policy network, and use the loss function to calculate the error in model parameter optimization. Its form is:

[0118]

[0119] in, This represents the probability ratio of the old and new policies on the current trajectory. The probability ratio is calculated from the output probabilities of the old and new policies on the same trajectory. The old policy network is used for relative comparison with the current policy. clip(...) is the truncation mechanism (ε=0.2). The truncation mechanism is used to limit the magnitude of policy parameter changes. The truncation threshold is set to 0.2 to avoid the policy update magnitude being too large, which would lead to training instability.

[0120] (θ) represents the policy entropy, and β is the entropy coefficient. Policy entropy measures the randomness of the policy, while the entropy coefficient adjusts the weight of randomness. Together, they encourage the policy to maintain a certain level of exploration ability and prevent the model from over-converging. Over-convergence can cause the model to generate too much homogeneous content, reducing generalization performance.

[0121] The optimizer employs the AdamW optimizer, which adds weight decay to the traditional adaptive gradient descent method, resulting in higher training stability. The initial learning rate is set to 3e-5, and gradient descent is used to update the parameters. The old policy network θ_old is synchronized to the current policy θ every 10 policy updates to ensure consistency in the relative evaluation basis. The value network is initialized using intermediate layer vectors from the supervised fine-tuning model and is used to predict the long-term expected value corresponding to the state, further improving the accuracy of advantage value estimation.

[0122] S2505: Iterative optimization and convergence verification.

[0123] The reinforcement training cycle consists of trajectory sampling and policy optimization, with a total of 3-5 iterations. After each training cycle, trajectory sampling is performed again based on the current policy to avoid relying on samples generated by the old policy during optimization, thus improving learning effectiveness and diversity. The model is considered converged when the average reward improvement over two consecutive training cycles is less than 5%, and the model achieves stability or an upward trend in the test set task metrics. Convergence indicates that the model performance no longer improves significantly and the parameters tend to stabilize. Finally, the network weights of the current policy are saved as the final model after adjusting the generalized reward policy optimization algorithm. A horizontal performance comparison is performed with the model from the supervised fine-tuning stage, and the version that performs best in real-world tasks is selected for deployment and evaluation.

[0124] S260. Perform a multi-dimensional performance evaluation on the initial training model based on the benchmark dataset. After the evaluation is passed, the large model in the field of offshore wind power and marine engineering is obtained.

[0125] Specifically, a multi-dimensional performance evaluation can be performed on the initially trained large model based on a benchmark dataset. Upon successful evaluation, a large model suitable for offshore wind power and marine engineering applications is obtained. The hierarchical training method of this invention allows the model to gradually acquire domain knowledge and task processing capabilities, ensuring the domain adaptability and professional performance of the large model.

[0126] In some embodiments, the step of performing a multi-dimensional performance evaluation on the initial training model based on the benchmark dataset, and obtaining the large model in the field of offshore wind power and marine engineering after passing the evaluation, includes: dividing the benchmark dataset into test subsets according to scenario complexity; configuring automated evaluation tools and professional evaluation indicators to perform task-based quantitative evaluation on the initial training model; analyzing the performance of the initial training model based on the evaluation results; performing iterative parameter optimization on modules whose performance does not meet the standards; re-executing the evaluation after optimization; obtaining the fifth model weight after passing the evaluation; and obtaining the large model in the field of offshore wind power and marine engineering based on the fifth model weight.

[0127] The scenario complexity can refer to different levels of complexity, such as simple, complex, or marginal, in the context of offshore wind power and marine engineering tasks. The test subset refers to a subset of benchmark test data divided according to scenario complexity. For example... Figure 5 The diagram shows a classification of the benchmark test set provided in this embodiment of the invention. Generation tasks (such as T10 and T3) focus on output quality and task adaptability, evaluating format compliance and technical feasibility; classification and recognition tasks (such as T6 and T9) emphasize classification accuracy and fault tolerance, evaluating recognition accuracy and the completeness of complex fault diagnosis; reasoning and analysis tasks (such as T1 and T2) focus on logical correctness and knowledge application ability, evaluating reasoning matching degree and knowledge depth; simultaneously, multilingual capability assessments are conducted for all tasks to verify cross-language consistency and the accuracy of terminology translation.

[0128] Automated evaluation tools can be used to quantitatively evaluate model performance. Professional evaluation indicators refer to model performance evaluation indicators that meet the requirements of the offshore wind power and marine engineering fields. Task-based quantitative evaluation refers to quantitative performance evaluation of the model for different domain tasks. Modules that fail to meet performance standards refer to functional modules in the model that fail to meet the preset standards in performance evaluation. Parameter iterative optimization refers to the process of repeatedly adjusting and optimizing model parameters. The fifth model weight refers to the model weight parameters obtained after the initial large model has been optimized and evaluated.

[0129] Specifically, the benchmark dataset is divided into test subsets based on scenario complexity. Automated evaluation tools and professional evaluation metrics are configured to perform task-based quantitative evaluations of the initially trained large model, achieving comprehensive performance testing across all scenarios and multiple tasks. Furthermore, based on the evaluation results, the performance of the initially trained large model is analyzed, and iterative parameter optimization is performed on modules that fail to meet performance standards, specifically improving the performance of the model's weak points. After optimization, the evaluation can be re-executed. Upon successful evaluation, a fifth model weight is obtained. Based on this fifth model weight, a large model for offshore wind power and marine engineering is obtained. In this embodiment, targeted iterative parameter optimization effectively improves the overall performance of the model, ensuring that the large model meets the requirements of practical applications in the field.

[0130] Compared with the prior art, the method provided by this invention has the following advantages:

[0131] By integrating domain-specific pre-training, supervised fine-tuning, and reward-based reinforcement learning strategies, a systematic and transferable large-scale model training process was constructed, significantly improving the language model's professional understanding and knowledge generation capabilities in complex vertical fields such as offshore wind power and marine engineering. This method fully utilizes high-quality domain corpora and diverse dialogue scenarios, combined with a structured reward feedback mechanism to guide the model to generate more accurate, reliable, and professionally in-depth responses, overcoming the limitations of existing general-purpose large-scale models such as insufficient mastery of professional terminology, poor adaptability to cross-disciplinary tasks, and incomplete reasoning chains. The training process incorporates low-resource, high-efficiency fine-tuning techniques and reinforcement learning strategies, ensuring model accuracy while significantly reducing computational overhead, demonstrating good engineering practicality and scalability. Furthermore, systematic evaluation using multilingual and multi-task benchmark sets validated the effectiveness and robustness of the proposed training method. The model exhibits superior performance compared to existing technologies in professional tasks, language generalization, and complex question-answering reasoning, demonstrating its potential and value for deployment in practical scenarios such as marine engineering, wind power operation and maintenance, and structural design assistance.

[0132] Figure 6 This is a schematic diagram of the structure of a large-scale model application device for offshore wind power and marine engineering provided in an embodiment of the present invention. Figure 6 As shown, the device includes:

[0133] The requirement data acquisition module 610 is used to acquire task requirement data in the field of offshore wind power and marine engineering.

[0134] The large model processing module 620 is used to input the task requirement data into a large model in the field of offshore wind power and marine engineering for processing, and to obtain professional processing results corresponding to the task requirement data.

[0135] The large model for offshore wind power and marine engineering is a domain-specific large model trained by integrating professional domain pre-training, supervised fine-tuning, and reinforcement learning algorithms based on reward models.

[0136] The technical solution of this invention acquires task requirement data in the field of offshore wind power and marine engineering, inputs this data into a large-scale model for processing, and obtains professional processing results corresponding to the task requirement data. This large-scale model is a domain-specific model trained by integrating professional domain pre-training, supervised fine-tuning, and a reward-based reinforcement learning algorithm. This technical solution constructs a domain-specific large-scale model through multi-stage targeted training. First, professional domain pre-training injects core domain knowledge into the basic model, addressing the lack of domain-specific accumulation in general models. Then, supervised fine-tuning adapts the model to actual task scenarios in the domain, compensating for the insufficient task adaptability of general models. Finally, a reward-based reinforcement learning algorithm optimizes the model's generation capabilities, addressing the problem that supervised learning models only learn fixed mapping relationships and cannot handle diverse and reasonable responses and dynamic context changes. Ultimately, this improves the professionalism and accuracy of domain task processing, reduces task processing costs, and enhances model generalization ability, providing efficient and reliable technical support for intelligent practices in the field of offshore wind power and marine engineering.

[0137] In some possible implementations, the device further includes a large model training module for training a large model in the field of offshore wind power and marine engineering.

[0138] The large model training module includes:

[0139] The target dataset construction submodule is used to collect raw text data in the field of offshore wind power and marine engineering. Multi-layer cleaning and filtering processes are performed on the raw text data to obtain the target dataset, which includes a pre-training dataset, a domain dialogue dataset, a reward model dataset, and a benchmark dataset.

[0140] The domain pre-training submodule is used to perform professional domain pre-training on the basic language model based on the pre-training dataset to obtain a domain pre-trained model.

[0141] The supervised fine-tuning submodule is used to perform supervised fine-tuning on the domain pre-trained model based on the domain dialogue dataset to obtain a supervised fine-tuned model.

[0142] The reward model training submodule is used to construct multi-level answer quality labeling rules based on the reward model dataset, and to perform supervised fine-tuning on the basic reward model based on the multi-level answer quality labeling rules to obtain the reward fine-tuning model;

[0143] The reinforcement learning optimization submodule is used to use the reward fine-tuning model as a quality evaluation model, and to perform reinforcement learning optimization training on the supervised fine-tuning model using a generalized reward policy optimization algorithm to obtain the initial large model.

[0144] The model evaluation submodule is used to perform multi-dimensional performance evaluation on the initial training large model based on the benchmark dataset. After passing the evaluation, the large model in the field of offshore wind power and marine engineering is obtained.

[0145] In some possible implementations, the target dataset construction submodule includes:

[0146] The preprocessing unit is used to perform basic format normalization and text feature quantization on the collected raw text data to obtain preprocessed text data;

[0147] The filtering unit is used to sequentially perform rule filtering and model quality scoring on the preprocessed text data, and obtain filtered text data after removing low-quality data.

[0148] The deduplication and quality inspection unit is used to perform deduplication and quality inspection on the screened text data to obtain qualified text data.

[0149] The data partitioning unit is used to partition the qualified text data according to its purpose to obtain the training dataset, the pre-training dataset, the domain dialogue dataset, the reward model dataset, and the benchmark dataset, and then merge them to obtain the target dataset.

[0150] In some possible implementations, the domain pre-training submodule includes:

[0151] The model initialization unit is used to load the basic language model and initialize the model parameters, word segmenter, and training framework of the basic language model.

[0152] The sample construction unit is used to construct the pre-training dataset into standardized training samples and divide the training batches according to text features.

[0153] The pre-training execution unit is used to set the pre-training objective that integrates autoregressive language modeling and domain instruction question answering, and to perform pre-training on the initialized basic language model using a training strategy that combines mixed precision training and gradient accumulation.

[0154] The weight derivation unit is used to dynamically adjust and optimize the pre-training process through training monitoring indicators, derive the first model weights after training convergence, and obtain the domain pre-trained model based on the first model weights.

[0155] In some possible implementations, the supervised fine-tuning submodule includes:

[0156] The weight loading unit is used to load the first model weights of the domain pre-trained model and construct a supervised fine-tuning training framework adapted to domain instruction-response data.

[0157] The data input unit is used to input the domain dialogue dataset into the supervised fine-tuning training framework in batches, using dialogue history and user questions as model inputs and standard reference answers as model outputs.

[0158] The fine-tuning execution unit is used to perform supervised fine-tuning of the domain pre-trained model using efficient parameter fine-tuning techniques, combined with loss function optimization and training stability guarantee mechanisms.

[0159] The fine-tuning weight derivation unit is used to derive the second model weights after training, and to obtain the supervised fine-tuning model based on the second model weights.

[0160] In some possible implementations, the reward model training submodule includes:

[0161] The annotation rule construction unit is used to perform quality level annotation on the answers in the reward model dataset based on the domain professional evaluation dimension, and to construct the multi-level answer quality annotation rules.

[0162] The reward model initialization unit is used to load the basic reward model, initialize the model parameters and word segmenter of the basic reward model, and adjust the model output layer to adapt to the multi-level answer quality labeling rules.

[0163] The reward model fine-tuning unit is used to input the labeled reward model dataset into the adjusted base reward model and perform supervised fine-tuning with the loss function as the optimization objective.

[0164] The reward weight derivation unit is used to monitor training metrics and dynamically adjust the training strategy through the validation set. After training converges, it derives the third model weights and obtains the reward fine-tuning model based on the third model weights.

[0165] In some possible implementations, the reinforcement learning optimization submodule includes:

[0166] The policy network initialization unit is used to load the second model weights of the supervised fine-tuning model, and use the second model weights as the initial policy network for reinforcement learning to construct a reinforcement learning dialogue prompt dataset.

[0167] The reward function construction unit is used to load the third model weights of the reward fine-tuning model, construct a custom reward function compatible with multi-dimensional quality scoring, and perform real-time quality evaluation on the content generated by the policy network.

[0168] The candidate response generation unit is used to configure the training hyperparameters of the generalized reward policy optimization algorithm, so that the initial policy network generates candidate responses based on the dialogue prompt dataset, performs quality scoring on the candidate responses through the custom reward function, and outputs reward values;

[0169] The parameter update unit is used to calculate the advantage value by combining the advantage estimation method, construct a loss function with a truncation mechanism and a policy entropy adjustment term, and perform parameter update and iterative training on the initial policy network based on the reward value and the advantage value.

[0170] The initial training model generation unit is used to introduce an early stopping mechanism to monitor the training effect, dynamically save the fourth model weights with the best training performance, and obtain the initial training large model based on the fourth model weights.

[0171] In some possible implementations, the model evaluation submodule includes:

[0172] The quantitative evaluation unit is used to divide the benchmark test dataset into test subsets according to scenario complexity, and configure automated evaluation tools and professional evaluation indicators to perform sub-task quantitative evaluation on the initial training large model.

[0173] The parameter optimization unit is used to analyze the performance of the initial training model based on the evaluation results and perform iterative parameter optimization for modules whose performance does not meet the standards.

[0174] The final model generation unit is used to re-perform the evaluation after optimization. After the evaluation is passed, the fifth model weight is obtained, and the large model of the offshore wind power and marine engineering field is obtained based on the fifth model weight.

[0175] The offshore wind power and marine engineering large-scale model application device provided in the embodiments of the present invention can execute the offshore wind power and marine engineering large-scale model application method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0176] Figure 7 This is a schematic diagram of the structure of an electronic device for implementing the large-scale offshore wind power and marine engineering application method according to embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0177] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0178] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0179] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as the large-scale application methods for offshore wind power and marine engineering.

[0180] In some embodiments, the offshore wind power and marine engineering large-scale model application method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the offshore wind power and marine engineering large-scale model application method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to execute the offshore wind power and marine engineering large-scale model application method by any other suitable means (e.g., by means of firmware).

[0181] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0182] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0183] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0184] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0185] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0186] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0187] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0188] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for applying large models of offshore wind power and ocean engineering, characterized in that, include: Acquire task requirement data in the field of offshore wind power and marine engineering; The task requirement data is input into a large model in the field of offshore wind power and marine engineering for processing to obtain professional processing results corresponding to the task requirement data. Among them, the large model in the field of offshore wind power and marine engineering is a domain-specific large model obtained by integrating professional domain pre-training, supervised fine-tuning and reinforcement learning algorithm based on reward model training; The training process of the large-scale model in the field of offshore wind power and marine engineering includes: Raw text data in the field of offshore wind power and marine engineering is collected, and multi-layer cleaning and filtering processes are performed on the raw text data to obtain the target dataset; wherein, the target dataset includes a training dataset, a pre-training dataset, a domain dialogue dataset, a reward model dataset, and a benchmark dataset; Based on the pre-trained dataset, the basic language model is pre-trained in a professional domain to obtain a domain-pre-trained model. Supervised fine-tuning is performed on the domain pre-trained model based on the domain dialogue dataset to obtain a supervised fine-tuned model. Based on the reward model dataset, a multi-level answer quality labeling rule is constructed, and a supervised fine-tuning is performed on the basic reward model based on the multi-level answer quality labeling rule to obtain the reward fine-tuning model; The reward fine-tuning model is used as a quality assessment model. The generalized reward strategy optimization algorithm is used to perform reinforcement learning optimization training on the supervised fine-tuning model to obtain the initial large model. Based on the benchmark dataset, a multi-dimensional performance evaluation is performed on the initial large model. After passing the evaluation, the large model in the field of offshore wind power and marine engineering is obtained. The step of performing reinforcement learning optimization training on the supervised fine-tuning model using a generalized reward strategy optimization algorithm includes: The reward model is configured and a reward function is constructed; the reward fine-tuning model is used to score the quality of the reply generated by the supervised fine-tuning model in reinforcement training; the scoring dimensions include logical rigor, professional accuracy and language expression integrity, and a comprehensive reward value r is output t The reward function adopts a batch calling mode and supports high-dimensional semantic alignment calculation; For each interaction sample corresponding state, based on the current strategy for random sampling, set the sampling temperature Temperature=0.8; each input state sampling generates 5~10 candidate replies, forms multiple trajectories , the trajectory is used to record the continuous sequence of state, action and reward in reinforcement learning The generalized advantage estimation is introduced to estimate the advantage value at each time point. The core formula of the generalized advantage estimation is: wherein the timing difference error γ is a discount factor, and λ is a GAE parameter. A loss function corresponding to the generalized reward policy optimization algorithm is constructed, and the policy network is updated accordingly; the form of the loss function is as follows: ,in This represents the probability ratio of the old and new strategies on the current trajectory, where clip(...) is the cutoff mechanism with ε=0.2, and β is the entropy coefficient. For policy entropy; In the generalized reward policy optimization algorithm, the optimizer adopts the AdamW optimizer, the initial learning rate is set to 3e-5, and the gradient descent strategy is used to update the parameters; the old policy network θ_old is synchronized to the current policy θ every 10 policy updates; the value network is initialized by the intermediate layer vectors of the supervised fine-tuning model. The total number of iterations for the generalized reward strategy optimization algorithm is set to 3 to 5 rounds; after each round of training, trajectory sampling is performed again based on the current strategy; when the average reward improvement of two consecutive rounds of training is less than 5%, the model is determined to have converged.

2. The method according to claim 1, characterized in that, The process involves collecting raw text data in the field of offshore wind power and marine engineering, performing multi-layer cleaning and filtering on the raw text data to obtain the target dataset, which includes: The collected raw text data is subjected to basic format normalization and text feature quantization to obtain preprocessed text data; The preprocessed text data is subjected to rule filtering and model quality scoring in sequence, and low-quality data is removed to obtain the filtered text data; The filtered text data is processed by an algorithm for deduplication and quality control to obtain qualified text data; The qualified text data is divided according to its purpose to obtain the training dataset, the pre-training dataset, the domain dialogue dataset, the reward model dataset, and the benchmark dataset, which are then merged to obtain the target dataset.

3. The method according to claim 1, characterized in that, The process of performing domain-specific pre-training on the basic language model based on the pre-trained dataset to obtain a domain-specific pre-trained model includes: Load the basic language model and initialize the model parameters, word segmenter, and training framework of the basic language model; The pre-training dataset is constructed into standardized training samples, and training batches are divided according to text features; Set a pre-training objective that integrates autoregressive language modeling and domain instruction question answering, and use a training strategy of mixed precision training and gradient accumulation to perform pre-training on the initialized basic language model; The pre-training process is dynamically adjusted and optimized by training monitoring indicators. After training convergence, the first model weights are derived, and the domain pre-trained model is obtained based on the first model weights.

4. The method according to claim 3, characterized in that, The process of performing supervised fine-tuning on the domain pre-trained model based on the domain dialogue dataset to obtain a supervised fine-tuned model includes: Load the first model weights of the domain pre-trained model to construct a supervised fine-tuning training framework adapted to domain instruction-response data; The domain dialogue dataset is input into the supervised fine-tuning training framework in batches, with dialogue history and user questions as model inputs and standard reference answers as model outputs. A parameter-efficient fine-tuning technique is employed, combined with loss function optimization and training stability assurance mechanisms, to perform supervised fine-tuning on the pre-trained model in the domain. After training is complete, the weights of the second model are exported, and the supervised fine-tuning model is obtained based on the weights of the second model.

5. The method according to claim 4, characterized in that, The step of constructing multi-level answer quality labeling rules based on the reward model dataset, and performing supervised fine-tuning on the basic reward model based on the multi-level answer quality labeling rules to obtain a reward fine-tuning model includes: Based on the domain-specific evaluation dimension, the answers in the reward model dataset are labeled with quality levels to construct the multi-level answer quality labeling rules. Load the basic reward model, initialize the model parameters and word segmenter of the basic reward model, and adjust the model output layer to adapt to the multi-level answer quality labeling rules; The labeled reward model dataset is input into the adjusted base reward model, and supervised fine-tuning is performed with the loss function as the optimization objective. By monitoring training metrics and dynamically adjusting the training strategy using a validation set, the weights of the third model are derived after training convergence, and the reward fine-tuning model is obtained based on the weights of the third model.

6. The method according to claim 5, characterized in that, The step involves using the reward-fine-tuned model as a quality assessment model, and employing a generalized reward policy optimization algorithm to perform reinforcement learning optimization training on the supervised fine-tuned model to obtain an initial large-scale model, including: Load the second model weights of the supervised fine-tuning model, use the second model weights as the initial policy network for reinforcement learning, and construct a reinforcement learning dialogue prompt dataset; Load the third model weights of the reward fine-tuning model, construct a custom reward function compatible with multi-dimensional quality scoring, and use it to perform real-time quality assessment on the content generated by the policy network; Configure the training hyperparameters of the generalized reward policy optimization algorithm so that the initial policy network generates candidate responses based on the dialogue prompt dataset, and performs quality scoring on the candidate responses and outputs reward values ​​through the custom reward function; The advantage value is calculated by combining the advantage estimation method, and a loss function with truncation mechanism and policy entropy adjustment term is constructed. Based on the reward value and the advantage value, the parameters of the initial policy network are updated and iteratively trained. An early stopping mechanism is introduced to monitor the training effect, and the weights of the fourth model with the best training performance are dynamically saved. The initial training large model is obtained based on the weights of the fourth model.

7. The method according to claim 1, characterized in that, The initial large model is subjected to a multi-dimensional performance evaluation based on the benchmark dataset. Upon passing the evaluation, the large model for the offshore wind power and marine engineering domain is obtained, including: The benchmark dataset is divided into test subsets according to scenario complexity, and automated evaluation tools and professional evaluation indicators are configured to perform task-based quantitative evaluation of the initial training model. Based on the evaluation results, the performance of the initial training model was analyzed, and the parameters of modules that did not meet the performance standards were iteratively optimized. After optimization, the evaluation is re-executed. If the evaluation passes, the weights of the fifth model are obtained. Based on the weights of the fifth model, the large model for the field of offshore wind power and marine engineering is obtained.

8. A large-scale model application device for offshore wind power and marine engineering, characterized in that, The apparatus for implementing the large-scale offshore wind power and marine engineering application method according to any one of claims 1-7 comprises: The requirement data acquisition module is used to acquire task requirement data in the field of offshore wind power and marine engineering. The large model processing module is used to input the task requirement data into a large model in the field of offshore wind power and marine engineering for processing, and to obtain professional processing results corresponding to the task requirement data. The large model for offshore wind power and marine engineering is a domain-specific large model trained by integrating professional domain pre-training, supervised fine-tuning, and reinforcement learning algorithms based on reward models.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to perform the offshore wind power and marine engineering large model application method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the application method of large-scale offshore wind power and marine engineering as described in any one of claims 1-7.