A Multi-Source Environmental Factor Suitability Assessment Method for Bluefin Tuna Aquaculture Based on LLM

CN120975285BActive Publication Date: 2026-09-01CHINESE ACAD OF FISHERY SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510952998.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2026-09-01
Estimated Expiration
2045-07-10

AI Technical Summary

Technical Problem

现阶段,相关研究在建模方法上仍存在较大空白,尤其是在如何有效捕捉环境数据中的时间依赖性、如何提升模型对数据噪声与异常的鲁棒性、以及如何实现适宜性趋势的高精度动态预测等方面,仍未得到有效解决

Benefits of technology

[0029] 1. The method of this invention is based on the large-scale artificial intelligence model LLaMA and combined with multi-source heterogeneous data of marine aquaculture environment. It can effectively capture the temporal change characteristics of environmental factors and realize intelligent assessment of the suitability level of bluefin tuna aquaculture environment. It establishes a closed-loop process from data collection, feature extraction, model training to prediction application, and has strong generalization ability and environmental adaptability. It breaks through the limitation of traditional methods that are difficult to handle dynamic environmental information and improves assessment efficiency and judgment accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120975285B_ABST
    Figure CN120975285B_ABST
Patent Text Reader

Abstract

This paper presents a multi-source environmental factor-based method for assessing the environmental suitability of bluefin tuna aquaculture, relating to the field of intelligent aquaculture technology in marine fisheries. The method involves collecting multi-source environmental factor data and labeling the aquaculture suitability under different environmental conditions into five levels, which serve as the raw data to construct a dataset. The raw data is preprocessed at a uniform scale, dividing it into training and test sets. A multi-class environmental suitability assessment model is constructed based on the LLaMA model, employing a LoRA fine-tuning strategy. The training set is input into the assessment model, and supervised learning is used for fine-tuning training. The performance of the fine-tuned assessment model is validated using the test set, evaluating its effectiveness. The fine-tuned and validated assessment model is then integrated into a management system. Based on the large-scale artificial intelligence model LLaMA and combined with multi-source environmental factor data, this method enables intelligent assessment of the environmental suitability levels for bluefin tuna aquaculture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent aquaculture technology in marine fisheries, specifically a method for assessing the environmental suitability of bluefin tuna aquaculture based on LLM (Multi-Source Environmental Factors). Background Technology

[0002] With the increasing demand from Chinese residents for high-quality aquatic products, especially high-value marine fish, promoting the transformation and upgrading of fisheries and achieving high-quality development has become the core direction of marine fisheries development. In recent years, with the increasing depletion of marine fishery resources and the continuous growth of global aquatic product market demand, artificial aquaculture has become an important path to achieve sustainable development of marine fisheries. As a large, high-value, pelagic migratory fish, the bluefin tuna, with its excellent meat quality, extremely high nutritional value, and widespread market recognition, has become an important target for many countries to develop deep-sea aquaculture. However, this species is extremely sensitive to its living environment and highly dependent on changes in the ecological factors of the aquaculture area. Its physiological processes, such as growth, feeding, immunity, and behavior, are all significantly affected by environmental factors. Therefore, scientifically assessing the suitability of the bluefin tuna aquaculture environment, rationally selecting aquaculture areas, and dynamically adjusting aquaculture strategies are key to ensuring the success rate of aquaculture, improving industry efficiency, and reducing risks.

[0003] Bluefin tuna exhibit distinct threshold characteristics in their environmental adaptability. Multiple environmental factors, including water temperature, dissolved oxygen, ammonia nitrogen concentration, and pH, interact to form a complex ecological response mechanism. Water temperature is a core factor influencing their feeding activity and metabolic rate, with an optimal range typically between 15°C and 24°C. Fluctuations in dissolved oxygen concentration directly determine the fish's respiratory efficiency and stress response. High concentrations of ammonia nitrogen and unstable pH conditions can trigger poisoning or immunosuppression. Furthermore, nonlinear coupling and spatiotemporal interactions exist between different factors; for example, fish tolerance significantly decreases under conditions of combined high temperature and low oxygen. Therefore, assessing the suitability of bluefin tuna farming environments requires not only considering static threshold determinations for single factors but also establishing a system model that integrates multiple environmental factors and incorporates multi-scale dynamic response mechanisms to comprehensively and accurately reflect the true state of the farming environment and its potential impact on fish physiology.

[0004] Traditional suitability assessment methods are mostly based on expert experience, the Analytic Hierarchy Process (AHP), or fuzzy comprehensive evaluation methods. While these methods have some practical value, they still have significant limitations. On the one hand, the determination of factor weights relies on subjective judgment and lacks adaptability. On the other hand, their modeling processes are mostly static analyses, making it difficult to characterize the dynamic changes of environmental factors over time, especially when facing complex and ever-changing marine environmental conditions, where their adaptability and generalization ability are poor. With the rapid development of marine observation technology and environmental sensors, massive amounts of multi-source, multi-dimensional, and time-series marine environmental data are constantly accumulating, providing a technological foundation for achieving more accurate environmental suitability modeling. However, how to extract effective features from complex data and model the nonlinear and dynamic relationship between the environment and biological responses remains a current research challenge and hot topic.

[0005] In recent years, artificial intelligence (AI) technology has been increasingly applied in fields such as environmental modeling, ecological prediction, and intelligent monitoring. Among these, deep learning methods, due to their powerful feature extraction capabilities and modeling flexibility, have demonstrated superior performance in climate change prediction, water quality assessment, and ecosystem response modeling. With the development of large-scale AI modeling (LLM) technology, representative models such as LLaMA, with their powerful language understanding and generation capabilities, are being widely explored for modeling and inference tasks involving unstructured environmental data. LLaMA models possess the ability to jointly model and understand unstructured text and structured data, and can uncover potential spatiotemporal relationships and semantic associations when analyzing complex environmental variables, demonstrating strong data analysis and reasoning capabilities.

[0006] Although the suitability assessment of aquaculture environments has received widespread attention in many fields, most current studies still rely on traditional methods, such as the analytic hierarchy process, fuzzy comprehensive evaluation, or static assessment methods based on empirical rules. These methods mainly focus on freshwater bodies or threshold determination of single environmental factors, making it difficult to adapt to the complex and ever-changing actual marine aquaculture environment, especially for warm-water migratory fish such as bluefin tuna.

[0007] In aquaculture suitability assessment research, there is currently a lack of systematic introduction of advanced data-driven methods for the fusion modeling of multi-source environmental factors. At present, there are still significant gaps in related research regarding modeling methods, particularly in how to effectively capture the time dependence of environmental data, how to improve the robustness of models to data noise and anomalies, and how to achieve high-precision dynamic prediction of suitability trends. These shortcomings limit the accuracy and practicality of the assessment results, making it difficult to meet the actual needs of modern intelligent aquaculture environmental monitoring and control. Therefore, constructing a multi-source environmental factor suitability assessment method will not only help improve the perception and adaptive prediction capabilities of bluefin tuna's living environment, but also provide theoretical support and technical assurance for achieving scientific site selection, dynamic control, and risk early warning. Summary of the Invention

[0008] To address the shortcomings of the existing technologies, this invention provides a multi-source environmental factor suitability assessment method for bluefin tuna aquaculture based on LLM. This method utilizes the large-scale artificial intelligence model LLaMA and combines multi-source environmental factor data to achieve intelligent assessment of the suitability level of the bluefin tuna aquaculture environment. It establishes a closed-loop process from data collection, feature extraction, model training to predictive applications, providing scientific decision support for bluefin tuna aquaculture.

[0009] To achieve the above objectives, the present invention adopts the following technical solution: a method for assessing the environmental suitability of bluefin tuna aquaculture based on LLM multi-source environmental factors, comprising the following steps:

[0010] Step 1: Collect multi-source environmental factor data of bluefin tuna farming areas, combine historical farming records and expert experience, and classify the farming suitability under different environmental conditions into five levels from high to low, and use this as the raw data to construct a dataset.

[0011] Step 2: Preprocess the acquired raw data to unify the scale of multi-source environmental factor data with different dimensions, and divide the dataset into training and test sets according to a set ratio. Then, perform word vector transformation on the constructed time sliding window samples.

[0012] Step 3: Construct a multi-class environmental suitability assessment model based on the LLaMA model. The LoRA fine-tuning strategy is adopted, and some trainable parameters are injected into the attention weights of LLaMA. The output layer is linearly mapped and connected to the Softmax activation function to form the LLaMA-LoRA model.

[0013] Step 4: Input the training set into the multi-class environmental suitability assessment model, fine-tune the training using supervised learning, use the cross-entropy loss function to measure the difference between the model's predicted output and the actual label, and update the trainable parameters injected by LoRA through backpropagation algorithm combined with gradient descent method until the loss function converges or the preset termination condition is met.

[0014] Step 5: Validate the performance of the fine-tuned multi-class environmental suitability assessment model using the test set, and evaluate the model's effectiveness through multi-dimensional evaluation metrics;

[0015] Step Six: Integrate the fine-tuned and validated multi-class environmental suitability assessment model into the Bluefin Tuna Intelligent Aquaculture Management System to achieve real-time environmental suitability prediction for the target aquaculture area.

[0016] Furthermore, in step one, the multi-source environmental factor data includes water temperature, salinity, dissolved oxygen, pH value, water flow velocity, light intensity, ammonia nitrogen concentration, and related meteorological conditions.

[0017] Furthermore, in step one, during the label setting process, bluefin tuna population ecological dynamics parameters are introduced to define the maximum sustainable catch (MSY) and catch rate multiple (C) under a sustainable fishing scenario. m The calculation formula is as follows:

[0018] ,

[0019] In the formula, r is the intrinsic growth rate of bluefin tuna, K is the carrying capacity, and Y is the intrinsic growth rate. actual This represents the actual catch per unit of time.

[0020] Furthermore, in step two, the preprocessing of the raw data includes missing value imputation, outlier removal or replacement, and data standardization.

[0021] Furthermore, in step two, during the word vector transformation process, variable type embedding and time position embedding are introduced to represent the physical meaning of different environmental factors and their relative / absolute position in the time dimension, respectively.

[0022] Furthermore, in step three, in the LoRA fine-tuning strategy, let W... frozen For the original weights, The update term is represented by the low-rank decomposition ΔW as follows: Then the forward propagation process becomes:

[0023]

[0024] In the formula, To reduce the dimension of the input matrix, To recover the matrix in the output dimension, the original weights W are used during training. frozen Keeping them unchanged, only optimize matrices A and B.

[0025] Furthermore, in step four, the fine-tuning training uses the Adam optimizer, with the initial learning rate set to 0.001. A learning rate scheduling strategy is used to avoid training oscillations, and an Early Stopping mechanism is set to prevent overfitting.

[0026] Furthermore, in step five, the multi-dimensional evaluation metrics include precision, recall, and F1 score.

[0027] Furthermore, in step six, during the real-time environmental suitability prediction period, the model is input according to the set time window for prediction, and the current time period's aquaculture environmental suitability level is output. The prediction results are presented in a visual format.

[0028] Compared with the prior art, the beneficial effects of the present invention are:

[0029] 1. The method of this invention is based on the large-scale artificial intelligence model LLaMA and combined with multi-source heterogeneous data of marine aquaculture environment. It can effectively capture the temporal change characteristics of environmental factors and realize intelligent assessment of the suitability level of bluefin tuna aquaculture environment. It establishes a closed-loop process from data collection, feature extraction, model training to prediction application, and has strong generalization ability and environmental adaptability. It breaks through the limitation of traditional methods that are difficult to handle dynamic environmental information and improves assessment efficiency and judgment accuracy.

[0030] 2. The method of this invention utilizes a data-driven learning mechanism to automatically mine the nonlinear relationship between multiple key environmental factors and aquaculture suitability. Unlike traditional suitability analysis methods that rely on expert experience and rule setting, this method reduces the reliance on human experience, enabling non-professional users to make scientific aquaculture decisions based on model prediction results. This improves the system's adaptability and ease of use, and can provide scientific decision support for the aquaculture of high-economic-value marine fish such as bluefin tuna.

[0031] 3. The method of this invention fully integrates multi-source environmental factor data such as water temperature, salinity, dissolved oxygen, pH value, water flow velocity, light intensity, ammonia nitrogen concentration and related meteorological conditions, and realizes suitability level classification through a unified standardized processing and classification mechanism. It not only has a stronger environmental perception capability, but also improves the generalization and robustness of the model, and is suitable for the assessment of bluefin tuna farming environment in different regions, different seasons and variable climatic conditions.

[0032] 4. The method of this invention is not only applicable to the aquaculture of high-economic-value deep-sea fish such as bluefin tuna, but also has the potential to be promoted to other marine aquaculture species. It has good versatility and scalability, and provides strong technical support for the realization of intelligent, refined and sustainable marine aquaculture environment management. It is expected to play an important role in multiple fields such as marine fisheries, smart aquaculture and ecological protection. Attached Figure Description

[0033] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0034] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the invention, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0035] like Figure 1 As shown, a method for assessing the environmental suitability of bluefin tuna aquaculture based on LLM multi-source environmental factors includes the following steps:

[0036] Step 1: Collect multi-source environmental factor data from the bluefin tuna farming area. Combine historical farming records and expert experience to label the suitability of farming under different environmental conditions. The labels are divided into five levels from high to low: very suitable, suitable, moderate, unsuitable, and very unsuitable. This data is used as the raw data to construct a dataset for subsequent model input. Specifically:

[0037] S101. Collect multi-source environmental factor data from bluefin tuna farming areas, including water temperature, salinity, dissolved oxygen, pH, water flow velocity, light intensity, ammonia nitrogen concentration, and related meteorological conditions. Combine this with historical farming records to comprehensively analyze the growth performance, physiological state, and farming efficiency of bluefin tuna under different environmental conditions. For example, collect production logs, real-time monitoring data, and relevant farming technology reports from typical deep-sea cage farming areas over several farming cycles. Extract key biological and economic indicators for the fish, including weight gain per unit time (e.g., daily weight gain), survival rate from fry to adult, feed conversion ratio, feeding activity score, and the frequency and severity of major diseases (e.g., gas bubble disease, gill disease, parasite infection, etc.). Time-align and sample-pair the above biological response indicators with the corresponding environmental factors to ensure the accuracy of the one-to-one mapping between environment and biological performance. Subsequently, statistical analysis methods (including Pearson correlation analysis, principal component analysis, hierarchical cluster analysis, and multiple linear regression) were used to quantitatively assess the correlation between environmental factors and aquaculture effects, and to identify key hard factors and their threshold characteristics that significantly affect bluefin tuna aquaculture output.

[0038] Simultaneously, by combining the geographical, hydrological, and climatic characteristics of the aquaculture area, the temporal variation patterns of various environmental factors at different aquaculture stages (such as the seedling stage, growth stage, and adult stage) and their impact on fish stress, growth inhibition, or disease outbreaks were identified. Further analysis of the synergistic or antagonistic relationships among various environmental factors, such as the interaction between dissolved oxygen and water temperature, and the correlation between pH and ammonia nitrogen concentration, was conducted to construct a candidate set of influencing variables. Finally, a set of representative, stable, and biologically relevant multi-source environmental factor variables was selected, providing a high-quality data foundation for subsequent model input feature setting and constructing key environmental factor feature dimensions affecting aquaculture suitability in the training dataset.

[0039] S102. Based on the pairing of environmental factors and biological response indicators and factor screening, suitability labeling is carried out to form high-quality supervised learning sample labels. First, based on the association analysis results constructed in S101, combined with existing marine fish farming research results, industry technical regulations, and expert experience, the physiological adaptation ranges of bluefin tuna to various environmental factors are comprehensively defined. For example, based on the thermal adaptability of bluefin tuna, water temperature is divided into ranges such as "<15°C (low, growth limited)," "15~29°C (suitable)," and ">29°C (high, may cause stress)," while dissolved oxygen is divided into standard levels such as ">6.5mg / L (sufficient)," "4.5~6.5mg / L (critical)," and "<4.5mg / L (risk of hypoxia)." Other environmental factors such as salinity, pH, water flow velocity, ammonia nitrogen concentration, and light intensity are also divided into ranges based on relevant literature and measured ranges to ensure that the division is reasonable and has biological interpretability.

[0040] Building upon this foundation, to enhance the ecological explanatory power and sustainability assessment capabilities of the labeling, ecological dynamics parameters of bluefin tuna populations are introduced, particularly a comprehensive assessment index of resource exploitation intensity and sustainable catch capacity. The intrinsic growth rate (r) of bluefin tuna measures the maximum population growth rate of this species under ideal conditions, while carrying capacity (K) represents the maximum population capacity that an ecosystem in a specific aquaculture or natural marine area can support. Together, these two parameters determine the Maximum Sustainable Yield (MSY) under sustainable fishing conditions, calculated as follows:

[0041]

[0042] Building on this, to measure the relationship between actual aquaculture or fishing intensity and the sustainable frontier, the Catch Multiplier (C) is introduced. m As an ecological parameter, it is defined as the actual catch per unit time (Y). actual The ratio between ) and MSY is calculated using the following formula:

[0043]

[0044] In setting the suitability level, this ecological parameter is incorporated into the multi-factor judgment logic along with environmental factors. For example, when all environmental factors are within the suitable range and C... m When C ≤ 1.0, it is considered very suitable; if the environmental factors are moderate but 1.0 < C m If C ≤ 1.2, the suitability is downgraded by one level; mWhen the value is greater than 1.5, meaning the actual catch exceeds the resource carrying capacity threshold, even if the environmental conditions are good, it is considered unsuitable or very unsuitable; conversely, if all environmental factors are in the disadvantageous range and C m If the levels are both too high, priority will be given to resource overdraft and systemic risk, and the suitability assessment will show a more conservative trend. Through the coupling mechanism of the above environmental factors and resource intensity, a context-driven label generation system is constructed to achieve a comprehensive judgment on the environmental suitability of samples for each time period.

[0045] The CatchMSY model and its expression are proprietary estimation formulas in the fisheries field, specifically designed for the irreversible, one-way estimation of the maximum possible catch based on historical catch data and environmental carrying capacity parameters on a purely mathematical scale. It represents a one-way reasoning method from environmental carrying capacity to expected yield. Environmental carrying capacity and environmental suitability are positively correlated. This invention uses this estimation method in conjunction with an artificial intelligence model to inversely estimate environmental carrying capacity from expected yield, and uses this as a parameter to evaluate environmental suitability, thus combining yield information with basic environmental conditions.

[0046] Finally, by combining a large number of labeled sample results, a dataset with clear and balanced label levels was formed. The labels were divided into five levels from high to low: very suitable, suitable, average, unsuitable, and very unsuitable. This labeling system not only reflects the physiological response of bluefin tuna to complex marine environments, but also possesses the clarity and practicality required for model training.

[0047] Step Two: Preprocess the acquired raw data, including missing value imputation, outlier removal or replacement, and data standardization. This ensures that multi-source environmental factor data with different scales are scaled uniformly for subsequent model input. After preprocessing, the dataset is divided into training and test sets according to a predetermined ratio to guarantee the diversity of training samples and the continuity of time series features. Specifically:

[0048] S201. For the collected raw data, a completeness and consistency check is first performed. For missing values ​​caused by sensor failure, network interruption, or manual recording errors, various strategies are used for imputation: when a single variable has discontinuous missing values ​​within a short time period, linear interpolation, forward fill, or moving average methods are used for restoration; when there are large-scale or structural missing values, K-nearest neighbors (KNN), multiple regression imputation, or deep imputation methods based on contextual semantics are used to improve imputation accuracy. For outliers in the data, observations that significantly deviate from the normal range are identified based on box plot method (IQR), Z-score method, or time series stationarity analysis, and outliers are removed or replaced to ensure the rationality of the sample distribution.

[0049] After handling missing and outlier data, environmental factors are standardized. Considering the significant differences in the physical dimensions and value ranges of various environmental factors (e.g., dissolved oxygen in mg / L, water flow velocity in m / s, light intensity in lux), Z-score standardization is uniformly used to convert each feature into a standard normal distribution with a mean of 0 and a standard deviation of 1. Alternatively, min-max normalization is used to scale the data to the [0,1] interval, depending on the specific model requirements. Standardization not only eliminates the scale effect caused by different units but also helps improve the model's training convergence speed and stability. Furthermore, for periodic time variables (e.g., months, hours), sine and cosine encoding can be used to map them into two-dimensional features, preserving their periodic structure information.

[0050] Based on data preprocessing, the input feature vectors for each time segment are reconstructed to ensure that all samples have consistent dimensions, accurate labels, and a data structure suitable for time series modeling. After preprocessing, a dataset with complete structure, numerical stability, and meeting the model input requirements is generated.

[0051] S202. Based on the preprocessed dataset, the training and test sets are divided according to the set time windows and sample ratios. While maintaining the time series structure of the data, a historical time window of a certain length is selected as the model input segment, and the output is the suitability label for the corresponding time point, constructing a multi-step time sliding window sample set. The time-series split method is used during the partitioning process, ensuring that the training set contains early data and the test set contains later data, preventing future information from being leaked into the training stage.

[0052] Specifically, the dataset is divided chronologically into an 80% training set and a 20% test set (or the ratio can be adjusted according to model validation needs). This ensures that the training samples cover a variety of typical environmental scenarios and suitability levels, improving the model's ability to identify and generalize different types of data. Simultaneously, a sliding window sample augmentation strategy (e.g., a 5-day step size and a 30-day window length) is used to construct more overlapping time-series samples, increasing the number of training samples and strengthening the model's temporal memory modeling ability. Ultimately, a training / test set with a time-dependent structure, accurate labels, and sufficient quantity is completed, providing a foundation for subsequent model construction and evaluation.

[0053] S203. The constructed time-sliding window samples undergo word vector transformation to adapt to the input requirements of the LLaMA model. First, the multi-source environmental factor data and their corresponding labels in each sample are converted into structured natural language descriptions using a predefined template, such as: "On day t, the water temperature is 22.4 degrees Celsius, the salinity is 35‰, the dissolved oxygen is 7.2 mg / L...", ensuring clear semantic relationships and sequential information between environmental variables. Then, the LLaMA model's built-in tokenizer is used to segment and encode the above text sequence, converting it into a token ID sequence, and adding special markers (such as...) according to model requirements. <bos> 、 <eos>(etc.) and positional encoding to preserve input semantics and temporal structure.

[0054] To further enhance semantic expressiveness, feature-type embedding and time-position embedding can be introduced before the word embedding layer. These embeddings are used to represent the physical meaning of different environmental factors and their relative / absolute positions in the time dimension, respectively. These word vectors will serve as input features in the model encoding stage, providing the LLaMA model with a cognitive foundation for multivariate dynamic change patterns and supporting effective modeling of subsequent classification tasks.

[0055] Step 3: Construct a multi-class environmental suitability assessment model based on the large-scale AI model LLaMA. Building upon the pre-trained LLaMA model, it employs Low-Rank Adaptation (LoRA) to inject some trainable parameters into the attention weights of LLaMA. The output layer uses a linear mapping followed by a Softmax activation function to form the LLaMA-LoRA model, enabling probabilistic prediction of suitability levels. Specifically:

[0056] S301. Select a suitable LLaMA version from the large models provided by Hugging Face, including "Tiny-LLaMA", "7B-LLaMA", etc. Introducing an LLaMA model can improve the effectiveness and accuracy of implementation.

[0057] S302. The Low-Rank Adaptation (LoRA) fine-tuning strategy is selected to initially set the hyperparameters, preparing for subsequent parameter tuning. LoRA aims to improve the parameter efficiency of LLM fine-tuning by training smaller low-rank factorization matrices. LoRA is based on the assumption that the intrinsic rank of weight updates is very low during task adaptation, therefore, two smaller matrices can be used to represent the weight changes, and these two trainable low-rank factorization matrices are injected into each layer of the Transformer architecture. Let W... frozen For the original weights in the pre-trained LLaMA model, the update term ΔW during the task adaptation process is represented by the following low-rank decomposition:

[0058]

[0059] Therefore, the forward propagation process becomes:

[0060]

[0061] In the formula, W frozen Represents the original weights. Indicates its update item, To reduce the dimension of the input matrix, To recover the output dimension matrix, r ≪ min(d,k), meaning the rank r is much smaller than the dimension of the original matrix. During the training initialization phase, matrix A is typically initialized randomly using a Gaussian distribution, while matrix B is initialized to all zeros, thus ensuring that ΔW = 0 at the initial time step. Throughout the training process, the original weights W... frozen Keeping them unchanged, only optimize matrices A and B. After fine-tuning, these two low-rank matrices can be merged with the frozen weights to simplify the model structure and optimize inference.

[0062] In practical fine-tuning, LoRA is used to efficiently adapt a pre-trained LLaMA model for sequence classification tasks. Unlike full-parameter fine-tuning, LoRA only injects trainable low-rank matrices into specific layers of the Transformer architecture, enabling efficient transfer learning while preserving the original knowledge. In LLaMA, LoRA is applied to the q_proj (query transformation) and v_proj (numerical transformation) layers of the attention mechanism, where the weight update term ΔW is factored into BA, significantly reducing the number of training parameters required by the model. This method decomposes the original weight matrix into two low-order matrices and updates only these two matrices during fine-tuning. Parameters such as "target_model" and "modules_to_save" can be adjusted during parameter tuning to improve performance.

[0063] Step 4: Input the training set into the constructed multi-class environmental suitability assessment model and fine-tune it using supervised learning. During training, the cross-entropy loss function is used to measure the difference between the model's predicted output and the actual labels. The trainable parameters injected by LoRA can be updated using backpropagation combined with gradient descent until the loss function converges or meets the preset termination condition. Appropriately setting the fine-tuning hyperparameters, including the learning rate and number of iterations, based on the sample size and task complexity, helps ensure the model has good convergence and generalization ability during training. Specifically:

[0064] S401. The pre-divided training set, after text template conversion and word segmentation, is used as the token sequence input to construct the completed LLaMA model. The time series data of multi-source environmental factors is encoded into structured text, with the input dimension dynamically adapting to the maximum sequence length of the LLaMA model. The corresponding labels are manually set aquaculture suitability levels. During training, a mini-batch strategy is adopted, with each batch containing a certain number of samples, and time-series order-preserving partitioning is combined to ensure the temporal continuity and representativeness of the training set data. Assuming the number of input multi-source environmental factor data T is z, then:

[0065]

[0066] The hidden state tensor corresponding to the output of the Transformer layer of the LLaMA model is , where 4096 is the hidden vector dimension at each position of the LLaMA model, reflecting the contextual embedding information of the environmental feature sequence after being encoded by the model.

[0067] The final hidden states output by the S402 and LLaMA models are projected onto a five-dimensional classification space via a fully connected linear mapping layer, corresponding to the predicted probabilities of the five fitness levels. This linear transformation result is then normalized using a softmax activation function to generate probability distribution vectors for each class label. The cross-entropy loss is calculated using the predicted probabilities and the true labels, and backpropagated for weight updates. To enhance training stability and generalization ability, input perturbation, label smoothing, and appropriate regularization techniques can be employed.

[0068]

[0069] Here, Result represents the model's final predicted probability output for the five suitability levels. During training, the cross-entropy loss between the predicted results and the true labels is used as the objective function for optimization, and the trainable parameters are updated in conjunction with the backpropagation mechanism.

[0070] S403. Fine-tuning training uses the Adam optimizer with an initial learning rate set to 0.001. A learning rate scheduling strategy (such as dynamic decay based on validation set performance) is employed to avoid training oscillations. During training, the loss changes on the training and validation sets are monitored in real time, and loss curves are plotted to aid in diagnosis. To prevent overfitting, an Early Stopping mechanism is implemented; training is terminated early if the validation set loss fails to improve for several consecutive rounds. The Result value is only compared to the pre-defined labels. Backpropagation is then used to update the model parameters.

[0071] Step 5: After training, the performance of the fine-tuned multi-class environmental suitability assessment model is validated using the test set to evaluate the environmental suitability level of different samples. The model's performance is quantitatively analyzed using multi-dimensional evaluation metrics such as precision, recall, and F1 score to verify its applicability, robustness, and generalization ability under complex multi-source environments. Specifically:

[0072] S501. Convert the test set into the input format, input it into the fine-tuned multi-class environmental suitability assessment model, obtain the prediction results of the aquaculture suitability level for each sample, and compare them with the actual labels. Use the overall accuracy as the main performance indicator to measure the overall prediction accuracy of the model for all test samples.

[0073] S502. Starting from each suitability level category dimension, calculate precision, recall, and F1 score respectively. Precision measures the proportion of samples predicted to be of a certain level that actually belong to that level. Recall reflects the proportion of samples that the model successfully identifies from those that actually belong to that level. The F1 score is the harmonic mean of the two and is used to comprehensively evaluate the model's classification stability and boundary learning ability at different levels.

[0074] Step Six: Integrate the fine-tuned and validated multi-class environmental suitability assessment model into the bluefin tuna intelligent aquaculture management system to achieve real-time environmental suitability prediction for the target aquaculture area, thereby improving the level of intelligent aquaculture and environmental risk response capabilities. Specifically:

[0075] S601. Deploy an environmental perception system: Deploy multiple types of sensor nodes in the actual bluefin tuna farming area to collect multi-source environmental factor data in real time, and realize data transmission and preliminary processing through a wireless transmission system or edge computing device.

[0076] S602. Construct a real-time prediction interface: Input real-time collected multi-source environmental factor data into a fine-tuned and validated multi-class environmental suitability assessment model. Input the data into the model according to a set time window (e.g., hourly or daily) for prediction, and output the aquaculture suitability level for the current period. The prediction results are visualized in five levels: very suitable, suitable, average, unsuitable, and very unsuitable. The dynamic distribution of regional suitability can be displayed using color coding or heatmaps.

[0077] This invention assists aquaculture managers in conducting intelligent suitability assessments of bluefin tuna farming environments. It utilizes LLM (Limited Learning Model) to automatically learn and predict suitability levels from multi-source environmental factor data. By using multi-source environmental factor data, such as water temperature, salinity, dissolved oxygen, pH, water flow velocity, light intensity, ammonia nitrogen concentration, and related meteorological conditions, as input, and combining this data with actual aquaculture tag data for training, it achieves accurate assessment and dynamic prediction of environmental conditions in target aquaculture areas. This method effectively replaces the traditional model relying on experience and manual judgment, reducing the risk of human error and improving the response efficiency of environmental monitoring and decision-making. To ensure the practicality and promotional value of the method, the assessment framework proposed in this invention can adapt to varying marine environmental conditions in different regions, seasons, and farming densities. Through the model's ability to uniformly process multi-source heterogeneous data and perform time-series modeling, the system possesses strong robustness and generalization capabilities, and can be widely applied to various scenarios such as deep-sea, bay, and shore-based aquaculture. This invention provides a scientific, replicable, and scalable environmental suitability assessment tool for the aquaculture of high-value marine fish such as bluefin tuna, which will help improve the level of intelligent, green, and efficient deep-sea aquaculture in my country.

[0078] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered illustrative and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

[0079] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.< / eos> < / bos>

Claims

1. A method for assessing the environmental suitability of bluefin tuna aquaculture based on LLM using multiple environmental factors, characterized in that: Includes the following steps: Step 1: Collect multi-source environmental factor data from bluefin tuna farming areas. Combining historical farming records and expert experience, classify the suitability of farming under different environmental conditions into five levels from high to low, and use this as the raw data to construct a dataset. In the labeling process, bluefin tuna population ecological dynamics parameters are introduced to define the maximum sustainable catch (MSY) and catch multiple (C) under a sustainable fishing scenario. m The calculation formula is as follows: , In the formula, r is the intrinsic growth rate of bluefin tuna, K is the carrying capacity, and Y is the intrinsic growth rate. actual This represents the actual catch per unit of time. Step 2: Preprocess the acquired raw data to unify the scale of multi-source environmental factor data with different dimensions, and divide the dataset into training and test sets according to a set ratio. Perform word vector transformation on the constructed time sliding window samples. In the word vector transformation process, introduce variable type embedding and time position embedding to represent the physical meaning of different environmental factors and their relative / absolute position in the time dimension, respectively. Step 3: Construct a multi-class environmental suitability assessment model based on the LLaMA model. Employ the LoRA fine-tuning strategy, injecting some trainable parameters into the attention weights of the LLaMA model. The output layer is linearly mapped and then connected to a Softmax activation function, forming an LLaMA-LoRA model. In the LoRA fine-tuning strategy, let W... frozen For the original weights, The update term is represented by the low-rank decomposition ΔW as follows: Then the forward propagation process becomes: In the formula, To reduce the dimension of the input matrix, To recover the matrix in the output dimension, the original weights W are used during training. frozen Keeping them unchanged, only optimize matrices A and B; Step 4: Input the training set into the multi-class environmental suitability assessment model, fine-tune the training using supervised learning, use the cross-entropy loss function to measure the difference between the model's predicted output and the actual label, and update the trainable parameters injected by LoRA through backpropagation algorithm combined with gradient descent method until the loss function converges or the preset termination condition is met. Step 5: Validate the performance of the fine-tuned multi-class environmental suitability assessment model using the test set, and evaluate the model's effectiveness through multi-dimensional evaluation metrics; Step Six: Integrate the fine-tuned and validated multi-class environmental suitability assessment model into the Bluefin Tuna Intelligent Aquaculture Management System to achieve real-time environmental suitability prediction for the target aquaculture area.

2. The method for assessing the environmental suitability of bluefin tuna aquaculture based on LLM-based multi-source environmental factors according to claim 1, characterized in that: In step one, the multi-source environmental factor data includes water temperature, salinity, dissolved oxygen, pH value, water flow velocity, light intensity, ammonia nitrogen concentration, and related meteorological conditions.

3. The method for assessing the environmental suitability of bluefin tuna aquaculture based on LLM-based multi-source environmental factors according to claim 1, characterized in that: In step two, the preprocessing of the raw data includes missing value imputation, outlier removal or replacement, and data standardization.

4. The method for assessing the environmental suitability of bluefin tuna aquaculture based on LLM-based multi-source environmental factors according to claim 1, characterized in that: In step four, the fine-tuning training uses the Adam optimizer, with the initial learning rate set to 0.

001. A learning rate scheduling strategy is used to avoid training oscillations, and an Early Stopping mechanism is set to prevent overfitting.

5. The method for assessing the environmental suitability of bluefin tuna aquaculture based on LLM-based multi-source environmental factors according to claim 1, characterized in that: In step five, the multi-dimensional evaluation metrics include precision, recall, and F1 score.

6. The method for assessing the environmental suitability of bluefin tuna aquaculture based on LLM-based multi-source environmental factors according to claim 1, characterized in that: In step six, during the real-time environmental suitability prediction period, the model is input according to the set time window for prediction, and the current time period's aquaculture environmental suitability level is output. The prediction results are presented in a visual format.

Citation Information

Patent Citations

  • Marine ecological multi-agent construction method and interaction system thereof

    CN119204861A

  • Aquaculture disease prevention and control text data enhancement method and device based on small model guidance and storage medium

    CN120277217A