An aquaculture disease prediction method based on multi-modal data fusion
The aquaculture disease prediction method based on multimodal data fusion utilizes TabTransformer and BERT models to fuse structured and unstructured data, solving the problems of low detection efficiency and weak generalization ability in existing technologies, and achieving efficient early warning of multiple pathogens.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NINGBO UNIV
- Filing Date
- 2025-06-19
- Publication Date
- 2026-04-24
AI Technical Summary
Existing technologies are inefficient in detecting parasitic diseases in fish farming, rely on professional personnel, and have a high false detection rate. Traditional machine learning models cannot handle unstructured information, resulting in insufficient coverage of environmental features and weak generalization ability, making it impossible to achieve simultaneous early warning of multiple pathogens.
A multimodal data fusion method is adopted, combining TabTransformer and BERT models to fuse structured water quality parameters with unstructured text data. Feature vectors are extracted through multi-layer Transformer encoder and BERT encoder, and the Sigmoid function is used to predict the probability of disease. The results are displayed through a visual interactive interface.
It improves the accuracy and generalization ability of disease prediction, reduces resource consumption, meets the real-time requirement of simultaneous early warning of multiple pathogens, and enhances the system's practicality and deployment efficiency.
Smart Images

Figure CN120782037B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of aquaculture technology, specifically to a method for predicting aquaculture diseases based on multimodal data fusion. Background Technology
[0002] In the field of parasitic disease early warning in fish farming, existing technologies mainly rely on three types of methods, but all of these methods have significant limitations. First, traditional manual inspection and microscopic diagnosis methods require aquaculture personnel to identify pathogens by observing fish or water samples under a microscope. This method is not only inefficient—a single sample test takes more than 30 minutes, and the inspection cycle is usually longer than 48 hours, resulting in a serious lack of real-time early warning capabilities; it also highly depends on the experience of professionals. According to statistics from the Fisheries Technology Platform of the Ministry of Agriculture and Rural Affairs, the false detection rate of manual inspections can reach 25%-40%, which is difficult to meet the monitoring needs of large-scale aquaculture farms with a capacity of thousands of tons. Second, although methods based on traditional machine learning models (such as support vector machines or random forests) can use structured water quality parameters (such as temperature, salinity, and pH) for modeling, they only support numerical variable inputs and cannot handle unstructured textual information such as "reduced feeding" or "white water color" in aquaculture logs, resulting in an environmental feature coverage rate of less than 60%; in addition, such models are sensitive to water quality fluctuations, and the accuracy drops by more than 35% when applied across aquaculture farms, indicating weak generalization ability. Finally, specialized models designed for single pathogens require the construction of a separate prediction system for each parasite, which not only increases the operation and maintenance costs by 300%, but also makes it impossible to achieve concurrent early warning for multiple pathogens due to the fragmentation between models. For example, it cannot provide a collaborative risk assessment when Cryptocaryon stimulans and Trypanosoma haematobium break out at the same time. Summary of the Invention
[0003] To address the low efficiency of current manual disease assessment in aquaculture and the inability of machine learning models to incorporate unstructured information, this invention proposes a method for predicting aquaculture diseases based on multimodal data fusion, comprising the following steps:
[0004] S1: Collect aquaculture data containing structured data and unstructured text data, wherein the structured data consists of various target water quality parameters, and the unstructured text data consists of aquaculture logs and environmental description text;
[0005] S2: Input structured data into the TabTransformer model, perform column embedding and feature context relationship modeling through multi-layer Transformer encoder, and output structured feature vectors;
[0006] S3: Input unstructured text data into the pre-trained BERT encoder and extract the output vector of CLS tags as text semantic features;
[0007] S4: Concatenate the structured feature vector with the textual semantic features to form a fused feature and input it into the fully connected layer. Then, use the Sigmoid function to predict the incidence probability of each target disease.
[0008] This invention integrates structured water quality parameters with unstructured text data, and combines a dual-channel modeling architecture of TabTransformer and BERT to fully capture the multidimensional correlations of pathogen induction factors in complex aquaculture environments. Compared with traditional single-modal models, the average prediction accuracy is improved, effectively solving the problems of single prediction dimension and weak generalization ability of existing technologies.
[0009] Furthermore, in step S1, the structured data includes at least three water quality parameters selected from temperature, salinity, pH value, dissolved oxygen, nitrite concentration, and silicate concentration, and the unstructured text data includes at least one of aquaculture logs, on-site descriptions, and expert records.
[0010] Furthermore, in step S1, the structured data is input in CSV or Excel format and includes continuous variables and discrete categorical variables.
[0011] Furthermore, in step S3, the extraction of the semantic vector of the CLS marker specifically involves: inputting the unstructured text data into the BERT encoder and taking the output vector at the corresponding position of the CLS marker as the text semantic feature.
[0012] Furthermore, in step S4, the fully connected layer is a multi-label output structure with an output dimension of... These correspond to the incidence rates of each type of target disease, among which... Let n be the number of samples and m be the total number of target disease categories.
[0013] Furthermore, in step S4, the model predicting the probability of disease onset using the sigmoid function is trained using the following loss function:
[0014]
[0015] In the formula, The total loss is represented by , and 'i' is the target disease number. The cross-entropy loss function is used to measure the predicted probability of the i-th type of target disease. Corresponding real tags The differences between them.
[0016] Furthermore, following step S4, the following step is also included:
[0017] S5: Divide risk levels based on the positional relationship between the incidence probability of the target disease and the preset threshold interval;
[0018] When the incidence rate of the target disease is below a preset threshold range, it is classified as low risk;
[0019] When the incidence rate of the target disease is within a preset threshold range, it is classified as medium risk;
[0020] When the incidence rate of the target disease is above a preset threshold range, it is classified as high risk.
[0021] Furthermore, it also includes the following steps:
[0022] S6: Display the prediction results and the variable factors that affect the current prediction results through a visual interactive interface.
[0023] Compared with the prior art, the present invention has at least the following beneficial effects:
[0024] (1) The present invention proposes a method for predicting aquatic diseases based on multimodal data fusion. By fusing structured water quality parameters with unstructured text data and combining the TabTransformer and BERT dual-channel modeling architecture, it fully captures the multidimensional association of pathogen inducing factors in complex aquaculture environments. Compared with traditional single-modal models, the average prediction accuracy is improved, effectively solving the problems of single prediction dimension and weak generalization ability of existing technologies.
[0025] (2) The multi-label Sigmoid output layer design is adopted, and the single model outputs the incidence probability of multiple parasites simultaneously. Compared with the traditional single pathogen independent modeling scheme, the resource consumption is reduced and the compatibility problem between models is avoided. It meets the real-time needs of the breeding site for concurrent early warning of multiple pathogens, and significantly improves the system's practicality and deployment efficiency. Attached Figure Description
[0026] Figure 1 This is a flowchart illustrating the steps of a global measurement and control task scheduling method based on resource conflict optimization. Detailed Implementation
[0027] The following are specific embodiments of the present invention, which are described in conjunction with the accompanying drawings. However, the present invention is not limited to these embodiments.
[0028] To enable those skilled in the art to more clearly understand the technical solution and advantages of the present invention, the implementation process of the present invention is described in detail below in conjunction with a specific aquaculture scenario. This embodiment takes the intelligent early warning of three types of parasites—Cryptocele irritans, Trypanosoma haematobium, and Benedenia rubescens (i.e., the total number of target disease categories m=3)—in a marine fish farm as its application target, demonstrating how the present invention achieves end-to-end prediction through multimodal data fusion, deep learning modeling, and collaborative interaction with an interactive platform. It should be noted that this embodiment is merely illustrative and does not constitute a limitation on the scope of protection of the present invention. Figure 1 As shown, this invention proposes a method for predicting aquaculture diseases based on multimodal data fusion, including the following steps:
[0029] S1: Collect aquaculture data containing structured data and unstructured text data, wherein the structured data consists of various target water quality parameters, and the unstructured text data consists of aquaculture logs and environmental description text;
[0030] S2: Input structured data into the TabTransformer model, perform column embedding and feature context relationship modeling through multi-layer Transformer encoder, and output structured feature vectors;
[0031] S3: Input unstructured text data into the pre-trained BERT encoder and extract the output vector of CLS tags as text semantic features;
[0032] S4: Concatenate the structured feature vector with the text semantic features to form a fusion feature and input it into the fully connected layer. Predict the incidence probability of each target disease through the Sigmoid function.
[0033] S5: Divide risk levels based on the positional relationship between the incidence probability of the target disease and the preset threshold interval;
[0034] S6: Display the prediction results and the variable factors that affect the current prediction results through a visual interactive interface.
[0035] This invention acquires multi-source heterogeneous information from the aquaculture environment through an integrated data acquisition process. Regarding structured data, the system extracts target water quality parameters from the aquaculture farm's sensor network, water quality monitoring equipment, and manual records. These parameters cover continuous variables such as temperature (°C), salinity (‰), pH, dissolved oxygen (mg / L), nitrite concentration, and silicate concentration, while also incorporating discrete categorical variables such as aquaculture unit number and feeding cycle. After this data is input in CSV or Excel spreadsheet format, it undergoes standardized preprocessing: continuous variables are normalized using Z-score to eliminate dimensional differences (the calculation formula is...). ),in The sample mean. Standard deviation Normalized value For categorical variables, one-hot encoding or embedding layers are used to convert them into parsable numerical vectors to preserve the potential correlations between discrete features.
[0036] For unstructured text data, the system simultaneously collects aquaculture logs, on-site observation records, and technician descriptions (such as "fish feeding has decreased sharply" and "pond water turbidity has increased"), and imports them in TXT or CSV format. The text processing workflow includes two stages: semantic cleaning and vectorization. First, special characters, stop words, and irrelevant symbols are removed to retain core descriptive phrases. Then, the semantic representation capabilities of the pre-trained BERT model are used to map the cleaned text sequence into fixed-dimensional semantic feature vectors. Specifically, the 768-dimensional vector corresponding to the [CLS] marker in the BERT output layer is extracted as the global semantic representation of the entire text, providing high-dimensional input for subsequent multimodal fusion.
[0037] To ensure the spatiotemporal consistency of multi-source data, the system establishes a dual indexing mechanism of timestamps and aquaculture unit IDs to align structured data with text data sources. Invalid samples with a missing rate exceeding 15% are removed, and outliers are detected and imputed based on box plot rules. Finally, a standardized dataset with spatiotemporal pairing is generated for model training and inference.
[0038] In this stage, the present invention overcomes the limitation of traditional models that only support numerical input by differentiating between discrete / continuous structured parameters and free text through a differentiated processing procedure. BERT's [CLS] vector extraction mechanism transforms unstructured text into computable semantic features, filling the gap in disease modeling using aquaculture logs. Furthermore, spatiotemporal alignment and outlier repair mechanisms ensure the reliability of multimodal inputs and avoid prediction bias caused by data misalignment.
[0039] The standardized continuous variables and the encoded categorical variables together constitute the dimension. The input matrix (where Let n be the number of samples and d be the feature dimension (in real number space). The input is then fed into the TabTransformer module for column-level embedding and contextual relationship modeling. This process first maps each feature column to a high-dimensional vector space through a learnable column embedding layer, generating a vector space with dimension d. The embedded representation (h is the hidden layer dimension) establishes semantic relationships between discrete categorical variables and continuous variables in a unified vector space. Subsequently, the embedded vectors are input into a multi-layer Transformer encoder stack, which utilizes self-attention to dynamically capture global dependencies between features, such as the synergistic effect of dissolved oxygen concentration and salinity changes, and the moderating effect of pH fluctuations on nitrite toxicity—complex nonlinear interaction patterns. Finally, the context-aware feature matrix output by the encoder stack is... As a deep representation of structured data, it provides high-information-density feature vectors for subsequent multimodal fusion.
[0040] In this stage, the dynamic relationships between parameters are explicitly learned through a self-attention mechanism (such as the physicochemical law that "rising temperature leads to a decrease in dissolved oxygen saturation"), breaking through the high dependence of traditional models on feature engineering. At the same time, the column embedding layer is compatible with continuous and categorical variables, avoiding the curse of dimensionality caused by one-hot encoding, and uses attention weights to visualize traceable key influencing factors (such as the contribution of salinity mutations to parasite disease incidence), meeting the transparency requirements of farm decision-making.
[0041] The preprocessed aquaculture logs and environmental descriptions (e.g., "Fish gather at the surface, feeding activity decreases," "Flocculated sediment appears at the bottom of the pond") are converted into token sequences, with [CLS] start tags and [SEP] separator tags added to adapt to BERT's input specifications. This sequence is then fed into BERT's multi-layer Transformer encoder, where a self-attention mechanism parses the contextual semantic relationships between tokens (e.g., the potential causal chain between "decreased feeding activity" and parasite stress). Finally, a hidden state vector for each token is generated at the output layer. The hidden state vector corresponding to the [CLS] tag at the beginning of the sequence is extracted as the global semantic representation of the entire text, with a fixed dimension. (Taking the BERT-Base model as an example). This vector condenses the key semantic information in the text (such as disease precursor behaviors and abnormal environmental phenomena), serving as a unified feature representation of unstructured data and providing high-dimensional semantic input for subsequent fusion with structured features.
[0042] Compared to traditional text processing methods, BERT's self-attention mechanism identifies implicit causal chains in text (such as "increased foam on the water surface → insufficient dissolved oxygen → susceptibility to parasites"), overcoming the limitation of the bag-of-words model, which only captures surface word frequencies. By dynamically parsing the true meaning of descriptive phrases (such as the specificity of "fish rubbing against the pool wall" in the context of parasite infection), it avoids misjudgments caused by semantic ambiguity in traditional methods. Furthermore, the fixed-dimensional CLS vector (768 dimensions) ensures compatibility with the TabTransformer output structure, achieving seamless alignment of heterogeneous data.
[0043] The structured feature matrix generated by TabTransformer Text semantic feature vectors extracted by BERT The features are concatenated along the feature dimension to form a fused feature matrix. This matrix is then input into the fully connected layer (output dimension is...). The algorithm learns cross-modal associations between structured parameters and text semantics through nonlinear transformations (e.g., the synergistic indicative meaning of "dissolved oxygen decreases" and "fish surfacing" in the log). Finally, the output layer activated by three nodes using sigmoid synchronously generates the symptomatic probabilities of stimulating Cryptocaryon, Trypanosoma haematobium, and Benedenia. .
[0044] To optimize the robustness of multi-label prediction, the model training employs a class imbalance-aware weighted loss function:
[0045]
[0046] in, The total loss is represented by , and 'i' is the target disease number. The cross-entropy loss function is used for binary classification. This represents the true label of the i-th type of pathogen (0 for no disease, 1 for disease). The class weights are dynamically adjusted. This design improves the sensitivity for identifying rare parasites (such as Trypanosoma haematobium, which has a small sample size) and avoids bias in the model towards classes with large sample sizes.
[0047] This invention identifies latent risk patterns that cannot be captured by a single data type by fusing complementary information from numerical parameters and textual descriptions (such as pH anomaly + "increased mucus in fish" text). It utilizes three sigmoid layers to independently calculate the incidence probability of various parasites, avoiding the maintenance burden of multiple independent models, and uses weights... Amplify the loss contribution of rare diseases to address modeling biases caused by uneven pathogen distribution in actual data.
[0048] Furthermore, this invention transforms the pathogen incidence probability output by the deep learning model into an operable three-tiered risk label. For each type of target parasite, the system automatically classifies it based on a preset dynamic probability threshold range: when the incidence probability value does not exceed 0.5, it is marked as low risk, corresponding to routine monitoring recommendations; when the probability value is between 0.5 and 0.8, it is marked as medium risk, triggering early warning instructions for water quality parameter optimization and increased observation frequency; when the probability value exceeds 0.8, it is marked as high risk, forcibly initiating emergency procedures such as isolation treatment or medicated bath intervention. This classification mechanism calibrates the threshold boundaries through a combination of historical incidence data and an expert knowledge base. For example, for Trypanosoma japonicum, which has a long incubation period, the high-risk threshold is appropriately relaxed to 0.85 to ensure the biological rationality of risk decisions. The final generated risk label, together with the original probability, constitutes a structured output, forming a direct mapping chain from prediction results to aquaculture operations.
[0049] Meanwhile, the system utilizes a visualization platform built on the Shiny framework to transform prediction results into multi-dimensional interactive views. The risk matrix heatmap uses aquaculture units as rows and pathogens as columns, visually presenting the overall risk distribution using red, yellow, and green color blocks. Users can click on color blocks to view detailed prediction data. The key influencing factor analysis module, based on TabTransformer's attention weights and backpropagation of gradients in fully connected layers, quantifies the contribution of each input variable (such as abnormal dissolved oxygen values or the log keyword "white spots on fish") to the prediction results, highlighting the top five risk drivers with a descending bar chart. Simultaneously, the fused feature space is projected onto a two-dimensional plane using the UMAP dimensionality reduction algorithm, revealing the spatial clustering patterns of high-risk samples through clustering scatter plots, aiding in the identification of potential disease outbreak trends. Users can dynamically adjust the confidence threshold using a slider and refresh all views in real time; a one-click export function generates a timestamped CSV report, structurally recording sample numbers, pathogen probabilities, risk levels, and major influencing factors, directly connecting to the aquaculture farm management system for intervention scheduling.
[0050] In summary, the aquaculture disease prediction method proposed in this invention, based on multimodal data fusion, fully captures the multidimensional correlation of pathogen inducing factors in complex aquaculture environments by fusing structured water quality parameters with unstructured text data and combining a TabTransformer and BERT dual-channel modeling architecture. Compared with traditional single-modal models, the average prediction accuracy is improved, effectively solving the problems of single prediction dimension and weak generalization ability of existing technologies.
[0051] The design employs a multi-label sigmoid output layer, which simultaneously outputs the incidence probabilities of multiple parasites in a single model. Compared with the traditional single-pathogen independent modeling scheme, it reduces resource consumption and avoids compatibility issues between models, meeting the real-time needs of concurrent early warning of multiple pathogens in aquaculture sites and significantly improving the system's practicality and deployment efficiency.
[0052] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.
[0053] Furthermore, in this invention, descriptions involving terms such as "first," "second," and "a" are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0054] In this invention, unless otherwise explicitly specified and limited, the terms "connection," "fixed," etc., should be interpreted broadly. For example, "fixed" can mean a fixed connection, a detachable connection, or an integral part; it can mean a mechanical connection or an electrical connection; it can mean a direct connection or an indirect connection through an intermediate medium; it can mean the internal communication of two components or the interaction between two components, unless otherwise explicitly limited. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0055] Furthermore, the technical solutions of the various embodiments of the present invention can be combined with each other, but only if they are based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.
Claims
1. A method for predicting aquaculture diseases based on multimodal data fusion, characterized in that, Including the following steps: S1: Collect aquaculture data containing structured data and unstructured text data, wherein the structured data consists of various target water quality parameters, and the unstructured text data consists of aquaculture logs and environmental description text; S2: Input structured data into the TabTransformer model, perform column embedding and feature context relationship modeling through multi-layer Transformer encoder, and output structured feature vectors; S3: Input unstructured text data into the pre-trained BERT encoder and extract the output vector of CLS tags as text semantic features; S4: Concatenate the structured feature vector with the textual semantic features to form a fused feature and input it into the fully connected layer. Then, use the Sigmoid function to predict the incidence probability of each target disease.
2. The aquaculture disease prediction method based on multimodal data fusion as described in claim 1, characterized in that, In step S1, the structured data includes at least three water quality parameters selected from temperature, salinity, pH, dissolved oxygen, nitrite concentration, and silicate concentration.
3. The aquaculture disease prediction method based on multimodal data fusion as described in claim 1, characterized in that, In step S1, the structured data is input in CSV or Excel format and includes continuous variables and discrete categorical variables.
4. The aquaculture disease prediction method based on multimodal data fusion as described in claim 1, characterized in that, In step S3, extracting the output vector of the CLS marker as the text semantic feature specifically involves: inputting unstructured text data into the BERT encoder and taking the output vector at the corresponding position of the CLS marker as the text semantic feature.
5. The aquaculture disease prediction method based on multimodal data fusion as described in claim 1, characterized in that, In step S4, the fully connected layer is a multi-label output structure with an output dimension of . These correspond to the incidence rates of each type of target disease, among which... Let n be the number of samples and m be the total number of target disease categories.
6. The aquaculture disease prediction method based on multimodal data fusion as described in claim 1, characterized in that, Following step S4, the following step is also included: S5: Divide risk levels based on the positional relationship between the incidence probability of the target disease and the preset threshold interval; When the incidence rate of the target disease is below a preset threshold range, it is classified as low risk; When the incidence rate of the target disease is within a preset threshold range, it is classified as medium risk; When the incidence rate of the target disease is above a preset threshold range, it is classified as high risk.
7. The aquaculture disease prediction method based on multimodal data fusion as described in claim 6, characterized in that, It also includes the following steps: S6: Display the prediction results and the variable factors that affect the current prediction results through a visual interactive interface.
Citation Information
Patent Citations
Intelligent management system for mariculture
CN102124964A
Intelligent diagnosis method for aquatic diseases
CN119626518A