Method, device, computer device and storage medium for monitoring fish in a net cage

By combining transducer arrays and deep learning models, the problem of accurate fish identification in cage aquaculture was solved, achieving high-resolution fish identification and health status assessment, thus overcoming the limitations of traditional methods.

CN122135398APending Publication Date: 2026-06-02ZHEJIANG OCEAN UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG OCEAN UNIV
Filing Date
2026-02-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately identify fish and debris disturbances in cage aquaculture environments, resulting in difficulties in real-time monitoring of fish distribution, population statistics, and behavioral dynamics.

Method used

Acoustic imaging is performed using a transducer array, combined with phased array technology and a deep learning classification model. By suppressing clutter and analyzing connected components, specific fish targets are identified, and their health status is assessed in conjunction with environmental parameters.

Benefits of technology

It enables high-resolution identification and health status assessment of fish in complex aquatic environments, overcoming the limitations of optical methods and improving identification accuracy and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122135398A_ABST
    Figure CN122135398A_ABST
Patent Text Reader

Abstract

This application relates to the field of aquaculture information technology, and discloses a method, device, computer equipment, and storage medium for monitoring fish in net cages. The method involves transmitting acoustic signals into the water within the net cage via a transducer array and receiving the echoes to generate an acoustic image. Clutter suppression and connected component analysis are performed on the image to extract candidate target regions. Pixel intensity statistical features, geometric shape features, and swimming motion features of each region are extracted to form a feature vector. The feature vector is input into a pre-trained deep learning classification model to obtain a confidence score. Specific fish targets are selected based on the confidence score. Finally, their swimming motion features are calculated and visualized. This application effectively overcomes the technical challenges of distinguishing different fish species and the significant interference from the net cage structure in mixed-culture models, achieving non-invasive and accurate identification and behavioral monitoring of fish.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology in aquaculture, and in particular to a method, apparatus, computer equipment, and storage medium for monitoring fish in net cages. Background Technology

[0002] Cage aquaculture is an important method of intensive aquaculture; however, monitoring fish in this aquaculture environment has always faced significant technical challenges.

[0003] Currently, aquaculture monitoring of fish in net cages mainly relies on two types of technologies. One is optical imaging, which uses underwater cameras to acquire visual images. This method is applicable in clear water, but net cage aquaculture areas are usually rich in plankton and suspended particles, and are significantly affected by weather conditions. Changes in water turbidity and lighting often lead to a severe decline in image quality, making stable identification difficult. The other type is acoustic detection technology, which utilizes the excellent propagation characteristics of sound waves in water for detection. Traditional sonar equipment mostly uses single-beam or fixed multi-beam designs, which have limited spatial resolution and low scanning efficiency. In acoustic images acquired in complex aquaculture environments, it is difficult to effectively distinguish specific fish from net cage structures, attached organisms, suspended debris, and other interfering objects due to their similar morphological characteristics.

[0004] Existing acoustic monitoring technologies have significant limitations in processing complex echo signals. Strong reflection interference from the metal frame of the net cage, sound wave scattering caused by attached organisms, and random echoes from debris in the water can all severely interfere with the identification of target fish. Current technologies lack effective target discrimination capabilities, cannot accurately distinguish between designated fish and various types of debris interference, and struggle to monitor the accurate distribution, quantity, and behavioral dynamics of designated fish in real time. Summary of the Invention

[0005] Based on this, it is necessary to address the technical problem of low accuracy in monitoring fish in net cage aquaculture using existing technologies, and propose a method, device, computer equipment, and storage medium for monitoring fish in net cages.

[0006] In a first aspect, a method for monitoring fish in a net cage is provided, the method comprising: The transducer array transmits acoustic signals into the water body of the net cage; the transducer array operates based on the phased array principle. The transducer array receives an echo signal based on the acoustic signal and generates a current frame acoustic image based on the echo signal, wherein the pixel value of the current frame acoustic image represents the echo intensity. Clutter suppression is performed on the current frame acoustic image to obtain a foreground target image, and connected component analysis is performed on the foreground target image to extract multiple candidate target regions; A feature vector is constructed based on the features of each candidate target region; the features include: pixel intensity statistical features, geometric shape features, and swimming motion features calculated from multiple consecutive acoustic images; The feature vectors are input into a pre-trained deep learning classification model, which then outputs the confidence scores for each target region. Based on the confidence level, a specified fish target is selected from the multiple candidate target regions; Calculate the swimming characteristics of the specified fish target and collect environmental parameters of the net cage water, including at least water temperature and dissolved oxygen. The environmental parameters are coupled with the swimming motion characteristics to obtain a health status score, and based on the health status score, a health status warning prompt for the fish group is generated and displayed.

[0007] Secondly, a device for monitoring fish in a net cage is provided, the device comprising: The transmitting module is used to transmit acoustic signals into the water body of the net cage through a transducer array; the transducer array operates based on the phased array principle. A receiving module is configured to receive an echo signal based on the acoustic signal through the transducer array, and to generate a current frame acoustic image based on the echo signal, wherein the pixel values ​​of the current frame acoustic image represent the echo intensity. The extraction module is used to perform clutter suppression on the current frame acoustic image to obtain a foreground target image, and to perform connected component analysis on the foreground target image to extract multiple candidate target regions; The calculation module is used to construct a feature vector based on the features of each candidate target region; the features include: pixel intensity statistical features, geometric shape features, and swimming motion features calculated from multiple consecutive acoustic images; The classification module is used to input the feature vector into a pre-trained deep learning classification model and output the confidence score of each target region. The filtering module is used to filter specified fish targets in the plurality of candidate target regions according to the confidence level; The monitoring module is used to calculate the swimming characteristics of the specified fish target and to collect environmental parameters of the cage water, including at least water temperature and dissolved oxygen; to couple the environmental parameters with the swimming characteristics to obtain a health status score, and to generate and display a health status warning for the fish based on the health status score.

[0008] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method for monitoring fish in a net cage.

[0009] Fourthly, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described method for monitoring fish in a net cage.

[0010] Beneficial effects: High-quality acoustic imaging of the water in the net cage is achieved using a phased array transducer array. Specifically, by controlling the transmission phase of each array element to form a directional beam, the spatial concentration of acoustic energy is significantly improved, enhancing the acoustic illumination intensity on the target area. At the receiving end, a high-resolution acoustic image of the current frame is generated through digital beamforming and pulse compression processing. This effectively penetrates turbid water, overcoming the limitations of optical methods that are greatly affected by water quality, and providing a reliable data foundation for subsequent identification.

[0011] This application addresses the complex acoustic environment of cage environments, including reflections from metal structures and interference from attached organisms. It effectively separates foreground targets from background interference by performing dynamic background modeling and clutter suppression on acoustic images. Furthermore, through connected component analysis combined with multi-dimensional feature extraction, a comprehensive feature vector is constructed, including pixel intensity statistics, geometric shape, and swimming motion features. This significantly reduces false alarms caused by fixed obstacles such as cage structures, improving the accuracy of target detection in mixed-species environments.

[0012] A deep learning classification model trained with physical model enhancement is used for classification, which enables accurate identification of specified fish species in mixed-species groups. This overcomes the limitations of traditional methods that rely on a single threshold or human experience, and significantly improves the accuracy and robustness of classification. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] in: Figure 1 This is an application environment diagram of a method for monitoring fish in a net cage in one embodiment; Figure 2 This is a flowchart of a method for monitoring fish in a net cage in one embodiment; Figure 3 This is a structural block diagram of a device for monitoring fish in a net cage in one embodiment; Figure 4 This is a structural block diagram of a computer device in one embodiment. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0016] The method for monitoring fish in net cages provided in this invention can be applied to, for example... Figure 1 The application environment includes: computer equipment 1, transducer array 2, and cage 3.

[0017] The transducer array consists of multiple acoustic transducer units (i.e., array elements in this application) arranged in a specific geometric structure. Each transducer unit uses piezoelectric ceramic material as the core transducer element, which has the characteristic of converting electrical energy into acoustic energy. The transducer unit is externally encapsulated with an acoustically transparent sealed structure, and internally includes a matching layer, a backing damping layer, and an acoustic isolation layer. The units in the array are arranged at equal intervals on a rigid substrate to form a linear array or planar array configuration. The entire array is protected by a corrosion-resistant metal shell, and the internal wiring uses coaxial cables for signal transmission. The connectors are designed to be waterproof and sealed.

[0018] For example, a typical transducer array contains 64 individual transducer units arranged in an 8×8 matrix on a 0.3 m × 0.3 m substrate. Each transducer unit has a diameter of 0.03 m and a center-to-center spacing of 0.0375 m. The unit core is made of PZT-4 piezoelectric ceramic material, with a polyurethane acoustic layer covering the front and an epoxy resin backing material filling the rear. The array housing is made of 316 stainless steel, and all cables are led out through waterproof connectors.

[0019] The transducer array operates based on the principle of a phased array. In transmission mode, the monitoring device applies an electrical excitation signal with a specific phase relationship to each element. The piezoelectric elements convert electrical energy into mechanical vibrations, generating sound waves. By precisely controlling the transmission delay of each element, the sound waves radiated by each element interfere and superimpose in space, forming a sound beam with a specific directionality. In reception mode, each element converts the received sound wave signal into an electrical signal. Through signal synthesis processing, the acoustic signal in the target direction is enhanced, while interference signals from other directions are suppressed. This beamforming mechanism enables the array to achieve electronic scanning rather than mechanical rotation.

[0020] In this application, the transducer array can be deployed in the cage in the following three ways.

[0021] Option 1: Vertical column installation.

[0022] The transducer array is installed along a vertical support column of the net cage. The array is fixed to the side of the column facing inwards from the net cage in a linear arrangement, with each element distributed vertically to cover different water layers. This installation method allows the array to form a horizontal scanning fan, achieving radial layered detection coverage of the net cage. The array housing is rigidly connected to the column structure, and cables are laid along the inside or surface of the column to the water surface control unit.

[0023] Option 2: Bottom frame installation.

[0024] The transducer array is installed flat or at a specific angle on the frame structure at the bottom of the net cage. The array faces the water above the net cage to transmit and receive sound waves. When installed flat, the array is parallel to the bottom net, forming a wide-angle upward detection coverage; when installed at an angle, the array is at a specific angle to the horizontal plane, forming a directional upward oblique detection coverage. A protective structure is provided at the bottom of the array to prevent sediment accumulation from affecting the acoustic performance.

[0025] Option 3: Distributed installation of the side network.

[0026] Multiple small transducer subarrays are distributed and installed on the side mesh structure of the cage. Each subarray faces the center of the cage's interior, working collaboratively to achieve multi-angle joint detection of the entire cage's water area. The subarrays are interconnected via underwater cables, and their operating timing is coordinated by a unified synchronization control unit. This distributed deployment method effectively reduces detection blind spots.

[0027] Specifically, computer equipment 1 controls transducer array 2 to perform a wide-area scan of the net cage water area using low-frequency acoustic signals to obtain low-frequency panoramic acoustic scan data; based on the low-frequency panoramic acoustic scan data, it identifies one or more suspected fish school areas within the net cage; it determines the core location of the fish school within the fish school area; it controls transducer array 2 to use high-frequency acoustic signals to focus and detect the core location of the fish school to obtain high-frequency acoustic echo data; and based on the high-frequency acoustic echo data, it analyzes the dynamic information of the fish school area, which includes at least the number of fish and the fish density.

[0028] The computer equipment can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will now be described in detail through specific embodiments.

[0029] Please see Figure 2 As shown, Figure 2 A flowchart illustrating a method for monitoring fish in a net cage according to an embodiment of the present invention includes the following steps: S1. Transmit acoustic signals into the water body of the net cage through the transducer array.

[0030] The computer equipment controls the signal transmission module to generate electrical signals with specific waveform parameters. These signals drive the transducer array through a multi-channel power amplifier unit. The transducer array consists of multiple acoustic transducer elements (array elements) arranged in a regular geometric pattern. Each element converts the input electrical signal into mechanical vibrations, generating sound waves in the water. The computer equipment controls the timing and phase difference of the excitation signals of each element, causing the sound waves radiated by each element to coherently superimpose during spatial propagation, ultimately forming a directional transmission beam in the target water area. Phased array beamforming technology enables spatially directional focusing of sound wave energy, effectively increasing the intensity of sound wave illumination on the target area. The transducer array operates in a time-division multiplexing manner, first transmitting sound signals and then switching to receiving mode. The transducer array is calibrated using a standard sound source to compensate for sound wave attenuation in water.

[0031] S2. Receive echo signals based on acoustic signals through a transducer array, and generate an acoustic image of the current frame based on the echo signals.

[0032] In this process, the computer equipment controls the transducer array to enter signal receiving mode. Each element in the array converts the received acoustic echo signal into a weak analog electrical signal. These signals are first amplified by a low-noise amplifier, and then pass through a bandpass filter to remove out-of-band noise and interference. The conditioned analog signal is then converted into a digital signal by a high-precision analog-to-digital converter. The computer equipment applies a digital beamforming algorithm to these digital signals, adjusting the phase and amplitude relationships of the signals in each channel to form multiple receiving beams with specific spatial directivity. Simultaneously, pulse compression processing is performed on the signal of each beam channel, utilizing the matching characteristics of the transmitted and received signals to improve range resolution. Finally, the processed beam data is arranged according to the correspondence between the azimuth and range dimensions to generate a two-dimensional acoustic image matrix, where the gray value of each pixel represents the acoustic echo intensity at the corresponding spatial location.

[0033] For example, a receiver array consisting of 64 elements is used, with 128 receiving beams covering a 90-degree sector. The bandpass filter is set to a passband range of 100 kHz to 140 kHz, and the analog-to-digital converter has a sampling rate of 1 MHz. A matched filtering algorithm is used to compress the echo signal from each beam, reducing the transmitted 1-millisecond linear frequency modulated pulse to a narrow 0.05-millisecond pulse. The resulting acoustic image has a resolution of 128 pixels * 512 pixels, corresponding to the azimuth and range dimensions, respectively. Each pixel value is quantized into an 8-bit grayscale value, representing the echo intensity of that spatial cell.

[0034] S3. Perform clutter suppression on the current frame acoustic image to obtain the foreground target image, and perform connected component analysis on the foreground target image to extract multiple candidate target regions.

[0035] The process involves clutter suppression of the current frame's acoustic image. First, a dynamically updated background model is established, and the distribution characteristics of background clutter are characterized by calculating the statistical properties of each pixel in multiple consecutive frames. The current frame image is then differenced from the background model to highlight foreground regions that differ from the background features. An adaptive thresholding method is applied to the difference results, determining the binarization threshold based on the distribution characteristics of pixel values ​​in local regions to generate a preliminary binary image of the foreground target. Morphological processing is then performed on this binary image: erosion is used to eliminate isolated noise points, followed by dilation to restore the contour of the effective target. Finally, connected component labeling is performed on the processed binary image, analyzing the adjacency relationships between pixels and identifying interconnected pixel sets as independent candidate target regions, while recording the geometric attribute parameters of each region. Through dynamic background modeling and morphological processing, fixed fish cage structures and moving fish schools can be effectively separated, reducing false alarms. Opening operations are used in clutter suppression to eliminate noise, and closing operations are used in connected component analysis to fill holes in the target region.

[0036] For example, a background model is built using 10 consecutive frames of images, with an update coefficient set to 0.95. After difference operations, an adaptive thresholding algorithm based on a local window is used, with the window size set to 15×15 pixels. Morphological processing uses a 3×3 circular structuring element, performing an erosion operation followed by a dilation operation. Connectivity analysis uses the 8-neighborhood criterion, retaining only regions with an area between 10 and 1000 pixels and an aspect ratio between 1.5 and 8.0 as valid candidate regions. Finally, 25 candidate regions meeting the criteria are extracted from the original image.

[0037] S4. Construct a feature vector based on the characteristics of each candidate target region.

[0038] The computer equipment performs multi-dimensional feature extraction operations on each candidate target region. Regarding pixel intensity statistics, it calculates the grayscale distribution characteristics of all pixels within the region, including the mean reflecting overall echo intensity, the standard deviation characterizing internal structural complexity, and skewness and kurtosis describing the distribution morphology. Regarding geometric features, it measures the basic dimensional parameters and contour characteristics of the region, including projected area, boundary perimeter, and the aspect ratio of the minimum bounding rectangle, while also calculating shape descriptors such as contour rectangularity, circularity, and the Hu invariant moment set. Regarding swimming motion features, by associating trajectory data of the same target in multiple consecutive frames, it calculates its displacement vector and time derivative, including instantaneous velocity, acceleration magnitude, and rate of change of orientation angle.

[0039] Specifically, pixel intensity statistical features refer to the numerical representation extracted from the grayscale value distribution of pixels within a candidate target region in an acoustic image. These features reflect the acoustic wave reflection characteristics of the target region, where the mean represents the average echo intensity of the region and is directly related to the acoustic reflectivity of the target. The standard deviation describes the dispersion of pixel values ​​within the region, reflecting the uniformity of the target's internal structure. Skewness measures the asymmetry of the pixel value distribution and can be used to identify targets with special reflection patterns. Kurtosis characterizes the concentration of pixel value distribution and helps distinguish target regions with different texture features.

[0040] Geometric features refer to a set of parameters that describe the spatial morphology and structural characteristics of candidate target regions in acoustic images. These features include basic dimensional measures such as region area and boundary perimeter, morphological proportion parameters such as aspect ratio and compactness, and high-level shape descriptors such as Hu invariant moments. Together, these features constitute a complete mathematical description of the target's shape, effectively distinguishing different categories of target objects.

[0041] Swimming motion features refer to dynamic parameters obtained by analyzing the position and shape changes of the same target in multiple consecutive acoustic images. These features include the target's displacement vector in the image sequence, the instantaneous velocity and acceleration derived from it, and the changing pattern of the direction of motion. Extracting swimming motion features requires first establishing the trajectory correlation of the target across different frames, creating a motion sequence of the same target, and then calculating quantitative indicators of its motion state.

[0042] For example, for candidate region number 15, its pixel grayscale mean is calculated to be 125, and its standard deviation is 32. Geometric features show that the region has an area of ​​86 pixels, a perimeter of 45 pixels, an aspect ratio of 3.2, and a contour rectangularity of 0.78. By associating trajectory data from the most recent 5 frames of images, the calculated motion speed is 0.3 meters per second, the direction of motion is northeast-northeast, and the acceleration is 0.1 meters per square second. These feature parameters together constitute a 23-dimensional feature vector for subsequent classification and recognition.

[0043] S5. Input the feature vectors into the pre-trained deep learning classification model and output the confidence scores of each target region.

[0044] The computer equipment inputs the constructed feature vectors into a pre-trained deep learning classification model. This model employs a multi-layer neural network structure, using a forward propagation algorithm to perform layer-by-layer non-linear transformations on the input features. The model first normalizes the input features, then progressively extracts and combines feature information through multiple fully connected layers. Each fully connected layer contains several neurons, with connections established between nodes via adjustable weight parameters. The network ultimately generates a classification result through an output layer, representing the confidence level that the current candidate target region belongs to a specified fish species as a probability value. The entire process is based on the parameter matrix learned during the model's training phase; these parameters are determined through optimization using a large number of labeled samples.

[0045] For example, a feedforward neural network with three hidden layers is used, with 64, 32, and 16 neurons in each layer, respectively. For the input 23-dimensional feature vector, the network first performs standardization, and then sequentially passes it through each hidden layer for feature transformation. The output layer uses the sigmoid activation function to generate a confidence score between 0 and 1. For instance, when the feature vector of a candidate region is input, the model outputs a confidence score of 0.87, indicating that there is an 87% probability that the region is the specified fish target. A decision threshold of 0.75 is set; candidate regions with values ​​higher than this will be retained for subsequent processing.

[0046] S6. Select the specified fish target from multiple candidate target regions based on confidence level.

[0047] Targeted fish species refer to individuals of a specific economic fish species identified in a net cage aquaculture environment. Examples include large yellow croaker, black sea bream, or basketfish.

[0048] The computer equipment filters and determines candidate target regions based on the confidence scores output by a deep learning classification model. First, candidate regions with spatial overlap are deduplicated using a non-maximum suppression algorithm to eliminate duplicate detections. This algorithm sorts all candidate regions from highest to lowest confidence score, selecting the region with the highest confidence score as a baseline and calculating its spatial overlap with the remaining regions. When the overlap between two regions exceeds a set threshold, the region with the higher confidence score is retained while the region with the lower confidence score is removed. After deduplication, a final determination is made based on a preset confidence score threshold, retaining only candidate regions with a confidence score not lower than the threshold as confirmed fish targets. Simultaneously, the azimuth coordinates and confidence score information of each confirmed target are recorded to create a target list for subsequent processing.

[0049] For example, an overlap threshold of 60% and a confidence threshold of 0.75 are set. When processing 50 candidate regions, spatial overlap is found in eight of them. Seven duplicate regions are removed using a non-maximum suppression algorithm, retaining the region with the highest confidence. From the remaining 43 regions, 28 regions with a confidence level of 0.75 or higher are selected as the final designated fish targets. The spatial location information of these confirmed targets is recorded in the target list, including the center point coordinates and bounding rectangle parameters of each target.

[0050] S7. Calculate the swimming characteristics of the specified fish target and collect environmental parameters of the water body in the net cage.

[0051] The study calculates swimming motion characteristics of a specific fish target based on a sequence of spatial locations within multiple consecutive acoustic images. First, a target trajectory tracking sequence is established, forming a complete trajectory by associating coordinate data of the same target in adjacent frames. The velocity vector of each target, including magnitude and direction, is calculated based on the trajectory data. Acceleration parameters are obtained through differentiation, and the consistency and spatial distribution characteristics of the group's movement are statistically analyzed. For environmental parameter acquisition, sensor arrays deployed at various water layers within the net cage acquire real-time aquatic environmental data, focusing on water temperature and dissolved oxygen levels. These environmental data are synchronized with the motion characteristic data to form a complete dataset linking fish behavior and the environment.

[0052] For example, by tracking and identifying 32 specific fish targets, and analyzing the location data from the most recent 10 frames (5-second time span), the average swimming speed of the group was calculated to be 0.28 meters per second, with the main direction of movement being southeast-southeast. Sudden acceleration behavior was also recorded for three targets, with acceleration reaching 0.5 meters per second. Data from environmental monitoring sensors showed that the current water temperature was 18.5 degrees Celsius, and the dissolved oxygen level in the mid-water was 6.8 mg / L. Correlation analysis between these movement characteristics and environmental parameters revealed that when the water temperature was below 16 degrees Celsius, the average swimming speed of the fish group decreased to approximately 0.15 meters per second, while when the dissolved oxygen level was below 5 mg / L, the fish group exhibited significant surfacing and aggregation behavior.

[0053] S8. Couple the environmental parameters with the swimming motion characteristics to obtain a health status score, and generate and display a health status warning for the fish based on the health status score.

[0054] The computer equipment establishes a coupled analysis model between environmental parameters and swimming characteristics. This model uses machine learning algorithms to learn the inherent correlation between environmental parameters and fish behavioral characteristics in historical data, establishing a quantitative assessment system for health status. First, the input parameters are standardized and preprocessed to eliminate dimensional differences. Then, a weighted fusion algorithm is used to calculate a comprehensive health status score. The scoring model comprehensively considers two dimensions: environmental suitability and behavioral normality. Environmental suitability reflects the degree to which current water temperature and dissolved oxygen parameters are suitable for the target fish, while behavioral normality assesses the degree to which the fish's movement characteristics match the expected pattern. Based on the numerical range and trend of the health status score, corresponding warning levels are automatically generated and displayed in a differentiated visual format on the human-computer interaction interface.

[0055] Furthermore, the computer equipment constructs a multi-dimensional coupled analysis model through a systematic training process. First, a training sample library is established, collecting historical monitoring data covering different seasons and aquaculture stages, including environmental parameters such as water temperature and dissolved oxygen, as well as swimming characteristics such as swimming speed, movement trajectory, and group distribution. Aquaculture experts are invited to professionally score the health status of the fish population at each time period based on video recordings and growth monitoring data, forming a labeled dataset.

[0056] In the data preprocessing stage, environmental and behavioral characteristics are normalized to eliminate dimensional differences. Correlation analysis and principal component extraction are used to screen key feature indicators that significantly impact health status. Machine learning algorithms such as random forest or gradient boosting decision tree are employed to train the nonlinear mapping relationship between environmental parameters, swimming motion characteristics, and health status scores. During training, k-fold cross-validation is used to optimize model parameters, prevent overfitting, and ensure the model's generalization ability.

[0057] During the model validation phase, independent test sets are used to evaluate prediction accuracy, and model performance is quantified using metrics such as mean squared error and coefficient of determination. After deployment, a model update mechanism is established to regularly incorporate new monitoring data and expert evaluation results, maintaining the model's accuracy and adaptability through incremental learning.

[0058] For example: The monitoring showed a current water temperature of 22 degrees Celsius and a dissolved oxygen level of 5.8 mg / L. Simultaneously, the average swimming speed of the fish school decreased to 0.15 meters per second, and significant grouping and dispersal behavior was observed. Based on preset evaluation rules, the coupled analysis model calculated an environmental suitability score of 65 points and a behavioral normality score of 45 points. Through weighted calculation, the overall health status score was 52 points. A score below 60 points was set as a warning zone, thus automatically generating a yellow warning signal. The monitoring interface displayed a dual alert indicating both abnormal water environment and abnormal behavior, and recommended oxygenation measures. When the score further decreased to 40 points, it was upgraded to a red warning, and the audible and visual alarm was activated.

[0059] In one possible embodiment of this application, S8, the step of coupling the environmental parameters with the swimming characteristics to obtain a health status score, includes: S81. The environmental parameters and swimming motion characteristics are fused at the data level to construct a multi-dimensional joint feature vector; wherein the environmental parameters include at least water temperature and dissolved oxygen, and the swimming motion characteristics include at least average swimming speed, group aggregation degree, and vertical distribution depth. S82. Input the joint feature vector into a pre-trained coupled analysis model and output a health status score. S83. Generate and display health status warning prompts based on the health status score.

[0060] In step S81, the average swimming speed refers to the average distance traveled by all individuals in the fish school per unit time. This parameter is calculated by analyzing the positional changes of the same target in multiple consecutive frames of acoustic images, reflecting the overall activity level of the fish school. Vertical distribution depth refers to the average depth of the fish school's activity in the middle layer of the water, obtained by analyzing the longitudinal coordinates of the target in the acoustic image. This parameter reflects the vertical utilization characteristics of the aquatic environment by the fish school and is closely related to environmental factors such as temperature stratification and dissolved oxygen distribution. The depth coordinates of each target can be determined using the distance dimension information of the acoustic image, and then the average depth distribution and distribution range of the group can be calculated. Group aggregation degree is an indicator that quantifies the spatial concentration of the fish school, characterized by calculating the average distance between individuals and the distribution density. This parameter reflects the social behavior patterns and spatial utilization characteristics of the fish school. First, the boundary range of the fish school distribution is determined, the number of individuals per unit area is counted, and the average distance from each individual to the centroid of the group is calculated. A higher aggregation degree indicates a more concentrated fish school, while a lower value indicates a more dispersed fish school. This parameter is highly sensitive to environmental changes and group status, and is an important basis for assessing the normality of fish school behavior.

[0061] The process of generating a joint feature vector includes: first, establishing a unified timestamp alignment mechanism to ensure all parameters originate from the same monitoring period; second, cleaning the data for various parameters, removing outliers and noisy data, and retaining valid observations; and third, performing standardization to convert the original parameters of different dimensions into dimensionless values. For water temperature parameters, a standardized range is set based on the suitable temperature range for fish; for dissolved oxygen parameters, a baseline value is determined based on the saturated dissolved oxygen concentration; and for motion characteristic parameters, a statistical normalization method is used.

[0062] During the vector construction phase, features are arranged and combined according to a pre-defined sequence. Environmental parameters are placed at the beginning of the vector as basic features, while swimming motion features are placed at the end as behavioral indicators. To enhance feature expressiveness, a set of derived features is also calculated, including gradient changes in environmental parameters and temporal stability indices for motion features. All features are arranged along a fixed dimension to form a structurally unified joint feature vector. A time stamp and quality control identifier are added to each feature vector to ensure data traceability and reliability.

[0063] For example, during a monitoring cycle, environmental data were collected showing a water temperature of 22 degrees Celsius and a dissolved oxygen level of 5.8 mg / L. Simultaneously, the average swimming speed (0.15 m / s), group aggregation degree (0.35), and vertical distribution depth (2.5 m) were calculated. First, the raw data were standardized: water temperature was converted to 0.60 (range 10-30 degrees Celsius), dissolved oxygen was converted to 0.45 (range 4-8 mg / L), average speed was converted to 0.30 (range 0-0.5 m / s), aggregation degree was taken directly from the original value of 0.35, and vertical distribution depth was converted to 0.50 (range 0-5 m).

[0064] Subsequently, derived features were calculated: water temperature change rate (difference between current value and previous period) was -0.02, dissolved oxygen change rate was 0.01, and velocity fluctuation coefficient (velocity standard deviation) was 0.08. The final constructed joint feature vector is [0.60, 0.45, 0.30, 0.35, 0.50, -0.02, 0.01, 0.08], where the first five dimensions are basic features, and the last three dimensions are derived features. This vector comprehensively records the coupling information between environmental conditions and fish behavior, providing standardized data input for health status assessment.

[0065] In step S82, the computer device inputs the constructed joint feature vector into the pre-trained coupled analysis model. This model is built on a deep neural network architecture and includes an input layer, multiple hidden layers, and an output layer. The number of neurons in the input layer corresponds exactly to the dimension of the joint feature vector, ensuring complete reception of all feature information. The hidden layers use non-linear activation functions to transform and extract features layer by layer from the input features, gradually abstracting the deep correlation between environmental parameters and swimming motion features. The output layer uses a specific activation function to map the final calculation result to a predetermined scoring range.

[0066] During model operation, the input feature vector is first propagated forward, undergoing linear transformation through the weight matrices and bias vectors of each layer, followed by nonlinear mapping via activation functions. The final output layer generates a quantitative score representing health status. Simultaneously, the confidence score for this prediction is calculated; if the confidence score falls below a set threshold, it is automatically flagged as requiring manual review. The entire calculation process employs batch processing, supporting parallel evaluation of feature vectors at multiple time points.

[0067] For example, an eight-dimensional joint feature vector is input into a neural network model containing three hidden layers. The first hidden layer contains thirty-two neurons using the ReLU activation function; the second hidden layer contains sixteen neurons, also using the ReLU activation function; the third hidden layer contains eight neurons; the output layer uses the Sigmoid activation function to map the result to the range of 0 to 100. For the input vector [0.60, 0.45, 0.30, 0.35, 0.50, -0.02, 0.01, 0.08], the model outputs a health status score of 68 after forward propagation, with a confidence level of 0.92. According to the preset scoring level standard, this result is judged as a good health status. When the output score for a certain input is 45 with a confidence level of 0.88, the result is automatically marked as a sub-healthy state warning.

[0068] In step S83, the computer device automatically generates corresponding early warning information based on the numerical range and trend of the health status score. Multiple score level thresholds are preset, each level corresponding to a different warning level and display scheme. When the score falls into a specific range, the corresponding warning template is invoked to generate complete early warning information including the specific score value, main abnormal parameters, and improvement suggestions. In terms of display scheme design, a multi-level color coding mechanism is adopted, with different colors representing different warning severity levels. Simultaneously, auxiliary information such as historical score trend charts and parameter comparison analysis are provided, and the process of health status change is intuitively displayed through visualization components.

[0069] A dynamic early warning update mechanism is established to continuously monitor score trends. When a score shows a continuous decline or abnormal fluctuation, a trend warning will be generated even if the next level of warning threshold has not been reached. All warning messages are timestamped and have an expiration date to ensure timeliness and accuracy. The system also supports categorized filtering and historical querying of warning messages, allowing users to track processing progress and analyze warning patterns.

[0070] For example, when the health status score drops from 72 to 58, it indicates that the score has moved from the "good" level to the "concern" level. A yellow warning is automatically generated, displaying the current score of 58 on a yellow background in the main monitoring interface, along with two main abnormal parameters: low dissolved oxygen and decreased swimming speed. Recommended measures are displayed in the suggestion section, including checking the operation of the aeration equipment and appropriately reducing the feeding amount.

[0071] The right side of the interface synchronously updates a health status trend chart, using a curve to display the score changes over the past 24 hours, with a yellow warning point marked at the 58-point level. When the score further drops to 45 points and persists for three monitoring cycles, it upgrades to an orange warning, the interface's main color changes to orange, and a flashing reminder is added. A detailed report is also generated, including historical comparison data for each parameter and an evaluation of improvement measures, which can be exported and shared by users. All warning information is logged, forming a complete warning processing archive.

[0072] This technical solution integrates environmental parameters and swimming characteristics at the data level to construct a joint feature vector that comprehensively reflects the health status of fish populations. Based on a pre-trained coupled analysis model, it achieves quantitative assessment of health status, overcoming the limitations of traditional single-parameter assessments. A tiered early warning mechanism and visual display provide intuitive decision support for aquaculture management, significantly improving the accuracy and timeliness of health risk identification.

[0073] In one possible embodiment of this application, S1, the transmission of acoustic signals to the water body of the net cage via the transducer array, includes: S11. Obtain the predetermined target beam transmission direction; S12. Based on the target beam transmission direction and the geometric layout of the transducer array, calculate the phase compensation amount of the transmitted signal of each element in the transducer array. S13. Generate multiple baseband signals and adjust the phase of each baseband signal according to the phase compensation amount or time delay amount. S14. Amplify the power of each baseband signal after phase adjustment; S15. The amplified baseband electrical signals are synchronously applied to the corresponding array elements of the transducer array, and the acoustic wave signal in the direction of the target beam is transmitted into the water body of the cage through the coordinated work of the array elements.

[0074] In step S11, the computer device reads preset beam control parameters from the configuration parameter library. These parameters define the expected coverage area of ​​the acoustic signal in the water body of the net cage. Based on the monitoring task requirements, the appropriate beam transmission mode is selected, including fixed-direction scanning mode and dynamic tracking mode. In fixed-direction scanning mode, the beam direction is adjusted sequentially according to a preset angle sequence. For example, the beam pointing is updated every 5 seconds based on the fish school behavior model. In dynamic tracking mode, the optimal beam pointing angle is calculated in real time by combining the distribution characteristics of the specified fish school in historical monitoring data. The computer device converts the spatial position of the target area into azimuth and elevation angles relative to the transducer array through coordinate transformation, forming a complete beam control command.

[0075] For example, using a fan-shaped scanning mode, the 90-degree monitoring area is divided into six 15-degree beam directions. The computer equipment sequentially generates beam pointing commands of 0 degrees, 15 degrees, 30 degrees, up to 75 degrees. In dynamic tracking mode, based on the location of the center of the specified fish school detected in the previous cycle, the spatial coordinates that require the beam to be pointed northeast at a 45-degree azimuth angle and a 10-degree downtilt angle are calculated.

[0076] In step S12, the computer equipment calculates the required phase compensation for each array element based on the target beam transmission direction and the geometric layout of the transducer array. First, a geometric model of the array is established to determine the precise position of each element in the array coordinate system. The theoretical wavefront corresponding to the desired beam direction is determined through beam steering vector calculation. The path difference between each element and the reference element is calculated; this path difference is closely related to the target direction. Finally, based on the operating frequency of the acoustic signal and the speed of sound in water, the path difference is converted into the corresponding phase compensation. The entire calculation process employs matrix operations to ensure that the phase relationship of the excitation signals for each element meets the beamforming requirements.

[0077] For example, for an 8×8 planar array consisting of 64 elements with an element spacing of 12.5 mm, when a beam pointing 30 degrees forward needs to be formed, the computer calculates the path difference between each element and the central element, with a maximum path difference of 5.2 mm. Based on the operating frequency of 120 kHz and the speed of sound in water of 1500 m / s, the path difference is converted into a corresponding phase compensation, ranging from 0 to 130 degrees.

[0078] In step S13, the computer equipment generates multiple baseband electrical signals with specific waveform parameters using digital frequency synthesis technology. The signal generation unit generates a reference digital waveform, typically a linear frequency modulated pulse or a single-frequency pulse waveform. Based on the phase compensation amount calculated in step S12, a digital phase rotation operation is performed on each baseband signal. Phase adjustment is achieved through complex multiplication, multiplying the baseband signal by the corresponding phase rotation factor. The adjusted multiple signals maintain a strict synchronization relationship, ensuring that the phase difference between each channel signal accurately meets the design requirements. Amplitude weighting can also be performed during signal generation to control the beam sidelobe level.

[0079] For example, direct digital frequency synthesis (DDS) technology is used to generate 64 baseband signals, each in the form of a 1-millisecond linear frequency modulated pulse, with the frequency linearly varying from 110 kHz to 130 kHz. Based on the calculated phase compensation, each signal undergoes digital phase rotation with a phase adjustment accuracy of 0.1 degrees. All signals are strictly synchronized in time, with a synchronization error of less than 1 nanosecond.

[0080] In step S14, the computer equipment controls a multi-channel power amplifier to boost the power of the phase-adjusted baseband signal. Each channel is equipped with an independent power amplifier unit, which employs a linear amplification architecture to maintain the signal's phase and waveform characteristics. Real-time monitoring and protection mechanisms are implemented during the power amplification process, including output power detection, reflected power monitoring, and temperature monitoring. Impedance matching optimization is performed based on the impedance characteristics of the transducer array elements to ensure power transmission efficiency. Simultaneously, digital predistortion technology is used to compensate for the nonlinear characteristics of the power amplifier, ensuring the quality of the output signal.

[0081] For example, a 64-channel power amplifier module can be configured, with each channel outputting a peak power of 50 watts. The amplified signal voltage reaches a peak of 200 volts, effectively driving the transducer elements. The power amplifier bandwidth covers 100 kHz to 150 kHz, with in-band ripple of less than 1 dB. The output status of each channel is monitored in real time, and when an abnormal reflected power is detected in a channel, the output power is automatically adjusted to protect the equipment.

[0082] In step S15, the computer equipment simultaneously applies the amplified baseband electrical signals of each channel to the corresponding array elements of the transducer array via a synchronization control circuit. The synchronization control circuit generates precise trigger signals to ensure that the signal transmission time deviation of all channels is within the required range. Under the excitation of the electrical signals, each array element generates mechanical vibration, converting electrical energy into acoustic energy and radiating sound waves into the water. Due to the specific phase relationship of the sound waves emitted by each array element, these sound waves undergo interference during spatial propagation, forming a reinforced sound beam in the target direction and weakening each other in other directions. The spatial directional radiation of sound wave energy is achieved through the superposition effect of this constructive and destructive interference.

[0083] For example, a high-precision clock source is used to generate a synchronization trigger signal, and the time synchronization error between channels is controlled within 2 nanoseconds. When all 64 array elements work simultaneously, a transmitted sound field with a beamwidth of 15 degrees is formed in a set 30-degree direction. The sound pressure level of the sound beam in the main lobe direction is about 30 dB higher than that of a single array element, while the sound pressure level in the side lobe direction is significantly reduced, achieving effective concentration of sound energy.

[0084] This technical solution achieves directional transmission of acoustic signals in the target direction through precise phase control and multi-channel synchronous transmission technology. This beamforming method significantly improves the utilization efficiency of acoustic energy and enhances the acoustic illumination intensity of a specific monitoring area. Simultaneously, the flexible beam control capability allows for rapid adjustment of the monitoring range to adapt to the dynamic distribution characteristics of a designated fish population.

[0085] In one possible embodiment of this application, it further includes: In one possible embodiment of this application, it further includes: A1. Obtain historical movement trajectory data and current environmental parameters of the core location of the fish school. The current environmental parameters include water temperature distribution, dissolved oxygen concentration, and water flow velocity.

[0086] The computer equipment reads historical movement trajectory records of the core location of the fish swarm from the data storage unit. These records contain spatial coordinate data over a continuous time series. Simultaneously, the device collects environmental parameters in real time through a sensor network deployed at various water layers in the cage, including water temperature distribution profiles at different depths, dissolved oxygen concentration gradients, and three-dimensional water flow velocity vectors. The computer equipment performs quality checks and outlier removal on the raw data to ensure the reliability and accuracy of the input data.

[0087] A2. Extract motion feature parameters from the historical motion trajectory data; the motion feature parameters include: motion speed, rate of change of direction, and spatial distribution pattern.

[0088] The computer equipment performs feature extraction and analysis on historical movement trajectory data. By calculating the displacement changes between continuous trajectory points, the average movement speed and acceleration characteristics of the fish school are obtained. By analyzing the time series changes in the movement direction, the rate of change of direction and turning frequency are calculated. At the same time, principal component analysis is used to quantitatively describe the spatial distribution pattern of the fish school, extracting morphological parameters such as the eccentricity and major axis direction of the distribution ellipse.

[0089] A3. Input the motion feature parameters and the current environment parameters into the pre-trained fish behavior model to obtain the probability distribution of the motion direction and displacement in the next time period.

[0090] The computer equipment inputs extracted motion feature parameters and real-time environmental parameters into a pre-trained fish behavior prediction model. This model, trained on a large amount of historical observation data, establishes a nonlinear mapping relationship between environmental factors and the fish's motion response. The model output is the joint probability distribution of the fish's motion direction and displacement in the next time period, expressed as a probability density function representing the likelihood of different motion states.

[0091] For example, after the computer device inputs the current motion characteristic parameters and environmental parameters into the prediction model, it obtains the probability distribution of the fish's movement direction in the next 5 minutes: the probability of moving northeast is 45%, the probability of moving due east is 30%, and the probability of other directions is 25%; the probability distribution of displacement shows that the most likely movement distance is 8 meters, with a standard deviation of 2 meters.

[0092] The training process of the fish school behavior model is explained below.

[0093] Model training first requires constructing a high-quality training dataset. A multi-source monitoring system is deployed in a typical net cage environment to collect synchronized data on fish movement trajectories and corresponding environmental parameters over a long period. Movement trajectory data is obtained through acoustic tag tracking or video analysis, recording the fish's position coordinate sequence in three-dimensional space. Environmental parameter data includes vertical water temperature profiles, dissolved oxygen concentration gradients, light intensity variations, water flow velocity vectors, and feeding activity records. Data collection continues across multiple aquaculture cycles, covering different seasons, weather conditions, and aquaculture operation scenarios to ensure data representativeness and diversity.

[0094] The collected raw data underwent rigorous quality control to remove outliers and noise interference. Motion trajectory data was smoothed and missing values ​​were imputed. Derived feature parameters, including instantaneous velocity, acceleration, motion direction angle, turning frequency, cluster density, and diffusion index, were calculated. Environmental parameter data underwent spatiotemporal alignment and standardization, and the gradient, extreme values, and trend features of parameter changes were extracted. Simultaneously, time window features were constructed to analyze the impact of historical motion sequences on current behavior.

[0095] A deep learning architecture based on Long Short-Term Memory (LSTM) networks is employed as the foundational model, which effectively captures the temporal dependence of fish movement. The model input includes historical movement feature sequences and real-time environmental parameter vectors, and the output is the probability distribution of movement states for future time periods. The training process employs a phased strategy: first, pre-training is performed using large-scale historical data to optimize the initial network weights; subsequently, online learning is used to incrementally train the model using the latest monitoring data, enabling the model to adapt to the specific environmental characteristics of the fish cage.

[0096] A4. Adjust the direction of the acoustic signal according to the probability distribution.

[0097] The computer equipment optimizes the acoustic signal pointing strategy of the distributed transducer array based on the predicted motion probability distribution. The device first identifies regions with high-probability motion directions and calculates the spatial coordinate range of these regions. Then, it recalculates the beamforming parameters of each transducer array to ensure that the main lobe direction of the synthesized beam covers the high-probability motion regions while maintaining sufficient beamwidth to cope with motion uncertainties.

[0098] In one possible embodiment of this application, S3, the clutter suppression of the current frame acoustic image to obtain the foreground target image, includes: S31. Based on the current frame acoustic image and the previous time step background model image, recursively update the current time step background model image using the following formula: B t (x, y)=α*B t-1 (x, y)+(1-α)*I t (x, y), B t (x, y) represents the background model intensity value at pixel position (x, y) at the current time t; B t-1 (x, y) represents the background model intensity value at the same pixel location at the previous time t-1; I t (x, y) represents the intensity value of the acoustic image at pixel position (x, y) in the current frame; α is the background attenuation factor, which is a constant between 0 and 1; S32, transfer the current frame acoustic image I t Compared with the updated background model image B at the current moment t Perform pixel-by-pixel difference operations to obtain the difference image D. t ; S33, for the difference image D t Binarization is performed, and morphological opening is applied to the binarized image to finally output a clutter-suppressed foreground target image.

[0099] In step S31, the computer device uses a recursive update algorithm to dynamically maintain the background model. This algorithm is based on a weighted combination of the intensity values ​​of each pixel in the current frame of the acoustic image and the intensity values ​​of the corresponding positions in the background model at the previous time point. The background attenuation factor controls the update speed of the background model; a larger attenuation factor results in slower changes in the background model, which is beneficial for maintaining background stability. A smaller attenuation factor allows the background model to quickly adapt to environmental changes. During initialization, the background model is set to the average value of several initial frames of acoustic images to ensure that the background model reflects the initial state of the monitoring environment. In subsequent operations, the background model value of each pixel is updated in real time according to this recursive formula. Opening operations are used in clutter suppression to eliminate noise, and closing operations are used in connected component analysis to fill holes in the target region.

[0100] The determination of the background attenuation factor is an optimization process based on the characteristics of the monitoring environment and the system performance requirements. The computer equipment first analyzes the environmental stability characteristics of the monitored water area, including water flow velocity, the rate of change of suspended solids concentration, and the temporal variation characteristics of background clutter. For monitoring areas with relatively slow environmental changes, the system tends to select a larger attenuation factor value to maintain the stability of the background model. For areas with rapid environmental changes, a smaller attenuation factor value needs to be selected so that the background model can adapt to environmental changes in a timely manner. Simultaneously, the target motion characteristics are considered, including the swimming speed and frequency of occurrence of specified fish, to ensure that the background model accurately reflects environmental changes without prematurely merging into the foreground targets.

[0101] The optimal range of attenuation factor values ​​was determined through experimental testing. During the initial debugging phase, comparative experiments were conducted with different attenuation factor values, and background modeling performance data was collected under various parameters. Evaluation metrics included the stability of the background model, the completeness of foreground target detection, and the false detection rate. By analyzing the relationship between these evaluation metrics and the attenuation factor values, the optimal parameter range for a specific environment was found. The background attenuation factor α was determined based on the stability of the aquatic environment and optimized through experimental testing, typically ranging from 0.95 to 0.99.

[0102] In step S32, the computer device compares the current frame acoustic image with the updated background model image pixel by pixel. This operation is achieved by calculating the intensity difference between corresponding pixels in the two images. Differential operations can effectively separate foreground targets from the background environment because foreground targets typically have significantly different acoustic reflection characteristics than the background model. Taking the absolute value of the difference result ensures that both positive and negative differences are effectively detected. High-intensity regions in the difference image correspond to foreground targets that differ significantly from the background environment, while low-intensity regions correspond to static environments consistent with the background.

[0103] In step S33, the computer device performs binarization segmentation on the difference image, converting the grayscale difference image into a black-and-white binary image. An adaptive thresholding method is used to determine the binarization threshold, which is dynamically adjusted based on the statistical characteristics of local regions of the image. After binarization, the image contains only black and white pixels, with white pixels representing potential foreground target regions. Subsequently, a morphological opening operation is performed on the binary image. This operation first performs erosion to eliminate isolated noise points and small false targets, and then performs dilation to restore the original size and connectivity of the real target, ultimately obtaining a clutter-suppressed foreground target image. The morphological processing uses structuring elements of specific shapes and sizes to ensure that the complete shape of the real target is preserved while removing noise.

[0104] This technical solution, through dynamic background modeling and recursive update mechanisms, can effectively adapt to the slow changes in the underwater monitoring environment and eliminate the influence of stationary clutter and slowly varying interference. Differential operations accurately extract foreground targets with acoustic characteristics different from the background environment, providing a high-quality data foundation for subsequent target identification. Morphological processing further purifies the foreground target image, eliminating noise interference while maintaining the integrity of the target's shape.

[0105] In one possible embodiment of this application, the connected component analysis of the foreground target image to extract multiple candidate target regions includes: S34. Perform morphological closing operation on the foreground target image to obtain an optimized binary image; S35. Based on predefined connectivity criteria, identify all interconnected foreground pixel sets in the binary image, and record each interconnected pixel set as a labeled region and assign a region label. S36. Based on the preset target organism body shape feature screening conditions, select candidate target areas that meet the conditions from all marked areas; S37. Generate a set of multiple candidate target regions based on the filtering results; each candidate target region is set with a region label and pixel coordinate information.

[0106] In step S34, the computer device performs morphological closing operations on the foreground target image obtained through clutter suppression. This operation aims to improve the morphological integrity of the foreground target region. First, a dilation operation is performed to expand the boundary of the target region outward, filling small holes inside the region and connecting adjacent broken parts. Then, an erosion operation is performed to shrink the expanded boundary inward, restoring the approximate original size of the target. This combination of dilation and erosion effectively maintains the shape of the target while eliminating voids and boundary breaks caused by noise or imperfect threshold segmentation, forming a structurally complete binary image.

[0107] For example, a morphological closing operation is performed using a circular structuring element with a size of 3 by 3 pixels. For a certain foreground target region, the original binary image shows a 2-pixel hole in the center of the region, and multiple 1-pixel-wide breaks at the boundary. After dilation, the hole is filled, and the breaks are connected to form a continuous region. The subsequent erosion operation, while maintaining the connectivity of the region, restores the over-dilated boundary to a state close to the original contour, ultimately resulting in a morphologically complete optimized binary image.

[0108] In step S35, the computer device performs region labeling on the optimized binary image based on predefined connectivity criteria. The system uses a scanline algorithm to traverse the entire image. When an unlabeled foreground pixel is found, region growing is performed starting from that pixel, grouping all connected foreground pixels into the same region. Connectivity criteria are typically defined using eight-neighborhood or four-neighborhood. Eight-neighborhood considers the top, bottom, left, right, and four diagonal directions of a pixel, while four-neighborhood only considers the top, bottom, left, and right directions. Each identified set of connected pixels is recorded as an independent labeled region and assigned a unique region label for subsequent identification and processing.

[0109] For example, the eight-neighbor connectivity criterion is used for region labeling. During the traversal, a group of 56 interconnected foreground pixels is found, and the system groups these pixels into the same region, assigning them the label number 15. Simultaneously, another group of 32 interconnected foreground pixels is found, assigned the label number 16. The entire image is ultimately labeled with 28 independent connected regions, each containing between 15 and 210 pixels.

[0110] In step S36, the target organism body shape feature screening criteria refer to a series of quantitative judgment standards set based on the biological morphological characteristics of the specified target fish. These criteria are derived from morphological statistical analysis of a large number of target fish samples, including their typical body proportions, contour features, and spatial size range. In acoustic image analysis, the system calculates the geometric parameters of each marked region and compares them with the preset screening criteria to determine whether the region conforms to the basic morphological characteristics of the target fish. The screening criteria typically include multiple parameter thresholds, such as upper and lower limits of region area, aspect ratio range, and contour complexity index. The area threshold ensures that the target has a reasonable physical size, the aspect ratio range reflects the body proportion characteristics of the fish, and the contour complexity is used to distinguish fish from regularly shaped debris. The system comprehensively evaluates each parameter through logical operations, and only regions that meet all conditions are retained as candidate target regions.

[0111] The computer system filters all marked regions based on preset target organism shape characteristics. First, it calculates the geometric parameters of each marked region, including region area, aspect ratio of the bounding rectangle, perimeter of the outline, and shape complexity. These parameters are compared with typical morphological characteristics of the specified fish, and reasonable threshold ranges are set to exclude regions that clearly do not conform to fish morphology. The filtering criteria are determined based on statistical analysis of a large number of samples, ensuring that noise and interference are effectively filtered out while retaining the true target.

[0112] For example, the filtering criteria were set as follows: the area of ​​the region must be between 10 and 1000 pixels, and the aspect ratio of the bounding rectangle must be between 1.5:1 and 8:1. Of the 28 marked regions, 5 were excluded because their area was less than 10 pixels, 3 were excluded because their aspect ratio exceeded the range, and 2 were judged as clutter due to their overly complex shapes. Ultimately, 18 regions meeting the criteria were retained as candidate target regions.

[0113] In step S37, the computer device constructs a set of candidate target regions based on the screening results. The system establishes a complete data structure for each selected candidate target region, recording its region label, pixel coordinate set, bounding rectangle parameters, and basic geometric features. This information is organized into a standardized data format to facilitate subsequent feature extraction and target recognition processing. Simultaneously, a region indexing mechanism is established to support quick access to the corresponding region data via region labels.

[0114] For example, generate a set containing 18 candidate target regions. For candidate region number 15, record the coordinate information of its 56 pixels. The coordinates of its upper left corner of the bounding rectangle are (120, 85), the coordinates of its lower right corner are (145, 105), the region area is 420 pixels, and the aspect ratio is 3.2:1. The data of all candidate regions are stored in order of region number, forming a complete list of candidate targets.

[0115] This technical solution optimizes the integrity of the target region through morphological closing operations, accurately identifies independent candidate targets through connected component analysis, and effectively eliminates non-target interference through a screening mechanism based on biological shape features. This systematic candidate target extraction method significantly improves the accuracy and reliability of target detection, providing high-quality input data for subsequent fine recognition.

[0116] In one possible embodiment of this application, S6, the step of filtering the specified fish target in the plurality of candidate target regions according to the confidence level, includes: S61. Obtain the classification confidence of multiple candidate target regions; S62. Calculate the intersection-union ratio between candidate target regions and remove redundant regions with an overlap higher than the preset overlap threshold and a low confidence level to obtain a deduplicated set of candidate target regions. S63. Remove candidate target regions from the candidate target region set whose classification confidence is lower than a preset confidence threshold; S64. The retained candidate target areas are used as the final designated fish targets.

[0117] In step S61, the computer device obtains the classification confidence data of each candidate target region from the output of the deep learning classification model. This confidence data is stored in floating-point form, with values ​​ranging from zero to one, representing the probability that the corresponding candidate region belongs to the specified fish target. A correspondence table between confidence scores and candidate regions is established to ensure that each candidate region is associated with the correct confidence score. Simultaneously, preliminary statistical analysis is performed on the confidence data to calculate the overall confidence distribution characteristics, providing a reference for subsequent threshold setting.

[0118] In step S62, the computer device implements a non-maximum suppression algorithm to process candidate regions with spatially overlapping locations. First, the intersection-union ratio (IUGR) between any two candidate regions is calculated; this value is obtained by dividing the area of ​​their intersection by the area of ​​their union. A reasonable overlap threshold is set; when the overlap between two regions exceeds this threshold, they are identified as duplicate detection regions. A confidence-first principle is adopted, retaining the candidate region with the highest confidence among the overlapping regions and removing the remaining regions with lower confidence. This process is iterated until all highly overlapping regions have been processed.

[0119] In step S63, the computer device performs a final screening of the deduplicated candidate target regions based on a preset confidence threshold. The confidence threshold, determined based on extensive experimental data statistics and practical application requirements, is read from the configuration parameters. The confidence score of each candidate region is compared with this threshold, and all regions with confidence scores below the threshold are removed. This step ensures that only regions highly identified by the classification model as the target fish are retained, effectively reducing the false detection rate.

[0120] In step S64, the computer device identifies the candidate target regions retained after overlap deduplication and confidence level filtering as the final designated fish targets. A complete target profile is created for each confirmed target, including its spatial location information, geometric feature parameters, and corresponding confidence score. This target data is organized into a standardized output format for easy subsequent swimming motion feature analysis and visualization. Simultaneously, the target tracking sequence is updated, and a unique identifier is assigned to newly detected targets.

[0121] This technical solution effectively addresses the common problem of duplicate target detection in acoustic images through a combined screening mechanism of confidence ranking and overlapping region deduplication. The non-maximum suppression algorithm based on cross-union ratio accurately identifies and removes spatially overlapping redundant detection results, ensuring that each true target is counted only once. Confidence threshold screening ensures that only high-confidence detection results are retained, significantly improving the accuracy and reliability of target detection.

[0122] In one possible embodiment of this application, it further includes: B1. Establish a physical model of acoustic scattering for a specified fish species; the input parameters of the physical model of acoustic scattering include: fish biological parameters, acoustic incident parameters, and fish posture parameters; B2. Randomize the input parameters of the acoustic scattering physical model to generate multiple synthetic training samples that simulate the acoustic response of a specified fish under different physical states, and label each synthetic sample with a category label. B3. The synthetic training samples are mixed with the real-collected labeled acoustic image samples to form a hybrid training dataset; B4. The final deep learning classification model is obtained by training using the hybrid training dataset.

[0123] In step B1, the computer equipment constructs a physical model of sound scattering based on the principle of interaction between sound waves and fish organisms. This model comprehensively considers the acoustic characteristics of the fish's body structure, including the acoustic impedance difference between the fish's tissues and the surrounding water, the strong reflectivity of the swim bladder, and the scattering contributions from various parts of the body. The model input parameters are divided into three main categories: fish biological parameters describe the physical characteristics of the fish, acoustic incident parameters define the characteristics of the probed sound waves, and fish posture parameters determine the spatial orientation of the fish in the sound field. The model uses numerical calculation methods to solve for the scattering field of sound waves on the complex organism, simulating the theoretical echo intensity distribution. The physical model of sound scattering uses the finite element method or boundary element method to solve for the sound wave scattering field, simulating the acoustic response of the fish body and swim bladder.

[0124] For example, the established sound scattering model includes biological parameters such as fish length, width, and swim bladder size; acoustic parameters such as sound wave frequency and incident angle; and attitude parameters such as pitch and yaw angles. For a given fish with a body length of 30 centimeters, the model calculates the echo intensity distribution when a 120 kHz sound wave is incident at a 45-degree angle, generating the corresponding theoretical acoustic image.

[0125] In step B2, the computer equipment randomly generates multiple sets of input parameter combinations within a reasonable parameter range, driving the acoustic scattering physical model to generate synthetic training samples in batches. These synthetic training samples are validated using real-world environmental parameters to ensure consistency with real-world data distribution. Reasonable value ranges and distribution patterns are set for each parameter to ensure that the generated samples cover the acoustic response characteristics of the specified fish under various physiological states and spatial postures. Each synthetic sample is automatically labeled as a specified fish category, and the specific parameter values ​​used during generation are recorded. Through large-scale parameter randomization, it is possible to create samples with rare postures and special sizes that are difficult to collect in reality.

[0126] For example, one thousand sets of biological parameters are randomly generated within the range of 20 to 40 centimeters in body length, and posture parameters are randomly generated within the range of 0 to 180 degrees. Combined with different acoustic incident conditions, a total of 100,000 synthetic training samples are generated. Each sample is saved as an acoustic image and labeled with a specified fish category.

[0127] In step B3, the computer equipment mixes the synthesized training samples with the real-collected labeled samples at a predetermined ratio. First, the real samples undergo quality screening and label verification to ensure their accuracy and reliability. During the mixing process, a reasonable ratio of the two types of samples is maintained to ensure both the diversity and authenticity of the training data. Simultaneously, statistical analysis is performed on the mixed dataset to ensure a balanced distribution of samples across all feature dimensions. Data augmentation operations are also implemented, further expanding the dataset size through methods such as rotation, scaling, and adding noise.

[0128] For example, five thousand real-world labeled acoustic image samples of a specified fish species are collected and mixed with one hundred thousand synthetic samples at a ratio of 1:20. Augmentation processes such as random rotation and brightness adjustment are then applied to the mixed dataset to ultimately obtain an augmented training set containing one million and fifty thousand samples.

[0129] In step B4, the computer equipment performs end-to-end training of the deep learning classification model using a hybrid training dataset. A phased training strategy is employed: first, pre-training is performed using synthetic samples to allow the model to initially learn the basic characteristics of acoustic scattering; then, fine-tuning is performed using real samples. Dynamic learning rate adjustment and early stopping mechanisms are used during training to prevent overfitting and ensure model performance. The training process is monitored using a validation set, and the model parameters that perform best on the validation set are selected as the final model.

[0130] For example, a convolutional neural network was used as the classification model, optimized using stochastic gradient descent. The initial learning rate was set to 0.01, decaying once every ten training epochs. After one hundred training epochs, the model achieved a classification accuracy of 95.6% on the test set, and the final model was saved in a deployable format.

[0131] This technical solution effectively addresses the challenge of scarce training samples in underwater acoustic monitoring by combining physical models with real data. The acoustic scattering physical model can generate training samples covering various rare scenarios, significantly improving the model's generalization ability. The hybrid training strategy maintains the regularity of the physical model while incorporating the characteristics of real data, enabling the deep learning model to learn both theoretical laws and practical features simultaneously. This training method significantly improves the robustness and accuracy of the classification model in complex underwater environments, providing core technical support for achieving reliable intelligent monitoring of specific fish species.

[0132] Furthermore, the computer equipment establishes a continuous update mechanism for the model, maintaining the accuracy and adaptability of the classification model by periodically incorporating new monitoring data. The system first collects new sample data validated during actual monitoring, including correctly identified target fish samples and confirmed interfering species samples. Newly collected samples undergo rigorous quality control to ensure accurate labeling and data validity. The system maintains a dynamically updated sample database, recording the collection time, environmental conditions, and identification results for each sample.

[0133] The model update process employs incremental learning, continuing training based on the original model parameters. The system determines the appropriate number of training epochs and learning rate parameters according to the quantity and distribution characteristics of the new samples. During training, the system retains some original training samples to prevent the model from over-adapting to new data and forgetting existing knowledge. Simultaneously, a dynamic weight adjustment strategy is used to increase the model's attention to rare samples. After the update, the system verifies the model's performance using an independent test dataset, ensuring that the updated model improves or at least maintains the original level across all metrics.

[0134] Please see Figure 3 As shown, in one embodiment, a device for monitoring fish in a net cage is provided, the device comprising: The transmitting module 301 is used to transmit acoustic signals into the water body of the net cage through a transducer array; the transducer array operates based on the phased array principle. The receiving module 302 is configured to receive an echo signal based on the acoustic signal through the transducer array, and to generate a current frame acoustic image based on the echo signal, wherein the pixel value of the current frame acoustic image represents the echo intensity. The extraction module 303 is used to perform clutter suppression on the current frame acoustic image to obtain a foreground target image, and to perform connected component analysis on the foreground target image to extract multiple candidate target regions; The calculation module 304 is used to construct a feature vector based on the features of each candidate target region; the features include: pixel intensity statistical features, geometric shape features, and swimming motion features calculated from multiple consecutive acoustic images; The classification module 305 is used to input the feature vector into a pre-trained deep learning classification model and output the confidence level of each target region. The filtering module 306 is used to filter specified fish targets in the plurality of candidate target regions according to the confidence level; The monitoring module 307 is used to calculate the swimming characteristics of the specified fish target and to collect environmental parameters of the net cage water, including at least water temperature and dissolved oxygen; to couple the environmental parameters with the swimming characteristics to obtain a health status score; and to generate and display a health status warning for the fish based on the health status score.

[0135] In one possible embodiment, the step of coupling the environmental parameters with the swimming characteristics to obtain a health status score includes: The environmental parameters and swimming motion characteristics are fused at the data level to construct a multi-dimensional joint feature vector; wherein the environmental parameters include at least water temperature and dissolved oxygen, and the swimming motion characteristics include at least average swimming speed, group aggregation degree, and vertical distribution depth. The joint feature vectors are simultaneously input into a pre-trained coupled analysis model to output a health status score. Based on the health status score, generate and display health status warning prompts.

[0136] In one possible embodiment, the transmission of acoustic signals into the water body of the cage via the transducer array includes: Obtain the predetermined target beam transmission direction; Based on the target beam transmission direction and the geometric layout of the transducer array, calculate the phase compensation amount of the transmitted signal of each element in the transducer array. Generate multiple baseband signals, and adjust the phase of each baseband signal according to the phase compensation amount or time delay amount; Power amplification is performed on each baseband signal after phase adjustment; The amplified baseband electrical signals are synchronously applied to the corresponding array elements of the transducer array, and the acoustic wave signal in the direction of the target beam is transmitted into the water body of the cage through the coordinated work of the array elements.

[0137] In one possible embodiment, it also includes: The adjustment module is used to acquire historical movement trajectory data and current environmental parameters of the core location of the fish school. The current environmental parameters include water temperature distribution, dissolved oxygen concentration, and water flow velocity. Motion feature parameters are extracted from the historical motion trajectory data; the motion feature parameters include: motion speed, rate of change of direction, and spatial distribution pattern. The motion feature parameters and the current environment parameters are input into a pre-trained fish behavior model to obtain the probability distribution of the motion direction and displacement in the next time period. The direction of the acoustic signal is adjusted according to the probability distribution.

[0138] In one possible embodiment, the clutter suppression of the current frame acoustic image to obtain the foreground target image includes: Based on the current frame acoustic image and the previous frame background model image, the current frame background model image is recursively updated using the following formula: B t (x, y)=α*B t-1 (x, y)+(1-α)*I t (x, y), B t (x, y) represents the background model intensity value at pixel position (x, y) at the current time t; B t-1 (x, y) represents the background model intensity value at the same pixel location at the previous time t-1; I t (x, y) represents the intensity value of the acoustic image at pixel position (x, y) in the current frame; α is the background attenuation factor, which is a constant between 0 and 1; The current frame acoustic image I t Compared with the updated background model image B at the current moment t Perform pixel-by-pixel difference operations to obtain the difference image D. t ; For the difference image D t Binarization is performed, and morphological opening is applied to the binarized image to finally output a clutter-suppressed foreground target image.

[0139] In one possible embodiment, the connected component analysis of the foreground target image to extract multiple candidate target regions includes: Perform a morphological closing operation on the foreground target image to obtain an optimized binary image; Based on predefined connectivity criteria, all interconnected foreground pixel sets are identified in the binary image, and each interconnected pixel set is recorded as a labeled region and assigned a region label. Based on the preset target organism body shape characteristics screening criteria, candidate target regions that meet the criteria are selected from all marked regions; The selection results generate a set of multiple candidate target regions; each candidate target region is assigned a region label and pixel coordinate information.

[0140] In one possible embodiment, the step of filtering the specified fish target from the plurality of candidate target regions based on the confidence level includes: Obtain the classification confidence scores of multiple candidate target regions; Calculate the intersection-union ratio between candidate target regions, and remove redundant regions with an overlap higher than a preset overlap threshold and a low confidence level to obtain a deduplicated set of candidate target regions; Candidate target regions in the candidate target region set whose classification confidence is lower than a preset confidence threshold are removed; The retained candidate target areas will be used as the final designated fish targets.

[0141] In one possible embodiment, it also includes: The training module is used to establish a physical model of acoustic scattering for a specified fish species; the input parameters of the physical model of acoustic scattering include: fish biological parameters, acoustic incident parameters, and fish posture parameters. The input parameters of the acoustic scattering physical model are randomized to generate multiple synthetic training samples that simulate the acoustic response of a specified fish under different physical states, and each synthetic sample is labeled with a category label. The synthetic training samples are mixed with real-world collected labeled acoustic image samples to form a hybrid training dataset; The final deep learning classification model is obtained by training using the hybrid training dataset.

[0142] In one embodiment, a computer device is provided, the internal structure of which can be shown as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements client-side functions or steps based on a method for monitoring fish in a net cage.

[0143] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, performs the following steps: The transducer array transmits acoustic signals into the water body of the net cage; the transducer array operates based on the phased array principle. The transducer array receives an echo signal based on the acoustic signal and generates a current frame acoustic image based on the echo signal, wherein the pixel value of the current frame acoustic image represents the echo intensity. Clutter suppression is performed on the current frame acoustic image to obtain a foreground target image, and connected component analysis is performed on the foreground target image to extract multiple candidate target regions; A feature vector is constructed based on the features of each candidate target region; the features include: pixel intensity statistical features, geometric shape features, and swimming motion features calculated from multiple consecutive acoustic images; The feature vectors are input into a pre-trained deep learning classification model, which then outputs the confidence scores for each target region. Based on the confidence level, a specified fish target is selected from the multiple candidate target regions; Calculate the swimming characteristics of the specified fish target and collect environmental parameters of the net cage water, including at least water temperature and dissolved oxygen. The environmental parameters are coupled with the swimming motion characteristics to obtain a health status score, and based on the health status score, a health status warning prompt for the fish group is generated and displayed.

[0144] This application achieves accurate identification and effective monitoring of designated fish species in a mixed-culture cage environment by organically combining a series of technologies such as phased array acoustic imaging, intelligent clutter suppression, multi-dimensional feature extraction, and deep learning recognition.

[0145] In one embodiment, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program that, when executed by a processor, performs the following steps: The transducer array transmits acoustic signals into the water body of the net cage; the transducer array operates based on the phased array principle. The transducer array receives an echo signal based on the acoustic signal and generates a current frame acoustic image based on the echo signal, wherein the pixel value of the current frame acoustic image represents the echo intensity. Clutter suppression is performed on the current frame acoustic image to obtain a foreground target image, and connected component analysis is performed on the foreground target image to extract multiple candidate target regions; A feature vector is constructed based on the features of each candidate target region; the features include: pixel intensity statistical features, geometric shape features, and swimming motion features calculated from multiple consecutive acoustic images; The feature vectors are input into a pre-trained deep learning classification model, which then outputs the confidence scores for each target region. Based on the confidence level, a specified fish target is selected from the multiple candidate target regions; Calculate the swimming characteristics of the specified fish target and collect environmental parameters of the net cage water, including at least water temperature and dissolved oxygen. The environmental parameters are coupled with the swimming motion characteristics to obtain a health status score, and based on the health status score, a health status warning prompt for the fish group is generated and displayed.

[0146] This application achieves accurate identification and effective monitoring of designated fish species in a mixed-culture cage environment by organically combining a series of technologies such as phased array acoustic imaging, intelligent clutter suppression, multi-dimensional feature extraction, and deep learning recognition.

[0147] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0148] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0149] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0150] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for monitoring fish in a net cage, characterized in that, The method includes: The transducer array transmits acoustic signals into the water body of the net cage; the transducer array operates based on the phased array principle. The transducer array receives an echo signal based on the acoustic signal and generates a current frame acoustic image based on the echo signal, wherein the pixel value of the current frame acoustic image represents the echo intensity. Clutter suppression is performed on the current frame acoustic image to obtain a foreground target image, and connected component analysis is performed on the foreground target image to extract multiple candidate target regions; A feature vector is constructed based on the features of each candidate target region; the features include: pixel intensity statistical features, geometric shape features, and swimming motion features calculated from multiple consecutive acoustic images; The feature vectors are input into a pre-trained deep learning classification model, which then outputs the confidence scores for each target region. Based on the confidence level, a specified fish target is selected from the multiple candidate target regions; Calculate the swimming characteristics of the specified fish target and collect environmental parameters of the net cage water, including at least water temperature and dissolved oxygen. The environmental parameters are coupled with the swimming motion characteristics to obtain a health status score, and based on the health status score, a health status warning prompt for the fish group is generated and displayed.

2. The method for monitoring fish in a net cage according to claim 1, characterized in that, The step of coupling the environmental parameters with the swimming characteristics to obtain a health status score includes: The environmental parameters and swimming motion characteristics are fused at the data level to construct a multi-dimensional joint feature vector; wherein the environmental parameters include at least water temperature and dissolved oxygen, and the swimming motion characteristics include at least average swimming speed, group aggregation degree, and vertical distribution depth. The joint feature vectors are simultaneously input into a pre-trained coupled analysis model to output a health status score. Based on the health status score, generate and display health status warning prompts.

3. The method for monitoring fish in a net cage according to claim 1, characterized in that, The transmission of acoustic signals into the water body of the net cage via the transducer array includes: Obtain the predetermined target beam transmission direction; Based on the target beam transmission direction and the geometric layout of the transducer array, calculate the phase compensation amount of the transmitted signal of each element in the transducer array. Generate multiple baseband signals, and adjust the phase of each baseband signal according to the phase compensation amount or time delay amount; Power amplification is performed on each baseband signal after phase adjustment; The amplified baseband electrical signals are synchronously applied to the corresponding array elements of the transducer array, and the acoustic wave signal in the direction of the target beam is transmitted into the water body of the cage through the coordinated work of the array elements.

4. The method for monitoring fish in a net cage according to claim 3, characterized in that, The method further includes: The historical movement trajectory data and current environmental parameters of the core location of the fish school are obtained, including water temperature distribution, dissolved oxygen concentration and water flow velocity. Motion feature parameters are extracted from the historical motion trajectory data; the motion feature parameters include: motion speed, rate of change of direction, and spatial distribution pattern. The motion feature parameters and the current environment parameters are input into a pre-trained fish behavior model to obtain the probability distribution of the motion direction and displacement in the next time period. The direction of the acoustic signal is adjusted according to the probability distribution.

5. The method for monitoring fish in a net cage according to claim 1, characterized in that, The step of obtaining the foreground target image by performing clutter suppression on the current frame acoustic image includes: Based on the current frame acoustic image and the previous frame background model image, the current frame background model image is recursively updated using the following formula: B t (x, y)=α*B t-1 (x, y)+(1-α)*I t (x, y), B t (x, y) represents the background model intensity value at pixel position (x, y) at the current time t; B t-1 (x, y) represents the background model intensity value at the same pixel location at the previous time t-1; I t (x, y) represents the intensity value of the acoustic image at pixel position (x, y) in the current frame; α is the background attenuation factor, which is a constant between 0 and 1; The current frame acoustic image I t Compared with the updated background model image B at the current moment t Perform pixel-by-pixel difference operations to obtain the difference image D. t ; For the difference image D t Binarization is performed, and morphological opening is applied to the binarized image to finally output a clutter-suppressed foreground target image.

6. The method for monitoring fish in a net cage according to claim 1, characterized in that, The connected component analysis of the foreground target image to extract multiple candidate target regions includes: Perform a morphological closing operation on the foreground target image to obtain an optimized binary image; Based on predefined connectivity criteria, all interconnected foreground pixel sets are identified in the binary image, and each interconnected pixel set is recorded as a labeled region and assigned a region label. Based on the preset target organism body shape characteristics screening criteria, candidate target regions that meet the criteria are selected from all marked regions; The selection results generate a set of multiple candidate target regions; each candidate target region is assigned a region label and pixel coordinate information.

7. The method for monitoring fish in a net cage according to claim 1, characterized in that, The method further includes: Establish a physical model of acoustic scattering for a specified fish species; the input parameters of the physical model of acoustic scattering include: fish biological parameters, acoustic incident parameters, and fish posture parameters; The input parameters of the acoustic scattering physical model are randomized to generate multiple synthetic training samples that simulate the acoustic response of a specified fish under different physical states, and each synthetic sample is labeled with a category label. The synthetic training samples are mixed with real-world collected labeled acoustic image samples to form a hybrid training dataset; The final deep learning classification model is obtained by training using the hybrid training dataset.

8. A device for monitoring fish in a net cage, characterized in that, The device includes: The transmitting module is used to transmit acoustic signals into the water body of the net cage through a transducer array; the transducer array operates based on the phased array principle. A receiving module is configured to receive an echo signal based on the acoustic signal through the transducer array, and to generate a current frame acoustic image based on the echo signal, wherein the pixel values ​​of the current frame acoustic image represent the echo intensity. The extraction module is used to perform clutter suppression on the current frame acoustic image to obtain a foreground target image, and to perform connected component analysis on the foreground target image to extract multiple candidate target regions; The calculation module is used to construct a feature vector based on the features of each candidate target region; the features include: pixel intensity statistical features, geometric shape features, and swimming motion features calculated from multiple consecutive acoustic images; The classification module is used to input the feature vector into a pre-trained deep learning classification model and output the confidence score of each target region. The filtering module is used to filter specified fish targets in the plurality of candidate target regions according to the confidence level; The monitoring module is used to calculate the swimming characteristics of the specified fish target and to collect environmental parameters of the cage water, including at least water temperature and dissolved oxygen; to couple the environmental parameters with the swimming characteristics to obtain a health status score, and to generate and display a health status warning for the fish based on the health status score.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method for monitoring fish in a net cage as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method for monitoring fish in a net cage as described in any one of claims 1 to 7.