Construction method of passive acoustic identification system for behavior and physiological state of sound-producing fish
By constructing a passive acoustic recognition system, utilizing the natural vocalization characteristics of vocal fish and combining them with deep learning algorithms, we can achieve non-interference, accurate monitoring and early warning of the reproductive and health status of vocal fish. This solves the problem of insufficient recognition in existing technologies and provides a cost-effective solution.
Patent Information
- Application Number
- CN202610342986.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-20
- Publication Date
- 2026-06-09
- Estimated Expiration
- 2046-03-20
AI Technical Summary
Existing technologies lack the ability to accurately identify reproductive and health status in fish farming, especially in terms of integrated differentiation. Furthermore, active sonar technology is expensive and disrupts the lives of fish.
A passive acoustic recognition system for the behavior and physiological state of vocal fish was constructed. Vocalization experiments were conducted under noise exposure, reproductive behavior, and disease conditions through a sound acquisition experimental platform. Acoustic signal features were trained using convolutional neural networks and stacked autoencoders to establish a mapping model between sound features and abnormal states.
It enables undisturbed and precise monitoring of fish that make sounds, allows for in-depth and integrated assessment of reproductive potential and health levels, provides early warnings, reduces hardware costs, and is easy to apply in fish farms for the long term.
Smart Images

Figure CN121905219B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of sound recognition, specifically to a method for constructing a passive acoustic recognition system for the behavior and physiological state of vocal fish. Background Technology
[0002] Some vocal fish (such as grouper and large yellow croaker) are high-value marine aquaculture species in my country, and their health and reproductive efficiency directly determine their economic benefits during the aquaculture process. Currently, fish farms mainly rely on manual observation and experience to judge the condition of fish schools, which suffers from problems such as strong subjectivity, low efficiency, inability to provide real-time early warnings, and difficulty in detecting early or underwater abnormalities (such as potential diseases or latent estrus). Therefore, accurately grasping the behavior (such as courtship, estrus, and spawning) and physiological state (such as diseases) of vocal fish is crucial for fish farms.
[0003] To address these needs, active sonar and high-definition camera technologies have been applied to fish identification, fish density inversion, and monitoring of abnormal fish behavior. For example, existing patent CN201210063831, "A Fish Identification Method and System Based on Segmented Temporal Qualitative Features," proposes an active acoustic fish identification technology that uses a transducer to emit sound waves and identifies fish based on the characteristics of the sound signals reflected by the fish. Patent CN201510054151, "A Fish Identification Method Based on Multi-Feature and Multi-Bit Data Fusion," also proposes an active acoustic fish identification technology that identifies fish by fusing the characteristics of the sound signals reflected by the fish and multi-bit data.
[0004] However, most existing technologies remain at the level of fish species identification or simple behavioral recording, lacking detailed information on the specific economic fish species, especially in-depth research that integrates and distinguishes between "reproductive behavior" and "healthy physiological state." Furthermore, existing technologies all employ active sonar, which is not only expensive but also disrupts the normal life of the fish that emit the sonar due to frequent emitting signals.
[0005] In summary, there is a need to design a method for constructing a passive acoustic recognition system for the behavior and physiological state of vocal fish to solve the problems in the existing technology. Summary of the Invention
[0006] To address the problems in the prior art, this invention provides a method for constructing a passive acoustic recognition system for the behavior and physiological state of vocal fish, which solves the problem that existing recognition methods lack depth in recognizing the "reproductive state" and "health state" of vocal fish and cannot distinguish them in an integrated manner.
[0007] To achieve the above objectives, the present invention adopts the following technical solution:
[0008] The method for constructing a passive acoustic recognition system for the behavior and physiological state of vocalizing fish includes the following steps:
[0009] S1. Construct a sound acquisition experimental platform for fish that emit sound;
[0010] S2. Conduct natural vocalization experiments, vocalization experiments of fish exhibiting different reproductive behaviors, and vocalization experiments of fish exhibiting disease states on the experimental platform under noise exposure.
[0011] S3. Acquire acoustic signals during each sound emission test in step S2, and extract features from the acoustic signals.
[0012] S4. Using convolutional neural networks and stacked autoencoders to train the features, a mapping model of "sound features - state anomalies" for the vocal fish is obtained.
[0013] In some embodiments of the present invention, the specific steps for feature extraction of the acoustic signal in step S3 include:
[0014] S31. Perform low-pass filtering on the acoustic signal to suppress background noise;
[0015] S32. The filtered acoustic signal is segmented to obtain a syllable scale map;
[0016] S33. Calculate the time-domain and frequency-domain characteristics of the separated syllables to obtain the linear frequency cepstral coefficients and Mel frequency cepstral coefficients; output the spectrogram of the syllable based on the linear frequency cepstral coefficients and Mel frequency cepstral coefficients.
[0017] In some embodiments of the present invention, step S4 specifically includes:
[0018] The spectrogram is trained using a convolutional neural network, and the syllable scale map is trained using a stacked autoencoder. The training results are then sent to a fusion module for decision-making, and a mapping model of the "sound features - state anomalies" of the vocal fish is output.
[0019] In some embodiments of the present invention, the step of obtaining the syllable scale diagram in step S32 is as follows:
[0020] S321. Determine the peak value of the highest amplitude and set the position of the nth syllable to tn. The amplitude calculation formula for this point is as follows:
[0021] ;
[0022] Where Y is the amplitude, f is the frequency index, and t is the time index;
[0023] S322. Tracing adjacent peaks between t>tn and t<tn, the start and end times of the nth syllable are defined as t. n -t s and t n+t e ;
[0024] S323, Time [t] n -t s, t n +t e The trajectory of ] is saved as the nth byte, and the following elements are deleted from the matrix:
[0025] ;
[0026] S324. Repeat the above steps until the spectrum is finished.
[0027] In some embodiments of the present invention, the calculation process of the linear frequency cepstral coefficients is as follows:
[0028] Frequency scaling from Hertz to Mel scale is achieved using multiple triangular bandpass filters:
[0029] ;
[0030] The linear frequency cepstral coefficients are derived from the logarithmic magnitude output Y of each triangular filter. i Obtained from the discrete cosine transform:
[0031] ;
[0032] Where N is the number of cepstral coefficients and M represents the number of triangular filters.
[0033] In some embodiments of the present invention, the Mel frequency cepstral coefficients are calculated from the logarithmic amplitude discrete Fourier transform:
[0034] ;
[0035] Where K represents the discrete Fourier transform amplitude coefficient (|X) i The quantity of |).
[0036] In some embodiments of the present invention, the sound acquisition test platform includes a test pool, a hydrophone, a hydroacoustic transducer, an electrocardiogram / vital signs monitoring system, a camera, a power amplifier, a signal generator, a data acquisition card, a controller, a life support system, and a programmable lighting system; the hydroacoustic transducer is located at one or both ends of the test pool and is connected to the power amplifier and the signal generator; the hydrophone is used to acquire the sound signals emitted by the sounding fish in different states and is connected to the data acquisition card and the controller for real-time recording of acoustic spectrum characteristics; the life support system includes a physical filtration module, a biological filtration module, a water quality conditioning and stabilization module, and an oxygenation and gas exchange module.
[0037] In some embodiments of the present invention, the natural vocalization experiment in step S2 includes the following steps:
[0038] The sound acquisition test platform was debugged to ensure that the water quality remained stable within the suitable temperature and salinity range for the fish that were making the sound.
[0039] Several healthy, non-sexually mature vocal fish were transferred to the experimental tank and left to stand for one day to allow them to adapt to the new environment.
[0040] Maintain standard environmental conditions in the experimental pool, use a hydroacoustic transducer to emit real environmental noise from the aquaculture farm, and simulate the real living conditions of the vocal fish; continuously record the spontaneous vocalization activities of the vocal fish in a stable environment without specific interference for n consecutive days.
[0041] Record the changes in sound during and before feeding behavior, repeating for at least 5 feeding cycles; during non-feeding times, simulate feeding actions but do not provide food, and record the fish's reactions and sounds.
[0042] In some embodiments of the present invention, the vocalization experiment of different reproductive behaviors in step S2 includes the following steps:
[0043] Multiple male and mature female vocal fish were introduced into the experimental pond; they were then kept in the experimental pond for n weeks to allow them to become familiar with the environment, establish a stable social hierarchy, and accept feeding.
[0044] Maintain standard environmental conditions in the experimental pool, use a hydroacoustic transducer to emit real environmental noise from the aquaculture farm, and simulate the real living conditions of the fish that make sounds; continuously record the acoustic and video data of the fish that make sounds for n days to establish a behavioral and acoustic baseline during the non-breeding season.
[0045] Adjust the life support system in the pool, activate the temperature regulation program, and set it to a suitable temperature for the vocal fish to spawn; activate the programmable lighting system to simulate the complete lunar cycle from new moon to full moon.
[0046] The vocalizations of male vocal fish during courtship displays, competition, pairing of male and female vocal fish, and spawning of female vocal fish were continuously recorded.
[0047] In some embodiments of the present invention, the fish vocalization test in the disease state in step S2 includes the following steps:
[0048] Fish infected with a specific pathogen that produce vocalizations were transferred to an experimental tank, and a standardized disease severity rating scale was used to score the infected fish; behavioral and physiological characteristics were collected, including subtle changes in respiratory rate, activity level, and appetite.
[0049] Continuously record the vocalization information and physiological indicators such as heart rate, respiratory rate, and metabolism of the vocalizing fish;
[0050] At least one fish is randomly selected each week for non-lethal comprehensive sampling for pathogen quantification, immune and stress index analysis; a comprehensive necropsy and histopathological examination are performed on the dead fish to determine the disease status of the individual.
[0051] The technical solution of the present invention has the following technical effects compared with the prior art:
[0052] This invention utilizes the unique acoustic characteristics of vocal fish to construct a passive acoustic monitoring mode, completely eliminating the interference and stress caused by external energy injection to the fish population, achieving truly zero-intrusion monitoring of the farmed organisms; enabling in-depth, integrated assessment and early warning of the reproductive potential and health level of vocal fish; and robustly extracting key acoustic features from complex environmental noise through advanced signal processing and deep learning algorithms, allowing this technology to operate and be promoted economically and reliably in actual aquaculture environments for a long time. Attached Figure Description
[0053] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1 This is a schematic diagram of the sound acquisition test platform structure shown in this embodiment.
[0055] Figure 2 This is a schematic diagram of the acoustic data processing flow shown in this embodiment.
[0056] Figure labels: 1. Test tank 1; 2. Sound-emitting fish; 3. Hydrophone; 4. Hydroacoustic transducer; 5. ECG / vital signs monitoring system; 6. Camera; 7. Power amplifier; 8. Signal generator; 9. Data acquisition card; 10. Controller; 11. Life support system; 12. Programmable lighting system. Detailed Implementation
[0057] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0058] In the description of this application, it should be understood that the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this application.
[0059] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0060] In the description of this application, it should be noted that, unless otherwise expressly specified and limited, the terms "installation," "connection," and "joining" should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections, direct connections, or indirect connections through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances.
[0061] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can include direct contact between the first and second features, or contact between the first and second features through another feature between them. Furthermore, "above," "over," and "on top" of the second feature includes the first feature directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature includes the first feature directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature.
[0062] The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the invention. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0063] Example 1: In this example, the grouper is used as the experimental sound-producing fish 2 object for construction. The specific method is as follows.
[0064] The method for constructing a passive acoustic recognition system for the behavior and physiological state of vocalizing fish includes the following steps:
[0065] S1. Construct a sound acquisition experimental platform for fish that emit sound;
[0066] Reference Figure 1 As shown, the sound acquisition test platform includes a test water tank 1, a hydrophone 3, a hydroacoustic transducer 4, an electrocardiogram / vital signs monitoring system 5, a camera 6, a power amplifier 7, a signal generator 8, a data acquisition card 9, a controller 10, a life support system 11, and a programmable lighting system 12.
[0067] Specifically, test pool 1 was made of acrylic or glass with a diameter of about 3 meters to ensure that the grouper could swim normally in the pool and that the equipment could easily track and record data. Before the test, the inner wall of the pool and all parts in contact with the water were thoroughly cleaned and disinfected. After standing for 15-30 minutes, it was rinsed thoroughly with plenty of clean water until there was no obvious chemical smell, and then placed in a ventilated place to air dry naturally.
[0068] The hydrophone 3 is installed in the test pool 1 to collect the acoustic signals emitted by the grouper in different states, and records the acoustic spectrum characteristics in real time in conjunction with the data acquisition card 9 and analysis software; the underwater acoustic transducer 4 is located at one or both ends of the test pool 1, and is connected to the power amplifier 7 and the signal generator 8. The underwater acoustic transducer 4 is used to generate noise of different frequencies, sound pressure levels and signal forms.
[0069] A high-speed camera 6 was installed at the top of the pool to acquire behavioral data such as the swimming trajectory, feeding behavior, and changes in group structure of the grouper during the experiment; an online electrocardiogram acquisition device was installed on the side wall or bottom of the pool to record physiological indicators such as heart rate, respiratory rate, and metabolism.
[0070] The sound acquisition test platform is controlled by the data acquisition card 9 and various devices in a unified manner, realizing the time synchronization and visualization of multi-source data such as behavior, vocalization, and electrocardiogram, and providing a standardized experimental environment for the vocalization of grouper in different states.
[0071] The life support system 11 includes a physical filtration module, a biological filtration module, a water quality conditioning and stabilization module, and an oxygenation and gas exchange module. The physical filtration module quickly removes suspended solids (feces, uneaten food, and mucus) from the water. The biological filtration module converts toxic ammonia nitrogen into less toxic nitrates, establishing and maintaining a strong nitrifying bacterial community. The water quality conditioning and stabilization module automatically maintains stable key parameters such as water temperature, salinity, pH, and alkalinity. The oxygenation and gas exchange module maintains high dissolved oxygen levels in the pool and removes carbon dioxide.
[0072] The programmable lighting system 12 includes a main lighting system and a moonlight simulation system. The main lighting system consists of high-power full-spectrum LED panels or strips, used to simulate the color temperature and spectrum of natural sunlight to promote the normal physiological rhythms and health of fish. The moonlight simulation system consists of independent arrays of low-intensity, cool white light or specific wavelength blue light LEDs, used to simulate moonlight at night, and is equipped with programmable functions to automatically adjust the intensity of nighttime illumination according to the set lunar phase (new moon, first quarter moon, full moon, last quarter moon).
[0073] S2. Conduct natural vocalization experiments, vocalization experiments of fish exhibiting different reproductive behaviors, and vocalization experiments of fish exhibiting disease states on the experimental platform.
[0074] S21. Measuring Natural Background Noise: Natural background noise is recorded in a fishless pond in the aquaculture farm using a hydrophone 3. The obtained audio file is preprocessed, and the data is read using audio processing software. The sampled data is output as a text file. Then, the text file is processed by the host computer software of the function signal generator 8 and converted into a recognizable RAF format. The RAF file is transmitted from the computer to the function signal generator 8, and after the amplitude is modified, it is sent to the power amplifier 7. The power amplifier 7 amplifies the signal and transmits it to the underwater acoustic transducer 4, thereby simulating the background noise in the pond under the real aquaculture farm environment.
[0075] Natural vocalization experiment:
[0076] This experiment aims to systematically monitor the spontaneous vocalization behavior of grouper in a controlled experimental tank (Pool 1) under simulated natural conditions. The focus is on the impact of non-breeding season daily behaviors (such as feeding, exploration, and territorial behavior) and basic environmental factors (such as photoperiod) on their vocalization activities. High-sensitivity hydrophones (Pool 3) deployed in the tank, combined with underwater high-definition video, are used to simultaneously record acoustic signals and behavioral performance, analyzing the correlation between acoustic activity and external stimuli.
[0077] The main steps of the experiment include:
[0078] S22. Debug the life support system 11 in the platform pool to ensure that the water quality is stable within the temperature and salinity range suitable for grouper; calibrate the time of all data recording devices (hydrophone 3, video camera 6) to ensure millisecond-level synchronization accuracy;
[0079] S23. Transfer 5-10 healthy, non-sexually mature groupers to experimental pond 1 and let them stand for 1 day to allow the vocal fish 2 to adapt to the new environment.
[0080] S24. Maintain standard environmental conditions in experimental pond 1 (constant temperature, standard light cycle, and routine feeding), and use hydroacoustic transducer 4 to emit real environmental noise from the aquaculture farm to simulate the real living conditions of the vocal fish 2; continuously record the spontaneous vocalization activities of the grouper in a stable environment without specific interference for n consecutive days (e.g., 7 consecutive days).
[0081] S25. Record the changes in sound during and before feeding behavior (fighting, swallowing, etc.) and repeat for at least 5 feeding cycles; during non-feeding times, simulate feeding actions (e.g., inserting a tool into the water) but do not provide food, and record the fish's reactions and sounds.
[0082] Grouper vocalization experiment under different reproductive behaviors:
[0083] This experiment aims to simulate and induce the complete grouper reproductive cycle in a platform pool, thereby systematically collecting acoustic data precisely synchronized with key behaviors such as courtship, competition, mating, and spawning. The goal is to establish an acoustic database of grouper during the reproductive period that is precisely synchronized with high-definition video behavioral tags, and to identify and quantify the acoustic differences in the sounds produced during different stages of reproductive behavior (courtship display, male competition, mating, and spawning).
[0084] The experimental steps are as follows:
[0085] S22. Introduce 2 male and 3-5 mature female grouper into experimental pond 1; preliminarily confirm that the grouper's gonads are well developed by ultrasound imaging or non-invasive hormone testing; conduct acclimatization feeding in experimental pond 1 for n weeks (e.g., 1 week) to familiarize them with the environment, establish a stable social hierarchy, and accept feeding.
[0086] S23. Maintain standard environmental conditions in experimental pond 1 (constant temperature, standard light cycle, and routine feeding), and use underwater acoustic transducer 4 to emit real environmental noise from the aquaculture farm to simulate the real living conditions of the vocal fish 2; continuously record acoustic and video data of the vocal fish 2 for n days to establish a behavioral and acoustic baseline during the non-breeding season.
[0087] S24. Adjust the life support system 11 of the water tank, start the temperature regulation program, and adjust it to the suitable temperature for grouper spawning; start the programmable lighting system 12 to simulate the complete lunar cycle from new moon to full moon.
[0088] S25. Continuously record the vocalizations during male grouper courtship displays, male competition, male-female pairing, and female spawning. Focus on acoustic activity before, during, and after these behaviors.
[0089] Grouper vocalization test under disease conditions:
[0090] Different pathogens (bacteria, viruses) possess varying tissue tropism, pathogenic mechanisms, and pathological processes, leading to specific physiological and behavioral changes in the host. These changes affect the function of vocal organs (such as inflammation of the swim bladder-muscle junction), respiratory patterns (gill damage), overall activity levels, and motivational states, ultimately leaving unique "fingerprints" in acoustic signals. This experiment established a database of dynamic changes in acoustic characteristics of grouper during different stages of disease development (incubation period, symptomatic period, and recovery / death period) in the infection of three typical diseases (bacterial, parasitic, and viral). The aim was to identify acoustic indicators with disease-specific and early warning value, providing a theoretical basis for developing a precision fish disease diagnostic system based on passive acoustics.
[0091] The experimental steps are as follows:
[0092] S22. Select approximately 10 healthy grouper and infect them with common grouper bacteria strains (such as Vibrio alginolyticus, Streptococcus agalactiae, Edwardsiella tarda, etc.) or common viruses (neurone necrosis virus, iridovirus disease) via injection. Use a pathogen load analyzer and a portable biochemical analyzer to determine the specific bacteria or virus infecting the grouper.
[0093] S23. Transfer the vocal fish 2 infected with a specific pathogen to the test tank 1, and score the infected vocal fish 2 using a standardized disease severity rating scale; closely monitor its behavior and physiological characteristics, including respiratory rate, activity level, and subtle changes in appetite; collect data on its behavior and physiological characteristics, including respiratory rate, activity level, and subtle changes in appetite.
[0094] S24. Continuously record the vocalization information and physiological indicators such as heart rate, respiratory rate, and metabolism of the grouper.
[0095] S25. Randomly select 1-2 fish each week for non-lethal comprehensive sampling (blood, mucus, feces) for pathogen quantification, immune and stress index analysis; perform a comprehensive necropsy and histopathological examination on the dead fish 2 to determine the disease status of the individual.
[0096] After the test, thoroughly clean and disinfect the inner wall of the platform pool and all parts that come into contact with the water: drain the pool, wipe the inner wall with a soft cloth or sponge to remove residue, add appropriate disinfectant (such as hydrogen peroxide, potassium permanganate, etc.) in proportion, let stand for 15-30 minutes, rinse thoroughly with plenty of clean water until there is no obvious chemical smell, and place in a ventilated place to air dry naturally.
[0097] S3. Acquire acoustic signals during each sound emission test in step S2, and extract features from the acoustic signals.
[0098] S31. The sounds produced by grouper are mostly low-frequency signals, usually below 500 Hz; therefore, the sound signal received by hydrophone 3 is first subjected to a 500 Hz low-pass filter to suppress high-frequency background noise.
[0099] S32. The filtered acoustic signal is segmented to obtain a syllable scale map; in the signal segmentation stage, the audio signal is automatically segmented into syllables to separate each acoustic sound. This process should be performed using a short-time Fourier transform to convert the audio signal into a frequency domain signal. ,
[0100] S321. Determine the peak value of the highest amplitude and set the position of the nth syllable to tn. The amplitude calculation formula for this point is as follows:
[0101] ;
[0102] Where Y is the amplitude, f is the frequency index, and t is the time index;
[0103] S322. Tracing adjacent peaks between t>tn and t<tn, the start and end times of the nth syllable are defined as t. n -t s and t n +t e ;
[0104] S323, Time [t] n -t s, t n +t e The trajectory of ] is saved as the nth byte, and the following elements are deleted from the matrix:
[0105] ;
[0106] S324. Repeat the above steps until the spectrum is finished.
[0107] S33. Calculate the time-domain and frequency-domain characteristics of the separated syllables to obtain the linear frequency cepstral coefficients and Mel frequency cepstral coefficients; output the spectrogram of the syllable based on the linear frequency cepstral coefficients and Mel frequency cepstral coefficients.
[0108] After segmentation, time-domain and frequency-domain features were computed for each syllable to generate information useful for taxonomic classification. Syllables were spectrally characterized using linear frequency cepstral coefficients (LFCC) and Mel frequency cepstral coefficients (MFCC) to preserve information in the low- and high-frequency regions of the signal. These coefficients were similarly computed based on short-time analysis, using a 25-millisecond Hamming window with 45% overlap for both features.
[0109] S331, The calculation process of the linear frequency cepstral coefficients is as follows:
[0110] Frequency scaling from Hertz to Mel scale is achieved using multiple triangular bandpass filters:
[0111] ;
[0112] The linear frequency cepstral coefficients are derived from the logarithmic magnitude output Y of each triangular filter. i Obtained from the discrete cosine transform:
[0113] ;
[0114] Where N is the number of cepstral coefficients and M represents the number of triangular filters.
[0115] S332, The Mel frequency cepstral coefficients are calculated from the logarithmic amplitude discrete Fourier transform:
[0116] ;
[0117] Where K represents the discrete Fourier transform amplitude coefficient (|X) i The quantity of |).
[0118] S333, Step S333 also includes calculating Shannon entropy, the specific calculation process is as follows:
[0119] Perform a Fast Fourier Transform on each frame of sound to obtain the power spectrum S(f) of that frame.
[0120] Divide the energy S(f) of each frequency point by the total energy ∑S(f) of the frame to obtain the "probability" P(f) of that frequency.
[0121] .
[0122] Calculate the Shannon entropy using the probabilities of all frequency points in the frame:
[0123] .
[0124] S4. Using convolutional neural networks and stacked autoencoders to train the features, a mapping model of "sound features - state anomalies" for vocal fish is obtained.
[0125] Convolutional Neural Networks (CNNs) have proven to be one of the most efficient deep learning algorithms for image classification and recognition. The three main components of a CNN are convolution, pooling, and activation. Convolutional layers convolve the input data with a set of convolutional kernels or filter impulse responses to produce feature maps; pooling layers operate independently on each feature map to reduce its size, with nonlinear max pooling being one of the most commonly used operations; activation layers consist of a nonlinear operation, similar to signal conditioning, controlling the range of its input. Rectified Linear Unified Activation (ReLU) is a commonly used activation function. The convolution-pooling-activation process is repeated until a set of features with sufficient discriminative power and conciseness is obtained. The feature vectors are then fed into a fully connected layer with an activation function (usually sigmoid) for decision-making.
[0126] Stacked autoencoders consist of multiple layers of unsupervised autoencoders followed by a fully connected layer using SoftMax or Sigmoid as the activation function. SAE training is a two-step process, involving unsupervised learning followed by supervised learning. Unlabeled samples are fed into the first layer of the SAE for unsupervised training. The autoencoder layers are stacked such that the parameter vector from the (k-1)th layer is used as the input to train the kth AE layer. Once the AE is trained, labeled data is fed into the fully connected layer to train its parameters.
[0127] S41. The spectrogram is trained using a convolutional neural network, and the syllable scale map is trained using a stacked autoencoder. The Shannon entropy is calculated for each frame of audio using a sliding window method to obtain a sequence of entropy values that change over time, and then the sequence is input into the convolutional neural network.
[0128] S42. The training results are sent to the fusion module for decision-making. The fusion module examines the output of each model to find locally consistent, discriminative, and representative patterns. The feature types used in the MMDL component employ the PatternNet fusion mechanism.
[0129] S43. Output the mapping model of "sound characteristics - abnormal state" of vocal fish.
[0130] This invention utilizes the unique acoustic characteristics of the sound-emitting fish 2 to construct a passive acoustic monitoring mode, which has the following technical advantages compared to existing technologies:
[0131] (1) Non-intrusive and precise monitoring: This invention utilizes the vocal characteristics of grouper to propose a "passive acoustic monitoring" technology. By collecting and analyzing the sounds actively emitted by the grouper itself, its state can be determined, thereby completely eliminating the interference and stress caused by external energy injection to the fish population, and achieving truly zero-intrusion monitoring of the farmed species.
[0132] (2) Integrated identification of behavior and psychological state: This invention utilizes the key biological characteristic of grouper as a vocal fish, and establishes a mapping model between its specific vocal characteristics (such as the pulse pattern of courtship calls and the spectral characteristics of pathological sounds) and its internal "behavioral states" (courtship, spawning, etc.) and "physiological states" (disease, stress, etc.). By analyzing the acoustic signals actively expressed by these organisms, we can directly understand their internal state and achieve in-depth, integrated assessment and early warning of reproductive potential and health level.
[0133] (3) Provides a cost-effective and easy-to-deploy practical solution: Based on the principle of passive listening, the core of the system consists only of hydrophone 3 and intelligent analysis module, which significantly reduces hardware costs and has low energy consumption, making it easy to install and maintain. Through advanced signal processing and deep learning algorithms, key acoustic features can be robustly extracted from complex environmental noise, enabling the technology to operate and be promoted economically and reliably in actual aquaculture environments for a long time.
[0134] In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.
[0135] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for constructing a passive acoustic recognition system for the behavior and physiological state of vocalizing fish, characterized in that, Includes the following steps: S1. Construct a sound acquisition experimental platform for vocal fish; S2. Conduct natural vocalization experiments, vocalization experiments of fish exhibiting different reproductive behaviors, and vocalization experiments of fish under disease conditions on the experimental platform. S3. Acquire acoustic signals during each sound emission test in step S2, and extract features from the acoustic signals. S4. Using convolutional neural networks and stacked autoencoders to train the features, a mapping model of "sound features - state anomalies" for vocal fish is obtained. The natural vocalization test in step S2 includes the following steps: The sound acquisition test platform was debugged to ensure that the water quality remained stable within the suitable temperature and salinity range for the fish that were making the sound. Several healthy, non-sexually mature vocal fish were transferred to the experimental tank and left to stand for one day to allow them to adapt to the new environment. Maintain standard environmental conditions in the experimental pool, use a hydroacoustic transducer to emit real environmental noise from the aquaculture farm, and simulate the real living conditions of the vocal fish; continuously record the spontaneous vocalization activities of the vocal fish in a stable environment without specific interference for n consecutive days. Record the changes in sound during and before feeding behavior, repeating at least 5 feeding cycles; during non-feeding times, simulate feeding actions but do not provide food, and record the fish's reactions and sounds; The vocalization test for different reproductive behaviors in step S2 includes the following steps: Multiple male and mature female vocal fish were introduced into the experimental pond; they were then kept in the experimental pond for n weeks to allow them to become familiar with the environment, establish a stable social hierarchy, and accept feeding. Maintain standard environmental conditions in the experimental pool, use a hydroacoustic transducer to emit real environmental noise from the aquaculture farm, and simulate the real living conditions of the fish that make sounds; continuously record the acoustic and video data of the fish that make sounds for n days to establish a behavioral and acoustic baseline during the non-breeding season. Adjust the life support system in the pool, activate the temperature regulation program, and set it to a suitable temperature for the vocal fish to spawn; activate the programmable lighting system to simulate the complete lunar cycle from new moon to full moon. The study continuously recorded the courtship displays of male vocal fish, male competition, pairing of male and female vocal fish, and vocalizations during the spawning and reproductive behaviors of female vocal fish. The fish vocalization test in the disease state in step S2 includes the following steps: Fish infected with a specific pathogen that produce vocalizations were transferred to an experimental tank, and a standardized disease severity rating scale was used to score the infected fish; behavioral and physiological characteristics were collected, including subtle changes in respiratory rate, activity level, and appetite. Continuously record the vocalization information, heart rate, respiratory rate, and metabolic physiological indicators of the vocalizing fish; At least one fish is randomly selected each week for non-lethal comprehensive sampling for pathogen quantification, immune and stress index analysis; a comprehensive necropsy and histopathological examination are performed on the dead fish to determine the disease status of the individual.
2. The construction method according to claim 1, characterized in that, The specific steps for feature extraction of the acoustic signal in step S3 include: S31. Perform low-pass filtering on the acoustic signal to suppress background noise; S32. The filtered acoustic signal is segmented to obtain a syllable scale map; S33. Calculate the time-domain and frequency-domain characteristics of the separated syllables to obtain the linear frequency cepstral coefficients and Mel frequency cepstral coefficients; output the spectrogram of the syllable based on the linear frequency cepstral coefficients and Mel frequency cepstral coefficients.
3. The construction method according to claim 2, characterized in that, Step S4 specifically includes: The spectrogram is trained using a convolutional neural network, and the syllable scale map is trained using a stacked autoencoder. The training results are then sent to a fusion module for decision-making, and a mapping model of "sound features - state anomalies" of the vocal fish is output.
4. The construction method according to claim 2, characterized in that, The steps for obtaining the syllable scale map in step S32 are as follows: S321. Determine the peak value of the highest amplitude and set the position of the nth syllable to tn. The amplitude calculation formula for this point is as follows: ; Where Y is the amplitude, f is the frequency index, and t is the time index; S322. Tracing adjacent peaks between t>tn and t<tn, the start and end times of the nth syllable are defined as t. n -t s and t n +t e ; S323, Time [t] n -t s, t n +t e The trajectory of ] is saved as the nth byte, and the following elements are deleted from the matrix: ; S324. Repeat the above steps until the spectrum is finished.
5. The construction method according to claim 2, characterized in that, The calculation process for the linear frequency cepstral coefficients is as follows: Frequency scaling from Hertz to Mel scale is achieved using multiple triangular bandpass filters: ; The linear frequency cepstral coefficients are derived from the logarithmic magnitude output Y of each triangular filter. i Obtained from the discrete cosine transform: ; Where N is the number of cepstral coefficients and M represents the number of triangular filters.
6. The construction method according to claim 2, characterized in that, The Mel frequency cepstral coefficients are calculated from the logarithmic amplitude discrete Fourier transform: ; Where K represents the discrete Fourier transform amplitude coefficient (|X) i The quantity of |).
7. The construction method according to claim 1, characterized in that, The sound acquisition test platform includes a test tank, a hydrophone, a hydroacoustic transducer, an electrocardiogram / vital signs monitoring system, a camera, a power amplifier, a signal generator, a data acquisition card, a controller, a life support system, and a programmable lighting system. The hydroacoustic transducer is located at one or both ends of the test tank and is connected to the power amplifier and the signal generator. The hydrophone is used to acquire the sound signals emitted by the fish in different states and is connected to the data acquisition card and the controller for real-time recording of the acoustic spectrum characteristics. The life support system includes a physical filtration module, a biological filtration module, a water quality regulation and stabilization module, and an oxygenation and gas exchange module.
Citation Information
Patent Citations
Fish identification method and system based on segmented time-domain centroid features
CN103308918A
Fish identification method with multi-feature and multidirectional data fused
CN104714237A
Statistical characteristic analysis method based on sound signal characteristics of large-scale net cage culture fish school
CN115394313A
Acoustic intelligent identification method for behavior state of penaeus vannamei boone and model building method
CN121054005A