Crab stress identification method, device, equipment, medium and product

By combining bio-voiceprint signals, vibration signals, and behavioral images, cross-modal multi-head attention alignment technology can accurately identify crab stress responses, solving the accuracy problem of crab stress identification in existing technologies and enabling timely early warning and control.

CN121834446APending Publication Date: 2026-04-10浙江云翎信息技术有限公司 +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing crab farming technologies struggle to accurately identify and regulate crab stress responses in complex environments, especially under low light or turbid water conditions. Visual and water quality monitoring methods are ineffective in identifying changes in crab stress behavior.

Method used

By acquiring the bioacoustic signals of crabs, the vibration signals of water bodies, and the behavioral images of crabs, and using cross-modal multi-head attention alignment technology, combined with a one-dimensional convolutional neural network and a SlowFast dual-time base network, features are extracted and stress indices are calculated to achieve accurate stress identification and control.

Benefits of technology

It enables accurate identification of crab stress responses under different environmental conditions, provides timely early warning and control, reduces the false negative rate, improves identification accuracy, and reduces the impact of stress responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834446A_ABST
    Figure CN121834446A_ABST
Patent Text Reader

Abstract

The invention discloses a crab stress recognition method and device, equipment, a medium and a product, and relates to the field of crab culture, and the method comprises the steps: obtaining a biological voiceprint signal of a crab, a vibration signal of a water body, and a behavior image of the crab; performing feature extraction according to the biological voiceprint signal to obtain a two-dimensional spectrogram and voiceprint behavior features; performing feature extraction according to the vibration signal to obtain a vibration fusion feature; performing behavior detection according to the behavior image to obtain an image behavior probability; performing cross-modal multi-head attention alignment according to the two-dimensional spectrogram and the behavior image to obtain a voiceprint image matching result; determining a behavior probability vector according to the voiceprint image matching result, the voiceprint behavior characteristics, the vibration fusion characteristics and the image behavior probability; and determining a stress index of the crabs according to the behavior probability vector, and regulating and controlling the environment according to the stress index. According to the invention, accurate stress identification can be realized, and accurate regulation and control can be realized in an early stage.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of crab breeding, and in particular to a crab stress recognition method, device, equipment, medium and product. BACKGROUND

[0002] In crab breeding, stress and fighting are common phenomena, especially in intensive breeding environments. These behaviors not only cause health problems in crabs, such as reduced appetite, difficulty molting, and growth retardation, but also can trigger group diseases, further exacerbating mortality rates. Stress reactions are often accompanied by behaviors such as escape and attack in crabs. If these changes cannot be identified and addressed in a timely manner, it may lead to deterioration of the entire breeding environment, and even affect the economic benefits and sustainable development of breeding.

[0003] Currently, crab behavior monitoring technology mainly relies on visual monitoring and water quality sensors, but these technologies have many limitations. Visual recognition technology, such as underwater cameras and image analysis methods, is often limited by environmental lighting, water turbidity, and intensive breeding. In low light or turbid water, the recognition accuracy of the camera is greatly reduced. In addition, even if advanced target detection networks such as YOLO and Transformer are introduced, it is still difficult to meet the monitoring needs of intensive breeding and small crabs, especially in detecting early stress reactions, fighting or molting, etc. Small changes often result in missed detection or false identification.

[0004] Water quality monitoring technology, although it can provide real-time environmental data such as pH, dissolved oxygen, and conductivity, cannot accurately determine changes in individual crab behavior. Even with multiple sensors, water quality monitoring can only provide macroscopic environmental data and cannot directly solve the problem of crab stress reaction recognition. Overall, existing technologies can only assist in breeding management to a certain extent and are difficult to accurately identify and regulate crab stress behavior, especially in complex breeding environments, real-time monitoring and intervention measures are particularly difficult. SUMMARY

[0005] The purpose of the present application is to provide a crab stress recognition method, device, equipment, medium and product, which can realize accurate stress recognition and accurate regulation in the early stage.

[0006] To achieve the above purpose, the present application provides the following solutions: In a first aspect, the present application provides a crab stress recognition method, comprising: obtaining biological voiceprint signals of crabs, vibration signals of water bodies, and behavior images of crabs; extracting features according to the biological voiceprint signals to obtain two-dimensional spectrograms and voiceprint behavior features; characteristic extraction is performed on the vibration signal to obtain vibration fusion characteristics; behavior detection is performed on the behavior image to obtain an image behavior probability; cross-modal multi-head attention alignment is performed on the two-dimensional acoustic spectrogram and the behavior image to obtain a voiceprint image matching result; A behavior probability vector is determined according to the voiceprint image matching result, the voiceprint behavior characteristics, the vibration fusion characteristics, and the image behavior probability. A stress index of the crabs is determined according to the behavior probability vector, and environmental regulation is performed according to the stress index.

[0007] In an embodiment, before the two-dimensional acoustic spectrogram and the voiceprint behavior characteristics are obtained by performing characteristic extraction on the biological voiceprint signal, the following steps are further included: The biological voiceprint signal is subjected to band-pass filtering and normalization processing; The vibration signal is subjected to high-pass filtering using a high-pass filter; The behavior image is subjected to resolution processing and time specification; The processed biological voiceprint signal and the processed vibration signal are aligned with the processed behavior image through a synchronization protocol of a clock pulse.

[0008] In an embodiment, the two-dimensional acoustic spectrogram and the voiceprint behavior characteristics are obtained by performing characteristic extraction on the biological voiceprint signal, specifically including: The two-dimensional acoustic spectrogram is obtained by processing the biological voiceprint signal using a Hamming window and a short-time Fourier transform; The voiceprint behavior characteristics are obtained by performing feature recognition on the biological voiceprint signal using a one-dimensional convolutional neural network.

[0009] In an embodiment, the vibration fusion characteristics are obtained by performing characteristic extraction on the vibration signal, specifically including: The long-term trend is extracted using the slow channel of the SlowFast dual-time base network; The detail feature is extracted using the fast channel of the SlowFast dual-time base network; The long-term trend and the detail feature are fused using a cross-channel feature fusion function of the SlowFast dual-time base network to obtain vibration fusion characteristics.

[0010] In an embodiment, the voiceprint image matching result is obtained by performing cross-modal multi-head attention alignment on the two-dimensional acoustic spectrogram and the behavior image, specifically including: The two-dimensional acoustic spectrogram and the behavior image are respectively represented by sequences; The two-dimensional sound spectrogram represented by the sequence table is used as a query vector, and the behavior image represented by the sequence table is used as a key-value pair, cross-modal matching is performed, and a weighted combination feature corresponding to each sound node is obtained; The weighted combination feature corresponding to each sound node is subjected to semantic relationship enhancement by using a multi-head attention mechanism, and a voiceprint image matching result is obtained.

[0011] In an embodiment, a behavior probability vector is determined according to the voiceprint image matching result, the voiceprint behavior feature, the vibration fusion feature, and the image behavior probability, specifically including: The voiceprint image matching result is represented as a semantic enhancement vector Z; The semantic enhancement vector Z, the voiceprint behavior feature , the vibration fusion feature , and the image behavior probability are input into a fusion reasoning module to determine the behavior probability vector according to the following formula: p wherein represents a feature concatenation operation, is a weight matrix, is a bias term, and Softmax is used for normalization to a probability distribution.

[0012] In a second aspect, the present application provides a crab stress recognition device, comprising: An acquisition module is configured to acquire a biological voiceprint signal of a crab, a vibration signal of a water body, and a behavior image of the crab; A first feature extraction module is configured to perform feature extraction according to the biological voiceprint signal to obtain a two-dimensional sound spectrogram and a voiceprint behavior feature; A second feature extraction module is configured to perform feature extraction according to the vibration signal to obtain a vibration fusion feature; A third feature extraction module is configured to perform behavior detection according to the behavior image to obtain an image behavior probability; A cross-modal multi-head attention alignment module is configured to perform cross-modal multi-head attention alignment according to the two-dimensional sound spectrogram and the behavior image to obtain a voiceprint image matching result; A reasoning module is configured to determine a behavior probability vector according to the voiceprint image matching result, the voiceprint behavior feature, the vibration fusion feature, and the image behavior probability; A recognition and regulation module is configured to determine a stress index of the crab according to the behavior probability vector, and to regulate the environment according to the stress index.

[0013] ​​In a third aspect, the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the crab stress recognition method.

[0014] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the crab stress recognition method.

[0015] In a fifth aspect, the present application provides a computer program product, comprising a computer program, wherein the computer program is executed by a processor to implement the crab stress recognition method.

[0016] According to the specific embodiments provided by the present application, the following technical effects are disclosed: The present application provides a crab stress recognition method, device, equipment, medium and product, respectively process the biological voiceprint signal, vibration signal and behavior image, and then align the cross-modal multi-head attention, so as to determine the stress index of crabs. Through the fusion of voiceprint, vibration and image data, the stress response of crabs in different environmental conditions is accurately identified. The behavior probability vector is determined through the voiceprint image matching result, voiceprint behavior feature, vibration fusion feature and image behavior probability, and finally the stress index is calculated, so as to realize the timely early warning and regulation and control of the behavior change of crabs. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0018] Figure 1 An application environment diagram of a crab stress recognition method in an embodiment of the present application; Figure 2 A flowchart of a crab stress recognition method provided in an embodiment of the present application; Figure 3 A crab stress recognition method diagram; Figure 4 A functional module diagram of a crab stress recognition device provided in an embodiment of the present application; Figure 5 A crab stress recognition device diagram; Figure 6 A structure diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0019] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0020] MEMS (hydrophone): a micro underwater microphone processed by micro-electro-mechanical system, high sensitivity, small size, arrayable, used to capture weak sound signals such as crab clawing, crawling, friction, etc.

[0021] 1D-CNN (one-dimensional convolutional neural network): a variant of convolutional neural network, specially used for processing one-dimensional sequence data, extracting local features by sliding one-dimensional convolution kernel on the sequence, such as time series, text, audio signals, etc.

[0022] SlowFast dual-time-base network: a neural network structure for video and time series data analysis, its core idea is to extract features of different time scales through two branch channels with different sampling rates.

[0023] The above purposes, features and advantages of the present application can be more obvious and easy to understand. The present application will be further described in detail below with reference to the drawings and specific embodiments.

[0024] The crab stress recognition method provided by the embodiments of the present application can be applied to, for example Figure 1The terminal 102 communicates with the server 104 through the network. The data storage system can store the data required by the server 104 for processing. The data storage system can be separately arranged, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the biological voiceprint signal to be processed, the vibration signal of the water body, and the behavior image of the crab to the server 104. After receiving the biological voiceprint signal to be processed, the vibration signal of the water body, and the behavior image of the crab, the server 104 performs feature extraction on the biological voiceprint signal to be processed, the vibration signal of the water body, and the behavior image of the crab, obtains a two-dimensional sound spectrum graph and a voiceprint behavior feature; performs feature extraction on the vibration signal to obtain a vibration fusion feature; performs behavior detection on the behavior image to obtain an image behavior probability; performs cross-modal multi-head attention alignment on the two-dimensional sound spectrum graph and the behavior image to obtain a voiceprint image matching result; determines a behavior probability vector according to the voiceprint image matching result, the voiceprint behavior feature, the vibration fusion feature, and the image behavior probability; determines a stress index of the crab according to the behavior probability vector, and performs environmental regulation according to the stress index. The server 104 can feed back the obtained stress index to the terminal 102. In addition, in some embodiments, the crab stress recognition method can also be implemented by the server 104 or the terminal 102 alone, for example, the terminal 102 can directly perform crab stress recognition on the biological voiceprint signal to be processed, the vibration signal of the water body, and the behavior image of the crab, or the server 104 can obtain the biological voiceprint signal to be processed, the vibration signal of the water body, and the behavior image of the crab from the data storage system and perform crab stress recognition.

[0025] The terminal 102 can be, but is not limited to, various desktop computers, notebook computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things device can be a smart speaker, a smart television, a smart air conditioner, a smart vehicle device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 104 can be implemented by a single server or a server cluster composed of multiple servers, and can also be a cloud server.

[0026] In an exemplary embodiment, as shown in Figure 2 and Figure 3 A crab stress recognition method is provided, which is executed by a computer device, specifically, can be executed by a terminal or a server, or both, in the embodiment of the present application, the method is applied to the server 104 in Figure 1 The method includes the following steps.

[0027] Step 201: Obtain the biological acoustic fingerprint signal of crabs, the vibration signal of the water body, and the behavior image of the crabs.

[0028] Step 202: Feature extraction is performed according to the biological acoustic fingerprint signal to obtain a two-dimensional acoustic spectrogram and an acoustic fingerprint behavior feature.

[0029] Step 203: Feature extraction is performed according to the vibration signal to obtain a vibration fusion feature.

[0030] Step 204: Behavior detection is performed according to the behavior image to obtain an image behavior probability.

[0031] Step 205: Cross-modal multi-head attention alignment is performed according to the two-dimensional acoustic spectrogram and the behavior image to obtain an acoustic fingerprint image matching result.

[0032] Step 206: A behavior probability vector is determined according to the acoustic fingerprint image matching result, the acoustic fingerprint behavior feature, the vibration fusion feature, and the image behavior probability.

[0033] Step 207: A stress index of the crabs is determined according to the behavior probability vector, and environmental regulation is performed according to the stress index.

[0034] The biological acoustic fingerprint signal, the vibration signal, and the behavior image are respectively processed, and then cross-modal multi-head attention alignment is performed, so as to determine the stress index of the crabs. Through the fusion of acoustic fingerprint, vibration, and image data, the stress response of the crabs under different environmental conditions is accurately identified. The behavior probability vector is determined according to the acoustic fingerprint image matching result, the acoustic fingerprint behavior feature, the vibration fusion feature, and the image behavior probability, and finally the stress index is calculated, so as to realize timely early warning and regulation of the behavior change of the crabs.

[0035] In an exemplary embodiment, before the feature extraction according to the biological acoustic fingerprint signal to obtain the two-dimensional acoustic spectrogram and the acoustic fingerprint behavior feature, it further includes: performing band-pass filtering and normalization processing on the biological acoustic fingerprint signal; using a high-pass filter to perform high-pass filtering on the vibration signal; performing resolution processing and time specification on the behavior image; and aligning the processed biological acoustic fingerprint signal and the processed vibration signal with the processed behavior image through a clock pulse synchronization protocol.

[0036] Specifically, the power hum below 50Hz and the high-frequency noise above 12kHz in the biological acoustic fingerprint signal for acoustic fingerprint recognition are removed by a band-pass filter through the formula: wherein, is the original signal of the sound wave, is the band-pass filter, is the convolution operation, is the filtered signal, irrelevant frequency band noise is removed, and only the target acoustic fingerprint frequency band is retained.

[0037] Then the bio-acoustic signal is peak normalized, and the maximum absolute amplitude of a sound is scaled to 1. The purpose is to uniformly adjust the amplitude of the original acoustic signal to a certain range, eliminate the interference between different signals due to intensity difference, and facilitate subsequent feature extraction and analysis. The formula is used to achieve is the original time domain sample, is the normalized sample, is the peak in the sound, is a very small positive number, and 10 -9 Prevent zero denominator when silent.

[0038] The 0Hz~ 5Hz slow drift of the tidal, water pump low frequency jitter collected by the 4th order high pass filter is removed, and the sampling rate is 1000Hz. The differential equation is: , represents the current time (the nth sampling point) vibration original signal input, represents the previous 1 sampling point, represents the current time (the nth sampling point) vibration signal output after high pass filter processing; represents the previous 1 sampling point vibration signal output after high pass filter processing, , , is the numerator coefficient of the filter (the weighting coefficient of the input signal), , is the denominator coefficient of the filter (the feedback coefficient of the output signal), to filter. The coefficients a and b are obtained by using the butter and filtfilt in the python library scipy.

[0039] The original resolution of the behavior image data collected by the dual-mode visual camera is 1080p, which is directly reduced to 720p while maintaining the aspect ratio. The local MCU sets a unified time source according to the system internal time, and the sound and vibration frames are synchronized with the image through the clock pulse synchronization protocol.

[0040] In an exemplary embodiment, according to the bio-acoustic signal, a two-dimensional acoustic spectrum and an acoustic behavior feature are extracted, specifically including: according to the bio-acoustic signal, a Hamming window and a short-time Fourier transform are used for processing to obtain a two-dimensional acoustic spectrum; according to the bio-acoustic signal, a one-dimensional convolutional neural network is used for feature recognition to obtain an acoustic behavior feature.

[0041] Specifically, a Hamming window is used for every 50 milliseconds of sound to reduce spectral leakage, and the formula is: , ​denotes the discrete samples of the original bio-acoustic signal in time domain, denotes the discrete samples of the Hamming window, which is a commonly used window function and its expression is: is used to weight the original acoustic signal to reduce the spectral leakage caused by signal truncation. denotes the acoustic signal samples after Hamming window processing, which is the result of point-by-point multiplication of the original acoustic signal samples and the Hamming window samples, denotes the total number of samples of each frame of acoustic signal.

[0042] Then do the short-time Fourier transform: , is the frequency spectrum point, is a complex vector reflecting the frequency components. Finally, calculate its amplitude spectrum, formula: , ε is a very small positive number, take 10 -9 , to prevent 0 in the logarithm, the amplitude spectrum is converted into an image, and a two-dimensional acoustic spectrogram is obtained by stacking multiple frames and using pseudo-color mapping.

[0043] The filtered and normalized bio-acoustic signal is further processed by 1D-CNN to extract features such as “clawing”, “crawling”, and “struggling”. The key algorithm of the acoustic fingerprint-CNN module is the feature extraction process, which specifically includes: , is the sequence of filtered and normalized bio-acoustic signals, denotes the acoustic behavior features, is the weight of the first layer of convolution, used to perform preliminary feature extraction on the acoustic signal after short-time Fourier transform, is the weight of the second layer of convolution, used to abstract and combine the features output by the first layer of convolution to a deeper level and extract more complex acoustic features, is a one-dimensional convolution operation; is a SiLU (Sigmoid Linear Unit) activation that combines the characteristics of linear and nonlinear functions.

[0044] In an exemplary embodiment, according to the vibration signal, a vibration fusion feature is extracted, specifically including: using a slow channel of a SlowFast dual-time base network to extract a long-time trend; using a fast channel of the SlowFast dual-time base network to extract a detail feature; and using a cross-channel feature fusion function of the SlowFast dual-time base network to fuse the long-time trend and the detail feature to obtain the vibration fusion feature.

[0045] Specifically, for vibration data, a SlowFast dual timebase network is used, one slow lane looks at long-term trends, formula: , is the original vibration signal sequence, is the time downsampling factor, is the convolution network extraction function on the slow lane, is the long-term trend feature vector. One fast lane grabs details and hits. is the convolution network extraction function on the slow lane, is the fast-changing time series feature vector.

[0046] Fuse the outputs of the two lanes for joint decision-making: , is the cross-channel feature fusion function, which can finally be sent to the classifier or used for cross-modal reasoning, is the vibration fusion feature.

[0047] In an exemplary embodiment, behavior detection is performed according to the behavior image, obtaining an image behavior probability, specifically including: using a lightweight YOLO-Nano-NIR to detect crab shape and posture, and outputting "normal / molting / fighting" probability. Input: infrared + visible light images obtained by a Sony IMX462 dual-mode camera, object frame coordinates , class probability vector: Each detection frame outputs a probability vector after softmax: , is the feature map of each frame , is the weight of the classification fully connected layer, d is the bias term, is the state probability distribution of the crab in the frame.

[0048] Softmax classification formula: , is the probability of the i-th behavior in the behavior probability vector, such as represents the probability of "feeding", represents "molting". is the embedding representation of the i-th modal feature (such as image, voiceprint, vibration), is the embedding representation of the j-th modal feature (such as another modal), j≠i, indicating "cross-modal" features.

[0049] In an exemplary embodiment, cross-modal multi-head attention alignment is performed according to the two-dimensional acoustic spectrogram and the behavior image to obtain a voiceprint image matching result, specifically including: respectively performing sequence representation on the two-dimensional acoustic spectrogram and the behavior image; performing cross-modal matching on the sequence represented two-dimensional acoustic spectrogram as a query vector and the sequence represented behavior image as a key-value pair to obtain a weighted combined feature corresponding to each sound node; and performing semantic relationship enhancement on the weighted combined feature corresponding to each sound node by using a multi-head attention mechanism to obtain the voiceprint image matching result.

[0050] wherein the cross-modal multi-head attention alignment represents the two-dimensional acoustic spectrogram after short-time Fourier transform as a sequence: , each represents a sound feature at time t. Each frame of image is also encoded into a feature vector sequence: , each represents a behavior image feature of all crabs in the picture at time t. Taking sound as a query vector and picture as a key-value pair, cross-modal matching is attempted at each time point, through the following formula: , is a query matrix after linear transformation of the sound feature; is a key matrix, is a value matrix, and are linear transformation matrices of the voiceprint sequence feature and the image frame sequence feature in the multi-head attention mechanism, which are used to map the original modal features to a unified attention subspace, thereby realizing cross-modal alignment and correlation modeling. d k is a vector dimension for scaling; the output is a picture weighted combined feature corresponding to each sound node, which is used to determine which crab in the picture is moving when the “plop” sound occurs. The multi-head mechanism is used to enhance the ability to capture different semantic relationships, and the calculation formula is: , the calculation formula of each head . Wherein each , , is a linear mapping matrix of the first i attention head, which acts on the voiceprint sequence (Query) and the image sequence (Key, Value) respectively, realizes cross-modal alignment in multiple groups of different semantic subspaces, and thereby enhances the understanding ability of the model to complex behaviors (such as stress, fighting, molting, etc.). The above two formulas are the core expressions of the multi-head attention mechanism in the Transformer architecture, and the main purpose is to let the model understand the relationship between the input data from multiple angles at the same time. MultiHead(Q, K, V) represents that the outputs of multiple single-head attentions (head1 to head h are spliced together (Concat), and then a linear transformation The final multi-head attention result is obtained.

[0051] In the calculation formula of the head, each head multiplies different linear transformation weight matrices to obtain different projection spaces, and "focuses" on the relevance in the input data from different angles. This way can help the model capture attention relationships at different semantic levels or feature dimensions.

[0052] 1D-CNN extracted deep semantic features (e.g., pinching, struggling) is another dimension of representation, which does not directly participate in "alignment matching", but will be integrated as a high semantic layer input in the behavior probability vector determination, used to enhance the inference model's understanding ability of complex behaviors (e.g., stress).

[0053] In an exemplary embodiment, according to the voiceprint image matching result, the voiceprint behavior feature, the vibration fusion feature, and the image behavior probability, the behavior probability vector is determined, specifically comprising: representing the voiceprint image matching result as a semantic enhancement vector Z.

[0054] The semantic enhancement vector , the voiceprint behavior feature , the vibration fusion feature , and the image behavior probability are input into the fusion inference module to determine the behavior probability vector p according to the following formula: .

[0055] Wherein represents a feature concatenation operation, is a weight matrix, d is a bias term, and Softmax is used for normalization to a probability distribution.

[0056] Specifically, the voiceprint image matching result is represented as a semantic enhancement vector Z, and the voiceprint behavior feature , the vibration fusion feature , and the image behavior probability are input into the fusion inference module; the fusion inference module includes a multi-layer perception or Transformer network structure, which is used to fuse multi-modal semantic features and output the final behavior probability vector p = [p 摄食 , p 脱壳 , p 争斗 , p 应激 ].

[0057] The image behavior probability vector As a high-confidence behavior prior, it provides constraints and guidance on behavior semantics for the fusion module, enhancing the discriminability and robustness of the final behavior reasoning.

[0058] The matched information is then fused into a more reliable behavior probability vector p = [p 摄食 , p 脱壳 , p 争斗 , p 应激 ], and the reasoning formula is: , f ac is the motion feature of the crab in the picture, f vi is the visual position information, f vd is the visual dynamic feature, and Z represents the matching semantic context vector output by the cross-modal alignment module, which integrates the temporal correspondence between the voiceprint and the image to guide the semantic enhancement of the multi-modal behavior feature. Z is a multi-modal semantic fusion vector, which is the output result from the cross-modal alignment module, used to describe the correspondence between the voiceprint and the image behavior. The behavior probability vector p is obtained by inputting Z and the image motion feature (f ac ), visual position feature (f vi ), and visual dynamic feature (f vd ) into the reasoning network. The introduction of Z makes the prediction of p have semantic consistency and context reasoning ability.

[0059] In an exemplary embodiment, determining a stress index of the crab according to the behavior probability vector and regulating the environment according to the stress index includes the following steps.

[0060] Stress index key formula: , the stress index is calculated according to the formula. The result is a real number between 0 and 1.

[0061] The system checks the stress index every second, and according to the set threshold, the stress state is divided into three levels: "no warning", "yellow light warning" and "red light strong intervention", and is mapped into the corresponding control matrix. If the stress index is low, no "warning" record is triggered; if the stress index is slightly high, "warning" is triggered, and the mild control is first done to speed up the oxygen pump by 10-20%; a small amount of "calming bait" is thrown to distract the crab; if the stress index is very high, the white light is turned off / the infrared light is turned on, the light stimulation is reduced; strong intervention is immediately executed and the message is pushed through WeChat / SMS. The specific environmental regulation is shown in Table 1.

[0062] Table 1 Environmental regulation parameter table

[0063] The system can dynamically adjust the alert threshold based on the SEI history distribution in the past period of time: set the initial early warning threshold θ0=0.45, and the strong intervention threshold θ1=0.65; the average value μ and the standard deviation σ of SEI are calculated every hour; if μ rises significantly, the system automatically adjusts θ0and θ1up by ξ (such as 0.05); if the SEI maintains a low level for 3 consecutive periods, the system gradually adjusts the threshold to avoid excessive regulation; the update strategy can be written as: where θ t is the early warning value at time t, λ is the learning rate, μ ref is the reference baseline.

[0064] The application can accurately identify the stress response of crabs under different environmental conditions through the fusion of voice prints, vibrations and visual signals. Through multi-modal data analysis, timely prediction and early warning of crab behavior changes are realized. While reducing energy consumption, high-precision real-time monitoring and regulation are maintained to achieve precise aquaculture environment control and reduce the stress response of crabs.

[0065] As shown in Figure 5 , the system applied by the crab stress identification method of the application is as follows. The system includes a MEMS hydrophone array, a piezoelectric film vibration sheet and a low-illumination dual-mode visual camera. The hydrophone array is used to capture the biological voice print signals of crabs (such as the sound of crab crawling, fighting, and pinching); the piezoelectric film vibration sheet captures the weak vibration changes in the water; and the visual camera is used to supplement the detection of the behavior characteristics of crabs. The local MCU transmits the data to the edge computing node AI gateway SoC through the LoRa-Mesh low-power long-distance network after simple processing of the original data, and performs real-time analysis and decision-making.

[0066] On the hardware structure, the application is a multi-modal sound-vibration-image integrated IP68 sensing node; the hydrophone array is arranged with a corrosion coupling member. On the method, the "voice print-vibration-image" three-mode attention collaborative network, the stress index calculation formula and the adaptive threshold updating strategy. The environmental parameter (light, oxygen, bait) closed-loop control logic. Thus, the following effects are achieved. The acoustic signal appears before the visible behavior occurs (for example, crab pinching friction), so the stress index can give a warning 2 to 3 times earlier than the pure visual scheme. Weak light and turbidity adaptability even if the vision fails, sound-vibration can still maintain more than 90% recognition rate; avoid turning on the light at night to cause disturbance and energy consumption. The power of the MEMS hydrophone and the LoRa-Mesh single node is less than or equal to 1.5W. Scalability, the same framework can be migrated to shrimp, swimming crab and other crustaceans; without changing the platform, only adjust the voice print model.

[0067] Based on the same inventive concept, the application further provides a crab stress recognition device for implementing the crab stress recognition method described above. The device provides a solution to the implementation scheme as described in the above method, so the specific limitations in one or more crab stress recognition device embodiments provided below can refer to the limitations of the crab stress recognition method described above, and will not be repeated here.

[0068] In an exemplary embodiment, as shown in Figure 4 a crab stress recognition device is provided, comprising: an acquisition module configured to acquire a biological voiceprint signal of a crab, a vibration signal of a water body, and a behavior image of the crab.

[0069] a first feature extraction module configured to perform feature extraction based on the biological voiceprint signal to obtain a two-dimensional sound spectrogram and a voiceprint behavior feature.

[0070] a second feature extraction module configured to perform feature extraction based on the vibration signal to obtain a vibration fusion feature.

[0071] a third feature extraction module configured to perform behavior detection based on the behavior image to obtain an image behavior probability.

[0072] a cross-modal multi-head attention alignment module configured to perform cross-modal multi-head attention alignment based on the two-dimensional sound spectrogram and the behavior image to obtain a voiceprint-image matching result.

[0073] an inference module configured to determine a behavior probability vector based on the voiceprint-image matching result, the voiceprint behavior feature, the vibration fusion feature, and the image behavior probability.

[0074] a recognition and regulation module configured to determine a stress index of the crab based on the behavior probability vector, and to perform environmental regulation based on the stress index.

[0075] In an exemplary embodiment, a computer device is provided, which can be a server or a terminal, and its internal structure diagram can be as shown in Figure 6As shown in the figure. The computer device includes a processor, a memory, an input / output interface and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capability. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store crab stress identification data. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through network connection. The computer program is executed by the processor to realize a crab stress identification method.

[0076] Those skilled in the art can understand that, Figure 6 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement. In one exemplary embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to realize each of the method embodiments.

[0077] In one exemplary embodiment, a computer readable storage medium is provided, storing a computer program, which is executed by a processor to realize each of the method embodiments.

[0078] In one exemplary embodiment, a computer program product is provided, including a computer program, which is executed by a processor to realize each of the method embodiments.

[0079] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.

[0080] In the present application, all actions of obtaining signals, information or data are carried out in compliance with the corresponding data protection regulations and policies of the country where the device is located, and with the authorization given by the corresponding device owner.

[0081] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to a memory, a database or other medium used in the embodiments provided in the present application can include at least one of a non-volatile and a volatile memory. The non-volatile memory can include a read-only memory (ROM), a magnetic tape, a floppy disk, a flash memory, an optical storage, a high-density embedded non-volatile memory, a resistive random access memory (ReRAM), a magnetoresistive random access memory (MRAM), a ferroelectric random access memory (FRAM), a phase change memory (PCM), a graphene memory, etc. The volatile memory can include a random access memory (RAM) or an external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as a static random access memory (SRAM) or a dynamic random access memory (DRAM), etc.

[0082] The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a blockchain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0083] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combinations of the technical features do not exist contradictory, they should be considered as the scope of the present application.

[0084] The principles and implementation modes of the present application are described by applying specific examples in the present application. The above-mentioned embodiments are only used to help understand the method and its core idea of the present application; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range can be changed. In conclusion, the content of the present application should not be understood as a limitation.

Claims

1. A method for recognizing crab stress, characterized in that, The crab stress recognition method includes: Acquire bioacoustic signals of crabs, vibration signals of water bodies, and behavioral images of crabs; Based on the bio-voiceprint signal, feature extraction is performed to obtain a two-dimensional acoustic spectrogram and voiceprint behavior features; Based on the vibration signal, feature extraction is performed to obtain vibration fusion features; Behavior detection is performed on the behavior image to obtain the image behavior probability; Cross-modal multi-head attention alignment is performed based on the two-dimensional spectrogram and the behavior image to obtain the voiceprint image matching result; A behavior probability vector is determined based on the voiceprint image matching result, the voiceprint behavior features, the vibration fusion features, and the image behavior probability. The stress index of crabs is determined based on the behavioral probability vector, and environmental regulation is carried out based on the stress index.

2. The crab stress recognition method according to claim 1, characterized in that, Before performing feature extraction based on the bio-voiceprint signal to obtain the two-dimensional spectrogram and voiceprint behavior features, the process also includes: The bio-voiceprint signal is subjected to bandpass filtering and normalization. The vibration signal is filtered using a high-pass filter. The behavioral images are then processed for resolution and time. The processed bio-voiceprint signal and the processed vibration signal are aligned with the processed behavioral image through a clock pulse synchronization protocol.

3. The crab stress recognition method according to claim 1, characterized in that, Based on the bio-voiceprint signal, feature extraction is performed to obtain a two-dimensional acoustic spectrogram and voiceprint behavior features, specifically including: The bio-voiceprint signal is processed using a Hamming window and short-time Fourier transform to obtain a two-dimensional acoustic spectrogram; Based on the bio-voiceprint signal, a one-dimensional convolutional neural network is used for feature recognition to obtain voiceprint behavioral features.

4. The crab stress recognition method according to claim 1, characterized in that, Based on the vibration signal, feature extraction is performed to obtain vibration fusion features, specifically including: Long-term trends are extracted using the slow channel of the SlowFast dual-time base network; Detail features are extracted using the fast channels of the SlowFast dual-temporal network; The long-term trend and the detailed features are fused using the cross-channel feature fusion function of the SlowFast dual-time base network to obtain the vibration fusion feature.

5. The crab stress recognition method according to claim 1, characterized in that, Cross-modal multi-head attention alignment is performed based on the two-dimensional spectrogram and the behavior image to obtain the voiceprint image matching result, specifically including: The two-dimensional spectrogram and the behavioral image are respectively represented sequentially; Using a two-dimensional spectrogram represented by a sequence as the query vector and a behavioral image represented by a sequence as the key-value pair, cross-modal matching is performed to obtain the weighted combined features corresponding to each sound node; The weighted combination features corresponding to each sound node are used to enhance semantic relationships using a multi-head attention mechanism to obtain the voiceprint image matching result.

6. The crab stress recognition method according to claim 1, characterized in that, The behavior probability vector is determined based on the voiceprint image matching result, the voiceprint behavior features, the vibration fusion features, and the image behavior probability, specifically including: The voiceprint image matching result is represented as a semantic enhancement vector Z; The semantic enhancement vector Z and the voiceprint behavior features are used. Vibration fusion characteristics and image behavior probability The behavior probability vector is determined using the fusion inference module according to the following formula. p : ; in, This indicates a feature concatenation operation. This is the weight matrix. As a bias term, Softmax is used to normalize to a probability distribution.

7. A crab stress recognition device, characterized in that, The crab stress recognition device includes: The acquisition module is used to acquire the biological acoustic signals of crabs, the vibration signals of water bodies, and the behavioral images of crabs; The first feature extraction module is used to extract features based on the biological voiceprint signal to obtain a two-dimensional spectrogram and voiceprint behavior features. The second feature extraction module is used to extract features based on the vibration signal to obtain vibration fusion features; The third feature extraction module is used to perform behavior detection based on the behavior image to obtain the image behavior probability; A cross-modal multi-head attention alignment module is used to perform cross-modal multi-head attention alignment based on the two-dimensional spectrogram and the behavior image to obtain a voiceprint image matching result; The inference module is used to determine a behavior probability vector based on the voiceprint image matching result, the voiceprint behavior features, the vibration fusion features, and the image behavior probability; The identification and control module is used to determine the stress index of crabs based on the behavior probability vector, and to carry out environmental control based on the stress index.

8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the crab stress recognition method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the crab stress recognition method according to any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the crab stress recognition method according to any one of claims 1-6.