Training method, recognition method and device of hairy crab behavior intelligent recognition model

By collecting and preprocessing behavior video data in hairy crab breeding scenarios, combining feature extraction of spatial stream networks and time stream networks, and using attention mechanisms and multi-scale information fusion strategies to train behavior recognition models, the accuracy of hairy crab behavior analysis in the existing technology is solved, and the accurate identification and analysis of hairy crab behavior is achieved.

CN120107671APending Publication Date: 2025-06-06SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510172858.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing hairy crab behavior analysis and identification methods cannot accurately identify the hairy crab behavior, mainly due to insufficient data sets, environmental complexity and behavior dynamics.

Method used

A training method of hairy crab behavior intelligent recognition model is adopted. The training data set is generated by collecting behavior video data in outdoor open air and indoor land-based breeding scenarios, preprocessing and annotating. Then, based on the spatial flow network and the temporal flow network, combining attention mechanisms and multi-scale information fusion strategies, the behavioral recognition model is trained.

Benefits of technology

It realizes accurate identification of hairy crab behavior in complex environments, improves the accuracy and robustness of behavior analysis, and can effectively identify the key behaviors of hairy crabs in different scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107671A_ABST
    Figure CN120107671A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of aquaculture, and discloses a training method, a recognition method and a device for a hairy crab behavior intelligent recognition model, and the method comprises the steps: collecting behavior video data of hairy crabs in two scenes of outdoor open-air breeding and indoor land-based breeding, carrying out the preprocessing of the behavior video data, marking the corresponding behaviors, and carrying out the recognition of the hairy crabs. Generating a corresponding training data set; determining a to-be-trained behavior recognition model based on the spatial flow network and the time flow network in combination with an attention mechanism and a multi-scale information fusion strategy; and inputting the training data set into a to-be-trained behavior recognition model for training until a preset stop condition is met, and obtaining a trained hairy crab behavior recognition model for recognizing hairy crab behaviors. According to the method, different time scales and space scales are combined, and the key behaviors of the hairy crabs in a complex environment can be accurately identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of aquaculture technology, and for example, to a training method, an identification method, and a device for an intelligent identification model of hairy crab behavior. Background Art

[0002] Aquaculture is an important part of my country's agricultural economy, and the health status and growth environment of hairy crabs, as an important breeding object, directly affect the breeding benefits. Traditional breeding management mainly relies on manual observation, which is inefficient and difficult to monitor the behavior of hairy crabs in real time. In recent years, with the development of computer vision and deep learning technology, behavior analysis technology based on image recognition has been gradually applied to aquaculture, but there are still some problems and limitations, such as insufficient data sets, environmental complexity, and dynamic behavior.

[0003] Therefore, in the existing hairy crab behavior analysis and identification methods, there is a problem that the hairy crab behavior cannot be accurately analyzed and identified.

[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of this application. Summary of the invention

[0005] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical components or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0006] The present disclosure provides a method for training a hairy crab behavior intelligent recognition model. The method:

[0007] Collect behavioral video data of hairy crabs in both outdoor and indoor land-based farming scenarios, pre-process the behavioral video data, and annotate the corresponding behaviors;

[0008] The preprocessed video data of hairy crabs is used as samples, and the behaviors corresponding to the samples are used as sample labels to generate corresponding training data sets;

[0009] Based on the spatial stream network and the temporal stream network, combined with the attention mechanism and multi-scale information fusion strategy, the behavior recognition model to be trained is determined;

[0010] Input the training data set into the behavior recognition model to be trained, and calculate the loss function based on the difference between the output result of the behavior recognition model to be trained and the corresponding sample label;

[0011] The behavior recognition model to be trained is trained based on the loss function until a preset stop condition is met, and a trained hairy crab behavior recognition model is obtained to be used for identifying hairy crab behavior.

[0012] In some embodiments, the behavior video data of hairy crabs is collected in two scenarios, outdoor open-air breeding and indoor land-based breeding, including:

[0013] Using cameras, in both outdoor open-air farming and indoor land-based farming scenarios, by changing the shooting angle and lighting conditions, the behavioral video data of hairy crabs under different angles and lighting conditions were obtained.

[0014] In some embodiments, the spatial flow network is used to extract the spatial features of the hairy crab in each frame of the image, and the spatial features include shape and texture features;

[0015] The spatial stream network includes an input layer, a convolution layer, and a pooling layer. The input layer of the spatial stream network is used to normalize the behavior video data of the hairy crab. Each convolution block of the spatial stream network includes multiple convolution layers and batch normalization layers. The convolution layer uses a 3×3 convolution kernel with a step size of 1 and a padding method to keep the size consistent. The batch normalization layer is used to accelerate the training process of the network. The pooling layer of the spatial stream network uses a maximum pooling operation with a pooling kernel size of 2×2 and a step size of 2 to reduce the spatial size of the feature map.

[0016] The time stream network is used to capture the temporal dynamic information of the hairy crab behavior, including the continuity and change trend of the action;

[0017] The temporal stream network includes an input layer, a 3D convolutional layer and a pooling layer. The convolution kernel size of the 3D convolutional layer is 3×3×3, the step size is 1, and the padding method is to keep the size consistent, which is used to capture the spatial and temporal information in the video data; the pooling layer of the temporal stream network uses a maximum pooling operation, the pooling kernel size is 2×2×2, and the step size is 2, which is used to reduce the spatial and temporal dimensions of the feature map.

[0018] In some embodiments, based on the spatial stream network and the temporal stream network, combined with the attention mechanism and the multi-scale information fusion strategy, the behavior recognition model to be trained is determined, including:

[0019] Based on the spatial stream network and the temporal stream network, a fusion layer is used to concatenate the feature vectors of the spatial stream network and the temporal stream network in the last dimension to form a comprehensive feature vector;

[0020] The comprehensive feature vector is transformed and compressed through one or more fully connected layers, and the hairy crab behavior is classified and identified based on the softmax layer combined with the attention mechanism and multi-scale information fusion strategy.

[0021] In some embodiments, the attention mechanism includes three parts: query, key, and value. The attention weight is obtained by calculating the dot product between the query and the key, and then the attention weight is multiplied by the value to obtain a weighted feature representation;

[0022] The multi-scale information fusion strategy includes using 3D convolution kernels of different sizes to perform convolution operations on the sequences corresponding to the video data, extracting spatiotemporal features of different scales, and fusing features of different scales.

[0023] In some embodiments, the training data set includes a training set and a test set, the training set is used for the training process of the model, and the test set is used to evaluate the performance and generalization ability of the model;

[0024] In the process of model training, the training data set is divided into multiple subsets, each subset is used as a test set in turn, and the remaining subsets are used as training sets. Multiple training and testing are performed, and the model performance evaluation result is the average of multiple test results.

[0025] The present disclosure provides a method for intelligently identifying hairy crab behaviors, the method comprising:

[0026] Collecting video data of hairy crab behaviors to be identified, and preprocessing the video data of hairy crab behaviors to be identified;

[0027] The pre-processed video data of hairy crab behavior to be identified is input into the trained hairy crab behavior identification model to obtain the corresponding hairy crab behavior, wherein the trained hairy crab behavior identification model is obtained based on the training method of the above-mentioned hairy crab behavior intelligent identification model.

[0028] The present disclosure provides a training device for a hairy crab behavior intelligent recognition model, the device comprising:

[0029] The acquisition module is used to collect the behavior video data of hairy crabs in two scenarios: outdoor open-air breeding and indoor land-based breeding, and to pre-process the behavior video data and annotate the corresponding behaviors;

[0030] A processing module, used to use the pre-processed video data of hairy crabs as samples and the behaviors corresponding to the samples as sample labels to generate corresponding training data sets;

[0031] The processing module is also used to determine the behavior recognition model to be trained based on the spatial stream network and the temporal stream network, combined with the attention mechanism and the multi-scale information fusion strategy;

[0032] The processing module is further used to input the training data set into the behavior recognition model to be trained, and calculate the loss function based on the difference between the output result of the behavior recognition model to be trained and the corresponding sample label;

[0033] The processing module is also used to train the behavior recognition model to be trained based on the loss function until a preset stop condition is met, thereby obtaining a trained hairy crab behavior recognition model for identifying hairy crab behavior.

[0034] An embodiment of the present disclosure provides an electronic device, the device comprising at least one processor;

[0035] and a memory communicatively coupled to the at least one processor;

[0036] The memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor so that the at least one processor can execute the above-mentioned training method and recognition method.

[0037] The embodiment of the present disclosure provides a storage medium storing program instructions, which execute the above-mentioned training method and recognition method when running.

[0038] The training method, recognition method, device, equipment and storage medium of the hairy crab behavior intelligent recognition model provided by the embodiments of the present disclosure can achieve the following technical effects:

[0039] The training method based on the intelligent recognition model of hairy crab behavior proposed in the embodiment of the present disclosure obtains the process of the recognition model for hairy crab behavior recognition, by collecting the behavior video data of hairy crabs in two scenarios of outdoor open-air breeding and indoor land-based breeding, and pre-processing the behavior video data and marking the corresponding behavior to generate a corresponding training data set; based on the spatial stream network and the temporal stream network, combined with the attention mechanism and the multi-scale information fusion strategy, the behavior recognition model to be trained is determined; the training data set is input into the behavior recognition model to be trained, and training is performed until the preset stop condition is met, and the trained hairy crab behavior recognition model is obtained for identifying the hairy crab behavior. The present disclosure combines different time scales and spatial scales to realize accurate recognition of the key behaviors of hairy crabs in complex environments.

[0040] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] One or more embodiments are exemplarily described by corresponding drawings, which do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements, and the drawings do not constitute a scale limitation, and wherein:

[0042] Figure 1 It is a flowchart of a method for training a hairy crab behavior intelligent recognition model provided by an embodiment of the present disclosure;

[0043] Figure 2 It is a flowchart of another method for training a hairy crab behavior intelligent recognition model provided by an embodiment of the present disclosure;

[0044] Figure 3 It is a flow chart of a method for intelligently identifying hairy crab behaviors provided by an embodiment of the present disclosure;

[0045] Figure 4 It is a structural schematic diagram of a training device for a hairy crab behavior intelligent recognition model provided by an embodiment of the present disclosure;

[0046] Figure 5 It is a structural schematic diagram of a hairy crab behavior intelligent identification device provided by an embodiment of the present disclosure;

[0047] Figure 6 It is a structural schematic diagram of a hairy crab behavior intelligent identification device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0048] In order to be able to understand the features and technical contents of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.

[0049] The terms "first", "second", etc. in the embodiments of the present disclosure are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so as to describe the embodiments of the present disclosure described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions.

[0050] Unless otherwise stated, the term "plurality" means two or more.

[0051] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B indicates: A or B.

[0052] The term "and / or" is a description of the association relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B.

[0053] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.

[0054] Aquaculture is an important part of my country's agricultural economy, and the health status and growth environment of hairy crabs, as an important breeding object, directly affect the breeding benefits. Traditional breeding management mainly relies on manual observation, which is inefficient and difficult to monitor the behavior of hairy crabs in real time. In recent years, with the development of computer vision and deep learning technology, behavior analysis technology based on image recognition has gradually been applied to aquaculture, but there are still some problems and limitations:

[0055] Insufficient datasets: There are few underwater image and video datasets on the behavior of hairy crabs, which limits the training and optimization of behavior recognition models.

[0056] Environmental complexity: The underwater environment is dark, the water is turbid, and there are phenomena such as refraction and scattering on the water surface, which causes problems such as color cast, blur, and low contrast in the video data, affecting the accuracy of behavior recognition.

[0057] Behavioral dynamics: The behavior of hairy crabs is diverse and complex, with obvious dynamic changes, but existing research mostly focuses on static image analysis, which makes it difficult to accurately portray dynamic behavior.

[0058] Poor algorithm robustness: In complex and changeable breeding environments, the accuracy and robustness of existing behavior recognition algorithms are poor and it is difficult to meet the needs of actual applications.

[0059] Therefore, in the existing hairy crab behavior analysis and identification methods, there is a problem that the hairy crab behavior cannot be accurately analyzed and identified.

[0060] In order to solve the above problems, the present disclosure provides a training method, recognition method, device, equipment and storage medium for an intelligent recognition model of hairy crab behavior.

[0061] The training method, recognition method, device, equipment and storage medium of the intelligent recognition model of hairy crab behavior provided by the embodiments of the present disclosure are described below in conjunction with the accompanying drawings.

[0062] Figure 1 It is a flow chart of a method for training a hairy crab behavior intelligent recognition model provided in an embodiment of the present disclosure.

[0063] Combination Figure 1 As shown, the training method of the hairy crab behavior intelligent recognition model may include:

[0064] S101, collecting behavioral video data of hairy crabs in two scenarios: outdoor open-air breeding and indoor land-based breeding, and preprocessing the behavioral video data and annotating corresponding behaviors;

[0065] S102, using the preprocessed video data of hairy crabs as samples, and using the behaviors corresponding to the samples as sample labels, to generate corresponding training data sets;

[0066] S103, based on the spatial stream network and the temporal stream network, combined with the attention mechanism and the multi-scale information fusion strategy, determine the behavior recognition model to be trained;

[0067] S104, inputting the training data set into the behavior recognition model to be trained, and calculating the loss function based on the difference between the output result of the behavior recognition model to be trained and the corresponding sample label;

[0068] S105, training the behavior recognition model to be trained based on the loss function until a preset stop condition is met, thereby obtaining a trained hairy crab behavior recognition model for use in recognizing hairy crab behavior.

[0069] In some embodiments, the behavior video data of hairy crabs is collected in two scenarios, outdoor open-air breeding and indoor land-based breeding, including:

[0070] Using cameras, in both outdoor open-air farming and indoor land-based farming scenarios, by changing the shooting angle and lighting conditions, the behavioral video data of hairy crabs under different angles and lighting conditions were obtained.

[0071] In some embodiments, the spatial flow network is used to extract the spatial features of the hairy crab in each frame of the image, and the spatial features include shape and texture features;

[0072] The spatial stream network includes an input layer, a convolution layer, and a pooling layer. The input layer of the spatial stream network is used to normalize the behavior video data of the hairy crab. Each convolution block of the spatial stream network includes multiple convolution layers and batch normalization layers. The convolution layer uses a 3×3 convolution kernel with a step size of 1 and a padding method to keep the size consistent. The batch normalization layer is used to accelerate the training process of the network. The pooling layer of the spatial stream network uses a maximum pooling operation with a pooling kernel size of 2×2 and a step size of 2 to reduce the spatial size of the feature map.

[0073] The time stream network is used to capture the temporal dynamic information of the hairy crab behavior, including the continuity and change trend of the action;

[0074] The temporal stream network includes an input layer, a 3D convolutional layer and a pooling layer. The convolution kernel size of the 3D convolutional layer is 3×3×3, the step size is 1, and the padding method is to keep the size consistent, which is used to capture the spatial and temporal information in the video data; the pooling layer of the temporal stream network uses a maximum pooling operation, the pooling kernel size is 2×2×2, and the step size is 2, which is used to reduce the spatial and temporal dimensions of the feature map.

[0075] In some embodiments, based on the spatial stream network and the temporal stream network, combined with the attention mechanism and the multi-scale information fusion strategy, the behavior recognition model to be trained is determined, including:

[0076] Based on the spatial stream network and the temporal stream network, a fusion layer is used to concatenate the feature vectors of the spatial stream network and the temporal stream network in the last dimension to form a comprehensive feature vector;

[0077] The comprehensive feature vector is transformed and compressed through one or more fully connected layers, and the hairy crab behavior is classified and identified based on the softmax layer combined with the attention mechanism and multi-scale information fusion strategy.

[0078] In some embodiments, the attention mechanism includes three parts: query, key, and value. The attention weight is obtained by calculating the dot product between the query and the key, and then the attention weight is multiplied by the value to obtain a weighted feature representation;

[0079] The multi-scale information fusion strategy includes using 3D convolution kernels of different sizes to perform convolution operations on the sequences corresponding to the video data, extracting spatiotemporal features of different scales, and fusing features of different scales.

[0080] In some embodiments, the training data set includes a training set and a test set, the training set is used for the training process of the model, and the test set is used to evaluate the performance and generalization ability of the model;

[0081] In the process of model training, the training data set is divided into multiple subsets, each subset is used as a test set in turn, and the remaining subsets are used as training sets. Multiple training and testing are performed, and the model performance evaluation result is the average of multiple test results.

[0082] Figure 2 is a flow chart of another method for training a hairy crab behavior intelligent recognition model provided by an embodiment of the present disclosure, combined with Figure 2 ,right Figure 1 The method is further described in .

[0083] 1. Data collection and preprocessing

[0084] ① Data collection: High-resolution optical cameras can be used to collect video data of hairy crab behavior in both outdoor open-air farming and indoor land-based farming. By changing the shooting angle and lighting conditions, hairy crab behavior images under different angles and lighting conditions can be obtained to improve the diversity and comprehensiveness of the data set.

[0085] ② Data preprocessing: In view of the particularity of underwater images, a variety of image processing technologies can be used to preprocess the collected video data. This includes operations such as denoising, contrast enhancement, and color cast correction to improve image quality and provide clear and accurate visual information for subsequent behavior recognition. At the same time, the preprocessed data is annotated with behaviors to establish a high-quality training data set.

[0086] Specifically, the main function of data collection and preprocessing is to automatically or manually collect, process and standardize the relevant behaviors of the identified object, and provide high-quality input data support for subsequent identification judgment. Specifically, it can include the following steps:

[0087] 1) Scene selection and equipment layout

[0088] Outdoor open-air breeding scene: You can select a representative outdoor open-air hairy crab breeding pond to ensure that the water quality, water temperature, light and other conditions of the pond meet the growth needs of hairy crabs. High-resolution optical cameras are arranged at several key locations in the pond, such as areas where hairy crabs are active frequently and near feeding points. The installation height of the camera should be slightly higher than the water surface to capture the behavior of hairy crabs underwater, while avoiding the influence of water surface reflection on the shooting effect. In order to obtain hairy crab behavior images from different angles and fields of view, multiple cameras can be used to work together, and the shooting angle and focal length of each camera should be carefully adjusted to ensure a wide coverage and moderate overlapping area.

[0089] Indoor land-based breeding scenarios: In indoor land-based breeding facilities, different breeding environmental conditions can be simulated, such as different water temperatures, light intensities, water quality, etc., to study the behavioral changes of hairy crabs in different environments. Multiple optical cameras can be installed around and above the breeding pond to record the behavior of hairy crabs in all directions. The layout of the camera should take into account the distribution of indoor light to avoid poor image quality due to excessive or weak light. At the same time, in order to simulate the changes in light in the natural environment, adjustable brightness lamps can be used to work synchronously with the camera to ensure that clear video data can be obtained under different lighting conditions.

[0090] 2) Collection time and frequency

[0091] Time selection: According to the biological clock and behavioral habits of hairy crabs, data can be collected during the time period when they are more active. Generally speaking, hairy crabs are more active in the early morning and evening, so these two time periods can be used as the main collection period. In addition, appropriate amounts of data can be collected during other time periods during the day and night to obtain behavioral data of hairy crabs in different time periods to ensure the comprehensiveness and representativeness of the data.

[0092] Frequency arrangement: In each collection cycle, the collection frequency can be reasonably arranged. For example, during the peak period of hairy crab molting, the collection frequency can be increased to collect data every 2 to 3 hours, each time for 1 hour, to capture the entire process of molting behavior; during the daily monitoring stage, the collection frequency can be appropriately reduced to collect data every 4 to 5 hours, each time for 1 hour, to reduce the amount of data and reduce interference with the breeding environment. At the same time, according to the growth stage and health status of hairy crabs, the collection frequency can be flexibly adjusted to ensure that key behaviors can be captured in a timely manner.

[0093] 3) Image denoising

[0094] Denoising algorithm selection: Select a suitable denoising algorithm for the noise types unique to underwater images, such as granular noise and stripe noise. The non-local mean denoising algorithm can be used. This algorithm can effectively remove random noise in the image while retaining the image details by finding similar image blocks in the image and replacing the noise pixels with the average value of these similar blocks. For stripe noise, the Fourier transform denoising method can be used to convert the image from the spatial domain to the frequency domain, identify and eliminate the frequency components corresponding to the stripe noise, and then convert the image back to the spatial domain to achieve denoising.

[0095] Denoising parameter optimization: In the actual denoising process, the parameters of the denoising algorithm can be adjusted according to the noise characteristics of the image. For example, in the non-local mean denoising algorithm, parameters such as the size of the search window and the weight of similar blocks will affect the denoising effect. The optimal combination of parameters can be determined through experiments or based on image quality evaluation indicators (such as signal-to-noise ratio, peak signal-to-noise ratio, etc.) to achieve the best denoising effect.

[0096] Contrast Enhancement

[0097] Histogram equalization: The histogram equalization method can be used to enhance the contrast of the image. First, the histogram of the image is calculated, that is, the number of pixels at each gray level; then, the cumulative distribution function is calculated based on the histogram and mapped to a new gray level range so that the histogram of the image is evenly distributed. This can stretch the gray level range of the image, improve the overall contrast of the image, and make the outline and details of the hairy crab more clearly visible.

[0098] Adaptive contrast enhancement: Considering the complexity of the underwater environment, the lighting conditions in different areas may be different. An adaptive contrast enhancement algorithm can be used. This algorithm divides the image into multiple small areas and performs contrast enhancement processing on each area separately. For example, methods such as local histogram equalization or adaptive gamma correction can be used to dynamically adjust the degree of contrast enhancement according to the local lighting conditions and image content of each area to avoid image distortion caused by over-enhancement.

[0099] 4) Color cast correction

[0100] White balance correction: Use the white balance correction algorithm to correct the color cast of the image. First, determine the white area or neutral gray area in the image, and calculate the average value of the red, green, and blue color channels in these areas; then, calculate the color gain based on these average values ​​to adjust the color balance of the entire image. For example, the gray world assumption method can be used to assume that the average color of the image is neutral gray. By adjusting the gain of the color channel, the average color of the image reaches neutral gray, thereby achieving color cast correction.

[0101] Correction based on physical models: A physical model of underwater light propagation is established, taking into account the absorption and scattering characteristics of water, as well as the spectral distribution of the light source. Based on this model, the color correction parameters under different depths and water quality conditions are calculated to accurately correct the color of the image. This method can more accurately restore the true color of underwater objects, improve the visual effect of the image and the accuracy of subsequent behavior recognition.

[0102] 4) Data annotation and enhancement

[0103] Annotation tools and methods: You can use professional image annotation tools, such as LabelImg, VGG ImageAnnotator, etc., to annotate the behavior of the preprocessed video data. During the annotation process, you need to carefully observe the behavior of the hairy crab in the video, and accurately annotate the corresponding behavior area and behavior type on the video frame according to the characteristics of behaviors such as molting, predation, and fighting. For example, in the annotation of molting behavior, it is necessary to mark the starting position and ending position of the hairy crab molting, as well as the body changes of the hairy crab during the molting process; in the annotation of predation behavior, it is necessary to mark the key frames such as the starting action of the hairy crab's predation, the moment of capturing prey, and the eating process.

[0104] Data enhancement technology: In order to expand the data set and improve the generalization ability of the model, a variety of data enhancement technologies can be used to process the annotated data. For example, geometric transformation methods such as image rotation, scaling, flipping, and cropping can be used to change the shape and size of the image and increase the diversity of the data; color jitter, Gaussian blur, noise injection and other methods can also be used to simulate different lighting conditions and image quality to improve the model's adaptability to complex environments. When performing data enhancement, it is necessary to ensure that the enhanced image still maintains the original behavioral characteristics and the accuracy of the annotation information.

[0105] 2. Behavior recognition algorithm design

[0106] ① Dual-stream network architecture: Based on the spatial stream and temporal stream networks, a behavior recognition algorithm that integrates spatiotemporal features is designed. The spatial stream network is responsible for extracting the spatial features of the hairy crab in each frame of the image, such as shape, texture, etc.; the temporal stream network captures the temporal dynamic information of the hairy crab's behavior, such as the continuity and change trend of the action. By fusing these two features, accurate recognition of the hairy crab's behavior is achieved.

[0107] ② Improved 3D convolutional neural network: Improvements are made to the traditional 3D CNN to enhance its ability to utilize temporal information in video data. By introducing the attention mechanism and multi-scale information fusion strategy, the network can better capture the local details and global features of the hairy crab's behavior, which can improve the accuracy of behavior recognition.

[0108] ③ Multi-scale information fusion: To address the problem of insufficient local information modeling of hairy crab behavior, a multi-scale information fusion algorithm is designed. This algorithm can focus on both the local area and overall characteristics of hairy crab behavior, and can achieve comprehensive recognition of hairy crab behavior by fusing feature information at different scales.

[0109] Specifically, the behavior recognition algorithm mainly extracts the behavior characteristics of the target object by analyzing the dynamic changes in the video or image sequence. It usually combines computer vision technology and deep learning methods, such as two-stream networks and 3D convolutional neural networks, to achieve accurate recognition and classification of complex behaviors. The specific implementation method is as follows:

[0110] 1) Spatial Stream Network Construction

[0111] Input layer: The input of the spatial stream network is a single frame image, and the image size is uniformly adjusted to 224×224 pixels to accommodate the processing of the subsequent convolutional layer. The input image is first normalized to scale the pixel values ​​to between 0 and 1 to improve the convergence speed and stability of the network.

[0112] Convolutional layer and pooling layer: ResNet-50 architecture is used as the basis, which contains multiple convolutional blocks and pooling layers. Each convolutional block consists of several convolutional layers and batch normalization layers. The convolutional layer uses a 3×3 convolution kernel, a step size of 1, and a padding mode of "same" to maintain the spatial size of the feature map. The batch normalization layer is used to accelerate the training process of the network and reduce the internal covariate shift phenomenon. The pooling layer uses the maximum pooling operation, the pooling kernel size is 2×2, and the step size is 2, which is used to reduce the spatial size of the feature map and extract more abstract features.

[0113] Feature extraction: After multiple layers of convolution and pooling operations, the spatial flow network can extract the spatial features of the hairy crab in a single frame image, such as shape, texture, color, etc. These features provide an important spatial information basis for subsequent behavior recognition.

[0114] 2) Construction of time stream network

[0115] Input layer: The input of the temporal stream network is a video sequence, each sequence contains 16 consecutive frames of images, and the image size is also adjusted to 224×224 pixels. After the input video sequence is normalized, it is sent to the network for processing.

[0116] 3D convolution layer and pooling layer: 3D CNN architecture is used, which uses 3D convolution kernels to perform convolution operations on video sequences. The size of the 3D convolution kernel is 3×3×3, the stride is 1, and the padding mode is "same", which can capture both spatial and temporal information in the video. The pooling layer also uses the maximum pooling operation, with a pooling kernel size of 2×2×2 and a stride of 2, which is used to reduce the spatial and temporal dimensions of the feature map and extract more abstract spatiotemporal features.

[0117] Feature extraction: The time stream network can capture the temporal dynamic information of the hairy crab's behavior, such as the continuity and change trend of the action, through the processing of 3D convolutional layers and pooling layers. These features provide an important temporal information basis for subsequent behavior recognition.

[0118] 3) Feature fusion and classification

[0119] Fusion layer: The features extracted by the spatial stream network and the temporal stream network are fused. The feature vectors of the two networks are concatenated in the last dimension to form a comprehensive feature vector. This comprehensive feature vector contains the spatiotemporal information of the hairy crab's behavior and can more comprehensively describe the behavioral characteristics of the hairy crab.

[0120] Fully connected layer and classification layer: The fused feature vector is further transformed and compressed through one or more fully connected layers, and finally the classification and recognition of hairy crab behavior is realized through the softmax layer. The number of neurons in the fully connected layer can be adjusted according to actual needs, for example, it can be set to 1024 or 512. The output of the softmax layer is the category probability distribution of hairy crab behavior, and the category with the highest probability is selected as the final recognition result.

[0121] 4) Introduction of attention mechanism

[0122] Self-attention module design: The self-attention module is introduced into the temporal stream network of 3D CNN. This module enhances the network's perception of temporal information by calculating the similarity between different frames in the video sequence. The self-attention module consists of three parts: query, key, and value. The attention weight is obtained by calculating the dot product between the query and the key, and then the attention weight is multiplied by the value to obtain the weighted feature representation. This mechanism can make the network pay more attention to the important temporal information in the video and improve the accuracy of behavior recognition.

[0123] Attention module position selection: In different layers of 3D CNN, select appropriate positions to insert self-attention modules. For example, inserting self-attention modules in the middle or deep layers of the network can make the network pay more attention to temporal information when extracting high-level features. At the same time, self-attention modules can also be inserted in multiple layers to enhance the network's perception of temporal information at different levels.

[0124] 5) Multi-scale information fusion strategy

[0125] Multi-scale feature extraction: In the temporal stream network of 3D CNN, 3D convolution kernels of different sizes can be used to perform convolution operations on video sequences to extract spatiotemporal features of different scales. For example, convolution kernels of different sizes such as 3×3×3, 5×5×5, and 7×7×7 can be used to extract local and global spatiotemporal features respectively.

[0126] Feature fusion method: To fuse features of different scales, methods such as splicing, weighted summation or feature pyramid can be used. The splicing method splices feature vectors of different scales in the last dimension, the weighted summation method performs weighted summation of features of different scales according to the importance of the features, and the feature pyramid method constructs a feature pyramid through upsampling and downsampling operations to achieve the fusion of features of different scales. The fused multi-scale features can focus on the local details and overall characteristics of the hairy crab behavior at the same time, improving the accuracy of behavior recognition.

[0127] 3. Model training and optimization

[0128] ① Model training: The designed recognition algorithm can be trained using the labeled hairy crab behavior dataset. Use deep learning frameworks, such as TensorFlow or PyTorch, to train and optimize the model. By adjusting the network structure and parameters, the model can accurately identify the key behaviors of hairy crabs, such as molting, predation, and fighting.

[0129] ②Model optimization: During the training process, regularization, early stopping and other techniques are used to prevent the model from overfitting. At the same time, the model is evaluated and optimized through methods such as cross-validation to ensure that the model has good generalization ability and robustness in different environments and scenarios.

[0130] Specifically, 1) Behavior dataset division

[0131] Training set and test set division: The labeled hairy crab behavior dataset can be divided into training set and test set according to a certain ratio, usually 7:3 or 8:2. The training set is used for the model training process, and the test set is used to evaluate the performance and generalization ability of the model.

[0132] Cross-validation: In order to further improve the stability and generalization ability of the model, the cross-validation method can be used. The data set is divided into multiple subsets, each subset is used as a test set in turn, and the remaining subsets are used as training sets, and multiple training and testing are performed. The final model performance evaluation result is the average of multiple test results.

[0133] 2) Training process settings

[0134] Loss function selection: The cross entropy loss function is used as the loss function of the model. This function can measure the difference between the category probability distribution predicted by the model and the true category probability distribution, and is suitable for multi-classification problems.

[0135] Optimizer selection: Use the Adam optimizer to update the model parameters. The Adam optimizer combines the advantages of the momentum and RMSProp optimization methods, has the characteristics of adaptive learning rate, can accelerate the convergence of the model and improve the training effect.

[0136] Learning rate setting: The initial learning rate is set to 0.001. As the training progresses, a learning rate decay strategy, such as exponential decay or cosine annealing, can be used to gradually reduce the learning rate to improve the convergence accuracy and stability of the model.

[0137] Batch size setting: According to the hardware resources and the complexity of the model, the batch size should be set reasonably. For example, it can be set to 32 or 64. Too large a batch size will increase memory consumption, but it can improve the stability and convergence speed of training; too small a batch size will increase the noise of training, but it can improve the generalization ability of the model.

[0138] 3) Training monitoring and adjustment

[0139] Monitoring indicators: During the training process, monitor the model's loss value, accuracy, precision, recall and other indicators. Use visualization tools (such as TensorBoard) to observe the changing trends of these indicators in real time and understand the training status of the model in a timely manner.

[0140] Adjustment strategy: According to the changes in monitoring indicators, timely adjust the parameters in the training process. For example, when the loss value of the model no longer decreases or the accuracy no longer increases, you can reduce the learning rate, increase the regularization strength, or adjust the network structure to promote further optimization of the model.

[0141] 4) Application of regularization technology

[0142] L2 regularization: Add L2 regularization term to the loss function of the model to penalize the weight of the model to prevent overfitting caused by excessive weight. The coefficient of L2 regularization term can be adjusted according to actual needs, usually set to 0.01 or 0.001, etc.

[0143] Dropout regularization: Dropout regularization technology is used in the fully connected layer of the model to randomly discard a certain proportion of neurons and their connections, so that the model cannot rely on a specific combination of neurons during training, thereby improving the generalization ability of the model. The Dropout ratio can be set according to actual conditions, for example, it can be set to 0.5 or 0.3.

[0144] 5) Early stopping strategy implementation

[0145] Early stopping condition setting: When the performance of the model on the validation set no longer improves, that is, when the validation set loss value no longer decreases or the accuracy no longer improves for multiple consecutive training cycles (for example, 5 or 10), the early stopping condition is triggered.

[0146] Early stopping operation: Stop the training process and keep the current best model parameters. This can prevent the model from overfitting on the training set while ensuring that the model has good performance on the validation set.

[0147] 6) Model evaluation and selection

[0148] Evaluation indicators: Use accuracy, precision, recall, F1-Score and other indicators to comprehensively evaluate the model. Accuracy measures the overall classification accuracy of the model, precision measures the proportion of samples predicted by the model as positive that are actually positive, recall measures the proportion of samples that are actually positive that are correctly predicted as positive, and F1-Score is the harmonic mean of precision and recall, which comprehensively considers the accuracy and comprehensiveness of the model.

[0149] Model selection: Based on the evaluation results, select the model with the best performance as the final behavior recognition model. In the selection process, we should not only pay attention to the performance of the model on the training set and the validation set, but also consider factors such as the model's complexity, training time, and inference speed. For practical applications, the real-time performance and resource consumption of the model are also important considerations.

[0150] Model optimization and iteration: After the model is selected, it can be continuously optimized and iterated based on feedback and new data from actual applications. For example, new behavior data can be collected regularly to retrain the model to improve its adaptability and robustness; different model architectures and parameter adjustments can also be tried to further improve the performance of the model.

[0151] 4. Practical applications and benefits

[0152] ① Optimization of breeding management: The present invention is applied to the actual breeding management of hairy crabs, and the behavior of hairy crabs can be monitored in real time, providing a scientific basis for decision-making for breeders. For example, by identifying the molting behavior of hairy crabs, the feeding strategy can be adjusted in time to promote the growth and development of hairy crabs; by monitoring predation and fighting behaviors, the breeding density and environmental conditions can be optimized to reduce competition and damage among hairy crabs.

[0153] When the system recognizes the molting behavior of hairy crabs, it automatically adjusts the feeding strategy and increases the feeding amount to meet the nutritional needs of hairy crabs during molting; when predation and fighting behavior is monitored, the breeding density is adjusted in time to avoid excessive competition and damage among hairy crabs.

[0154] ② Improvement of breeding efficiency: Improve the growth rate and survival rate of hairy crabs, reduce the incidence of diseases, and thus improve breeding efficiency. At the same time, reduce the workload and cost of manual monitoring, and realize intelligent and automated breeding management.

[0155] The process of obtaining the recognition model based on the training method of the intelligent recognition model of hairy crab behavior proposed in the embodiment of the present disclosure can monitor the key behaviors of hairy crabs, such as molting, predation and fighting, in real time through the underwater visual system to optimize aquaculture management and improve the breeding efficiency. The technology includes data collection and preprocessing, behavior recognition algorithm design, model training and optimization, etc. In the data collection stage, a high-resolution optical camera is used to collect the behavior video data of hairy crabs in two breeding scenes, outdoor open-air and indoor land-based, and preprocessing operations such as denoising, contrast enhancement, and color cast correction are performed to establish a high-quality training data set. The behavior recognition algorithm adopts a two-stream network architecture. The spatial stream network extracts the spatial features in a single frame image, and the temporal stream network captures the temporal dynamic information in the video sequence, and achieves accurate recognition through feature fusion; at the same time, an improved 3D convolutional neural network is introduced, combined with the attention mechanism and multi-scale information fusion strategy, to further improve the recognition accuracy. In the model training process, the cross entropy loss function and Adam optimizer are used to optimize the model parameters through cross validation and early stopping strategy to ensure that the model has good generalization ability and robustness. Focusing on the two main scenarios of outdoor open-air farming and indoor land-based farming, the system can accurately identify the key behaviors of hairy crabs in complex underwater environments through scientific visual acquisition and model architecture, combined with different time and space scales. It can provide a scientific basis for aquaculture decision-making, help intelligent farming management, and have significant economic and social benefits.

[0156] Figure 3 is a flow chart of a method for intelligently identifying hairy crab behaviors provided by an embodiment of the present disclosure, such as Figure 3 As shown, the intelligent identification method of hairy crab behavior may include:

[0157] S301, obtaining video data of hairy crab behavior to be identified, and preprocessing the video data of hairy crab behavior to be identified;

[0158] S302, inputting the pre-processed video data of the hairy crab behavior to be identified into the trained hairy crab behavior recognition model to obtain the corresponding hairy crab behavior, wherein the trained hairy crab behavior recognition model is based on Figure 1 Obtained by the method in .

[0159] and Figure 1 Corresponding to the training method of the hairy crab behavior intelligent recognition model, the present disclosure also provides a training device for the hairy crab behavior intelligent recognition model, such as Figure 4 As shown, the device may specifically include:

[0160] The acquisition module 401 is used to collect the behavior video data of hairy crabs in two scenarios: outdoor open-air breeding and indoor land-based breeding, and pre-process the behavior video data and mark the corresponding behaviors;

[0161] Processing module 402, used to use the pre-processed video data of hairy crabs as samples, and use the behaviors corresponding to the samples as sample labels to generate corresponding training data sets;

[0162] The processing module 402 is further used to determine the behavior recognition model to be trained based on the spatial stream network and the temporal stream network in combination with the attention mechanism and the multi-scale information fusion strategy;

[0163] The processing module 402 is further used to input the training data set into the behavior recognition model to be trained, and calculate the loss function based on the difference between the output result of the behavior recognition model to be trained and the corresponding sample label;

[0164] The processing module 402 is also used to train the behavior recognition model to be trained based on the loss function until a preset stop condition is met, thereby obtaining a trained hairy crab behavior recognition model for identifying hairy crab behavior.

[0165] In some embodiments, the behavior video data of hairy crabs is collected in two scenarios, outdoor open-air breeding and indoor land-based breeding, including:

[0166] Using cameras, in both outdoor open-air farming and indoor land-based farming scenarios, by changing the shooting angle and lighting conditions, the behavioral video data of hairy crabs under different angles and lighting conditions were obtained.

[0167] In some embodiments, the spatial flow network is used to extract the spatial features of the hairy crab in each frame of the image, and the spatial features include shape and texture features;

[0168] The spatial stream network includes an input layer, a convolution layer, and a pooling layer. The input layer of the spatial stream network is used to normalize the behavior video data of the hairy crab. Each convolution block of the spatial stream network includes multiple convolution layers and batch normalization layers. The convolution layer uses a 3×3 convolution kernel with a step size of 1 and a padding method to keep the size consistent. The batch normalization layer is used to accelerate the training process of the network. The pooling layer of the spatial stream network uses a maximum pooling operation with a pooling kernel size of 2×2 and a step size of 2 to reduce the spatial size of the feature map.

[0169] The time stream network is used to capture the temporal dynamic information of the hairy crab behavior, including the continuity and change trend of the action;

[0170] The temporal stream network includes an input layer, a 3D convolutional layer and a pooling layer. The convolution kernel size of the 3D convolutional layer is 3×3×3, the step size is 1, and the padding method is to keep the size consistent, which is used to capture the spatial and temporal information in the video data; the pooling layer of the temporal stream network uses a maximum pooling operation, the pooling kernel size is 2×2×2, and the step size is 2, which is used to reduce the spatial and temporal dimensions of the feature map.

[0171] In some embodiments, based on the spatial stream network and the temporal stream network, combined with the attention mechanism and the multi-scale information fusion strategy, the behavior recognition model to be trained is determined, including:

[0172] Based on the spatial stream network and the temporal stream network, a fusion layer is used to concatenate the feature vectors of the spatial stream network and the temporal stream network in the last dimension to form a comprehensive feature vector;

[0173] The comprehensive feature vector is transformed and compressed through one or more fully connected layers, and the hairy crab behavior is classified and identified based on the softmax layer combined with the attention mechanism and multi-scale information fusion strategy.

[0174] In some embodiments, the attention mechanism includes three parts: query, key, and value. The attention weight is obtained by calculating the dot product between the query and the key, and then the attention weight is multiplied by the value to obtain a weighted feature representation;

[0175] The multi-scale information fusion strategy includes using 3D convolution kernels of different sizes to perform convolution operations on the sequences corresponding to the video data, extracting spatiotemporal features of different scales, and fusing features of different scales.

[0176] In some embodiments, the training data set includes a training set and a test set, the training set is used for the training process of the model, and the test set is used to evaluate the performance and generalization ability of the model;

[0177] In the process of model training, the training data set is divided into multiple subsets, each subset is used as a test set in turn, and the remaining subsets are used as training sets. Multiple training and testing are performed, and the model performance evaluation result is the average of multiple test results.

[0178] and Figure 3 Corresponding to the method for intelligently identifying the behavior of hairy crabs, the present disclosure also provides an intelligent device for identifying the behavior of hairy crabs, such as Figure 5 As shown, the device may specifically include:

[0179] The acquisition module 501 is used to acquire the video data of the hairy crab behavior to be identified, and pre-process the video data of the hairy crab behavior to be identified;

[0180] Processing module 502 is used to input the pre-processed video data of the hairy crab behavior to be identified into the trained hairy crab behavior recognition model to obtain the corresponding hairy crab behavior, wherein the trained hairy crab behavior recognition model is based on Figure 1 Obtained by the method in .

[0181] Combination Figure 6As shown, the embodiment of the present disclosure also provides a hairy crab behavior intelligent identification device 600, including a processor 604 and a memory (memory) 601. Optionally, the system may also include a communication interface (CommunicationInterface) 602 and a bus 603. Among them, the processor 604, the communication interface 602, and the memory 601 can communicate with each other through the bus 603. The communication interface 602 can be used for information transmission. The processor 604 can call the logic instructions in the memory 601 to execute the training method or identification method of the hairy crab behavior intelligent identification model of the above embodiment.

[0182] In addition, the logic instructions in the memory 601 described above can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.

[0183] The memory 601 is a computer-readable storage medium that can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the method in the embodiment of the present disclosure. The processor 604 executes the program instructions / modules stored in the memory 601 to perform functional applications and data processing, that is, to implement the training method or recognition method of the hairy crab behavior intelligent recognition model in the above embodiment.

[0184] The memory 601 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory 601 may include a high-speed random access memory and may also include a non-volatile memory.

[0185] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured as a training method or an identification method for an intelligent identification model of hairy crab behavior.

[0186] The computer-readable storage medium mentioned above may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.

[0187] The technical solution of the embodiment of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of the embodiment of the present disclosure. The aforementioned storage medium may be a non-transient storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a disk or an optical disk, and other media that can store program codes, or a transient storage medium.

[0188] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent possible changes only. Unless explicitly required, separate components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. As used in the description of the embodiments, unless the context clearly indicates otherwise, the singular forms of "a", "an" and "the" are intended to include plural forms as well. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of one or more associated listings. In addition, when used in this application, the term "comprise" and its variants "comprises" and / or including (comprising) refer to the existence of stated features, wholes, steps, operations, elements, and / or components, but do not exclude the existence or addition of one or more other features, wholes, steps, operations, elements, components and / or these groups. In the absence of further restrictions, the elements defined by the sentence "comprising a ..." do not exclude the existence of other identical elements in the process, method or device comprising the elements. In this article, each embodiment may focus on the differences from other embodiments, and the same or similar parts between the embodiments may refer to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can refer to the description of the method part.

[0189] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. Technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. Technicians can clearly understand that for the convenience and simplicity of description, the specific working process of the systems, devices and units described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0190] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units can be only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, each functional unit in the embodiment of the present disclosure may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.

[0191] The flowchart and block diagram in the accompanying drawings show the possible architecture, functions and operations of the system, method and computer program product according to the embodiment of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and a part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. In the description corresponding to the flowchart and block diagram in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in a different order from the order disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

[0192] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0193] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0194] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0195] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0196] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0197] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0198] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of this disclosure can be achieved, and this document is not limited here.

[0199] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A method for training a hairy crab behavior intelligent recognition model, characterized in that: The method comprises: Collecting behavioral video data of hairy crabs in two scenarios, outdoor open-air breeding and indoor land-based breeding, and preprocessing the behavioral video data and annotating corresponding behaviors; The preprocessed video data of hairy crabs is used as samples, and the behaviors corresponding to the samples are used as sample labels to generate corresponding training data sets; Based on the spatial stream network and the temporal stream network, combined with the attention mechanism and multi-scale information fusion strategy, the behavior recognition model to be trained is determined; Inputting the training data set into the behavior recognition model to be trained, and calculating a loss function based on the difference between the output result of the behavior recognition model to be trained and the corresponding sample label; The behavior recognition model to be trained is trained based on the loss function until a preset stop condition is met, thereby obtaining a trained hairy crab behavior recognition model for identifying hairy crab behavior.

2. The method according to claim 1, characterized in that The behavior video data of hairy crabs collected in two scenarios of outdoor open-air breeding and indoor land-based breeding include: Using cameras, in both outdoor open-air farming and indoor land-based farming scenarios, by changing the shooting angle and lighting conditions, the behavioral video data of hairy crabs under different angles and lighting conditions were obtained.

3. The method according to claim 1, characterized in that The spatial flow network is used to extract the spatial features of the hairy crab in each frame of the image, and the spatial features include shape and texture features; The spatial stream network includes an input layer, a convolution layer and a pooling layer. The input layer of the spatial stream network is used to normalize the behavior video data of the hairy crab; each convolution block of the spatial stream network includes multiple convolution layers and batch normalization layers. The convolution layer uses a 3×3 convolution kernel, a step size of 1, and a padding method to keep the size consistent. The batch normalization layer is used to accelerate the training process of the network; the pooling layer of the spatial stream network uses a maximum pooling operation, the pooling kernel size is 2×2, and the step size is 2, which is used to reduce the spatial size of the feature map; The time stream network is used to capture the time dynamic information of the hairy crab behavior, wherein the time dynamic information includes the continuity and change trend of the action; The temporal stream network includes an input layer, a 3D convolution layer and a pooling layer. The convolution kernel size of the 3D convolution layer is 3×3×3, the step size is 1, and the padding method is to keep the size consistent, which is used to capture the spatial and temporal information in the video data; the pooling layer of the temporal stream network uses a maximum pooling operation, the pooling kernel size is 2×2×2, and the step size is 2, which is used to reduce the spatial and temporal dimensions of the feature map.

4. The method according to claim 1, characterized in that: The behavior recognition model to be trained is determined based on the spatial stream network and the temporal stream network, combined with the attention mechanism and the multi-scale information fusion strategy, including: Based on the spatial stream network and the temporal stream network, a fusion layer is used to concatenate the feature vectors of the spatial stream network and the temporal stream network in the last dimension to form a comprehensive feature vector; The comprehensive feature vector is transformed and compressed through one or more fully connected layers, and the hairy crab behavior is classified and identified based on the softmax layer combined with the attention mechanism and multi-scale information fusion strategy.

5. The method according to claim 4, characterized in that The attention mechanism consists of three parts: query, key and value. The attention weight is obtained by calculating the dot product between the query and the key, and then the attention weight is multiplied by the value to obtain the weighted feature representation; The multi-scale information fusion strategy includes using 3D convolution kernels of different sizes to perform convolution operations on the sequences corresponding to the video data, extracting spatiotemporal features of different scales, and fusing features of different scales.

6. The method according to claim 1, characterized in that The training data set includes a training set and a test set, the training set is used for the training process of the model, and the test set is used to evaluate the performance and generalization ability of the model; In the process of model training, the training data set is divided into multiple subsets, each subset is used as a test set in turn, and the remaining subsets are used as training sets. Multiple training and testing are performed, and the model performance evaluation result is the average of multiple test results.

7. A method for intelligently identifying hairy crab behavior, characterized in that: The method comprises: Collecting video data of hairy crab behavior to be identified, and preprocessing the video data of hairy crab behavior to be identified; The pre-processed video data of hairy crab behavior to be identified is input into a trained hairy crab behavior identification model to obtain the corresponding hairy crab behavior, wherein the trained hairy crab behavior identification model is obtained based on the method described in any one of claims 1-6.

8. A training device for an intelligent recognition model of hairy crab behavior, characterized in that: The device comprises: A collection module is used to collect the behavior video data of hairy crabs in two scenarios: outdoor open-air breeding and indoor land-based breeding, and to pre-process the behavior video data and mark the corresponding behaviors; A processing module, used to use the pre-processed video data of hairy crabs as samples and the behaviors corresponding to the samples as sample labels to generate corresponding training data sets; The processing module is further used to determine the behavior recognition model to be trained based on the spatial stream network and the temporal stream network in combination with the attention mechanism and the multi-scale information fusion strategy; The processing module is further used to input the training data set into the behavior recognition model to be trained, and calculate the loss function based on the difference between the output result of the behavior recognition model to be trained and the corresponding sample label; The processing module is also used to train the behavior recognition model to be trained based on the loss function until a preset stop condition is met, so as to obtain a trained hairy crab behavior recognition model for identifying hairy crab behavior.

9. An electronic device, characterized in that: include: at least one processor; and a memory communicatively coupled to the at least one processor; It is characterized in that the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.