Fish detection methods, devices, equipment, and storage media based on underwater robots

By combining dynamic convolutional attention networks and latent game models, intelligent identification and collision risk assessment of underwater fish are achieved, solving the problem of high difficulty in fish identification in underwater environments and improving identification accuracy and efficiency.

CN120108003BActive Publication Date: 2025-10-31SHENZHEN CHASING INNOVATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510571022.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-10-31
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

In underwater environments, light attenuation and water currents increase the difficulty of fish identification and location, affecting the accuracy and reliability of detection equipment.

Method used

A dynamic convolutional attention network is used to identify fish categories, and a potential game model is used to assess collision risk. Fish information is obtained by combining multi-sensor data fusion, providing intelligent identification and early warning.

Benefits of technology

It improves the accuracy and efficiency of fish identification, adapts to different underwater environments, provides comprehensive and accurate fish information, and ensures the safe operation of underwater robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108003B_ABST
    Figure CN120108003B_ABST
Patent Text Reader

Abstract

This application relates to the field of intelligent detection technology, and in particular to a method, apparatus, device, and storage medium for fish detection based on an underwater robot. The method includes: acquiring an image to be processed; the image to be processed containing objects to be detected in an underwater region; identifying the category of the objects to be detected using a dynamic convolutional attention network; determining the relative position between the objects to be detected and the underwater robot based on the image to be processed; using a latent game model to detect whether the collision risk of the relative position falls within a warning range; and issuing fish alert information to the user based on the detection results and the object's category. This application can achieve intelligent fish identification in underwater areas, improving the accuracy and efficiency of fish identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent detection and processing, and in particular to a fish detection method, apparatus, equipment and storage medium based on an underwater robot. Background Technology

[0002] In related technologies, underwater light weakens rapidly with increasing depth, and the scattering and absorption of light in water lead to a decrease in image quality. This makes it difficult for visual inspection systems to acquire clear images of fish, affecting fish identification and localization. Furthermore, water currents can cause changes in the posture and position of fish, increasing the difficulty of detection. Simultaneously, water currents can also cause instability in underwater robots, affecting the accuracy and reliability of inspection equipment.

[0003] Therefore, there is an urgent need to design a technical solution to improve the accuracy and efficiency of fish identification in complex environments. Summary of the Invention

[0004] This application addresses the technical problems existing in the prior art by providing a fish detection method, device, equipment, and storage medium based on an underwater robot, which enables intelligent fish identification in underwater areas, improves the accuracy of fish identification in complex environments, and increases the efficiency of fish identification.

[0005] In a first aspect, embodiments of this application provide a fish detection method based on an underwater robot, the method comprising at least:

[0006] Acquire an image to be processed; the image to be processed contains the object to be detected in the underwater region;

[0007] The category to which the object to be detected belongs is identified by a dynamic convolutional attention network;

[0008] The relative position between the object to be detected and the underwater robot is determined based on the image to be processed;

[0009] A potential game theory model is used to detect whether the collision risk of the relative position falls within the warning range. Based on the detection results and the fish category, a fish alert is issued to the user. The fish alert is used to indicate the fish categories that the underwater robot can observe and the fish collision risk.

[0010] Secondly, embodiments of this application provide a fish detection system based on an underwater robot, the system comprising at least:

[0011] The acquisition module is used to acquire an image to be processed; the image to be processed contains objects to be detected in the underwater area.

[0012] The recognition module is used to identify the category to which the object to be detected belongs through a dynamic convolutional attention network;

[0013] The positioning module is used to determine the relative position between the object to be detected and the underwater robot based on the image to be processed;

[0014] The early warning module is used to detect whether the collision risk of the relative position is within the warning range using a potential game model, and to issue fish alert information to the user based on the detection results and the fish category; the fish alert information is used to indicate the fish categories that the underwater robot can observe and the fish collision risk.

[0015] Thirdly, embodiments of this application provide an electronic device, which includes a memory for storing computer software programs;

[0016] A processor is used to read and execute the computer software program, thereby implementing the first aspect of the fish detection method based on an underwater robot.

[0017] Fourthly, a computer-readable storage medium is provided, comprising instructions that, when executed on a computer, cause the computer to perform the underwater robot-based fish detection method of the first aspect.

[0018] The beneficial effects of this application are: it provides a fish detection method, apparatus, device, and storage medium based on an underwater robot. In this technical solution, firstly, an image to be processed is acquired; the image to be processed contains objects to be detected in the underwater area. Then, a dynamic convolutional attention network is used to identify the category to which the objects to be detected belong. Next, the relative position between the objects to be detected and the underwater robot is determined based on the image to be processed. Finally, a latent game model is used to detect whether the collision risk of the relative position falls within the warning range, and fish alert information is issued to the user based on the detection results and the category; the fish alert information is used to indicate the fish categories observable by the underwater robot and the fish collision risk. In this embodiment, intelligent fish identification in underwater areas can be achieved, improving the accuracy and efficiency of fish identification. Attached Figure Description

[0019] Figure 1 This is a schematic flowchart of a fish detection method based on an underwater robot according to an embodiment of this application;

[0020] Figure 2 This is a schematic diagram of the structure of a fish detection system based on an underwater robot according to an embodiment of this application;

[0021] Figure 3 This is a schematic diagram of the structure of a medium according to an embodiment of this application. Detailed Implementation

[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] This application provides a method, apparatus, device, and storage medium for fish detection based on an underwater robot. In this embodiment, firstly, an image to be processed is acquired; the image to be processed contains objects to be detected in an underwater region. Then, a dynamic convolutional attention network is used to identify the category to which the objects to be detected belong. Next, the relative position between the objects to be detected and the underwater robot is determined based on the image to be processed. Finally, a latent game model is used to detect whether the collision risk of the relative position falls within the warning range, and fish alert information is issued to the user based on the detection results and the category; the fish alert information is used to indicate the fish categories observable by the underwater robot and the fish collision risk. This application embodiment can achieve intelligent fish identification in underwater areas, improving the accuracy and efficiency of fish identification.

[0024] Specifically, firstly, the dynamic convolution module in the dynamic convolutional attention network dynamically generates convolutional kernels based on the fish pose estimation results. Due to the diverse poses of underwater fish, such as sideways swimming and sharp turns, traditional fixed convolutional kernels struggle to effectively extract features from different poses. This method, however, can adjust the convolutional kernel weights in real time according to the pose. For example, when a fish is sideways, it focuses on extracting tail fin features; when swimming upright, it focuses on body texture features. This allows for more accurate feature extraction for fish in different poses, reducing recognition errors caused by pose changes and improving the classification accuracy for closely related species (such as different species of tuna with similar appearances) and juvenile and adult fish (with morphological differences). Compared to traditional CNNs (such as ResNet), this improvement is 15%-20%.

[0025] Due to the complex underwater environment, factors such as light attenuation and water turbidity can lead to problems such as localized image blurring and low contrast. The multi-scale attention fusion module processes shallow (edge ​​details) and deep (semantic features) features in parallel and adaptively fuses multi-scale information through a self-attention mechanism. In blurred areas, it enhances the capture of edge details while combining deep semantic features for a better understanding of image content. For example, in dimly lit areas, adjusting attention weights to highlight the edge contours of fish helps in accurately identifying fish species. This is particularly advantageous in fish identification of low-quality underwater images, reducing the misclassification rate caused by image quality issues.

[0026] Secondly, the Dynamic Convolutional Attention Network (DCA) reduces model parameters by 30% compared to traditional CNNs, resulting in a corresponding reduction in computational cost. This reduces the computational resources required for processing images and accelerates inference speed. Taking embedded devices commonly used in underwater robots as an example, it can complete image feature extraction and fish category recognition in a short time, meeting real-time requirements. The inference speed can reach approximately 25 FPS, enabling rapid underwater fish identification. Compared to traditional models, this significantly improves recognition efficiency, allowing underwater robots to process more images and detect fish in more areas per unit time.

[0027] From acquiring the image to completing fish category identification, relative position determination, and collision risk detection, the entire process is seamlessly integrated. Through reasonable algorithm design and model architecture, unnecessary intermediate steps and data processing time are reduced. For example, when determining the relative position, calculations are performed directly based on the image, avoiding redundant data processing and transmission, improving overall processing efficiency, and enabling underwater robots to complete fish detection tasks more efficiently.

[0028] Third, dynamic convolutional attention networks can automatically learn key distinguishing features of fish without the need for manually designing complex feature extraction rules. As training data increases and the model is optimized, its understanding and recognition capabilities for various fish features continuously improve, enabling it to adapt to newly emerging fish species and different underwater environments. For example, when encountering a new fish species, the model can gradually and accurately identify it by learning its features, achieving intelligent and adaptive fish identification.

[0029] The latent game theory model treats the underwater robot and the target object (fish) as two parties in a game, comprehensively considering factors such as their relative positions, motion states, and the fish's behavioral patterns. By analyzing the payoffs and risks for both parties under different strategies, it dynamically determines whether a collision risk falls within the warning range. For example, for aggressive fish, the model can predict their potential aggressive behavior in advance based on changes in their trajectory and speed, and issue timely warnings; for harmless fish swimming normally, it can accurately determine that they will not pose a threat to the underwater robot, avoiding unnecessary warnings. This decision-making approach based on game theory models enables underwater robots to more intelligently cope with complex underwater environments and fish behavior, making reasonable decisions.

[0030] Fourth, the fish alerts sent to users based on the detection results and their categories not only include the types of fish observable by the underwater robot but also accurately indicate the risk of fish collisions. This provides users (such as underwater workers and researchers) with more comprehensive and accurate information, helping them better understand the underwater environment and fish populations. For example, in marine scientific monitoring, researchers can accurately record fish species and distribution based on the alerts, and rationally plan the underwater robot's route based on the collision risk information to ensure the smooth progress of monitoring tasks. In underwater operations, operators can take timely measures based on the alerts to avoid collisions between the underwater robot and fish, ensuring equipment safety and the normal operation of the work.

[0031] In this embodiment, regardless of whether it's in clear shallow water or turbid deep water, or under strong or low light conditions, the dynamic convolutional attention network and the latent game model can adapt to different environmental conditions through their own characteristics and algorithm adjustments, maintaining high accuracy in fish identification and collision risk detection. For example, in turbid water, the multi-scale attention fusion module can better capture fish features, and the latent game model can comprehensively consider the impact of the aquatic environment on fish movement, accurately judging collision risks. This makes the technology highly practical and reliable in various underwater environments.

[0032] By using a dynamic convolutional attention network to identify the category of the object to be detected, this approach better addresses the challenges of diverse fish species, varied morphologies, and individual differences. Compared to related technologies, it improves the accuracy and adaptability of fish category identification, effectively solving the problem of difficulty in accurately distinguishing different fish species. Determining the relative position between the object to be detected and the underwater robot based on the processed image provides an accurate data foundation for subsequent collision risk detection, overcoming the difficulty of accurately obtaining the relative position of fish and the underwater robot in existing technologies. This helps to more accurately assess the spatial relationship between fish and the underwater robot. Employing a latent game model to detect whether the collision risk based on the relative position falls within the warning range allows for a more comprehensive and accurate assessment of collision risk, considering multiple factors. This solves the problems of unscientific and untimely collision risk assessment in related technologies, ensuring the safe operation of the underwater robot. Based on the detection results and the fish category, the system provides users with fish alerts that include both the fish category and the fish collision risk. This allows users to have a more comprehensive understanding of the fish situation around the underwater robot, solving the problem of single and incomplete information feedback in related technologies, and helping users make more informed decisions.

[0033] The underwater robot-based fish detection solution provided in this application can also be executed by an electronic device, such as a server, server cluster, or cloud server. This electronic device can also be a terminal device such as a mobile phone, computer, tablet computer, wearable device, or dedicated device (such as a dedicated terminal device with an underwater robot-based fish detection system). These electronic devices can also incorporate the chips described in the above embodiments. Alternatively, these electronic devices can also install a service program for executing the underwater robot-based fish detection solution.

[0034] Figure 1 This is a schematic diagram illustrating a fish detection method based on an underwater robot, provided as an embodiment of this application. Figure 1 As shown, the method includes:

[0035] 101. Acquire the image to be processed; the image to be processed contains the object to be detected in the underwater area;

[0036] 102. Identify the category to which the object to be detected belongs through a dynamic convolutional attention network;

[0037] 103. Determine the relative position between the object to be detected and the underwater robot based on the image to be processed;

[0038] 104. A potential game theory model is used to detect whether the collision risk of the relative position is within the warning range, and fish alert information is issued to the user based on the detection results and the category to which it belongs.

[0039] In this embodiment, the fish alert information is used to indicate the types of fish that the underwater robot can observe and the risk of fish collision. For example, the fish alert information may include the following: Tuna detected (high-speed swimming fish), 1.2 meters away from the underwater robot, collision risk: high; it is recommended to immediately slow down and change course. For example, the fish alert information may include the following: Clownfish detected, 3 meters away from the underwater robot, collision risk: low; normal operation is possible. For example, the fish alert information may include the following: Shark detected, 2.5 meters away from the underwater robot, and accelerating towards it, collision risk: extremely high; please immediately control the robot to evacuate the area.

[0040] In step 101, the image to be processed is acquired. Further, the image to be processed contains the objects to be detected in the underwater region.

[0041] Specifically, underwater robots are equipped with sensors such as cameras and sonar to collect image data in the underwater environment. Cameras directly capture visual images of the underwater scene; sonar transmits and receives sound waves, converting reflected signals into image information, making it suitable for low-light or murky environments. Data collected by multiple sensors can be fused to provide more comprehensive information for subsequent processing. Thus, multi-source data acquisition ensures that the acquired images contain rich underwater scene information, covering the objects to be detected under different lighting and water quality conditions, laying a data foundation for subsequent fish detection and identification. At the same time, multi-sensor fusion overcomes the limitations of a single sensor, improving the reliability and completeness of the data.

[0042] In step 102, a dynamic convolutional attention network is used to identify the category to which the object to be detected belongs.

[0043] Specifically, the dynamic convolution module dynamically generates convolutional kernels based on fish pose estimation results using a conditional parameter generator. Pose estimation obtains parameters such as the fish's body rotation angle and tail fin swing amplitude, generating convolutional kernels with corresponding weights to extract fish features under different poses. The multi-scale attention fusion module processes shallow edge details and deep semantic features in parallel, adaptively fusing them through a self-attention mechanism to address feature imbalance caused by uneven lighting in underwater images. This effectively addresses fish pose variations and underwater image quality issues, accurately extracting key fish identification features and significantly improving the accuracy of fish category recognition. The classification accuracy for closely related species and juvenile versus adult fish is improved by 15%-20% compared to traditional CNNs, maintaining a high recognition rate even in complex underwater environments.

[0044] As an optional embodiment, in step 102, identifying the category to which the object to be detected belongs through a dynamic convolutional attention network includes:

[0045] Extract the identification information of the object to be detected; the identification information includes at least: real-time contour information, real-time surface information, and real-time pose information; perform dynamic convolution processing on the key identification features to obtain the dynamic shape features of the object to be detected under different activity states; the dynamic shape features include at least: scale texture features and tail fin shape features; set weight parameters for different activity states, and use the weight parameters to predict the category of the dynamic shape features to obtain the category to which the object to be detected belongs.

[0046] Understandably, in step 102, real-time contour information, real-time surface information, and real-time pose information of the target object are extracted from the image to be processed. This information can comprehensively describe the appearance and state of the fish. Contour information outlines the overall shape of the fish, surface information contains details of the fish's body surface such as color and markings, and pose information reflects the fish's movement state and body posture, providing rich raw data for subsequent accurate identification of fish categories. Furthermore, using a dynamic convolution module, based on the fish's pose estimation results, a conditional parameter generator dynamically generates convolution kernels to process key discriminative features. For fish in different activity states, by dynamically adjusting the weights of the convolution kernels, their dynamic shape features, such as scale texture features and tail fin shape features, can be extracted more effectively. This dynamic convolution method can adapt to the feature changes of fish in different poses, overcoming the problem that traditional fixed convolution kernels cannot cope with pose diversity.

[0047] Furthermore, weighting parameters are set for dynamic shape features under different activity states, comprehensively considering the importance of various features in fish category identification. These weighting parameters are used to weight and fuse the dynamic shape features, which are then input into a classifier for category prediction, thereby obtaining the category to which the object to be detected belongs. This method can reasonably weight the features according to their contribution to category judgment, improving the accuracy of identification.

[0048] For example, suppose the object to be detected is a fish. At a certain moment, the information to be identified is as follows: Real-time contour information: It presents a relatively slender body and a pointed head. Real-time surface information: The body surface has a silvery sheen and black spots. Real-time posture information: The body is slightly tilted, and the tail fin is swinging at a large angle, indicating a fast swimming posture.

[0049] After dynamic convolution processing, the extracted dynamic shape features are scale texture features and caudal fin shape features. Scale texture features include, for example, the scales exhibit a fine and tightly packed texture with a certain directionality. Caudal fin shape features include, for example, the caudal fin is forked with relatively smooth edges.

[0050] Furthermore, based on preset weight parameters for different activity states, these dynamic physical features are comprehensively evaluated and their categories predicted. For example, features such as a slender body outline, a silver surface with black spots, a fast swimming posture, and specific scale textures and tail fin shapes, combined, conform to the characteristic pattern of a bass, thus predicting that the object to be detected is a bass.

[0051] In this way, by extracting multi-dimensional information to be identified and performing dynamic convolution processing on key distinguishing features, the unique characteristics of different fish in various activity states can be captured more comprehensively and accurately. Setting weight parameters for category prediction further optimizes the comprehensive consideration of different features, thereby significantly improving the accuracy of fish category identification and effectively distinguishing easily confused categories such as juvenile fish, adult fish, and closely related species. The dynamic convolutional attention network can dynamically adjust the convolutional kernel weights and attention focus according to the real-time posture and activity state of the fish, adapting to the diversity of fish postures in underwater images and feature changes caused by environmental factors. This allows the model to maintain good performance in different underwater scenes and fish behavior patterns, exhibiting stronger generalization ability and adaptability.

[0052] Compared to traditional convolutional neural networks, this method achieves high-precision recognition while reducing the number of model parameters through dynamically generated convolutional kernels and an adaptive feature fusion mechanism. For example, compared to some classic CNN models (such as ResNet), the number of parameters can be reduced by about 30%, which is beneficial for deploying and running the model in resource-constrained environments such as edge devices, and reduces the requirements for hardware devices.

[0053] In practical applications, the training process of dynamic convolutional attention networks can include steps such as data preparation, model initialization, loss function definition, training loop, and model evaluation.

[0054] During data preparation, a large number of underwater images containing various fish species were collected as training data. At the same time, the fish in the images were categorized and labeled with key feature information, such as the fish's outline, posture, scale texture, and tail fin shape. This labeled information will serve as supervision signals for model training.

[0055] To increase data diversity and improve the model's generalization ability, data augmentation operations are performed on the original images, such as random cropping, flipping, rotating, and adjusting brightness, contrast, and color saturation. This allows the model to be exposed to fish images from more different perspectives and appearances during training, reducing the risk of overfitting.

[0056] Following the architecture design of a dynamic convolutional attention network, a network model is constructed that includes components such as a dynamic convolution module and a multi-scale attention fusion module. The number of parameters and connection methods of each layer in the network are determined to prepare for model training.

[0057] Use appropriate initialization methods to initialize the model's parameters. For example, random initialization can be used to assign initial values ​​to convolutional kernel weights, bias terms, and other learnable parameters. Some commonly used initialization strategies include Xavier initialization and Kaiming initialization. Good initialization can help the model converge faster.

[0058] In defining the loss function, based on the characteristics of fish classification tasks, the cross-entropy loss function is typically chosen as the model's loss function. The cross-entropy loss function measures the difference between the model's predictions and the true labels, exhibiting excellent performance for multi-class classification problems. During training, the model adjusts its parameters by minimizing the loss function, making the predictions as close as possible to the true labels.

[0059] To prevent overfitting and improve the model's generalization ability, a regularization term, such as L1 or L2 regularization, can be added to the loss function. The regularization term constrains the model's parameters, preventing them from becoming too large and thus avoiding overfitting due to model overcomplexity.

[0060] Then, the prepared training data is input into the model, and the data undergoes forward propagation within the network. In the dynamic convolution module, convolution kernels are dynamically generated based on the fish pose estimation results, and convolution operations are performed on the input data to extract fish features from different perspectives. The multi-scale attention fusion module processes shallow and deep features in parallel and adaptively fuses multi-scale information through a self-attention mechanism to obtain the final feature representation. Finally, the feature representation is input into the classifier to obtain the model's prediction result for the fish category. Based on the model's prediction result and the true label, a predefined loss function is used to calculate the loss value, which reflects the gap between the model's current prediction result and the actual situation. The gradient of the loss function with respect to the model parameters is calculated using the backpropagation algorithm; the gradient represents the rate of change of the loss function with respect to the current parameter values. Based on the gradient information, optimization algorithms (such as stochastic gradient descent, Adagrad, Adadelta, etc.) are used to update the model parameters, causing the loss function value to gradually decrease. When updating parameters, it is necessary to update them according to the corresponding rules based on different parameter types and network structures to ensure that the model can converge towards the optimal solution.

[0061] Repeat the forward propagation, loss calculation, and backpropagation process described above to iterate through the entire training dataset multiple times. As training progresses, the model parameters are continuously adjusted, the loss function value gradually decreases, and the model's performance gradually improves. During training, the model parameters can be saved periodically for recovery and evaluation when training is interrupted or when the model is needed again.

[0062] A subset of data is allocated from the original dataset as a validation set to evaluate the model's performance during training. This validation set data is not used in model training but rather to simulate the model's performance in real-world applications, allowing for timely detection of overfitting or underfitting. Appropriate evaluation metrics are chosen to measure model performance, such as accuracy, recall, and F1 score. For fish classification tasks, accuracy is a crucial metric, representing the proportion of fish categories correctly predicted by the model. Recall reflects the proportion of a particular fish category correctly identified by the model among all samples of that category. The F1 score, the harmonic mean of accuracy and recall, provides a comprehensive evaluation of the model's performance.

[0063] Based on the evaluation results on the validation set, optimize and adjust the model. If the model's accuracy on the validation set is low, it may be necessary to increase training data, adjust the network structure, and optimize hyperparameters to improve model performance. If the model exhibits overfitting (i.e., the loss function value on the training set is low, but the accuracy on the validation set decreases), consider methods such as adding regularization terms, reducing network complexity, or stopping training early. Through continuous evaluation and optimization, the model should achieve optimal performance on the validation set.

[0064] The trained model is then subjected to final performance testing using a separate test set. The test set data has not been used during model training and validation, thus accurately reflecting the model's generalization ability in real-world applications. Various evaluation metrics, such as accuracy, recall, and F1 score, are calculated on the test set to assess the model's final performance. If the model's performance on the test set meets the requirements, it can be deployed to real-world applications. If the performance does not meet the requirements, further analysis of the causes is needed, and the model should be improved and optimized until satisfactory performance is achieved.

[0065] Further, optionally, in the above steps, dynamic convolution processing is performed on the key identification features to obtain the dynamic shape features of the object to be detected under different activity states, including:

[0066] Extract the features to be processed at different scales from the key identification features; project the features to be processed at different scales onto the behavior prediction space corresponding to different activity states through a conditional parameter generator to obtain the projected features of the object to be detected in different activity states; the conditional parameter generator is used to indicate the activity states corresponding to different fish and the types of projected features to be extracted in various activity states; classify and enhance the projected features in different activity states to obtain the dynamic shape features of the object to be detected in different activity states.

[0067] In this embodiment of the application, the classification enhancement branch includes: a first enhancement branch for enhancing the caudal fin features in the sideways state, and a second enhancement branch for enhancing the scale texture features in the forward swimming state.

[0068] Specifically, in the steps described above, features at different scales can capture information at different levels of detail of the object to be detected. For example, small-scale features may contain finer texture information, while large-scale features can reflect the overall shape and structure information. By extracting features at different scales, key discriminative features can be comprehensively described, providing a rich data foundation for subsequent analysis.

[0069] The conditional parameter generator projects the features to be processed at different scales onto the behavior prediction space corresponding to different activity states, based on the activity characteristics of different fish species. This means that the model will selectively extract features related to the possible activity states of the fish. For example, for fast-swimming states, the model will focus more on features that reflect changes in speed and posture; while for stationary or slow-swimming states, it will focus more on relatively stable features such as shape and texture. This allows the model to better adapt to changes in the fish's shape under different activity states, improving the accuracy of extracting dynamic shape features.

[0070] The classification enhancement branch further enhances the projected features under different activity states, highlighting key features related to specific activity states. In this embodiment, the first enhancement branch enhances the caudal fin features in the sideways swimming state, because the caudal fin plays a crucial role in the fish's turning and propulsion when swimming sideways; enhancing the caudal fin features can better identify fish in the sideways swimming state. The second enhancement branch enhances the scale texture features in the upright swimming state, as the scale texture can serve as one of the important features for distinguishing different fish in the upright swimming state; enhancing this feature helps improve the accuracy of fish identification in the upright swimming state.

[0071] For example, suppose we have a goldfish as the object to be detected. First, we extract the features to be processed at different scales from the key distinguishing features of the goldfish. For instance, at a small scale, we can capture the fine texture of the goldfish scales, while at a large scale, we can see the overall body outline and general posture of the goldfish. Then, the conditional parameter generator projects these features to be processed at different scales onto the corresponding behavior prediction space based on the common activity states of the goldfish, such as fast swimming, slow swimming, and stillness. For example, when the goldfish is in a fast swimming state, the conditional parameter generator will pay more attention to speed-related features, such as the amplitude of body swaying and the frequency of tail fin swaying, and project the corresponding features to be processed onto the behavior prediction space of the fast swimming state, obtaining the projected features of the fast swimming state. Finally, the projected features are processed through a classification reinforcement branch. For the side-lying position, the first enhancement branch will focus on enhancing the tail fin features, such as highlighting the shape and swing angle of the tail fin; for the swimming position, the second enhancement branch will enhance the scale texture features, so that the model can more clearly identify the unique scale texture pattern when the goldfish is swimming, thereby accurately determining the type of goldfish and its activity state.

[0072] Therefore, by extracting features at different scales and projecting them into the behavior prediction space corresponding to different activity states, the key features of the detected object in different activity states can be captured more comprehensively and accurately, avoiding the limitations of single-scale or fixed feature extraction methods. For example, for fish of different sizes and species, feature extraction at different scales can adapt to their differences in shape and detail, while projection for activity states can better handle the dynamic changes in the shape of fish during movement, making the extracted dynamic shape features more representative and discriminative. The classification enhancement branch performs specialized enhancement processing on the projected features in different activity states, enabling the model to better adapt to the feature changes of various fish in different activity states, improving the model's adaptability to fish detection in different scenarios and conditions. For example, in complex underwater environments, fish activity states are diverse. By enhancing key features in specific activity states, the model can more accurately identify fish in different activity states. Even when encountering specific situations not seen in the training data, it can make reasonable judgments based on the enhanced features, thereby enhancing the model's generalization ability.

[0073] After the above processing, the model can more accurately extract the dynamic shape features of the object under different activity states. These features are of great significance for distinguishing different types of fish. Therefore, it can significantly improve the accuracy of fish recognition, reduce the occurrence of false positives and false negatives, and lay a solid foundation for providing accurate fish prompts to users in the future, making the entire fish detection system more reliable and practical.

[0074] In step 103, the relative position between the target object and the underwater robot is determined based on the image to be processed. Specifically, image processing techniques, such as target detection algorithms, are used to locate the fish's position in the image. This information is combined with the underwater robot's own position and attitude information (provided by an inertial navigation system, etc.) and the image's depth information (obtainable through sonar data or stereo vision calculations) to calculate the fish's three-dimensional spatial position (X, Y, Z coordinates) and orientation information relative to the underwater robot. This accurately determines the relative position between the fish and the underwater robot, providing accurate data support for subsequent collision risk assessment. It also enables the underwater robot to perceive the distribution and distance of surrounding fish, providing a basis for decision-making.

[0075] In step 104, a latent game theory model is used to detect whether the collision risk at the relative position falls within the warning range. Specifically, the latent game theory model treats the underwater robot and the fish as two parties in a game, considering factors such as their movement strategies, relative positions, speeds, and fish species. By constructing a game theory model, the payoffs and risks for both parties under different strategies are analyzed, and the collision risk probability is calculated. When the risk probability exceeds a set threshold, it is determined to enter the warning range. Based on the fish's species, corresponding fish alert information is generated and sent to the user. This allows for intelligent assessment of collision risk, dynamically adjusting the warning strategy based on fish behavior patterns and interactions with the robot, reducing false alarms and missed alarms. Compared to traditional fixed threshold methods, the accuracy of collision risk detection is significantly improved, ensuring the safety of underwater robot operations.

[0076] As an optional embodiment, in step 104, a potential game theory model is used to detect whether the collision risk of the relative positions falls within the warning range, including:

[0077] A construction layer employing a latent game model is used to construct a corresponding mathematical model based on the relative position in the three-dimensional space of the underwater region; a risk game layer employing a latent game model is used to predict the collision risk of the relative position according to the mathematical model; the collision risk of the relative position includes at least: the individual collision risk of the object to be detected at the relative position and the global collision risk; an early warning layer employing a latent game model is used to determine whether the collision risk of the relative position falls within the early warning range.

[0078] In the three-dimensional space of the underwater region, a mathematical model is constructed based on the relative positions of the object to be detected and the underwater robot. This model takes into account factors such as their coordinates, distance, direction of motion, and velocity in space. By transforming the actual physical positional relationship into mathematical expressions, a quantitative basis is provided for subsequent risk analysis. For example, a Cartesian coordinate system can be used to represent their positions, and vectors can be used to describe the direction of motion and velocity, thereby establishing a mathematical model that accurately reflects their relative motion state in three-dimensional space.

[0079] Based on the constructed mathematical model, the risk game layer considers the behavioral strategies of the detected object and the underwater robot, as well as their interactions. Individual collision risk primarily focuses on the probability of the detected object colliding with the underwater robot, which depends on factors such as its own motion state and distance from the underwater robot. Global collision risk, on the other hand, considers the probability of multiple detected objects colliding with the underwater robot and with each other in the entire underwater environment from a more macroscopic perspective. This requires comprehensively considering the position, motion state, and potential mutual influence of all relevant objects, and predicting collision risk through the analysis and calculation of these factors. For example, when multiple fish approach the underwater robot simultaneously, it is necessary to consider not only the individual collision risk of each fish with the underwater robot, but also the potential increase in global collision risk due to interference between the fish.

[0080] The early warning layer compares the predicted collision risks with pre-set thresholds to determine whether the collision risk falls within the warning range. These thresholds are determined based on factors such as the underwater robot's safety requirements, working environment, and actual application needs. If the calculated individual or global collision risk exceeds the corresponding threshold, the collision risk is considered to be within the warning range, and a notification is sent to the user so that they can take appropriate measures to avoid a collision.

[0081] Suppose an underwater robot operates in a rectangular underwater region with dimensions of 10 meters long, 8 meters wide, and 6 meters high, with its initial position at the origin (0, 0, 0). A fish (the object to be detected) is located at coordinates (3, 4, 2) and swims at a speed of 0.5 meters per second along the direction of vector (1, 1, 1). The underwater robot moves at a speed of 1 meter per second along the positive x-axis.

[0082] Mathematical models can be constructed based on their position and motion information. For example, the position of a fish at time t can be represented as (3 + 0.5t, 4 + 0.5t, 2 + 0.5t), and the position of an underwater robot at time t is (t, 0, 0). Their relative positions and distances at any given time can be calculated using these expressions.

[0083] Risk Game Theory Layer: When calculating individual collision risk, the change in distance between the fish and the underwater robot over time is considered. When the distance between the fish and the underwater robot at a certain moment is less than the safe distance (assumed to be 1 meter), the individual collision risk increases. For global collision risk, it is assumed that there are several other fish in the area, and their individual positions and movements also affect the overall system's collision risk. For example, the movement of other fish may cause their distances to this fish or underwater robot to decrease, thus increasing the global collision risk. Global collision risk is predicted by comprehensively considering the movement trajectories and mutual distances of all fish and the underwater robot.

[0084] Early warning layer: If the calculated individual collision risk or global collision risk exceeds the set threshold, such as the estimated shortest distance between the fish and the underwater robot being less than 0.8 meters (threshold) in the individual collision risk, or the degree of proximity between multiple objects exceeding a certain danger level in the global collision risk assessment, then the early warning layer will determine that the collision risk is within the warning range and trigger the issuance of fish alert information to the user, informing the user of the collision risk and related fish information.

[0085] Therefore, by constructing a mathematical model in three-dimensional space and comprehensively considering individual and global collision risks, the collision risk between the object to be detected and the underwater robot can be accurately assessed. This comprehensive and detailed analysis method avoids the problem of inaccurate risk assessment caused by considering only a single factor or simple distance relationship, providing a more reliable guarantee for the safe operation of underwater robots. It can promptly determine whether the collision risk is within the warning range, providing users with early warnings. Users can take corresponding measures in a timely manner based on the prompts, such as adjusting the underwater robot's movement direction and speed, or pausing the task, thereby effectively avoiding collision accidents, protecting the safety of the underwater robot and fish, and ensuring the smooth progress of the detection task. This model can adapt to complex underwater environments, including situations where multiple objects to be detected coexist with different motion states. By considering global collision risk, it can handle the interactions and interference between multiple objects, enabling accurate assessment of collision risk even in multi-object environments, improving the robustness and adaptability of the system in practical applications.

[0086] Further optionally, in the above steps, the risk game layer employing the latent game model predicts the collision risk of the relative positions based on the mathematical model, including:

[0087] Using the mathematical model of the relative positions as the game subjects, the finite improvement characteristics of the latent game model are used to predict the individual collision risks of the game subjects, thereby obtaining the individual collision risk function of the game subjects in three-dimensional space; the individual collision risk function is mapped to the latent game function; the global collision risk function between the game subjects and surrounding objects is determined based on the convergence of Nash equilibrium; and the collision risk estimate corresponding to the game subjects is obtained based on the global collision risk function.

[0088] It is worth noting that in the above steps, the mathematical model of relative positions is used as the game subject, utilizing the finite improvement property of the latent game model. The finite improvement property means that during the game, the participants (here referring to the object to be detected and the underwater robot) will continuously adjust their strategies (such as direction of movement, speed, etc.) to improve their own situation. By analyzing the impact of this strategy adjustment on collision risk, an individual collision risk function is obtained. This function describes the probability that the object to be detected, based on its relative position and motion state with the underwater robot in three-dimensional space, will collide with the underwater robot alone.

[0089] Mapping the individual collision risk function to the latent game function allows for the consideration of individual behavior within a broader game context. The latent game function comprehensively considers the behaviors of all participants and their mutual influences, thus providing a foundation for analyzing global collision risk. This mapping links individual-level collision risk to the game structure of the entire system.

[0090] The global collision risk function is determined based on the convergence of Nash equilibrium. Nash equilibrium refers to a game in which all participants have chosen their optimal strategies, and without changing the strategies of other participants, no participant can achieve a better outcome by changing their own strategy. In this state, the system reaches a stable state. By analyzing the interactions between the game subject (the object to be detected) and surrounding objects (including other objects to be detected and the underwater robot) during the Nash equilibrium convergence process, the global collision risk function can be obtained. This function describes the collision risk faced by the object to be detected in the entire underwater environment, considering the motion and interactions of all relevant objects.

[0091] Finally, the estimated collision risk for each player is calculated based on the global collision risk function. This estimate is a quantitative indicator that comprehensively considers both individual collision risk and the global collision risk resulting from interactions with surrounding objects, providing an accurate basis for determining whether to issue a warning.

[0092] For example, suppose an underwater robot operates in a three-dimensional space, with two fish, denoted as fish A and fish B, as the objects to be detected. Taking fish A as an example, its relative position mathematical model with the underwater robot considers factors such as its coordinates, velocity, and direction of motion. Based on the finite improvement characteristics, the influence of fish A adjusting its velocity and direction on the probability of collision with the underwater robot is analyzed, resulting in an individual collision risk function for fish A. For example, when fish A accelerates towards the underwater robot, the value of this function increases, indicating an increased collision risk.

[0093] Furthermore, the individual collision risk functions of fish A and fish B are mapped into a latent game function. This function considers not only the individual collision risks of fish A and fish B with the underwater robot, but also the mutual influence between fish A and fish B. For example, the movement trajectories of fish A and fish B may interfere with each other, affecting their respective collision risks with the underwater robot.

[0094] Next, the global collision risk function is determined based on the convergence of Nash equilibrium. It is assumed that at a certain moment, the motion states of fish A, fish B, and the underwater robot reach an approximate Nash equilibrium, meaning that their respective policy adjustments will not further reduce the overall collision risk under the current circumstances. By analyzing their interactions in this stable state, the global collision risk function is obtained, which takes into account the position and velocity information of fish A, fish B, and the underwater robot.

[0095] Finally, the estimated collision risks for fish A and fish B are calculated based on the global collision risk function. For example, using a specific algorithm and parameter settings, the estimated collision risk for fish A is 0.6, and for fish B it is 0.4 (the range of values ​​can be set according to the actual situation; 0 represents no risk, and 1 represents extremely high risk). If the set warning threshold is 0.5, then the collision risk of fish A is within the warning range, while that of fish B has not yet reached the warning level.

[0096] Therefore, by constructing individual collision risk functions and global collision risk functions, this embodiment of the application can accurately analyze the collision risk of the object to be detected in complex underwater environments. It not only considers the direct collision probability between the object to be detected and the underwater robot, but also comprehensively considers the impact of other surrounding objects on the collision risk, making the risk assessment more accurate and comprehensive. Utilizing the theoretical foundation of latent game models and Nash equilibrium, a scientific decision-making framework is provided for collision risk prediction. This method can reasonably describe and analyze the interactions and strategy choices among multiple participants (the object to be detected and the underwater robot), thereby more accurately predicting the system's behavior and risk state, providing users with reliable decision-making basis. This embodiment of the application has strong adaptability and flexibility, capable of adapting to different numbers of objects to be detected in different motion states and various complex underwater environments. Whether it is a single object or multiple objects present simultaneously, accurate risk assessment can be performed through corresponding mathematical models and functions, and model parameters and warning thresholds can be adjusted according to actual conditions to meet the needs of different application scenarios.

[0097] Further, optionally, in the above steps, the early warning layer of the potential game model is used to determine whether the collision risk of the relative position falls within the early warning range, including:

[0098] The warning range is dynamically configured based on the category to which the object to be detected belongs; wherein the shape of the warning range matches the outline shape of the object to be detected; it is determined whether there is an overlapping area between the collision risk estimate and the pre-configured warning range; if the area of ​​the overlapping area exceeds a set threshold, it is determined that the relative position belongs to the warning range.

[0099] In the above steps, different categories of objects to be detected have different outline shapes and behavioral characteristics. Dynamically configuring the warning range based on their category allows for a more accurate fit to the actual situation. For example, larger fish require more safety space, so their warning range is correspondingly larger; while smaller fish require a relatively smaller warning range. This allows for personalized risk assessment based on the characteristics of different objects.

[0100] The risk level is determined by calculating the overlap between the estimated collision risk and a pre-configured warning range. The estimated collision risk is considered a region or numerical range, and the warning range is also considered a specific region. The system determines whether there is overlap and the size of the overlapping area. If the area of ​​the overlapping region exceeds a set threshold, it indicates a high collision risk, and the relative position is within the warning range, requiring appropriate measures to avoid a collision.

[0101] For example, suppose the objects to be detected are a shark and a small fish. For the shark, its warning range is a large elliptical area, determined by its species, because sharks are large and swim fast, requiring a large safety space. Suppose the area corresponding to the shark's estimated collision risk is a fan-shaped area in front of it. If this fan-shaped area overlaps with the elliptical warning range, and the overlapping area exceeds a set threshold (e.g., 30% of the warning range area, and the actual overlap reaches 40%), then the shark's current relative position is determined to be within the warning range, and there is a potential collision risk. For the small fish, its warning range is a smaller circular area. If the small fish's estimated collision risk corresponds to a smaller annular area around it, and the overlap area between the annular area and the circular warning range exceeds a set threshold (e.g., 20%), then the small fish's relative position is also determined to be within the warning range.

[0102] Therefore, by dynamically configuring the warning range according to the category of the object to be detected, the warnings are more in line with the actual situation, avoiding misjudgments or omissions caused by a one-size-fits-all warning approach, thus improving the accuracy of collision risk assessment. It can adapt to different types, sizes, and behavioral patterns of objects to be detected, whether large marine organisms or small aquatic animals, allowing for reasonable warning range configuration and risk assessment based on their own characteristics, enhancing the system's adaptability and versatility. This provides a more precise basis for subsequent decision-making. When the relative position is determined to be within the warning range, corresponding measures can be taken in a timely manner, such as adjusting the detection strategy and issuing alarms, helping to better protect the object to be detected and the surrounding environment and facilities, reducing losses from potential collision accidents.

[0103] In the above embodiments, the underwater environment includes objects to be detected (such as various fish) and underwater robots. The state information of each participant, such as position, speed, and direction of movement, is clearly defined. For each object to be detected, based on its relative position and motion state with the underwater robot and other potentially interacting objects (including other objects to be detected), the probability of it colliding with other objects individually is analyzed using the finite improvement characteristics of the latent game model, resulting in an individual collision risk function. For example, if an object to be detected moves rapidly towards the underwater robot, and the distance gradually decreases, its individual collision risk increases. Then, the individual collision risk functions of each object to be detected are mapped to the latent game function. This function comprehensively considers the behavior of all participants and their mutual influence. For example, the trajectories of two objects to be detected may interfere with each other, thus affecting their respective collision risks with the underwater robot; the latent game function needs to take into account this complex interrelationship.

[0104] Nash equilibrium refers to a stable state in a game where all participants have chosen their optimal strategies, and no participant can improve their outcome by changing their strategy if the strategies of other participants remain unchanged. This is achieved by analyzing the interactions between the game players (the objects to be detected) and surrounding objects during the Nash equilibrium convergence process to determine the global collision risk function. This process needs to consider the overall collision risk when the behaviors of all participants tend to stabilize. For example, when the motion states of multiple objects to be detected and an underwater robot reach an approximate Nash equilibrium, the impact of factors such as distance and velocity changes between them on the global collision risk is analyzed.

[0105] Taking all the above factors into account, a global collision risk function is calculated using specific algorithms and mathematical models. This process may involve weighted summation of various factors, integral operations, or the use of complex mathematical formulas and algorithms. The specific calculation method will vary depending on the underlying game model adopted and the specific problem setting. For example, different weights may be assigned based on factors such as the distance and speed between different objects, and then a series of calculations are used to quantify the global collision risk.

[0106] The calculation of the global collision risk function is a complex process based on a potential game model, which comprehensively considers various factors such as the individual behavior, mutual influence, and Nash equilibrium state of each participant in the system. The specific calculation method needs to be determined according to the specific application scenario and model settings.

[0107] In an optional example, suppose there is an underwater robot and two fish, fish A and fish B, in an underwater environment. To calculate the global collision risk function, three important factors are considered: relative distance, which is the straight-line distance between the fish and the underwater robot; the shorter the distance, the higher the risk of collision. Relative velocity includes the difference in velocity between the fish and the underwater robot, and the angle formed by their directions of motion. The larger the velocity difference, and the closer their directions of motion are to face-to-face, the higher the risk of collision. The influence of surrounding objects refers to the distance between fish A and fish B and their motion states. If the two fish are very close and their directions of motion interfere with each other, it may increase the risk of each of them colliding with the underwater robot.

[0108] To facilitate calculation, each factor is assigned a corresponding weight. Assume the weight of relative distance is 0.4, relative speed is 0.4, and the influence of surrounding objects is 0.2. First, let's calculate the individual collision risk factors for fish A: Relative distance between fish A and the underwater robot: Measurements or calculations show the distance between fish A and the underwater robot is 5 meters. To better measure the risk, we normalize the distance. Assuming the maximum safe distance is 10 meters, the normalized distance is calculated by subtracting the actual distance between fish A and the underwater robot from 1 and dividing by the maximum safe distance, i.e., 1 - 5 ÷ 10 = 0.5. In other words, the closer the distance, the closer the normalized value is to 1, representing a higher risk. Relative speed between fish A and the underwater robot: Assume fish A's speed is 2 meters per second, the underwater robot's speed is 1 meter per second, and the angle between their directions of motion is 60 degrees. Calculating the relative speed using a speed synthesis method, we obtain a relative speed of approximately 1.73 meters per second. Similarly, the relative velocity is normalized. Assuming the maximum relative velocity is 3 meters per second, the normalized relative velocity is the actual value of the relative velocity divided by the maximum relative velocity, i.e., 1.73 ÷ 3 ≈ 0.58. Assuming the distance between fish A and fish B is 3 meters, we set a condition that when the distance between the fish is less than 4 meters, they will influence each other, and the closer the distance, the greater the influence. Therefore, the normalized influence factor of fish A on fish B is calculated as 1 minus the actual distance between fish A and fish B divided by the maximum distance set for the influence, i.e., 1 - 3 ÷ 4 = 0.25. Then, the individual collision risk function value of fish A is calculated by multiplying the normalized value of the relative distance by the relative distance weight, adding the normalized value of the relative velocity multiplied by the relative velocity weight, and adding the influence factor of fish A on fish B multiplied by the weight of the influence of surrounding objects. That is, the individual collision risk function value of fish A = 0.4 × 0.5 + 0.4 × 0.58 + 0.2 × 0.25 = 0.2 + 0.232 + 0.05 = 0.482. The same method is used to calculate the individual collision risk function value of fish B. For example, if the distance between fish B and the underwater robot is 6 meters, the normalized value is 1 - 6 ÷ 10 = 0.4. Assuming fish B's speed is 1.5 meters per second and the angle between its speed and the underwater robot's direction of motion is 45 degrees, the calculated relative speed is approximately 1.12 meters per second, and the normalized value is 1.12 ÷ 3 ≈ 0.37. Because the distance between fish A and fish B is 3 meters, the influence factor of fish B on fish A is the same as the influence factor of fish A on fish B, which is 0.25. The individual collision risk function value for fish B is calculated as follows: 0.4 × 0.4 + 0.4 × 0.37 + 0.2 × 0.25 = 0.16 + 0.148 + 0.05 = 0.358. Finally, the global collision risk function value is calculated. An exemplary calculation method is used here: the individual collision risk function values ​​for fish A and fish B are added together, and then divided by 2 (in actual calculations, more complex models and algorithms may be used).That is: Global collision risk function value = (0.482 + 0.358) ÷ 2 = 0.84 ÷ 2 = 0.42. The obtained global collision risk function value of 0.42 is a relative index used to evaluate the degree of collision risk between the fish and the underwater robot in the entire system.

[0109] As an optional embodiment, in step 104, a fish alert is issued to the user based on the detection results and the fish's category, including:

[0110] If the overlapping area exceeds a set threshold, fish-related alert information is generated based on the fish category. This fish alert information is categorized into several dimensions: motion alerts, voice alerts, text alerts, and image alerts. The fish alert information is classified into alert priorities based on the overlapping area. The larger the overlapping area, the higher the alert priority of the fish alert information. The fish alert information is then pushed to the user in order of alert priority to inform the user of the currently appearing fish category and the risk of fish collision.

[0111] When the overlapping area exceeds a set threshold, it indicates a high risk of fish collision. At this point, relevant alerts are generated based on the fish's category. Because different fish categories have different characteristics and potential danger levels, targeted alerts are needed to ensure users understand the specific situation. Fish alerts are categorized into multiple dimensions, including sensory perception, voice, text, and images, to deliver information to users through different sensory channels, enabling them to receive and understand the alerts more comprehensively and promptly. Different users may have different preferences and perceptual abilities regarding different information presentation methods; multi-dimensional information can meet the needs of more users. Prioritizing fish alerts based on overlapping area area is crucial because a larger overlapping area signifies a higher collision risk, requiring users to pay more priority and urgent attention to this type of information. In this way, users can quickly assess the urgency of the situation based on alert priority and take appropriate measures.

[0112] For example, suppose the estimated collision risk of both a shark and a small fish overlaps with a warning area exceeding a set threshold. For the shark, being a large and highly aggressive fish, the generated warning information would be as follows: a haptic warning simulating water vibrations when a shark approaches; a voice warning stating, "Shark detected! Danger! Please be aware of your surroundings and avoid approaching!"; and a text warning displayed on the screen as, "Shark sighted, high collision risk, please be careful!" accompanied by an image of a shark. Because the overlapping area of ​​the shark is large, assuming it reaches 70% of the warning area, its warning priority would be set to high.

[0113] For small fish, the generated alerts might be: a haptic alert of a slight vibration simulating water flow changes caused by the fish swimming; a voice alert of "Small fish spotted, risk of collision, please be careful."; a text alert of "Small fish appeared, please avoid collision."; and an image alert of a small fish. The overlapping area of ​​the small fish is assumed to be 30% of the warning range, with a low priority. The system will first push shark alerts to the user because of their high priority, allowing the user to focus on the more dangerous situation first.

[0114] Therefore, by providing multi-dimensional prompts, the user's senses are stimulated from different angles, making it easier for them to notice the prompts and increasing their awareness of the presence of fish and the risk of collision, thus reducing the possibility of accidents caused by user negligence. Prioritizing warnings based on the area of ​​overlapping regions and pushing information accordingly allows users to quickly understand the urgency of the situation and rationally plan their response. This personalized and targeted information delivery method optimizes the user experience, enabling users to receive and process information more efficiently. Timely and accurate notifications of the types of fish currently present and the risk of collision help users take preventative measures in advance to avoid collisions, thereby enhancing user safety in relevant scenarios (such as underwater operations and diving).

[0115] In this embodiment, firstly, an image to be processed is acquired; the image to be processed contains objects to be detected in the underwater area. Then, a dynamic convolutional attention network is used to identify the category to which the objects to be detected belong. Next, the relative position between the objects to be detected and the underwater robot is determined based on the image to be processed. Finally, a latent game model is used to detect whether the collision risk of the relative position falls within the warning range, and fish alert information is issued to the user based on the detection results and the category; the fish alert information is used to indicate the types of fish observable by the underwater robot and the fish collision risk. This embodiment can achieve intelligent fish identification in underwater areas, improving the accuracy and efficiency of fish identification.

[0116] Figure 2 A schematic diagram of a fish detection system based on an underwater robot provided in this application embodiment, the system including at least:

[0117] The acquisition module is used to acquire an image to be processed; the image to be processed contains objects to be detected in the underwater area.

[0118] The recognition module is used to identify the category to which the object to be detected belongs through a dynamic convolutional attention network;

[0119] The positioning module is used to determine the relative position between the object to be detected and the underwater robot based on the image to be processed;

[0120] The early warning module is used to detect whether the collision risk of the relative position is within the warning range using a potential game model, and to issue fish alert information to the user based on the detection results and the fish category; the fish alert information is used to indicate the fish categories that the underwater robot can observe and the fish collision risk.

[0121] The system provided in this application embodiment can realize intelligent fish identification in underwater areas, improve the accuracy of fish identification, and increase the efficiency of fish identification.

[0122] This application provides a computer-readable storage medium according to its embodiments. For example... Figure 3 As shown, this embodiment provides a computer-readable storage medium 600 on which a computer program 611 is stored, which implements the aforementioned embodiment when executed by a processor.

[0123] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0124] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0125] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0126] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0127] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0128] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0129] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A fish detection method based on an underwater robot, characterized in that, The method includes: Acquire an image to be processed; the image to be processed contains the object to be detected in the underwater region; The category to which the object to be detected belongs is identified by a dynamic convolutional attention network; The relative position between the object to be detected and the underwater robot is determined based on the image to be processed; A potential game model is used to detect whether the collision risk of the relative position falls within the warning range. Based on the detection results and the fish category, a fish alert is issued to the user. The fish alert is used to indicate the fish categories that the underwater robot can observe and the fish collision risk. The step of using a potential game theory model to detect whether the collision risk of the relative positions falls within the warning range includes: A construction layer of a potential game model is adopted to construct a corresponding mathematical model in the three-dimensional space of the underwater region based on the relative position; Using the mathematical model of the relative positions as the game subjects, the individual collision risk of the game subjects is predicted by using the finite improvement characteristics of the potential game model, and the individual collision risk function of the game subjects in three-dimensional space is obtained. The individual collision risk function is mapped to the potential game function; The global collision risk function between the game player and surrounding objects is determined based on the convergence of Nash equilibrium. The collision risk estimate of the game subject is obtained based on the global collision risk function; the collision risk of the relative position includes at least: the individual collision risk of the object to be detected at the relative position and the global collision risk, the global collision risk including the collision risk caused by mutual interference between fish; An early warning layer employing a potential game theory model is used to determine whether the collision risk at the relative positions falls within the warning range.

2. The fish detection method based on an underwater robot according to claim 1, characterized in that, The step of identifying the category of the object to be detected using a dynamic convolutional attention network includes: Extract the identification information of the object to be detected; the identification information includes at least: real-time contour information, real-time surface information, and real-time pose information; The key identification features are subjected to dynamic convolution processing to obtain the dynamic shape features of the object to be detected under different activity states; the dynamic shape features include at least: scale texture features and tail fin shape features. Weight parameters are set for different activity states, and the category of the dynamic shape feature is predicted using the weight parameters to obtain the category to which the object to be detected belongs.

3. The fish detection method based on an underwater robot according to claim 2, characterized in that, The dynamic convolution processing of key identification features to obtain the dynamic shape features of the object to be detected under different activity states includes: Extract the features to be processed at different scales from the key identification features; The conditional parameter generator projects the features to be processed at different scales onto the behavior prediction space corresponding to different activity states to obtain the projected features of the object to be detected in different activity states. The conditional parameter generator is used to indicate the activity states corresponding to different fish and the types of projected features to be extracted in various activity states. The projected features under different activity states are classified and enhanced to obtain the dynamic shape features of the object under test under different activity states; the classification enhancement branch includes: a first enhancement branch for enhancing the tail fin features in the side-lying state, and a second enhancement branch for enhancing the scale texture features in the front-swimming state.

4. The fish detection method based on an underwater robot according to claim 1, characterized in that, The early warning layer employing a potential game theory model determines whether the collision risk at the relative position falls within the early warning range, including: The warning range is dynamically configured based on the category to which the object to be detected belongs; wherein the shape of the warning range matches the outline shape of the object to be detected. Determine whether there is an overlap between the estimated collision risk value and the pre-configured warning range; If the area of ​​the overlapping region exceeds a set threshold, then the relative position is determined to be within the warning range.

5. The fish detection method based on an underwater robot according to claim 4, characterized in that, The process of issuing fish alerts to users based on detection results and fish species includes: If the area of ​​the overlapping region exceeds a set threshold, fish-related prompts are generated based on the category; the fish prompts are divided into multiple dimensions: haptic prompts, voice prompts, text prompts, and image prompts. The fish alert information is classified into warning priority categories based on the area of ​​the overlapping region; wherein, the larger the area of ​​the overlapping region, the higher the warning priority of the fish alert information. The fish alerts are pushed to users in order of warning priority to inform them of the types of fish currently appearing and the risk of fish collisions.

6. A fish detection system based on an underwater robot, characterized in that, The system includes at least: The acquisition module is used to acquire an image to be processed; the image to be processed contains objects to be detected in the underwater area. The recognition module is used to identify the category to which the object to be detected belongs through a dynamic convolutional attention network; The positioning module is used to determine the relative position between the object to be detected and the underwater robot based on the image to be processed; The early warning module is used to detect whether the collision risk of the relative position falls within the warning range using a potential game model, and to issue fish alert information to the user based on the detection results and the fish category; the fish alert information is used to indicate the fish categories that the underwater robot can observe and the fish collision risk; The step of using a potential game theory model to detect whether the collision risk of the relative positions falls within the warning range includes: A construction layer of a potential game model is adopted to construct a corresponding mathematical model in the three-dimensional space of the underwater region based on the relative position; Using the mathematical model of the relative positions as the game subjects, the individual collision risk of the game subjects is predicted by using the finite improvement characteristics of the potential game model, and the individual collision risk function of the game subjects in three-dimensional space is obtained. The individual collision risk function is mapped to the potential game function; The global collision risk function between the game player and surrounding objects is determined based on the convergence of Nash equilibrium. The collision risk estimate of the game subject is obtained based on the global collision risk function; the collision risk of the relative position includes at least: the individual collision risk of the object to be detected at the relative position and the global collision risk, the global collision risk including the collision risk caused by mutual interference between fish; An early warning layer employing a potential game theory model is used to determine whether the collision risk at the relative positions falls within the warning range.

7. An electronic device, characterized in that, Includes memory used to store computer software programs; A processor is configured to read and execute the computer software program, thereby implementing the fish detection method based on an underwater robot as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The instructions include, when executed on a computer, causing the computer to perform the fish detection method based on an underwater robot as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Vehicle collision risk early warning method and system

    CN111785023A

  • Underwater robot operation risk assessment system based on multi-dimensional information calculation

    CN116362544A

  • Intelligent fish identifying and monitoring method and system based on multi-sensor data

    CN117214904A

  • Crab detection and counting method and device based on instance segmentation, medium and product

    CN118799716A