Fish detection method, device and equipment based on underwater robot and storage medium

By integrating dynamic convolutional attention networks and potential game models on underwater robots, the difficulties of fish identification and collision risk assessment in complex underwater environments are solved, and high accuracy and high efficiency of fish detection and early warning are achieved.

CN120108003AActive Publication Date: 2025-06-06SHENZHEN CHASING INNOVATION TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510571022.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-06
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

In complex underwater environments, it is difficult for the prior art to accurately identify fish, and the water flow causes changes in fish posture and position, increasing the difficulty of detection and affecting the accuracy and reliability of the equipment.

Method used

The fish detection method based on underwater robots is used to identify fish categories through dynamic convolutional attention networks, and the relative position between fish and underwater robots is determined based on images. The potential game model is used to detect collision risks and send fish prompt information to users.

Benefits of technology

It realizes intelligent fish identification in underwater areas, improves the accuracy and efficiency of fish identification, can adapt to different underwater environments and fish behaviors, and ensures the safe operation of underwater robots.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108003A_ABST
    Figure CN120108003A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent detection, in particular to a fish detection method, device and equipment based on an underwater robot and a storage medium. The method comprises the following steps: acquiring a to-be-processed image; the to-be-processed image comprises a to-be-detected object in an underwater area; identifying the category of the to-be-detected object through a dynamic convolution attention network; determining a relative position between the to-be-detected object and the underwater robot based on the to-be-processed image; and detecting whether the collision risk of the relative position belongs to an early warning range or not by adopting a potential game model, and sending fish prompt information to a user based on a detection result and a category to which the collision risk belongs. According to the invention, intelligent fish identification in the underwater area can be realized, the accuracy of fish identification is improved, and the fish identification efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent detection and processing, and in particular to a fish detection method, device, equipment and storage medium based on an underwater robot. Background Art

[0002] In the related art, underwater light weakens rapidly with increasing depth, and the scattering and absorption of light in water will cause image quality to deteriorate, which makes it difficult for the visual detection system to obtain clear fish images, affecting the identification and positioning of fish. In addition, water flow can change the posture and position of fish, increasing the difficulty of detection. At the same time, water flow may also cause instability of underwater robots, affecting the accuracy and reliability of detection equipment.

[0003] Therefore, it is urgent to design a technical solution to improve the accuracy and efficiency of fish identification in complex environments. Summary of the invention

[0004] In response to the technical problems existing in the prior art, the present application provides a fish detection method, device, equipment and storage medium based on an underwater robot, which are used to realize intelligent identification of fish in underwater areas, improve the accuracy of fish identification in complex environments, and improve the efficiency of fish identification.

[0005] In a first aspect, an embodiment of the present application provides a fish detection method based on an underwater robot, the method at least comprising: Acquire an image to be processed; the image to be processed includes an object to be detected in an underwater area; Identify the category to which the object to be detected belongs through a dynamic convolutional attention network; Determine the relative position between the object to be detected and the underwater robot based on the image to be processed; A potential game model is used to detect whether the collision risk of the relative position is within the warning range, and fish prompt information is sent to the user based on the detection result and the category to which it belongs; the fish prompt information is used to indicate the fish category observable by the underwater robot and the fish collision risk.

[0006] In a second aspect, an embodiment of the present application provides a fish detection system based on an underwater robot, the system comprising at least: An acquisition module, used for acquiring an image to be processed; the image to be processed includes an object to be detected in an underwater area; An identification module, used to identify the category to which the object to be detected belongs through a dynamic convolutional attention network; A positioning module, used to determine the relative position between the object to be detected and the underwater robot based on the image to be processed; The early warning module is used to use a potential game model to detect whether the collision risk of the relative position is within the early warning range, and to send fish prompt information to the user based on the detection result and the category to which it belongs; the fish prompt information is used to indicate the fish category observable by the underwater robot and the fish collision risk.

[0007] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device comprising a memory for storing a computer software program; The processor is used to read and execute the computer software program, thereby implementing the fish detection method based on the underwater robot of the first aspect.

[0008] In a fourth aspect, a computer-readable storage medium is provided, which includes instructions, and when the instructions are executed on a computer, the computer executes the underwater robot-based fish detection method of the first aspect.

[0009] The beneficial effects of the present application are: providing a fish detection method, device, equipment and storage medium based on an underwater robot. In this technical solution, first, an image to be processed is obtained; the image to be processed contains an object to be detected in an underwater area. Then, the category to which the object to be detected belongs is identified through a dynamic convolutional attention network. Next, the relative position between the object to be detected and the underwater robot is determined based on the image to be processed. Finally, a potential game model is used to detect whether the collision risk of the relative position is within the warning range, and fish prompt information is issued to the user based on the detection result and the category to which it belongs; the fish prompt information is used to indicate the fish category observable by the underwater robot and the risk of fish collision. In the embodiment of the present application, intelligent identification of fish in underwater areas can be achieved, the accuracy of fish identification can be improved, and the efficiency of fish identification can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 It is a flow chart of a fish detection method based on an underwater robot according to an embodiment of the present application; Figure 2 It is a structural schematic diagram of a fish detection system based on an underwater robot according to an embodiment of the present application; Figure 3 It is a structural schematic diagram of a medium in an embodiment of the present application. DETAILED DESCRIPTION

[0011] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0012] The embodiments of the present application provide a method, device, equipment and storage medium for fish detection based on an underwater robot. In the embodiments of the present application, first, an image to be processed is obtained; the image to be processed contains an object to be detected in an underwater area. Then, the category to which the object to be detected belongs is identified through a dynamic convolutional attention network. Next, the relative position between the object to be detected and the underwater robot is determined based on the image to be processed. Finally, a potential game model is used to detect whether the collision risk of the relative position is within the warning range, and fish prompt information is issued to the user based on the detection result and the category to which it belongs; the fish prompt information is used to indicate the fish category observable by the underwater robot and the risk of fish collision. The embodiments of the present application can realize intelligent identification of fish in underwater areas, improve the accuracy of fish identification, and improve the efficiency of fish identification.

[0013] Specifically, first, the dynamic convolution module in the dynamic convolution attention network dynamically generates convolution kernels based on the fish posture estimation results. Due to the diverse postures of underwater fish, such as swimming sideways and turning sharply, it is difficult for traditional fixed convolution kernels to effectively extract features in different postures. This method can adjust the weights of the convolution kernel in real time according to the posture. For example, when the fish is sideways, it focuses on extracting the tail fin features, and when it is swimming forward, it focuses on the body texture features, so as to extract features of fish in different postures more accurately, reduce recognition errors caused by posture changes, and improve the classification accuracy of closely related species (such as different types of tuna with similar appearances) and juveniles and adult fish (with different morphologies), which can be improved by 15%-20% compared to traditional CNN (such as ResNet).

[0014] Due to the complex underwater environment, factors such as light attenuation and turbid water can cause local image blur and low contrast. The multi-scale attention fusion module processes shallow (edge ​​details) and deep (semantic features) features in parallel, and adaptively fuses multi-scale information through the self-attention mechanism. In blurred areas, it can enhance the capture of edge details and combine deep semantic features to better understand the image content. For example, in dark areas, by adjusting the attention weights, the edge contours of fish can be highlighted, which helps to accurately identify the fish category. It has obvious advantages in fish recognition in low-quality underwater images and reduces the misjudgment rate caused by image quality problems.

[0015] Secondly, the dynamic convolutional attention network reduces the model parameters by 30% compared to the traditional CNN, and the amount of calculation is also reduced accordingly. This reduces the computing resources required by the model when processing the image to be processed, and speeds up the inference speed. Taking the embedded device carried by a common underwater robot as an example, it can complete the feature extraction of the image and the identification of the fish category in a relatively short time, meeting the real-time requirements, and the inference speed can reach about 25FPS, realizing the rapid identification of underwater fish. Compared with the traditional model, the recognition efficiency is greatly improved, and the underwater robot can process more images per unit time and detect fish in more areas.

[0016] From acquiring the image to be processed to completing fish category identification, relative position determination, and collision risk detection, the entire process is closely connected. Through reasonable algorithm design and model architecture, unnecessary intermediate links and data processing time are reduced. For example, when determining the relative position, calculations are performed directly based on the image to be processed, avoiding repeated processing and transmission of data, improving overall processing efficiency, and enabling underwater robots to complete fish detection tasks more efficiently.

[0017] Third, the dynamic convolutional attention network can automatically learn the key identification features of fish, without the need to manually design complex feature extraction rules. With the increase of training data and the optimization of the model, the model's ability to understand and identify the characteristics of various fish continues to improve, and it can adapt to new fish species and different underwater environments. For example, when encountering a new fish species, the model can gradually and accurately identify the species by learning its characteristics, realizing the intelligent and adaptive capabilities of fish identification.

[0018] The potential game model regards the underwater robot and the object to be detected (fish) as the two parties in the game, and comprehensively considers factors such as their relative position, motion state, and the behavior pattern of the fish. By analyzing the benefits and risks of both parties under different strategies, it dynamically determines whether the collision risk falls within the warning range. For example, for aggressive fish, the model can predict their possible attack behavior in advance based on their movement trajectory and speed changes, and issue a warning in time; for the normal swimming of harmless fish, it can accurately judge that they will not pose a threat to the underwater robot, avoiding unnecessary warnings. This decision-making method based on the game model enables the underwater robot to respond to complex underwater environments and fish behaviors more intelligently and make reasonable decisions.

[0019] Fourth, the fish prompt information sent to users based on the detection results and the categories they belong to not only includes the types of fish that can be observed by the underwater robot, but also accurately indicates the risk of fish collision. This provides users (such as underwater workers, scientific researchers, etc.) with more comprehensive and accurate information, helping them to better understand the underwater environment and fish conditions. For example, in marine scientific research monitoring, researchers can accurately record the types and distribution of fish based on the prompt information, and reasonably plan the underwater robot's route of action based on the collision risk information to ensure the smooth progress of the monitoring task; in underwater operations, operators can take timely measures based on the prompt information to avoid collisions between underwater robots and fish, and ensure the safety of equipment and the normal development of operations.

[0020] In the embodiments of the present application, whether in clear shallow sea areas or turbid deep sea areas, or under strong light or low light conditions, the dynamic convolutional attention network and the potential game model can adapt to different environmental conditions through their own characteristics and algorithm adjustments, and maintain high fish recognition accuracy and collision risk detection capabilities. For example, in turbid water bodies, the multi-scale attention fusion module can better capture the characteristics of fish, and the potential game model can also comprehensively consider the impact of the water environment on fish movement and accurately judge the collision risk, making this technology highly practical and reliable in various underwater environments.

[0021] By using a dynamic convolutional attention network to identify the category of the object to be detected, it can better handle the problems of a wide variety of fish species, different shapes, and individual differences. Compared with related technologies, it improves the accuracy and adaptability of fish category recognition, and effectively solves the problem that it is difficult to accurately distinguish different types of fish in related technologies. The relative position between the object to be detected and the underwater robot is determined based on the image to be processed, which provides an accurate data basis for subsequent collision risk detection, overcomes the problem that the relative position of fish and underwater robots may not be accurately obtained in the prior art, and helps to more accurately evaluate the spatial relationship between fish and underwater robots. The potential game model is used to detect whether the collision risk of the relative position is within the warning range, which can comprehensively consider multiple factors and make a more comprehensive and accurate assessment of the collision risk, solve the problem that the collision risk assessment in related technologies is unscientific and untimely, and provide a guarantee for the safe operation of underwater robots. Based on the detection results and the category to which they belong, fish prompt information is sent to the user, which includes both the fish category and the fish collision risk, so that the user can have a more comprehensive understanding of the fish situation around the underwater robot, solves the problem of single and incomplete information feedback in related technologies, and helps users make more reasonable decisions.

[0022] The fish detection scheme based on underwater robots provided in the embodiments of the present application can also be executed by electronic devices, which can be servers, server clusters, or cloud servers. The electronic device can also be a terminal device such as a mobile phone, a computer, a tablet computer, a wearable device, or a dedicated device (such as a dedicated terminal device with a fish detection system based on an underwater robot, etc.). These electronic devices can also be equipped with the chips introduced in the above embodiments. Alternatively, these electronic devices can also be installed with a service program for executing the fish detection scheme based on an underwater robot.

[0023] Figure 1 A schematic diagram of a fish detection method based on an underwater robot provided in an embodiment of the present application. Figure 1 As shown, the method includes: 101, obtaining an image to be processed; the image to be processed includes an object to be detected in an underwater area; 102, identifying the category to which the object to be detected belongs through a dynamic convolutional attention network; 103, determining the relative position between the object to be detected and the underwater robot based on the image to be processed; 104 , using a potential game model to detect whether the collision risk of the relative position is within a warning range, and sending fish prompt information to the user based on the detection result and the category.

[0024] In an embodiment of the present application, the fish prompt information is used to indicate the types of fish that can be observed by the underwater robot and the risk of fish collision. For example, the fish prompt information may include the following content: Tuna (high-speed swimming fish) is detected, 1.2 meters away from the underwater robot, collision risk: high; it is recommended to slow down and change course immediately. For example, the fish prompt information may include the following content: Clownfish is detected, 3 meters away from the underwater robot, collision risk: low; normal operation is possible. For example, the fish prompt information may include the following content: Shark is detected, 2.5 meters away from the underwater robot, and is accelerating closer, collision risk: extremely high; please control the robot to evacuate the area immediately.

[0025] In 101, an image to be processed is obtained. Further, the image to be processed includes an object to be detected in an underwater area.

[0026] Specifically, underwater robots are equipped with sensors such as cameras and sonars to collect image data in underwater environments. The camera directly captures the underwater scene to obtain visual images; the sonar converts the reflected signal into image information by emitting and receiving sound waves, which is suitable for environments with insufficient light or turbidity. The data collected by multiple sensors can be fused to provide more comprehensive information for subsequent processing. Therefore, multi-source data collection ensures that the acquired images to be processed contain rich underwater scene information, covering the objects to be detected under different lighting and water quality conditions, laying a data foundation for subsequent fish detection and identification. At the same time, multi-sensor fusion makes up for the limitations of a single sensor and improves the reliability and integrity of the data.

[0027] In 102, the category to which the object to be detected belongs is identified through a dynamic convolutional attention network.

[0028] Specifically, the dynamic convolution module uses a conditional parameter generator to dynamically generate convolution kernels based on the fish posture estimation results. Posture estimation obtains parameters such as the fish body deflection angle and tail fin swing amplitude, generates convolution kernels with corresponding weights, and specifically extracts fish features in different postures. The multi-scale attention fusion module processes shallow edge details and deep semantic features in parallel, and adaptively fuses them through the self-attention mechanism to solve the feature imbalance problem caused by uneven illumination of underwater images. In this way, it can effectively cope with fish posture changes and underwater image quality problems, accurately extract key identification features of fish, and significantly improve the accuracy of fish category recognition. The classification accuracy of closely related species and juvenile and adult fish is improved by 15%-20% compared with traditional CNN, and a high recognition rate can still be maintained in complex underwater environments.

[0029] As an optional embodiment, in 102, identifying the category to which the object to be detected belongs by using a dynamic convolutional attention network includes: Extract the information to be identified of the object to be detected; the information to be identified includes at least: real-time contour information, real-time surface information, and real-time posture information; perform dynamic convolution processing on the key identification features to obtain dynamic shape features of the object to be detected in different activity states; the dynamic shape features include at least: scale texture features and tail fin shape features; set weight parameters under different activity states, use the weight parameters to predict the category of the dynamic shape features, and obtain the category to which the object to be detected belongs.

[0030] It can be understood that in 102, the real-time contour information, real-time surface information and real-time posture information of the object to be detected are extracted from the image to be processed. This information can fully describe the appearance and state of the fish. The contour information outlines the overall shape of the fish, the surface information contains the details of the fish's body surface such as color, stripes, etc., and the posture information reflects the movement state and body posture of the fish, providing rich raw data for the subsequent accurate identification of fish categories. Then, using the dynamic convolution module, according to the fish's posture estimation results, the conditional parameter generator dynamically generates a convolution kernel to process the key identification features. For fish in different activity states, by dynamically adjusting the convolution kernel weights, its dynamic appearance features, such as scale texture features and tail fin shape features, can be more effectively extracted. This dynamic convolution method can adapt to the characteristic changes of fish in different postures, overcoming the problem that traditional fixed convolution kernels are difficult to cope with posture diversity.

[0031] Furthermore, weight parameters are set for dynamic appearance features under different activity states, and the importance of various features in fish category recognition is comprehensively considered. The dynamic appearance features are weighted and fused through these weight parameters, and then input into the classifier for category prediction, thereby obtaining the category to which the object to be detected belongs. This method can reasonably weight different features according to their contribution to category judgment, thereby improving the accuracy of recognition.

[0032] For example, suppose the object to be detected is a fish, and at a certain moment, the following information is obtained: Real-time contour information: The body is relatively slender and the head is relatively pointed. Real-time surface information: The body surface has a silver luster with black spots. Real-time posture information: The body is slightly tilted, and the tail fin swings at a large angle, showing a fast swimming posture.

[0033] After dynamic convolution processing, the dynamic shape features extracted are scale texture features and tail fin shape features. Scale texture features include small and closely arranged scales with a certain directionality. Tail fin shape features include a forked tail fin with a smooth edge.

[0034] Then, based on the preset weight parameters under different activity states, these dynamic appearance features are comprehensively evaluated and predicted. For example, the slender body outline, silver surface with black spots, fast swimming posture, and specific scale texture and tail fin shape are all consistent with the characteristic pattern of sea bass, so the object to be detected is predicted to be sea bass.

[0035] In this way, by extracting multi-dimensional information to be identified and performing dynamic convolution processing on key identification features, the unique characteristics of different fish in various activity states can be captured more comprehensively and accurately. Setting weight parameters for category prediction further optimizes the comprehensive consideration of different features, thereby significantly improving the accuracy of fish category recognition, and can effectively distinguish between juvenile fish and adult fish, closely related species and other easily confused categories. The dynamic convolutional attention network can dynamically adjust the convolution kernel weights and attention focus according to the real-time posture and activity status of the fish, and adapt to the diversity of fish postures in underwater images and feature changes caused by environmental factors. This enables the model to maintain good performance in different underwater scenes and fish behavior patterns, with stronger generalization and adaptability.

[0036] Compared with traditional convolutional neural networks, this method can reduce the number of model parameters while achieving high-precision recognition by dynamically generating convolution kernels and adaptive feature fusion mechanisms. For example, compared with some classic CNN models (such as ResNet), the parameters can be reduced by about 30%, which is conducive to the deployment and operation of the model in resource-constrained environments such as edge devices, and reduces the requirements for hardware devices.

[0037] In practical applications, the training process of the dynamic convolutional attention network can include steps such as data preparation, model initialization, loss function definition, training loop, and model evaluation.

[0038] During the data preparation process, a large number of underwater images containing various fish are collected as training data. At the same time, the fish in the images are labeled with categories and some key feature information, such as the fish's outline, posture, scale texture, tail fin shape, etc. These labeled information will serve as supervision signals for model training.

[0039] In order to increase the diversity of data and improve the generalization ability of the model, data augmentation operations are performed on the original images, such as random cropping, flipping, rotation, and adjustment of brightness, contrast, and color saturation. This allows the model to be exposed to more fish images with different perspectives and appearances during training, reducing the risk of overfitting.

[0040] According to the architecture design of the dynamic convolutional attention network, a network model including dynamic convolution module, multi-scale attention fusion module and other components is constructed. The number of parameters and connection methods of each layer in the network are determined to prepare for model training.

[0041] Use appropriate initialization methods to initialize the model parameters. For example, you can use random initialization methods to assign initial values ​​to convolution kernel weights, bias terms, and other learnable parameters. Some common initialization strategies include Xavier initialization and Kaiming initialization. Good initialization can help the model converge faster.

[0042] In the process of defining the loss function, the cross entropy loss function is usually selected as the loss function of the model according to the characteristics of the fish classification task. The cross entropy loss function can measure the difference between the model prediction result and the true label, and has good performance for multi-classification problems. During the training process, the model will adjust the parameters by minimizing the loss function so that the prediction result is as close to the true label as possible.

[0043] In order to prevent the model from overfitting and improve the generalization ability of the model, regularization terms such as L1 regularization or L2 regularization can be added to the loss function. The regularization term can constrain the parameters of the model so that they will not be too large, thereby avoiding overfitting caused by the model being too complex.

[0044] Then, the prepared training data is input into the model, and the data is forward propagated in the network. In the dynamic convolution module, the convolution kernel is dynamically generated according to the posture estimation results of the fish, and the input data is convolved to extract the features of the fish from different perspectives. The multi-scale attention fusion module processes shallow and deep features in parallel, and adaptively fuses multi-scale information through the self-attention mechanism to obtain the final feature representation. Finally, the feature representation is input into the classifier to obtain the model's prediction results for the fish category. According to the model's prediction results and the true label, the loss value is calculated using the defined loss function, which reflects the gap between the model's current prediction results and the actual situation. The gradient of the loss function to the model parameters is calculated by the back-propagation algorithm, and the gradient represents the rate of change of the loss function under the current parameter value. According to the gradient information, the optimization algorithm (such as stochastic gradient descent, Adagrad, Adadelta, etc.) is used to update the model parameters so that the loss function value gradually decreases. When updating the parameters, it is necessary to update according to the corresponding rules according to different parameter types and network structures to ensure that the model can converge to the optimal solution.

[0045] Repeat the above forward propagation, loss calculation and back propagation process to perform multiple iterations of training on the entire training data set. As the training progresses, the model parameters are constantly adjusted, the loss function value gradually decreases, and the model performance gradually improves. During the training process, the model parameters can be saved regularly so that they can be restored and evaluated when the training is interrupted or the model needs to be used.

[0046] A portion of the data is divided from the original data set as a validation set to evaluate the performance of the model during the training process. The data in the validation set does not participate in the training of the model, but is used to simulate the performance of the model in actual applications, so as to promptly detect whether the model is overfitting or underfitting. Select appropriate evaluation indicators to measure the performance of the model, such as accuracy, recall, F1 value, etc. For the task of fish classification, accuracy is an important indicator, which indicates the proportion of fish categories correctly predicted by the model. The recall rate reflects the proportion of a certain type of fish that the model can correctly identify among all samples of that type. The F1 value is the harmonic mean of accuracy and recall, which can comprehensively evaluate the performance of the model.

[0047] Optimize and adjust the model based on the evaluation results on the validation set. If the model has low accuracy on the validation set, you may need to increase training data, adjust the network structure, optimize hyperparameters, and other methods to improve the model's performance. If the model is overfitting, that is, the loss function value on the training set is very low, but the accuracy on the validation set decreases, you can consider adding regularization terms, reducing network complexity, stopping training early, and other methods to solve the problem. Through continuous evaluation and optimization, the model can achieve the best performance on the validation set.

[0048] Use an independent test set to perform a final performance test on the trained model. The data in the test set has not been used in the model training and verification process, and can truly reflect the generalization ability of the model in practical applications. Calculate various evaluation indicators of the model on the test set, such as accuracy, recall, F1 value, etc., to evaluate the final performance of the model. If the performance of the model on the test set meets the requirements, the model can be deployed in practical applications. If the performance does not meet the requirements, it is necessary to further analyze the reasons and improve and optimize the model until satisfactory performance is achieved.

[0049] Further optionally, in the above steps, the key identification features are subjected to dynamic convolution processing to obtain dynamic appearance features of the object to be detected in different activity states, including: Extract the features to be processed at different scales from the key identification features; project the features to be processed at different scales to the behavior prediction space corresponding to different activity states through a condition parameter generator, so as to obtain the projection features of the object to be detected in different activity states; the condition parameter generator is used to indicate the activity states corresponding to different fish and the types of projection features to be extracted under various activity states; classify and strengthen the projection features under different activity states, and obtain the dynamic appearance features of the object to be detected under different activity states.

[0050] In the embodiment of the present application, the classified strengthening branches include: a first strengthening branch for strengthening the tail fin characteristics in the sideways state, and a second strengthening branch for strengthening the scale texture characteristics in the swimming state.

[0051] Specifically, in the above steps, features of different scales can capture information at different levels of detail of the object to be detected. For example, small-scale features may contain finer texture information, while large-scale features can reflect the overall shape and structure information. By extracting features to be processed at different scales, key identification features can be fully described, providing a rich data foundation for subsequent analysis.

[0052] The conditional parameter generator projects the features to be processed at different scales to the behavior prediction space corresponding to different activity states according to the activity characteristics of different fish. This means that the model will extract features related to the activity state in a targeted manner according to the possible activity state of the fish. For example, for fast swimming states, the model will pay more attention to features that can reflect changes in speed and posture; while for static or slow swimming states, it will pay more attention to relatively stable features such as shape and texture. This allows the model to better adapt to changes in the appearance of fish in different activity states and improve the accuracy of extracting dynamic appearance features.

[0053] The classification enhancement branch further enhances the projection features in different activity states, highlighting the key features related to specific activity states. In the embodiment of the present application, the first enhancement branch is used to enhance the tail fin features in the sideways state, because the tail fin plays a key role in the turning and propulsion of the fish when it swims sideways. By enhancing the tail fin features, the fish in the sideways state can be better identified; the second enhancement branch is used to enhance the scale texture features in the normal swimming state. The scale texture in the normal swimming state can be used as one of the important features to distinguish different fish. Strengthening this feature helps to improve the accuracy of fish identification in the normal swimming state.

[0054] Exemplarily, suppose there is a goldfish as the object to be detected. First, extract the features to be processed at different scales in the key identification features of the goldfish, such as the fine texture of the goldfish scales can be captured at a small scale, and the overall body outline and general posture of the goldfish can be seen at a large scale. Then, the conditional parameter generator projects these features to be processed at different scales to the corresponding behavior prediction space according to the common activity states of the goldfish, such as fast swimming, slow swimming, and stillness. For example, when the goldfish is in a fast swimming state, the conditional parameter generator will pay more attention to speed-related features, such as the swing amplitude of the body, the swing frequency of the tail fin, etc., and project the corresponding features to be processed into the behavior prediction space of the fast swimming state to obtain the projected features in the fast swimming state. Finally, the projected features are processed by the classification reinforcement branch. For the sideways state, the first reinforcement branch will focus on enhancing the tail fin features, such as highlighting the shape of the tail fin, the swing angle, etc.; for the swimming state, the second reinforcement branch will enhance the scale texture features, so that the model can more clearly identify the unique scale texture pattern of the goldfish when swimming, thereby accurately judging the type of goldfish and its activity state.

[0055] Therefore, by extracting the features to be processed at different scales and projecting them to the behavior prediction space corresponding to different activity states, the key features of the object to be detected in different activity states can be captured more comprehensively and accurately, avoiding the limitations of a single scale or fixed feature extraction method. For example, for fish of different sizes and types, feature extraction at different scales can adapt to their differences in appearance and details, while projection for the activity state can better handle the dynamic changes in the appearance of fish during movement, making the extracted dynamic appearance features more representative and distinguishable. The classification reinforcement branch performs special reinforcement processing on the projection features under different activity states, so that the model can better adapt to the feature changes of various fish under different activity states, and improves the adaptability of the model to fish detection in different scenarios and conditions. For example, in a complex underwater environment, the activity states of fish are diverse. By strengthening the key features of a specific activity state, the model can more accurately identify fish in different activity states. Even if it encounters some specific situations that have not appeared in the training data, it can make reasonable judgments based on the strengthened features, thereby enhancing the generalization ability of the model.

[0056] After the above processing, the model can more accurately extract the dynamic appearance features of the object to be detected in different activity states, which are of great significance for distinguishing different types of fish. Therefore, it can significantly improve the accuracy of fish identification, reduce the occurrence of misjudgment and missed judgment, and lay a solid foundation for providing users with accurate fish prompt information in the future, making the entire fish detection system more reliable and practical.

[0057] In 103, the relative position between the object to be detected and the underwater robot is determined based on the image to be processed. Specifically, the position of the fish in the image is located by image processing technology, such as the target detection algorithm, and the position and posture information of the underwater robot itself (provided by the inertial navigation system, etc.) and the depth information of the image (which can be obtained through sonar data or stereo vision calculation) are combined to calculate the three-dimensional spatial position (X, Y, Z coordinates) and direction information of the fish relative to the underwater robot. Thus, the relative position of the fish and the underwater robot is accurately determined, providing accurate data support for subsequent collision risk assessment, enabling the underwater robot to perceive the distribution and distance of the surrounding fish, and providing a basis for decision-making.

[0058] In 104, a potential game model is used to detect whether the collision risk of the relative position is within the warning range. Specifically, the potential game model regards the underwater robot and the fish as the two parties in the game, and considers factors such as the movement strategies, relative positions, speeds, and fish categories of both parties. By constructing a game model, the benefits and risks of both parties under different strategies are analyzed, and the collision risk probability is calculated. When the risk probability exceeds the set threshold, it is determined to enter the warning range, and combined with the category of the fish, the corresponding fish prompt information is generated and sent to the user. In this way, the collision risk can be intelligently assessed, and the warning strategy can be dynamically adjusted according to the fish behavior pattern and the interaction with the robot to reduce false alarms and missed alarms. Compared with the traditional fixed threshold method, the detection accuracy of the collision risk is greatly improved, ensuring the safety of underwater robot operations.

[0059] As an optional embodiment, in 104, using a potential game model to detect whether the collision risk of the relative position is within the warning range includes: A construction layer of a potential game model is adopted to construct a corresponding mathematical model in the three-dimensional space of an underwater area based on the relative position; a risk game layer of the potential game model is adopted to predict the collision risk of the relative position according to the mathematical model; the collision risk of the relative position includes at least: the individual collision risk and the global collision risk of the object to be detected at the relative position; and a warning layer of the potential game model is adopted to determine whether the collision risk of the relative position is within the warning range.

[0060] In the three-dimensional space of the underwater area, a mathematical model is constructed based on the relative position of the object to be detected and the underwater robot. This model takes into account factors such as the coordinates, distance, movement direction and speed of the two in space. By converting the actual physical position relationship into a mathematical expression, a quantitative basis is provided for subsequent risk analysis. For example, a Cartesian coordinate system can be used to represent their position, and a vector can be used to describe the movement direction and speed, thereby establishing a mathematical model that can accurately reflect their relative motion state in three-dimensional space.

[0061] Based on the constructed mathematical model, the risk game layer will take into account the behavioral strategies of the objects to be detected and the underwater robot and the interactions between them. The individual collision risk mainly focuses on the possibility of collision between the object to be detected and the underwater robot, which depends on its own motion state, distance from the underwater robot and other factors. The global collision risk will consider the possibility of collision between multiple objects to be detected and the underwater robot in the entire underwater environment and between them from a more macro perspective. This requires comprehensive consideration of the position, motion state and potential mutual influence of all relevant objects, and predicting the collision risk through analysis and calculation of these factors. For example, when multiple fish approach the underwater robot at the same time, not only the individual collision risk between each fish and the underwater robot should be considered, but also the increase in global collision risk caused by mutual interference between fish.

[0062] The early warning layer will compare the predicted collision risk with the pre-set threshold to determine whether the collision risk is within the early warning range. This threshold is determined based on factors such as the underwater robot's safety requirements, working environment, and actual application needs. If the calculated individual collision risk or global collision risk exceeds the corresponding threshold, then the collision risk is considered to be within the early warning range, and a prompt message needs to be issued to the user so that the user can take appropriate measures to avoid collision.

[0063] Assume that the underwater robot is working in a rectangular underwater area with a length, width, and height of 10 meters, 8 meters, and 6 meters respectively, and its initial position is at the origin of the coordinate system (0, 0, 0). There is a fish (the object to be detected) at the coordinates (3, 4, 2), swimming along the direction of the vector (1, 1, 1) at a speed of 0.5 meters per second, and the underwater robot moves along the positive direction of the x-axis at a speed of 1 meter per second.

[0064] Based on their position and motion information, a mathematical model can be constructed. For example, the position of the fish at time t can be expressed as (3 + 0.5t, 4 + 0.5t, 2 + 0.5t), and the position of the underwater robot at time t is (t, 0, 0). Through these expressions, their relative positions and distances at any time can be calculated.

[0065] Risk game layer: When calculating the individual collision risk, consider how the distance between the fish and the underwater robot changes over time. When the distance between the fish and the underwater robot at a certain moment is less than the safe distance (assuming it is 1 meter), the individual collision risk will increase. For the global collision risk, assume that there are several other fish in the area, and their respective positions and motion states will also affect the collision risk of the entire system. For example, the movement of other fish may cause the distance between them and this fish or underwater robot to become smaller, thereby increasing the global collision risk. The global collision risk is predicted by comprehensively considering the movement trajectories and mutual distances of all fish and underwater robots.

[0066] Warning layer: If the calculated individual collision risk or global collision risk exceeds the set threshold, for example, the estimated shortest distance between the fish and the underwater robot in the individual collision risk is less than 0.8 meters (threshold), or the degree of mutual proximity between multiple objects in the global collision risk assessment exceeds a certain danger level, then the warning layer will determine that the collision risk is within the warning range and trigger a fish prompt message to the user, informing the user of the collision risk and the relevant fish information.

[0067] Therefore, by constructing a mathematical model in three-dimensional space and comprehensively considering individual and global collision risks, the collision risk between the object to be detected and the underwater robot can be accurately evaluated. This comprehensive and detailed analysis method can avoid the problem of inaccurate risk assessment caused by considering only a single factor or a simple distance relationship, and provide a more reliable guarantee for the safe operation of the underwater robot. It can timely determine whether the collision risk is within the warning range and provide early warning to the user. The user can take corresponding measures in time according to the prompt information, such as adjusting the movement direction and speed of the underwater robot or pausing the task, so as to effectively avoid the occurrence of collision accidents, protect the safety of the underwater robot and fish, and also ensure the smooth progress of the detection task. The model can adapt to complex underwater environments, including the situation where multiple objects to be detected exist at the same time and have different motion states. By considering the global collision risk, the interaction and interference between multiple objects can be handled, so that the collision risk can be accurately evaluated in a multi-target environment, which improves the robustness and adaptability of the system in practical applications.

[0068] Further optionally, in the above steps, using the risk game layer of the potential game model to predict the collision risk of the relative position according to the mathematical model includes: The mathematical model of the relative position is used as the game subject, and the limited improvement characteristics of the potential game model are used to predict the individual collision risks of the game subject to obtain the individual collision risk function of the game subject in the three-dimensional space; the individual collision risk function is mapped to the potential game function; based on the convergence of Nash equilibrium, the global collision risk function between the game subject and the surrounding objects is determined; based on the global collision risk function, the collision risk estimation value corresponding to the game subject is obtained.

[0069] It is worth noting that in the above steps, the mathematical model of relative position is used as the game subject, and the limited improvement characteristics of the potential game model are used. The limited improvement characteristics mean that in the game process, the participants (here refers to the object to be detected and the underwater robot) will constantly adjust their strategies (such as movement direction, speed, etc.) to improve their own conditions. By analyzing the impact of this strategy adjustment on the collision risk, the individual collision risk function is obtained. This function describes the possibility of the object to be detected colliding with the underwater robot alone in three-dimensional space based on factors such as its relative position and movement state with the underwater robot.

[0070] The purpose of mapping the individual collision risk function to the potential game function is to consider individual behavior in a broader game environment. The potential game function can comprehensively consider the behavior of all participants and their mutual influence, thus providing a basis for analyzing the global collision risk. Through this mapping, the collision risk at the individual level is linked to the game structure of the entire system.

[0071] The global collision risk function is determined based on the convergence of Nash equilibrium. Nash equilibrium means that in a game, all participants have chosen their own optimal strategies, and no participant can change their strategy to obtain a better result when the strategies of other participants remain unchanged. In this case, the system reaches a stable state. By analyzing the interaction between the game subject (the object to be detected) and the surrounding objects (including other objects to be detected and underwater robots) during the convergence of Nash equilibrium, the global collision risk function can be obtained. This function describes the collision risk faced by the object to be detected in the entire underwater environment, considering the movement and mutual influence of all relevant objects.

[0072] Finally, the estimated collision risk value corresponding to the game subject is calculated based on the global collision risk function. This estimated value is a quantitative indicator that comprehensively considers the individual collision risk and the global collision risk after interaction with surrounding objects, and can provide an accurate basis for determining whether to issue an early warning.

[0073] For example, suppose that an underwater robot is operating in a three-dimensional space, and there are two fish in the space as objects to be detected, denoted as fish A and fish B. Taking fish A as an example, the mathematical model of its relative position with the underwater robot takes into account factors such as their coordinates, speed, and direction of movement. According to the limited improvement characteristics, the influence of fish A's adjustment of speed and direction on the possibility of collision with the underwater robot is analyzed, and the individual collision risk function of fish A is obtained. For example, when fish A accelerates towards the underwater robot, the value of this function will increase, indicating an increase in the risk of collision.

[0074] Then, the individual collision risk functions of fish A and fish B are mapped into the potential game function. This function not only considers the collision risk of fish A and fish B with the underwater robot, but also considers the mutual influence between fish A and fish B. For example, the movement trajectories of fish A and fish B may interfere with each other, affecting their respective collision risks with the underwater robot.

[0075] Next, the global collision risk function is determined based on the convergence of Nash equilibrium. Assume that at a certain moment, the motion states of fish A, fish B, and the underwater robot reach a state that is close to Nash equilibrium, that is, their respective strategy adjustments will not further reduce the overall collision risk under the current circumstances. By analyzing the interaction between them in this stable state, the global collision risk function is obtained, which takes into account the position and velocity information of fish A, fish B, and the underwater robot.

[0076] Finally, the collision risk estimates for fish A and fish B are calculated based on the global collision risk function. For example, through a specific algorithm and parameter setting, the collision risk estimates for fish A are calculated to be 0.6, and for fish B to be 0.4 (the value range can be set according to the actual situation, 0 means no risk, 1 means extremely high risk). If the warning threshold is set to 0.5, then the collision risk of fish A is within the warning range, while that of fish B has not yet reached the warning level.

[0077] Thus, in the embodiment of the present application, by constructing individual collision risk function and global collision risk function, the collision risk of the object to be detected in a complex underwater environment can be accurately analyzed. Not only the direct collision possibility between the object to be detected and the underwater robot is considered, but also the influence of other surrounding objects on the collision risk is fully considered, making the risk assessment more accurate and comprehensive. Using the theoretical basis of potential game model and Nash equilibrium, a scientific decision-making framework is provided for collision risk prediction. This method can reasonably describe and analyze the interaction and strategy selection between multiple participants (objects to be detected and underwater robots), so as to more accurately predict the behavior and risk state of the system and provide users with a reliable decision-making basis. The embodiment of the present application has strong adaptability and flexibility, and can adapt to objects to be detected of different numbers and different motion states and various complex underwater environments. Whether it is a single object or multiple objects at the same time, accurate risk assessment can be performed through the corresponding mathematical model and function, and the model parameters and warning thresholds can be adjusted according to the actual situation to meet the needs of different application scenarios.

[0078] Further optionally, in the above steps, the warning layer of the potential game model is used to determine whether the collision risk of the relative position is within the warning range, including: Dynamically configure a warning range based on the category to which the object to be detected belongs; wherein the shape of the warning range matches the contour shape of the object to be detected; determine whether there is an overlapping area between the collision risk estimate and the pre-configured warning range; if the area of ​​the overlapping area exceeds a set threshold, determine that the relative position is within the warning range.

[0079] In the above steps, different categories of objects to be detected have different contour shapes and behavioral characteristics. Dynamically configuring the warning range based on their categories can more accurately fit the actual situation. For example, larger fish require a larger safety space, so their warning range is correspondingly larger; while for smaller fish, the warning range is relatively small. In this way, personalized risk assessment can be performed based on the characteristics of different objects.

[0080] The risk condition is determined by calculating the overlap area between the collision risk estimate and the pre-configured warning range. The collision risk estimate is considered as an area or value range, and the warning range is also considered as a specific area. It is determined whether the two overlap and the size of the overlapping area. If the area of ​​the overlapping area exceeds the set threshold, it means that the collision risk is high and the relative position is within the warning range, and appropriate measures need to be taken to avoid collision.

[0081] For example, suppose the objects to be detected are a shark and a small fish. For the shark, its warning range is determined to be a larger elliptical area according to its category, because sharks are large and swim fast, requiring a larger safety space. Assuming that the area corresponding to the estimated value of the shark's collision risk is a fan-shaped area in front of it, when this fan-shaped area overlaps with the elliptical warning range, and the area of ​​the overlapping area exceeds the set threshold, for example, the set threshold is 30% of the warning range area, and the actual overlapping area reaches 40%, then it is determined that the shark's current relative position is within the warning range, and there may be a risk of collision. For the small fish, its warning range is a smaller circular area. If the estimated value of the small fish's collision risk corresponds to a smaller annular area around it, when the overlapping area of ​​the annular area and the circular warning range exceeds the set threshold, such as 20%, it is determined that the relative position of the small fish is also within the warning range.

[0082] Therefore, the warning range is dynamically configured according to the category of the object to be detected, making the warning more in line with the actual situation, avoiding misjudgment or missed judgment caused by a one-size-fits-all warning method, and thus improving the accuracy of collision risk judgment. It can adapt to objects to be detected of different types, different sizes and different behavior patterns, whether large marine organisms or small aquatic animals, and can reasonably configure the warning range and judge the risk according to their own characteristics, enhancing the adaptability and versatility of the system. It provides a more accurate basis for subsequent decision-making. When it is determined that the relative position is within the warning range, corresponding measures can be taken in time, such as adjusting the detection strategy, issuing an alarm, etc., which will help to better protect the object to be detected and the surrounding environment and facilities, and reduce the losses caused by potential collision accidents.

[0083] In the above embodiment, the underwater environment includes objects to be detected (such as various fish) and underwater robots. The state information of each participant, such as position, speed, direction of movement, etc., is clarified. For each object to be detected, according to its relative position, motion state and other factors with the underwater robot and other objects that may interact (including other objects to be detected), the limited improvement characteristics of the potential game model are used to analyze the possibility of collision with other objects alone, and obtain an individual collision risk function. For example, if an object to be detected moves quickly toward the underwater robot and the distance gradually shortens, its individual collision risk will increase. Then, the individual collision risk functions of each object to be detected are mapped to the potential game function. This function can comprehensively consider the behavior of all participants and their mutual influence. For example, the motion trajectories of two objects to be detected may interfere with each other, thereby affecting their respective collision risks with the underwater robot. The potential game function needs to take into account this complex relationship.

[0084] Nash equilibrium means that in a game, all participants have chosen their own optimal strategies, and when the strategies of other participants remain unchanged, no participant can change their strategy to obtain better results. At this time, the system reaches a stable state. The global collision risk function is determined by analyzing the interaction between the game subject (the object to be detected) and the surrounding objects during the convergence of Nash equilibrium. In this process, it is necessary to consider the overall collision risk when the behavior of all participants tends to be stable. For example, when the motion states of multiple objects to be detected and underwater robots reach a state that is close to Nash equilibrium, the influence of factors such as the distance and speed change between them on the global collision risk is analyzed.

[0085] Taking all the above factors into consideration, the global collision risk function is calculated through specific algorithms and mathematical models. This process may involve weighted summation of various factors, integral operations, or the use of some complex mathematical formulas and algorithms. The specific calculation method will vary depending on the underlying game model adopted and the actual problem setting. For example, different weights may be assigned based on factors such as the distance and speed between different objects, and then a series of calculations may be performed to quantify the global collision risk.

[0086] The calculation of the global collision risk function is a complex process based on a potential game model that comprehensively considers multiple factors such as the individual behavior of each participant in the system, mutual influence, and Nash equilibrium state. The specific calculation method needs to be determined based on the specific application scenario and model settings.

[0087] In an optional example, suppose there is an underwater robot and two fish, fish A and fish B, in an underwater environment. To calculate the global collision risk function, three important factors are considered: Relative distance is the straight-line distance between the fish and the underwater robot. The shorter the distance, the higher the risk of collision. Relative speed includes the speed difference between the fish and the underwater robot, and the angle formed by their movement directions. The greater the speed difference and the closer the movement directions are to face-to-face movement, the higher the risk of collision. The influence of surrounding objects refers to the distance between fish A and fish B and their movement state. If the two fish are very close and their movement directions interfere with each other, it may increase the risk of collision between them and the underwater robot.

[0088] For the convenience of calculation, set a corresponding weight for each factor. Assume that the weight of the relative distance factor is 0.4, the weight of the relative speed factor is 0.4, and the weight of the influence of surrounding objects is 0.2. First, calculate the factors related to the individual collision risk of fish A: The relative distance between fish A and the underwater robot: After measurement or calculation, the distance between fish A and the underwater robot is 5 meters. In order to better measure the risk, we normalize the distance. Assuming that the maximum safe distance is 10 meters, the normalized distance calculation method is 1 minus the actual distance between fish A and the underwater robot divided by the maximum safe distance, that is, 1-5÷10=0.5. In other words, the closer the distance, the closer the normalized value is to 1, which means the higher the risk. The relative speed of fish A and the underwater robot: Assume that the speed of fish A is 2 meters per second, the speed of the underwater robot is 1 meter per second, and the angle between their movement directions is 60 degrees. The magnitude of the relative speed is calculated by the speed synthesis method, and the relative speed is about 1.73 meters per second. Similarly, the relative speed is normalized. Assuming that the maximum relative speed is 3 meters per second, the normalized relative speed is the actual value of the relative speed divided by the maximum relative speed, that is, 1.73÷3≈0.58. Assuming that the distance between fish A and fish B is 3 meters, we set that when the distance between the fish is less than 4 meters, they will have mutual influence, and the closer the distance, the greater the influence. Therefore, the calculation method of the influence factor of fish A on fish B after normalization is 1 minus the actual distance between fish A and fish B divided by the maximum distance set to have an influence, that is, 1-3÷4=0.25. Then, the individual collision risk function value of fish A is calculated by multiplying the normalized value of the relative distance by the weight of the relative distance, adding the normalized value of the relative speed multiplied by the weight of the relative speed, and adding the influence factor of fish A on fish B multiplied by the weight of the surrounding objects. That is: the individual collision risk function value of fish A = 0.4×0.5+0.4×0.58+0.2×0.25=0.2+0.232+0.05=0.482. The same method is used to calculate the individual collision risk function value of fish B. For example, the distance between fish B and the underwater robot is 6 meters, and the normalized value is 1-6÷10=0.4. Assuming that the speed of fish B is 1.5 meters per second, and the angle with the direction of movement of the underwater robot is 45 degrees, the relative speed is calculated to be about 1.12 meters per second, and the normalized value is 1.12÷3≈0.37. Because the distance between fish A and fish B is 3 meters, the influence factor of fish B on fish A is the same as the influence factor of fish A on fish B, which is 0.25. The individual collision risk function value of fish B = 0.4 × 0.4 + 0.4 × 0.37 + 0.2 × 0.25 = 0.16 + 0.148 + 0.05 = 0.358. Finally, calculate the global collision risk function value. Here is an exemplary calculation method, which is to add the individual collision risk function values ​​of fish A and fish B, and then divide by 2 (in actual calculations, more complex models and algorithms may be used for calculation).That is: global collision risk function value = (0.482 + 0.358) ÷ 2 = 0.84 ÷ 2 = 0.42. The global collision risk function value of 0.42 is a relative indicator used to evaluate the degree of collision risk between the fish and the underwater robot in the entire system.

[0089] As an optional embodiment, in 104, sending fish prompt information to the user based on the detection result and the category to which it belongs includes: If the area of ​​the overlapping region exceeds the set threshold, fish prompt information related to the fish is generated based on the category to which it belongs; the fish prompt information is divided into multiple dimensions: somatosensory prompt information, voice prompt information, text prompt information, and image prompt information; the fish prompt information is classified according to the warning priority based on the area of ​​the overlapping region; wherein, the larger the area of ​​the overlapping region, the higher the warning priority of the fish prompt information; the fish prompt information is pushed to the user in sequence according to the warning priority, so as to prompt the user of the current fish category and the risk of fish collision.

[0090] When the area of ​​the overlapping area exceeds the set threshold, it indicates that there is a high risk of fish collision. At this time, relevant prompt information is generated according to the category of the fish. Because different categories of fish have different characteristics and potential danger levels, it is necessary to generate prompt information in a targeted manner so that users can understand the specific situation. The fish prompt information is divided according to multiple dimensions such as somatosensory, voice, text, and image in order to convey information to users through different perception channels, so that users can receive and understand the prompt content more comprehensively and timely. Different users may have different preferences and perception abilities for different information presentation methods, and information in multiple dimensions can meet the needs of more users. The warning priority of fish prompt information is classified based on the area of ​​the overlapping area because the larger the area of ​​the overlapping area, the higher the risk of collision, and users need to pay more priority and more urgent attention to such information. In this way, users can quickly judge the urgency of the situation according to the warning priority and take corresponding measures.

[0091] For example, suppose that the collision risk estimate of a shark and a small fish is detected, and the overlapping area of ​​the warning range exceeds the set threshold. For sharks, since they are large and highly aggressive fish, the generated fish prompt information is as follows: the somatosensory prompt information can be a vibration device that simulates the vibration of the water flow when the shark approaches; the voice prompt information is "Shark found, danger! Please pay attention to the surrounding environment immediately and avoid approaching!"; the text prompt information is displayed on the screen "Shark appears, high risk of collision, please pay attention to safety!", and is accompanied by a picture of a shark as an image prompt information. Since the overlapping area of ​​the shark is large, assuming that it reaches 70% of the warning range, its warning priority is set to high.

[0092] For small fish, the generated prompt information may be: the somatosensory prompt information is a slight vibration, simulating the change of water flow caused by the swimming of small fish; the voice prompt information is "Small fish found, there is a certain risk of collision, please be careful."; the text prompt information shows "Small fish appears, please be careful to avoid collision.", and the image prompt information is a picture of the small fish. The overlapping area of ​​the small fish is assumed to be 30% of the warning range, and the warning priority is low. The system will push the shark prompt information to the user first because its warning priority is high, allowing the user to pay attention to more dangerous situations first.

[0093] Therefore, through the prompt information of multiple dimensions, the user's senses can be stimulated from different aspects, making it easier for users to notice the prompt content, increasing users' attention to the appearance of fish and the risk of collision, and reducing the possibility of accidents caused by user negligence. The warning priority is divided according to the area of ​​overlapping areas, and information is pushed according to the priority, so that users can quickly understand the urgency of the situation and arrange countermeasures reasonably. This personalized and targeted information push method optimizes the user experience and enables users to receive and process information more efficiently. Prompting users in a timely and accurate manner about the current types of fish and the risk of collision can help users take preventive measures in advance to avoid collisions with fish, thereby enhancing the safety of users in related scenarios (such as underwater operations, diving, etc.).

[0094] In an embodiment of the present application, first, an image to be processed is obtained; the image to be processed contains an object to be detected in an underwater area. Then, the category to which the object to be detected belongs is identified through a dynamic convolutional attention network. Next, the relative position between the object to be detected and the underwater robot is determined based on the image to be processed. Finally, a potential game model is used to detect whether the collision risk of the relative position is within the warning range, and fish prompt information is sent to the user based on the detection result and the category; the fish prompt information is used to indicate the category of fish that can be observed by the underwater robot and the risk of fish collision. The embodiment of the present application can realize intelligent identification of fish in underwater areas, improve the accuracy of fish identification, and improve the efficiency of fish identification.

[0095] Figure 2 A schematic diagram of a fish detection system based on an underwater robot provided in an embodiment of the present application, the system at least includes: An acquisition module, used for acquiring an image to be processed; the image to be processed includes an object to be detected in an underwater area; An identification module, used to identify the category to which the object to be detected belongs through a dynamic convolutional attention network; A positioning module, used to determine the relative position between the object to be detected and the underwater robot based on the image to be processed; The early warning module is used to use a potential game model to detect whether the collision risk of the relative position is within the early warning range, and to send fish prompt information to the user based on the detection result and the category to which it belongs; the fish prompt information is used to indicate the fish category observable by the underwater robot and the fish collision risk.

[0096] The system provided in the embodiment of the present application can realize intelligent identification of fish in underwater areas, improve the accuracy of fish identification, and improve the efficiency of fish identification.

[0097] The present application provides a computer-readable storage medium. Figure 3 As shown, this embodiment provides a computer-readable storage medium 600 on which a computer program 611 is stored. When the computer program 611 is executed by a processor, the aforementioned embodiment is implemented.

[0098] It should be noted that in the above embodiments, the description of each embodiment has its own emphasis, and for parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0099] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0100] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0101] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0102] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0103] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.

[0104] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A fish detection method based on an underwater robot, characterized in that: The method comprises: Acquire an image to be processed; the image to be processed includes an object to be detected in an underwater area; Identify the category to which the object to be detected belongs through a dynamic convolutional attention network; Determine the relative position between the object to be detected and the underwater robot based on the image to be processed; A potential game model is used to detect whether the collision risk of the relative position is within the warning range, and fish prompt information is sent to the user based on the detection result and the category to which it belongs; the fish prompt information is used to indicate the fish category observable by the underwater robot and the fish collision risk.

2. The fish detection method based on underwater robot according to claim 1, characterized in that: The step of identifying the category to which the object to be detected belongs by using a dynamic convolutional attention network includes: Extracting the information to be identified of the object to be detected; the information to be identified at least includes: real-time contour information, real-time surface information, and real-time posture information; Performing dynamic convolution processing on the key identification features to obtain dynamic appearance features of the object to be detected in different activity states; the dynamic appearance features at least include: scale texture features and tail fin shape features; The weight parameters under different activity states are set, and the category of the dynamic shape feature is predicted using the weight parameters to obtain the category to which the object to be detected belongs.

3. The fish detection method based on underwater robot according to claim 2 is characterized in that: The step of performing dynamic convolution processing on the key identification features to obtain dynamic appearance features of the object to be detected in different activity states includes: Extracting features to be processed at different scales from the key identification features; Through a conditional parameter generator, the features to be processed at different scales are projected to the behavior prediction space corresponding to different activity states, so as to obtain the projection features of the object to be detected in different activity states; the conditional parameter generator is used to indicate the activity states corresponding to different fish and the types of projection features to be extracted under various activity states; The projection features under different activity states are classified and enhanced to obtain the dynamic shape features of the object to be detected under different activity states; the classification and enhancement branches include: a first enhancement branch for enhancing the tail fin features in the sideways state, and a second enhancement branch for enhancing the scale texture features in the swimming state.

4. The fish detection method based on underwater robot according to claim 1, characterized in that: The adopting of a potential game model to detect whether the collision risk of the relative position is within the warning range includes: Using a construction layer of a potential game model, a corresponding mathematical model is constructed in a three-dimensional space of an underwater area based on the relative position; The risk game layer of the potential game model is adopted to predict the collision risk of the relative position according to the mathematical model; the collision risk of the relative position at least includes: the individual collision risk and the global collision risk of the object to be detected at the relative position; The early warning layer of the potential game model is used to determine whether the collision risk of the relative position is within the early warning range.

5. The fish detection method based on underwater robot according to claim 4 is characterized in that: The risk game layer using the potential game model predicts the collision risk of the relative position according to the mathematical model, including: Taking the mathematical model of the relative position as the game subject, using the limited improvement characteristics of the potential game model to predict the individual collision risk of the game subject, and obtaining the individual collision risk function of the game subject in the three-dimensional space; Mapping the individual collision risk function into a potential game function; Determine a global collision risk function between the game subject and surrounding objects based on the convergence of Nash equilibrium; A collision risk estimation value corresponding to the game subject is obtained based on the global collision risk function.

6. The fish detection method based on underwater robot according to claim 5, characterized in that: The early warning layer using the potential game model determines whether the collision risk of the relative position is within the early warning range, including: Dynamically configure the warning range based on the category to which the object to be detected belongs; wherein the shape of the warning range matches the contour shape of the object to be detected; Determining whether there is an overlapping area between the collision risk estimation value and a pre-configured warning range; If the area of ​​the overlapping region exceeds a set threshold, it is determined that the relative position is within the warning range.

7. The fish detection method based on underwater robot according to claim 6, characterized in that: The sending of fish prompt information to the user based on the detection result and the category thereof includes: If the area of ​​the overlapping region exceeds the set threshold, fish prompt information related to the fish is generated based on the category to which it belongs; the fish prompt information is divided into: somatosensory prompt information, voice prompt information, text prompt information, and image prompt information according to multiple dimensions; Classifying the fish prompt information according to the warning priority based on the area of ​​the overlapping region; wherein the larger the area of ​​the overlapping region, the higher the warning priority of the fish prompt information; The fish prompt information is pushed to the user in sequence according to the warning priority, so as to prompt the user with the current fish types and the fish collision risks.

8. A fish detection system based on an underwater robot, characterized in that: The system comprises at least: An acquisition module, used for acquiring an image to be processed; the image to be processed includes an object to be detected in an underwater area; An identification module, used to identify the category to which the object to be detected belongs through a dynamic convolutional attention network; A positioning module, used to determine the relative position between the object to be detected and the underwater robot based on the image to be processed; The early warning module is used to use a potential game model to detect whether the collision risk of the relative position is within the early warning range, and to send fish prompt information to the user based on the detection result and the category to which it belongs; the fish prompt information is used to indicate the fish category observable by the underwater robot and the fish collision risk.

9. An electronic device, characterized in that: including a memory for storing a computer software program; A processor is used to read and execute the computer software program, thereby implementing the fish detection method based on the underwater robot as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The method comprises instructions, which, when executed on a computer, enable the computer to execute the fish detection method based on an underwater robot as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Vehicle collision risk early warning method and system

    CN111785023A

  • Underwater robot operation risk assessment system based on multi-dimensional information calculation

    CN116362544A

  • Intelligent fish identifying and monitoring method and system based on multi-sensor data

    CN117214904A

  • Crab detection and counting method and device based on instance segmentation, medium and product

    CN118799716A

  • Fish behavior identification method and device

    CN119625816A