Method and system for feature recognition of water targets based on multiple types of mobile devices

By using heterogeneous data acquisition from multiple types of mobile devices and an adaptive sliding window detection mechanism, combined with a fully convolutional neural network, high-precision identification and intelligent status analysis of water targets were achieved. This solved the problems of small target detection and environmental adaptation, and met the precise needs of graded rescue.

CN122493324APending Publication Date: 2026-07-31THE NAVAL MEDICAL UNIV OF PLA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
THE NAVAL MEDICAL UNIV OF PLA
Filing Date
2026-04-08
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing technologies for water target feature recognition suffer from poor performance in detecting small targets, lack of environmental adaptability, and lack of intelligent assessment of the injury status of people who have fallen into the water, thus failing to meet the precise needs of tiered rescue.

Method used

A water target feature recognition method based on multiple types of mobile devices is adopted. Multimodal image data is collected through a heterogeneous mobile acquisition module. Combined with an adaptive sliding window detection mechanism and a multi-side output fully convolutional neural network, the water target is identified and located, and its motion status and injury level are intelligently assessed.

Benefits of technology

It achieves high-precision identification and status analysis of water targets, has environmental adaptive perception capabilities, and enhances the intelligent judgment function of target actions and injury levels, meeting the needs of precise and graded rescue in complex water scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493324A_ABST
    Figure CN122493324A_ABST
Patent Text Reader

Abstract

This application relates to the field of water target recognition technology, and discloses a feature recognition method and system for water targets based on multiple types of mobile devices. The method includes: S1, acquiring multimodal water image data through a heterogeneous mobile acquisition module, and performing standardized preprocessing on the image data; S2, using an adaptive sliding window detection mechanism to perform analysis block processing on the preprocessed high-resolution image. This invention acquires multimodal image data through a heterogeneous mobile acquisition module, and combines an adaptive sliding window mechanism with a multi-side output fully convolutional network to achieve high-precision recognition and state analysis of water targets. It effectively solves the problem of feature loss due to downsampling in the detection of small targets in traditional methods, has adaptive perception capability for environmental changes, and enhances the intelligent judgment function of target actions and injury levels, thereby meeting the needs of precise and graded rescue for search and rescue and safety monitoring of people falling into the water in complex water scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of waterborne target recognition technology, and in particular to a method and system for feature recognition of waterborne targets based on multiple types of mobile devices. Background Technology

[0002] Target feature recognition on water is a core technology in maritime search and rescue, water safety monitoring, and other fields. Its main task is to quickly detect, accurately locate, and analyze the status of targets such as people and vessels that have fallen into the water by analyzing water image data. In recent years, with the rapid development of mobile sensing platforms such as drones and unmanned surface vessels, water monitoring systems based on collaborative operations of multiple types of mobile devices have become an important research direction. Such systems typically use heterogeneous sensors such as visible light and infrared to collect multimodal image data and use artificial intelligence technologies such as deep learning to automatically extract and identify target features. Especially in the scenario of searching for people who have fallen into the water, the system needs to overcome the complexity of the water environment, achieve accurate detection of small-sized, low-pixel-ratio targets, and provide key status information support for subsequent rescue decisions.

[0003] However, existing technologies still have significant shortcomings in the identification of features of targets on water. First, in the detection of small targets, traditional image processing methods directly adopt downsampling strategies, which leads to the further loss of already limited target features and poor recognition results. Second, the system lacks the ability to adapt to environmental conditions and is difficult to intelligently adjust the perception strategy according to day and night and weather changes. Furthermore, most existing systems only realize the target positioning function and lack the ability to intelligently assess the injury status of people who have fallen into the water, thus failing to meet the precise needs of graded rescue. Summary of the Invention

[0004] To overcome the shortcomings of existing technologies in small target detection, adaptive environmental perception, and injury assessment, this application provides a feature recognition method and system for waterborne targets based on multiple types of mobile devices.

[0005] Firstly, this application provides a feature recognition method for waterborne targets based on multiple types of mobile devices, employing the following technical solution: A feature recognition method for waterborne targets based on multiple types of mobile devices, the method comprising: S1. Collect multimodal waterborne image data through a heterogeneous mobile acquisition module, and perform standardized preprocessing on the image data; S2. An adaptive sliding window detection mechanism is used to perform block analysis on the preprocessed high-resolution image. S3. Use a multi-side output fully convolutional neural network to extract the depth features of each image patch, and generate fusion feature parameters through a weighted fusion strategy; S4. Based on the fused feature parameters, identify and locate the water target, and intelligently assess its action status and injury level. S5. Output the recognition results and analysis information.

[0006] By adopting the above technical solution, the full-process feature recognition of water targets under multiple types of mobile devices is realized. First, the heterogeneous mobile acquisition module acquires multimodal water images and performs standardized preprocessing. Then, the adaptive sliding window detection block and multi-side output fully convolutional neural network extracts depth features and performs weighted fusion. Finally, the identification and positioning of water targets and the intelligent judgment of their action status and injury level are completed, and relevant results and information are output.

[0007] Optionally, in step S1, during the acquisition of multimodal waterborne image data, the dominant sensor is dynamically selected based on real-time weather and lighting conditions: During clear days, the system primarily uses visible light sensors, while at night or under low light conditions, it primarily uses infrared sensors. In foggy or rainy weather, it activates a collaborative acquisition mode using both visible light and infrared sensors.

[0008] By adopting the above technical solution, the dominant sensor can be dynamically selected according to real-time weather and lighting conditions. During sunny days, the visible light sensor is the main sensor, while at night or under low light conditions, the infrared sensor is the main sensor. In severe weather conditions such as fog and rain, dual-sensor collaborative acquisition is activated to ensure the adaptability and effectiveness of multimodal water image data acquisition.

[0009] Optionally, in step S2, the adaptive process of the sliding window mechanism includes: Based on the estimated size of the water target, image resolution, and real-time computing resources, the sliding window size is dynamically adjusted. and sliding step size S; The adjustment of the sliding window size follows the following constraints: in, The width of the input image. The height of the input image. To estimate the area of ​​the target region, The total area of ​​the image. This is the size matching factor; The sliding step size S is obtained by analyzing and calculating using the above formula; in, This is the step size adjustment factor. To calculate resource indices in real time.

[0010] By adopting the above technical solution, the sliding window size is dynamically adjusted according to specific constraints based on the estimated size of the water target, image resolution, and real-time computing resources. The sliding step size is calculated by combining the step size adjustment factor and the real-time computing resource index, thereby realizing the adaptive optimization of the sliding window detection mechanism.

[0011] Optionally, in step S3, the process of obtaining the fusion feature parameters includes: The fusion feature parameters are obtained through the above formula analysis and calculation. ; in, The feature map output by the i-th side branch. For the transformation function used to unify the feature scale, is the weight coefficient corresponding to the i-th branch.

[0012] By adopting the above technical solution, the feature maps output by each side branch are processed by a transformation function with a unified feature scale, and combined with the weight coefficients corresponding to each branch, the fused feature parameters are calculated by a specific formula, thereby improving the comprehensiveness and accuracy of feature representation.

[0013] Optionally, in step S4, the process of intelligently assessing the patient's action state and injury level includes: Human keypoint estimation technology is used to extract limb movement information of people who have fallen into the water and calculate their movement activity index. The sports activity index The calculation formula is: Where T is the total number of frames within the observation time window. N is the number of key points to be tracked. , Let J be the two-dimensional coordinates of the j-th keypoint at frame t. Let J be the two-dimensional coordinates of the j-th keypoint at frame t-1. The sports activity index Compared with the preset motion activity judgment threshold Perform a comparison; like If so, the person who fell into the water is deemed to have the ability to move. like If so, it is determined that the person who fell into the water has limited or lost the ability to move.

[0014] By adopting the above technical solution, the limb movement information of the person who fell into the water is extracted with the help of human key point estimation technology. The motion activity index is calculated based on the total number of frames in the observation time window, the number of tracked key points, and the coordinate changes of each key point between frames. After comparing it with the preset threshold, it is determined whether the person who fell into the water has the ability to move.

[0015] Optionally, for those determined to be capable of movement after falling into the water, a more refined vital sign assessment process should be initiated, including: Core body temperature is estimated by measuring the highest temperature in the facial region based on infrared image sequences. ; Respiratory rate is calculated by analyzing the periodic movements of the thoracic region based on visible light image sequences. ; Heart rate values ​​are extracted by detecting minute color changes on the facial skin surface based on visible light image sequences. ; The vital signs assessment values ​​of the drowning person were obtained through the above formula analysis and calculation. ; in, , , As the first weighting coefficient, , , This is a standardized function corresponding to the physical characteristics; when When vital signs fall below a preset threshold, a potential risk warning is generated.

[0016] By adopting the above technical solution, a refined vital sign assessment is initiated for people who have fallen into the water and are capable of movement. The core body temperature estimate is obtained based on infrared image sequences, and respiratory rate and heart rate are extracted by analyzing visible light image sequences. The vital sign assessment value is calculated by weighting coefficient allocation and standardization. When the value is lower than the preset threshold, a potential risk warning is generated.

[0017] Optionally, for individuals who have fallen into the water and are determined to have limited or lost mobility, an emergency injury assessment procedure should be initiated immediately, including: The respiratory function index is calculated and obtained by analyzing the amplitude and frequency of movement in the chest and abdominal regions. Based on the Glasgow Coma Scale, a consciousness status score is calculated and analyzed by examining eye-opening response, verbal response, and motor response. The body surface temperature distribution is obtained based on infrared images, and the core temperature deviation index is calculated. Based on the above assessment results, an injury level report is generated, and a corresponding level of rescue alert is triggered.

[0018] By adopting the above technical solutions, an emergency injury assessment is initiated for drowning personnel whose mobility is limited or lost. The respiratory function index is calculated, a consciousness status score is obtained based on the Glasgow Coma Scale, and a core temperature deviation index is obtained. The comprehensive assessment results generate an injury level report and trigger a corresponding level of rescue alarm.

[0019] Optionally, the method further includes: S6. Based on the real-time location, battery life, and sensor status of each mobile platform, the search area and task allocation are dynamically optimized through the central scheduler to achieve full coverage of water areas and maximize recognition efficiency under multi-device collaboration.

[0020] By adopting the above technical solution, a new multi-device collaborative scheduling mechanism is added. The central scheduler dynamically optimizes the search area and task allocation based on the real-time location, battery life, and sensor status of each mobile platform, thereby achieving full coverage of the water area and maximizing identification efficiency.

[0021] Optionally, in the adaptive sliding window detection mechanism, the total number of image blocks D is dynamically calculated as follows: Where Q is the image edge padding width.

[0022] By adopting the above technical solution, and combining the input image width and height, sliding window size, sliding step size and image edge filling width, the total number of image blocks D in the adaptive sliding window detection mechanism is dynamically calculated using a specific formula, providing a quantitative basis for block processing.

[0023] Secondly, this application provides a feature recognition system for waterborne targets based on multiple types of mobile devices, employing the following technical solution: A feature recognition system for waterborne targets based on multiple types of mobile devices, used to implement the feature recognition method for waterborne targets based on multiple types of mobile devices described in any one of the above, the system comprising: The heterogeneous mobile acquisition module consists of a drone and an unmanned surface vessel, equipped with visible light and infrared sensors, and is used to acquire multimodal waterborne image data. The central processing module is used to standardize and preprocess the acquired image data, and to analyze high-resolution images by using an adaptive sliding window detection mechanism. It extracts image block depth features through a multi-side output fully convolutional neural network, and generates fusion feature parameters by using a weighted fusion strategy. Based on the fusion feature parameters, it realizes the identification and positioning of water targets, and intelligently judges the action status and injury level. The decision support module is used to output identification results and analysis information, and generate rescue alarms.

[0024] By adopting the above technical solution, the water target feature recognition system utilizes a heterogeneous mobile acquisition module composed of UAVs and unmanned surface vessels, combined with visible light and infrared sensors to achieve comprehensive acquisition of multimodal water image data. After the image data is standardized and preprocessed by the central processing module, feature extraction and fusion are completed using adaptive sliding window detection, multi-side output fully convolutional neural networks, and weighted fusion strategies. This enables intelligent assessment of water target identification, positioning, movement status, and injury level. Finally, the decision support module outputs relevant results and generates rescue alarms, thus constructing a complete water target recognition and rescue support system from data acquisition and intelligent analysis to decision output.

[0025] In summary, this application includes at least one of the following beneficial technical effects: This invention acquires multimodal image data through a heterogeneous mobile acquisition module, and combines an adaptive sliding window mechanism with a multi-side output fully convolutional network to achieve high-precision identification and state analysis of targets on water. It effectively solves the problem of feature loss caused by downsampling in the detection of small targets in traditional methods, and has the ability to adaptively perceive environmental changes. At the same time, it enhances the intelligent judgment function of target actions and injury levels, thereby meeting the needs of precise and graded rescue for search and rescue and safety monitoring of people who have fallen into the water in complex water scenarios. Attached Figure Description

[0026] Figure 1 This is a flowchart of the steps of the feature recognition method for water targets based on multiple types of mobile devices proposed in this invention; Figure 2 This is a schematic diagram of the micro-target identification method for waterborne targets based on the feature recognition of multiple types of mobile devices proposed in this invention. Figure 1 ; Figure 3 This is a schematic diagram of the micro-target identification method for waterborne targets based on the feature recognition of multiple types of mobile devices proposed in this invention. Figure 2 ; Figure 4 This is a flowchart of the rapid identification of human targets at sea using multiple types of mobile devices in the feature recognition method for water targets based on multiple types of mobile devices proposed in this invention. Figure 5 This is a schematic diagram illustrating the personnel injury assessment method in the feature recognition method for waterborne targets based on multiple types of mobile devices proposed in this invention. Detailed Implementation

[0027] The embodiments of this application are described in detail below, and examples of the embodiments are shown in the accompanying drawings.

[0028] In the description of this specification, the references to "certain embodiments," "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples" refer to specific features, structures, materials, or characteristics described in connection with the described embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0029] This application discloses a feature recognition method for waterborne targets based on multiple types of mobile devices, referring to... Figure 1 - Figure 5 The method includes: S1. Multimodal waterborne image data is acquired through a heterogeneous mobile acquisition module, and the image data is subjected to standardized preprocessing. The heterogeneous mobile acquisition module includes devices such as drones and unmanned surface vessels equipped with visible light and infrared sensors, which are used to acquire multimodal images under different environmental conditions. The standardized preprocessing includes image denoising, color correction, and resolution unification operations. S2. An adaptive sliding window detection mechanism is used to perform analysis block processing on the preprocessed high-resolution image. The adaptive sliding window detection mechanism can dynamically adjust the sliding window size and step size according to the image content and environmental conditions, thereby balancing the target detection rate and computational efficiency. S3. The deep features of each image patch are extracted using a multi-side output fully convolutional neural network, and fusion feature parameters are generated through a weighted fusion strategy. The multi-side output structure can extract semantic information at different levels in the image. The weighted fusion strategy fuses different side outputs according to feature confidence, thereby enhancing the feature expression ability of targets with low pixel ratio. S4. Based on the fused feature parameters, the water target is identified and located, and its motion state and injury level are intelligently assessed. The identification and location include target type judgment and position regression. The motion state and injury level assessment is based on multi-dimensional analysis of target posture, heat map features and motion patterns. S5. Output recognition results and analysis information, including target category, location coordinates, action status, and injury level assessment results.

[0030] Through the above technical solution, this embodiment provides a feature recognition method for water targets based on multiple types of mobile devices. The method acquires multimodal image data through a heterogeneous mobile acquisition module, and combines an adaptive sliding window mechanism with a multi-side output fully convolutional network to achieve high-precision recognition and state analysis of water targets. It effectively solves the problem of feature loss caused by downsampling in the detection of small targets in traditional methods, has adaptive perception capability for environmental changes, and enhances the intelligent judgment function of target actions and injury levels, thereby meeting the needs of precise and graded rescue for search and rescue and safety monitoring of people who have fallen into the water in complex water scenarios.

[0031] In one embodiment, during step S1, in the process of acquiring multimodal waterborne image data, the dominant sensor is dynamically selected based on real-time weather and lighting conditions: Under clear daytime conditions, the visible light sensor is prioritized as the dominant sensor to take advantage of its rich color and texture information; At night or in low light conditions, it automatically switches to the infrared thermal imaging sensor as the primary sensor, and detects based on the difference in thermal radiation between the target and the water area. In adverse visibility conditions such as fog and rain, a collaborative acquisition mode using visible light and infrared sensors is activated to enhance overall robustness through multi-source data complementarity.

[0032] In step S2, the adaptive process of the sliding window mechanism includes: Based on the estimated size of the water target, image resolution, and real-time computing resources, the sliding window size is dynamically adjusted. and sliding step size S; The adjustment of the sliding window size follows the following constraints: in, The width of the input image is obtained directly from the input image to be processed, and is generally the pixel width of the image. The height of the input image is obtained directly from the input image to be processed, and is generally the pixel height of the image. The specific value depends on the resolution setting of the camera sensor, such as 1920x1080. To estimate the area of ​​a target region, it is usually based on prior knowledge or statistics on the size of the target bounding boxes in the training dataset. For example, a person falling into the water typically occupies about 0.5% to 5% of the total image area at a typical shooting distance. The total area of ​​the image can be obtained by multiplying the width and height of the input image. Size matching coefficients are hyperparameters determined through experimental verification. They are usually obtained by debugging on a validation set before model training and are used to fine-tune the baseline level of the sliding window size relative to the image size and the target proportion. The sliding step size S is obtained by analyzing and calculating using the above formula; in, The step size adjustment factor, a hyperparameter determined through experimental verification, controls the density of sliding window movements. The smaller the step size, the smaller the scan density, and the higher the detection accuracy may be, but the greater the computational load. The larger the step size, the faster the processing speed, but it may miss small targets. To calculate resource indices in real time, they can be dynamically calculated based on performance indicators such as CPU / GPU / memory usage monitored in real time by the system. When computing resources are sufficient, a smaller step size can be used for fine-grained searching; when resources are scarce, a larger step size is used to ensure the real-time performance of the system.

[0033] In step S3, the process of obtaining the fusion feature parameters includes: The fusion feature parameters are obtained through the above formula analysis and calculation. ; in, The feature map output by the i-th side branch is directly generated by the i-th side output layer of the fully convolutional neural network. These side outputs typically come from different depths of the convolutional neural network, such as shallow, medium, and deep layers. The transformation function used to unify feature scales is typically an upsampling or downsampling operation, such as bilinear interpolation or deconvolution. The weight coefficient corresponding to the i-th branch can be obtained by setting it based on experience.

[0034] Through the above technical solution, this embodiment provides an adaptive and high-precision method for identifying features of waterborne targets. This method employs an environment-adaptive sensor scheduling strategy, a dynamic sliding window mechanism based on target scale and computing resources, and weighted multi-scale deep feature fusion. First, the adaptive sliding window and multi-scale feature fusion avoid the problem of small target feature loss caused by traditional downsampling, improving the detection rate of small targets. Second, dynamically selecting sensors based on weather and lighting conditions and dynamically adjusting the step size based on resources endow the system with adaptive capabilities to cope with complex environments and ensure real-time performance. Finally, the extracted highly robust fusion feature parameters provide a solid data foundation for intelligent assessment of target movement status and injury level in subsequent steps, thereby meeting the urgent need for accurate information in tiered rescue operations.

[0035] In one embodiment, step S4, the process of intelligently assessing the patient's action state and injury level, includes: Human keypoint estimation technology is used to extract and track the limb joints (such as shoulders, elbows, hips, knees, etc.) of people who have fallen into the water, and calculate their motor activity index to quantify their voluntary motor ability. The sports activity index The calculation formula is: Where T is the total number of frames within the observation time window, preset by the system, referring to the number of continuously acquired image frames used for a single analysis and evaluation. A longer T results in more stable evaluation results and stronger resistance to transient jitter interference; a shorter T results in a faster system response, typically corresponding to a 1 to 3 second video segment. If the video frame rate is 25fps, then T is usually taken as 25 to 75. N represents the number of keypoints to be tracked, which is determined by the human keypoint estimation model used. The more keypoints there are, the more detailed the description of limb movement. Common values ​​are 13 (mainly the trunk and limbs), 18, or 25 (including hand and facial details). , Let J be the two-dimensional coordinates of the j-th keypoint at frame t. Let J be the two-dimensional coordinates of the j-th keypoint at frame t-1. , All of these can be obtained by directly outputting each frame of the image after processing it using keypoint detection models (such as HRNet and OpenPose); The sports activity index Compared with the preset motion activity judgment threshold The preset motion activity judgment threshold is compared. This can be obtained through statistical analysis of a large amount of normal and abnormal behavior data; like If so, the person who fell into the water is deemed to have the ability to move. like If so, it is determined that the person who fell into the water has limited or lost the ability to move.

[0036] For those determined to be capable of movement after falling into the water, a more refined vital sign assessment process will be initiated, including: Based on infrared image sequences, the facial regions of the person who fell into the water (such as the inner corner of the eye and forehead) were located, and the highest temperature in these regions was measured as an estimate of the core body temperature. ; Based on visible light image sequences, the respiratory rate is obtained by tracking the periodic undulations of the thoracic region (e.g., using optical flow or feature point tracking) and calculating the number of undulations per unit time. ; Based on visible light image sequences, and through signal processing techniques such as blind source separation, heart rate signals are extracted from the micro-color changes on the facial skin surface caused by changes in blood volume (principle of photoplethysmography), thus obtaining heart rate values. ; The vital signs assessment values ​​of the drowning person were obtained through the above formula analysis and calculation. ; in, , , The first weighting coefficient, derived from prior medical knowledge or trained on historical data using a machine learning model, is used to characterize the relative importance of each vital sign in the overall assessment. , , The standardization function for the corresponding vital signs is a function that maps the raw vital sign values ​​to a uniform scale and standard score range. For example, the Z-score standardization function or the Min-Max normalization function can be used. when When the vital signs fall below the preset threshold for assessment, it indicates that although the vital signs have not reached a critical level, they have deviated from the normal range, generating a potential risk warning.

[0037] For individuals who have fallen into the water and are determined to have limited or lost mobility, an emergency injury assessment process should be initiated immediately, including: By analyzing the amplitude and frequency of movement in the chest and abdominal regions, if the amplitude is consistently below the threshold and the frequency is abnormal (too fast, too slow, or irregular), the low respiratory function index is calculated. Based on the Glasgow Coma Scale, the system analyzes eye-opening response (whether the eyes open spontaneously) using visible light images, analyzes motor response (response to painful stimuli or unconscious movements) using a combination of infrared and visible light images, and combines audio analysis (if any) to determine language response (whether sounds are made). The system then calculates the consciousness status score. The surface temperature distribution is obtained based on infrared images. Combined with the ambient temperature, the gradient between the surface temperature and the estimated core body temperature is calculated. If the gradient is abnormal or the core body temperature is significantly lower than the normal value, the high core temperature deviation index is obtained. Based on the combined assessment results of breathing, consciousness, and body temperature, an injury level report is generated according to preset rules (such as decision trees or scoring cards), such as minor injury, serious injury, critical injury, etc., and the corresponding level of rescue alarm is automatically triggered to optimize the allocation of rescue resources.

[0038] Through the above technical solution, this embodiment provides a refined and graded intelligent assessment method for the status of people who have fallen into the water. The method achieves a rapid preliminary assessment of mobility by calculating the motion activity index, and initiates a differentiated assessment process based on the preliminary assessment results: non-contact vital sign screening is conducted for those who are capable to identify potential risks, and multiple emergency injury assessments are conducted for those who are incapacitated to determine the severity level. This method effectively addresses the problem of the lack of intelligent assessment capabilities for the injury status of people who have fallen into the water, which fails to meet the precise needs of graded rescue. By combining computer vision with medical assessment models, it achieves a leap from target discovery to status assessment, providing key information support for rescue decisions, from where to rescue, who to prioritize, and how to rescue, greatly improving the intelligence level and rescue efficiency of water search and rescue.

[0039] In one embodiment, the method further includes: S6. Based on the real-time location, battery life, and sensor status of each mobile platform, the search area and task allocation are dynamically optimized through the central scheduler to achieve full coverage of water areas and maximize recognition efficiency under multi-device collaboration.

[0040] In the adaptive sliding window detection mechanism, the total number of image blocks D is dynamically calculated as follows: Where Q is the image edge fill width, which is a preset parameter of the system, usually set during algorithm initialization. It is usually set to 0 (no fill) or a small positive number, depending on the importance attached to the edge target.

[0041] Through the above technical solution, this embodiment provides a method for identifying waterborne target features that integrates collaborative scheduling and resource optimization. The method dynamically integrates the status information of multiple mobile platforms through a central scheduler, achieving optimal allocation of search tasks in space and time. This effectively solves the problems of limited field of view and insufficient endurance of a single device, thereby enabling long-term and efficient monitoring of large-scale water areas. At the same time, by accurately calculating the total number of image blocks D, the method concretizes the adaptive sliding window mechanism into a quantifiable and controllable processing flow, ensuring that high-resolution images can be analyzed efficiently and without omissions. This avoids the shortcomings of traditional fixed block methods in terms of resource utilization and target coverage. It optimizes from both the system and computational levels, significantly improving the overall efficiency, coverage, and resource utilization of target search and identification in complex water environments.

[0042] This application also discloses a feature recognition system for waterborne targets based on multiple types of mobile devices, used to implement the feature recognition method for waterborne targets based on multiple types of mobile devices described in any one of the above embodiments. The system includes: The heterogeneous mobile acquisition module consists of a drone and an unmanned surface vessel, equipped with visible light and infrared sensors, and is used to acquire multimodal waterborne image data. The central processing module is used to standardize and preprocess the acquired image data, and to analyze high-resolution images by using an adaptive sliding window detection mechanism. It extracts image block depth features through a multi-side output fully convolutional neural network, and generates fusion feature parameters by using a weighted fusion strategy. Based on the fusion feature parameters, it realizes the identification and positioning of water targets, and intelligently judges the action status and injury level. The decision support module is used to output identification results and analysis information, and generate rescue alarms.

[0043] Through the above technical solution, this embodiment provides a feature recognition system for waterborne targets based on multiple types of mobile devices. The system achieves all-weather, multi-view data acquisition through a heterogeneous mobile acquisition module; through the adaptive analysis engine and intelligent judgment unit integrated in the central processing module, it realizes fully automated intelligent processing from image preprocessing to target recognition, and then to in-depth analysis of status and injury; finally, through the decision support module and collaborative scheduling module, the analysis results are transformed into operable command information and the overall search efficiency is optimized. This effectively solidifies the methodological innovation into an efficient and reliable physical solution, fundamentally solving the problems of difficulty in detecting small targets, poor environmental adaptability, and lack of intelligent judgment capabilities in existing technologies, providing an end-to-end solution for waterborne search and rescue and safety monitoring.

[0044] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A method for feature recognition of an object on water based on multiple types of mobile devices, characterized in that, The method includes: S1. Collect multimodal waterborne image data through a heterogeneous mobile acquisition module, and perform standardized preprocessing on the image data; S2. An adaptive sliding window detection mechanism is used to perform block analysis on the preprocessed high-resolution image. S3. Use a multi-side output fully convolutional neural network to extract the depth features of each image patch, and generate fusion feature parameters through a weighted fusion strategy; S4. Based on the fused feature parameters, identify and locate the water target, and intelligently assess its action status and injury level. S5. Output the recognition results and analysis information. 2.The method of claim 1, wherein, In step S1, during the acquisition of multimodal waterborne image data, the dominant sensor is dynamically selected based on real-time weather and lighting conditions. During clear days, the system primarily uses visible light sensors, while at night or under low light conditions, it primarily uses infrared sensors. In foggy or rainy weather, it activates a collaborative acquisition mode using both visible light and infrared sensors.

3. The feature recognition method for waterborne targets based on multiple types of mobile devices according to claim 2, characterized in that, In step S2, the adaptive process of the sliding window mechanism includes: Based on the estimated size of the water target, image resolution, and real-time computing resources, the sliding window size is dynamically adjusted. and sliding step size S; The adjustment of the sliding window size follows the following constraints: in, The width of the input image. The height of the input image. To estimate the area of ​​the target region, The total area of ​​the image. This is the size matching factor; The sliding step size S is obtained by analyzing and calculating using the above formula; in, This is the step size adjustment factor. To calculate resource indices in real time.

4. The feature recognition method for waterborne targets based on multiple types of mobile devices according to claim 3, characterized in that, In step S3, the process of obtaining the fusion feature parameters includes: The fusion feature parameters are obtained through the above formula analysis and calculation. ; in, The feature map output by the i-th side branch. For the transformation function used to unify the feature scale, is the weight coefficient corresponding to the i-th branch.

5. The feature recognition method for waterborne targets based on multiple types of mobile devices according to claim 4, characterized in that, In step S4, the process of intelligently assessing the patient's action state and injury level includes: Human keypoint estimation technology is used to extract limb movement information of people who have fallen into the water and calculate their movement activity index. The sports activity index The calculation formula is: Where T is the total number of frames within the observation time window. N is the number of key points to be tracked. , Let J be the two-dimensional coordinates of the j-th keypoint at frame t. Let J be the two-dimensional coordinates of the j-th keypoint at frame t-1. The sports activity index Compared with the preset motion activity judgment threshold Perform a comparison; like If so, the person who fell into the water is deemed to have the ability to move. like If so, it is determined that the person who fell into the water has limited or lost the ability to move.

6. The feature recognition method for waterborne targets based on multiple types of mobile devices according to claim 5, characterized in that, For those determined to be capable of movement after falling into the water, a more refined vital sign assessment process will be initiated, including: Core body temperature is estimated by measuring the highest temperature in the facial region based on infrared image sequences. ; Respiratory rate is calculated by analyzing the periodic movements of the thoracic region based on visible light image sequences. ; Heart rate values ​​are extracted by detecting minute color changes on the facial skin surface based on visible light image sequences. ; The vital signs assessment values ​​of the drowning person were obtained through the above formula analysis and calculation. ; in, , , As the first weighting coefficient, , , This is a standardized function corresponding to the physical characteristics; when When vital signs fall below a preset threshold, a potential risk warning is generated.

7. The feature recognition method for waterborne targets based on multiple types of mobile devices according to claim 6, characterized in that, For individuals who have fallen into the water and are determined to have limited or lost mobility, an emergency injury assessment process should be initiated immediately, including: The respiratory function index is calculated and obtained by analyzing the amplitude and frequency of movement in the chest and abdominal regions. Based on the Glasgow Coma Scale, a consciousness status score is calculated and analyzed by examining eye-opening response, verbal response, and motor response. The body surface temperature distribution is obtained based on infrared images, and the core temperature deviation index is calculated. Based on the above assessment results, an injury level report is generated, and a corresponding level of rescue alert is triggered.

8. The feature recognition method for waterborne targets based on multiple types of mobile devices according to claim 7, characterized in that, The method further includes: S6. Based on the real-time location, battery life, and sensor status of each mobile platform, the search area and task allocation are dynamically optimized through the central scheduler to achieve full coverage of water areas and maximize recognition efficiency under multi-device collaboration.

9. The feature recognition method for waterborne targets based on multiple types of mobile devices according to claim 8, characterized in that, In the adaptive sliding window detection mechanism, the total number of image blocks D is dynamically calculated as follows: Where Q is the image edge padding width.

10. A feature recognition system for waterborne targets based on multiple types of mobile devices, characterized in that, For implementing the feature recognition method for waterborne targets based on multiple types of mobile devices as described in any one of claims 1-9, the system comprises: The heterogeneous mobile acquisition module consists of a drone and an unmanned surface vessel, equipped with visible light and infrared sensors, and is used to acquire multimodal waterborne image data. The central processing module is used to standardize and preprocess the acquired image data, and to analyze high-resolution images by using an adaptive sliding window detection mechanism. It extracts image block depth features through a multi-side output fully convolutional neural network, and generates fusion feature parameters by using a weighted fusion strategy. Based on the fusion feature parameters, it realizes the identification and positioning of water targets, and intelligently judges the action status and injury level. The decision support module is used to output identification results and analysis information, and generate rescue alarms.