Miniaturized underwater target real-time monitoring system based on software and hardware collaborative optimization and data interaction

By constructing a lightweight deep learning model and a three-layer computing architecture with hardware and software co-optimization, combined with a foldable platform, the problems of large size, high power consumption and poor real-time performance of traditional underwater target monitoring systems are solved, and efficient real-time monitoring and accurate identification of underwater targets are achieved.

CN120932678APending Publication Date: 2025-11-11STATE OCEAN TECH CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511090233.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Traditional underwater target monitoring systems suffer from problems such as large size, high power consumption, and poor real-time performance, making it difficult to meet the needs of modern marine monitoring. Furthermore, deep learning models are difficult to deploy on resource-constrained embedded platforms, and there is limited hardware and software co-design and data interaction, resulting in low data security.

Method used

A lightweight deep learning model and a software-hardware co-optimization design are adopted to construct a three-layer computing architecture based on TinyDL for detection, fusion, and recognition. Combined with a foldable test bench, software and hardware co-optimization is achieved. Through the collaborative work of microcontroller and microprocessor, multimodal data fusion and real-time target recognition are performed.

Benefits of technology

It achieves miniaturization, real-time performance improvement, and accuracy enhancement of underwater target monitoring systems, making them suitable for scenarios such as marine environmental monitoring, marine resource exploration, and underwater vehicle safety assurance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932678A_ABST
    Figure CN120932678A_ABST
Patent Text Reader

Abstract

The invention discloses a miniaturized underwater target real-time monitoring system based on software and hardware collaborative optimization and data interaction, and relates to the field of marine acoustic monitoring, and the system comprises a detection device, a fusion device and an identification device. The detection device comprises a rack, an acoustic array and a data acquisition module, the rack is used for supporting each component; the acoustic array is used for collecting underwater acoustic signals in real time; the data acquisition module is used for converting underwater acoustic signals into digital signals and extracting audio features; lightweight deep learning models are deployed in the fusion device and the recognition device, and the fusion device is used for dynamically fusing digital signals and audio features by adopting the lightweight deep learning models based on professional domain knowledge; and the recognition device is used for carrying out target recognition by adopting a lightweight deep learning model according to the fusion features so as to carry out real-time monitoring on the underwater target. According to the invention, the size of the underwater target monitoring system is reduced, and the real-time performance and accuracy of underwater target detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of marine acoustic monitoring, specifically to a miniaturized real-time underwater target monitoring system based on software and hardware co-optimization and data interaction. Background Technology

[0002] With the increasing awareness of marine resource development and marine environmental protection, higher demands are being placed on the real-time monitoring and identification of underwater targets. Traditional underwater acoustic monitoring systems often suffer from problems such as large size, high power consumption, and poor real-time performance, making it difficult to meet the needs of modern marine monitoring.

[0003] Based on the TinyDL (Tiny Deep Learning) model and software-hardware collaborative technology, it can meet the requirements of future terminal devices for low power consumption, high standards, strong professionalism and high security.

[0004] TinyDL models face performance bottlenecks in underwater noise recognition, primarily due to challenges in balancing model compression and accuracy, real-time performance and energy consumption, and data volume limitations. Model compression requires sophisticated strategies to maintain high recognition accuracy and avoid information loss; real-time processing necessitates efficient inference engines and hardware acceleration technologies to meet the real-time requirements of low-power platforms; simultaneously, acquiring high-quality data is difficult, necessitating research into data augmentation and new learning techniques to alleviate data scarcity. Furthermore, current embedded hardware platforms have limitations in supporting TinyDL, including limited computing power of low-frequency MCUs, limited memory bandwidth, power constraints, insufficient storage space, and compatibility issues, making direct deployment of deep learning models difficult. There is an urgent need to optimize the model structure, leverage hardware acceleration to meet real-time processing requirements, and develop cross-platform deployment tools and frameworks to improve model standardization. Therefore, collaborative optimization of software and hardware is essential to overcoming these performance bottlenecks.

[0005] In summary, traditional underwater target monitoring systems face three core problems: First, deep learning models have a large number of parameters, making them difficult to deploy on resource-constrained embedded platforms; second, multimodal data fusion processing consumes a lot of energy, and there is little hardware and software co-design and data interaction; and third, there are insufficient real-time guarantee mechanisms and low data security. Summary of the Invention

[0006] The purpose of this application is to provide a miniaturized underwater target real-time monitoring system based on software and hardware co-optimization and data interaction, which can reduce the size of the underwater target monitoring system and improve the real-time performance and accuracy of underwater target detection.

[0007] To achieve the above objectives, this application provides the following solution:

[0008] This application provides a miniaturized real-time underwater target monitoring system based on software and hardware co-optimization and data interaction, including: a detection device, a fusion device, and an identification device;

[0009] The detection device includes a stand, an acoustic array, and a data acquisition module; the stand supports the acoustic array, the data acquisition module, the fusion device, and the recognition device; the acoustic array is used to acquire underwater acoustic signals in real time; the data acquisition module converts the underwater acoustic signals into digital signals and extracts the audio features of the underwater acoustic signals.

[0010] The fusion device is equipped with a lightweight deep learning model. The fusion device is used to dynamically fuse the digital signal and the audio features based on professional domain knowledge and the lightweight deep learning model to obtain fused features.

[0011] The identification device is equipped with a lightweight deep learning model, which is used to identify targets based on the fused features, so as to monitor underwater targets in real time.

[0012] In one embodiment, the acoustic array includes multiple hydrophones arranged in a tetrahedral array on the platform to collect underwater acoustic signals from different directions.

[0013] In one embodiment, the platform includes an annular disk, a support, an electronic instrument compartment, and a hydrophone frame. The bottom of the support is fixedly connected to the annular disk, and the support is fixedly sleeved outside the electronic instrument compartment. The annular disk is positioned below the electronic instrument compartment, and the annular disk can support the support and the electronic instrument compartment. The electronic instrument compartment can accommodate the data acquisition module, the fusion device, and the identification device. The hydrophone frame is fixedly mounted on the top of the support. The hydrophone frame has a first mounting portion, a second mounting portion, a third mounting portion, and a fourth mounting portion. At least one hydrophone can be fixedly mounted on the first mounting portion, at least one hydrophone can be fixedly mounted on the second mounting portion, at least one hydrophone can be fixedly mounted on the third mounting portion, and at least one hydrophone can be fixedly mounted on the fourth mounting portion. The first mounting portion, the second mounting portion, the third mounting portion, and the fourth mounting portion are respectively located at the four vertices of a regular tetrahedron.

[0014] In one embodiment, the platform includes an annular disk, a support, an electronic instrument compartment, two first folding rods, two second folding rods, and several support rods. The bottom of the support is fixedly connected to the annular disk, and the support is fixedly sleeved outside the electronic instrument compartment. The annular disk is positioned below the electronic instrument compartment. The annular disk can support the support and the electronic instrument compartment, and the electronic instrument compartment can accommodate the data acquisition module, the fusion device, and the identification device. One end of each of the two first folding rods is hinged to the top of the support. The other end of one first folding rod has a first mounting portion, and the other end of the other first folding rod has a second mounting portion. One end of each of the two second folding rods is hinged to the bottom of the support. The other end of one second folding rod has a third mounting portion, and the other end of the other second folding rod has a fourth mounting portion. Mounting parts: at least one hydrophone can be fixedly mounted on the first mounting part, at least one hydrophone can be fixedly mounted on the second mounting part, at least one hydrophone can be fixedly mounted on the third mounting part, and at least one hydrophone can be fixedly mounted on the fourth mounting part; both first folding rods and both second folding rods can be folded to attach to the bracket, and both first folding rods and both second folding rods can be unfolded and locked; after both first folding rods and both second folding rods are unfolded and locked, the first mounting part, the second mounting part, the third mounting part, and the fourth mounting part can be respectively placed on the four vertices of a regular tetrahedron; one end of the support rod is fixedly connected to the bracket, and the other end of the support rod is used to fixally connect to the unfolded first folding rod or the unfolded second folding rod.

[0015] In one embodiment, the interior of the electronic instrument compartment is equipped with a microelectromechanical system acoustic sensor.

[0016] In one embodiment, the data acquisition module includes an analog-to-digital converter, a filter, an operational amplifier, and a first microcontroller connected in sequence; the analog-to-digital converter is used to convert the underwater acoustic signal into a digital signal using a successive approximation analog-to-digital conversion algorithm; the filter is used to filter the digital signal using a finite impulse response filtering algorithm to obtain a filtered signal; the operational amplifier is used to amplify the filtered signal to obtain an amplified signal; and the first microcontroller is used to extract the audio features of the amplified signal.

[0017] In one embodiment, the first microcontroller is model MCX-N947.

[0018] In one embodiment, the audio features include Mel frequency cepstral coefficients and Mel spectrograms; the process of the first microcontroller extracting Mel frequency cepstral coefficients is as follows: the amplified signal is sequentially processed by pre-emphasis, framing, windowing, fast Fourier transform, Mel filtering, logarithmic transform, and discrete cosine transform; the process of the first microcontroller extracting Mel spectrograms is as follows: the amplified signal is sequentially processed by framing, fast Fourier transform, Mel filtering, and logarithmic transform.

[0019] In one embodiment, the fusion device is a second microcontroller; the recognition device includes a third microcontroller and a memory; both the second microcontroller and the third microcontroller are equipped with lightweight deep learning models; the memory is used to store the target recognition results.

[0020] In one embodiment, the second microcontroller is an STM32H757 or an STM32N657, and the third microcontroller is... i.MX 8M Mini or NXP i.MX 93.

[0021] In one embodiment, the system further includes a communication module; the communication module is used to transmit the target identification results to a remote monitoring center or cloud platform so that users can perform remote monitoring.

[0022] According to the specific embodiments provided in this application, this application has the following technical effects:

[0023] This application provides a miniaturized underwater target real-time monitoring system based on software and hardware co-optimization and data interaction. Based on a lightweight deep learning model, software and hardware co-design, and data interaction technology, it carries out hierarchical functional division and tiered hardware configuration of "detection-fusion-recognition". Through the synergy of hardware design of test bench, acoustic array, data acquisition module, fusion device, and recognition device and software design of lightweight deep learning model, it achieves triple improvement in latency, energy efficiency and accuracy in edge computing scenarios, reduces the size of underwater target monitoring system, and improves the real-time performance and accuracy of underwater target detection. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a structural block diagram of a miniaturized underwater target real-time monitoring system based on software and hardware co-optimization and data interaction in one embodiment of this application;

[0026] Figure 2 This is a schematic diagram of the simplified platform in this application;

[0027] Figure 3 This is a schematic diagram of the middle part of the support structure of the simplified platform in this application;

[0028] Figure 4 for Figure 3 Top view of the structure;

[0029] Figure 5 This is a schematic diagram of the bottom structure of the support frame of the simplified platform in this application;

[0030] Figure 6 This is a top view of the support top portion structure of the simplified platform in this application;

[0031] Figure 7 This is a schematic diagram of the annular disk of the simplified platform in this application;

[0032] Figure 8 This is a schematic diagram of a simple hydrophone frame for the stand-up device in this application.

[0033] Figure 9 This is a schematic diagram of another structure of the hydrophone frame of the simplified platform in this application;

[0034] Figure 10 This is another structural schematic diagram of the hydrophone frame of the simplified platform in this application;

[0035] Figure 11 This is a schematic diagram of the folding platform in this application;

[0036] Figure 12 for Figure 11 A schematic diagram of the unfolded second folding rod in the folding frame;

[0037] Figure 13 for Figure 11 A simplified structural diagram of the first and second folding rods of the folding frame after they are fully unfolded.

[0038] Reference numerals: 101-stand, 102-acoustic array, 103-data acquisition module, 104-fusion module, 105-identification module, 106-communication module, 107-power management module, 108-housing and packaging module, 1011-ring-shaped disk, 1012-bracket, 1013-electronic instrument compartment, 1014-hydrophone frame, 1015-first mounting part, 1016-second mounting part, 1017-third mounting part, 1018-fourth mounting part, 1019-first folding rod, 1020-second folding rod. Detailed Implementation

[0039] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0040] Traditional deep learning models face challenges such as high computational complexity, high energy consumption, and poor real-time performance when deployed on resource-constrained edge devices. Existing solutions mostly focus on single-dimensional optimization such as model compression or hardware acceleration, lacking cross-level software and hardware collaborative design methods and runtime software and hardware data interaction technologies.

[0041] This application constructs a three-layer computing architecture based on TinyDL, which integrates detection, fusion, and recognition through hardware and software collaboration. By combining this with the innovative mechanical structure of a foldable platform, the algorithm and hardware are deeply optimized to achieve end-to-end intelligent processing from edge nodes to edge servers. This addresses the technical bottlenecks of traditional underwater monitoring systems, such as large model parameters, high energy consumption, and poor real-time performance. As a result, it enables efficient real-time monitoring of underwater targets and is applicable to scenarios such as marine environmental monitoring, marine resource exploration, underwater vehicle safety assurance, and underwater structure health monitoring.

[0042] This application, based on lightweight deep learning models and software-hardware co-design and data interaction technologies, conducts hierarchical functional division and tiered hardware configuration of "detection-fusion-recognition," and promotes data interaction and collaborative development between layers of software and hardware. Through the co-design and data interaction technologies of hardware such as microcontrollers and microprocessors, and software such as physical models, deep learning technologies, signal processing, and automated control algorithms, a three-layer computing architecture for multi-task and multi-modal fusion is established, including a detection layer, a fusion layer, and a recognition layer. This reduces the performance impact between sensors, software / hardware platforms, and physical models within the underwater target real-time monitoring system, enhances the software / hardware collaboration capability under data interaction, and improves the recognition accuracy under complex system tasks.

[0043] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0044] In one exemplary embodiment, such as Figure 1 As shown, a miniaturized real-time underwater target monitoring system based on hardware and software co-optimization and data interaction is provided, including a detection device, a fusion device 104, and an identification device 105. The detection device includes a test bench 101, an acoustic array 102, and a data acquisition module 103.

[0045] Among them, the detection device serves as the detection layer in the software and hardware co-optimization architecture, the fusion device 104 serves as the fusion layer in the software and hardware co-optimization architecture, and the identification device 105 serves as the identification layer in the software and hardware co-optimization architecture.

[0046] In one embodiment, the detection layer has automatic acoustic signal preprocessing and feature extraction functions. Its main tasks are: first, to set sensor parameters and collect complete sensor data, while sending sensor status information to the fusion layer and the recognition layer, and improving the quality of sensor data in the event of missing sensor data; second, to switch between standby, continuous acquisition, timing, and response states. The hardware mainly uses a 100-200MHz microcontroller, such as the Arm Cortex-M33 microcontroller (MCX-N947), and the software mainly uses deep learning models such as ResNet (Residual Neural Network) and MCUNet (Microcontroller Neural Network).

[0047] In one implementation, the fusion layer possesses multimodal data dynamic fusion and feature joint representation capabilities. Its main tasks are: first, to fully integrate the effects of multi-source information, organically combining information from different sources and with different modes of representation to obtain a more accurate description of the identified object; and second, to optimize feature-level fusion, employing a joint representation system for heterogeneous data combinations and filtering and recombining multiple feature information from multiple sensors. The hardware primarily uses a 500MHz–800MHz microprocessor (STM32H757 or STM32N657), such as a Cortex-M7 microcontroller. The software primarily employs deep learning models such as CNN (Convolutional Neural Networks), MobileNet (Mobile Network), and YOLO (You Only Look Once). The heterogeneity of sensors (e.g., a hybrid architecture of multiple sensors) requires the model design to be deeply adapted to the hardware characteristics.

[0048] In one embodiment, the recognition layer has autonomous target recognition and decision output functions. Firstly, based on fused features, it combines feature learning and domain knowledge to use DNN (Deep Neural Networks) to process multi-label classification tasks, achieving high-resolution underwater target classification and recognition. Secondly, it strictly controls input and output information, isolating it from the detection and fusion layers to improve information security and device reliability. The hardware mainly uses a microprocessor with a main frequency of 1GHz to 1.5GHz, for example, based on... microprocessor( The software primarily uses deep learning frameworks such as DNN and TensorFlow Lite for MCUs (TFLM), and is based on the i.MX 8M Mini or NXP i.MX 93.

[0049] In one implementation, lightweight deep learning models, as a cutting-edge research area deeply integrating edge computing and embedded artificial intelligence, focus on breaking through the dependence of traditional deep learning models on high-performance computing resources. Through algorithmic innovation and hardware co-optimization, they achieve efficient deployment and real-time inference of deep neural networks on resource-constrained micro-devices (such as microcontroller units and low-power sensor nodes). This field focuses on resolving three key technical contradictions: first, the contradiction between the high computational complexity of deep learning models and the limited computing power of edge devices; second, the contradiction between model storage requirements and the memory capacity limitations of micro-devices; and third, the contradiction between cloud-based inference modes and the low latency and high privacy requirements of edge scenarios. Its technical path mainly covers three dimensions: first, model lightweighting technology, reducing model computation and storage overhead through parameter quantization (such as INT8 quantization), structural pruning, and knowledge distillation; second, dedicated lightweight framework design, such as TensorFlow Lite Micro and MCUNet, providing efficient model deployment and inference support for micro-devices; and third, hardware-algorithm co-optimization, combining dedicated accelerators to achieve model deployment under strict constraints of low memory consumption and low inference latency.

[0050] Although deep learning technology has made significant progress in the field of target detection, its high computational complexity and the resource constraints of underwater equipment (such as power consumption, computing power and storage) create a sharp contradiction. Lightweight deep learning models have become a key technical path to solve this contradiction by optimizing network structure and reducing the number of parameters and computation. The application value of lightweight models in underwater monitoring is mainly reflected in three aspects: First, by reducing the computational requirements through model compression technology (such as quantization and pruning), the algorithm can run in real time on edge devices; Second, the hierarchical architecture design (detection-fusion-recognition) can improve the robustness in complex environments, such as YOLOv8 achieving multi-scale feature extraction through backbone network and path fusion; Finally, hardware collaborative optimization (such as MCU and MPU collaboration) further reduces energy consumption and meets the requirements of long-term deployment. (1) Hierarchical architecture design of detection-fusion-recognition. The hierarchical architecture decomposes the target detection task into three levels: detection, fusion and recognition, realizing the rational allocation of computing resources and efficient information transmission. The advantages of this design concept are: 1) Modular design facilitates customized optimization for the special characteristics of the underwater environment; 2) Each level can be independently lightweighted, reducing the overall computational burden; 3) Reuse of features between levels reduces redundant calculations and improves real-time performance. (2) Hardware co-design and data interaction mechanism. Hardware co-design is a key link in achieving high-efficiency and low-power operation of the underwater target real-time monitoring system. Through the collaborative work of the microcontroller (MCU) and microprocessor (MPU), combined with optimized data transmission protocols and tiered hardware configuration, the system can achieve high-performance computing and real-time response in resource-constrained underwater environments. For example, the microcontroller and microprocessor work together. The collaborative working mechanism of MCU and MPU achieves efficient resource utilization through hierarchical task allocation. The MCU is responsible for low-level control tasks with high real-time requirements (such as sensor data acquisition and equipment status monitoring), while the MPU handles computationally intensive tasks (such as feature extraction and multimodal data fusion).

[0051] The underwater target real-time monitoring system is an integrated mechanical dynamic coordination underwater system. By constructing a three-layer computing architecture of "detection-fusion-recognition" and a tiered hardware configuration, combined with the TinyDL lightweight deep learning model and automated control algorithms, it achieves efficient interaction of multimodal sensor data and autonomous adaptation of the mechanical structure. The system simultaneously possesses a lightweight deep learning model and mechanical collaborative control capabilities. Through the tight coupling of the "detection-fusion-recognition" three-level architecture, it realizes multimodal data interaction. The collaborative mechanism between the model architecture and dynamic control establishes a feedback loop between the mechanical system and the recognition model, achieving closed-loop control of mechanical attitude adjustment and noise data fusion recognition. This effectively solves the contradiction between real-time performance, accuracy, and resource constraints in underwater monitoring.

[0052] The structure and function of each component are described below.

[0053] (a) Acoustic array 102: Acoustic array 102 is used to acquire underwater acoustic signals in real time. Acoustic array 102 is connected to data acquisition module 103 via a waterproof cable and converts underwater acoustic signals into electrical signals for transmission to data acquisition module 103.

[0054] The acoustic array 102 includes multiple hydrophones. These hydrophones are arranged in a tetrahedral array on the mounting platform 101 to collect underwater acoustic signals from different directions. The hydrophones are located at the four vertices of the tetrahedron and are omnidirectional within their operating frequency bandwidth, enabling them to receive acoustic signals from different directions and ensuring the omnidirectional and accurate acquisition of acoustic signals. Each hydrophone is carefully designed and calibrated to ensure its stability and sensitivity in the underwater environment. In this application, the tetrahedral arrangement of the acoustic array 102 enhances the stereoscopic and directional nature of signal acquisition, providing accurate data support for subsequent sound source localization and identification.

[0055] The support frame of the acoustic array 102 is made of HDPE material with a frame diameter of 15mm ± 0.2mm and an M6 threaded interface for fixing the hydrophone.

[0056] In a specific application example, multiple hydrophones are high-performance broadband voltage-type hydrophones (linear frequency response range 1Hz-100kHz). These high-performance broadband voltage-type hydrophones are mounted externally to the housing. The acoustic array 102 employs a tetrahedral topology, with one high-performance broadband voltage-type hydrophone positioned at each of the four vertices. The vertex spacing is precisely controlled to 200mm ± 0.05mm to ensure optimal spatial gain for the beamforming algorithm.

[0057] In another specific application example, due to the high cost of high-performance broadband voltage hydrophones, in order to reduce the cost of the underwater target real-time monitoring system, MEMS (Micro-Electro-Mechanical Systems) acoustic sensors are installed in the electronic instrument compartment on the inner wall of the shell as alternative acoustic acquisition devices. MEMS acoustic sensors have low cost and low power consumption, which can reduce the cost and power consumption of the underwater target real-time monitoring system.

[0058] (ii) Stand 101: Stand 101 is used to support acoustic array 102, data acquisition module 103, fusion device 104 and identification device 105.

[0059] In a specific application example, such as Figures 2 to 10As shown, the stand 101 is a simple stand, including an annular disk 1011, a support 1012, an electronic instrument compartment 1013, and a hydrophone frame 1014. The bottom of the support 1012 is fixedly connected to the annular disk 1011, and the support 1012 is fixedly fitted outside the electronic instrument compartment 1013. The annular disk 1011 is placed below the electronic instrument compartment 1013 and can support the support 1012 and the electronic instrument compartment 1013. The electronic instrument compartment 1013 can accommodate the data acquisition module 103, the fusion device 104, and the identification device 105. The hydrophone frame 1014 is fixedly mounted on the top of the support 1012 and has a first mounting part 1015, a second mounting part 1016, a third mounting part 1017, and a fourth mounting part 1018. Figures 8 to 10 Three different mounting methods are demonstrated for the first mounting part 1015, the second mounting part 1016, the third mounting part 1017, and the fourth mounting part 1018. At least one hydrophone can be fixedly mounted on the first mounting part 1015, the second mounting part 1016, the third mounting part 1017, and the fourth mounting part 1018. The first mounting part 1015, the second mounting part 1016, the third mounting part 1017, and the fourth mounting part 1018 are respectively positioned at the four vertices of a regular tetrahedron for easy deployment and maintenance. The structure is compact and provides stable support, ensuring the reliability and accuracy of the underwater target real-time monitoring system. It should be noted that the fixed connections between the structural components can be, but are not limited to, using bolts or clips to achieve a secure connection, ensuring the overall stability and reliability of the structure.

[0060] In a preferred embodiment of this invention, the annular disk 1011 is made of a high-strength material (such as ABS resin or stainless steel), possessing good corrosion resistance and mechanical strength. The annular disk 1011 has an inner diameter of 180mm, an outer diameter of 230mm, and a suitable thickness to provide sufficient mechanical strength. It is provided with multiple M6 threaded holes for connection to the bracket 1012. The annular disk 1011 is securely fixed to the center of the bottom of the bracket 1012 using M6 long screws, providing stable bottom support for the entire platform 101. The bracket 1012 is also made of a high-strength material (such as ABS resin). The diameter of the space inside the bracket 1012 used to accommodate the electronic instrument compartment 1013 is 250mm, and the height of the bracket 1012 is between 350mm and 500mm to ensure the stability of the entire platform 101. To ensure compactness and stability, the bracket 1012 is divided into multiple parts, each manufactured using 3D printing technology. These parts are then connected together through precision machining and assembly processes to form a complete bracket 1012 structure, improving structural accuracy and reliability. The top and sides of the bracket 1012 are equipped with multiple M8 and M6 threaded holes for connection to the annular disk 1011, the hydrophone frame 1014, or the electronic instrument compartment 1013. The hydrophone frame 1014, made of high-density polyethylene (HDPE) with a diameter of 15mm and featuring M6 threaded holes, is manufactured using 3D printing and possesses excellent corrosion resistance and mechanical strength. The hydrophone frame 1014 connects to the corresponding positions of the bracket 1012 via the M6 ​​threaded holes, ensuring that the acoustic array 102 can be precisely installed in the predetermined location.

[0061] In another specific application example, such as Figures 11 to 13As shown, the test stand 101 is a folding type, including an annular disk 1011, a support 1012, an electronic instrument compartment 1013, two first folding rods 1019, two second folding rods 1020, and several support rods. The bottom of the support 1012 is fixedly connected to the annular disk 1011, and the support 1012 is fixedly sleeved on the outside of the electronic instrument compartment 1013. The annular disk 1011 is placed below the electronic instrument compartment 1013. The annular disk 1011 can support the support 1012 and the electronic instrument compartment 1013. The electronic instrument compartment 1013 can accommodate the data acquisition module 103, the fusion device 104, and the identification device 105; the two first folding rods One end of each of the first folding rods 1019 is hinged to the top of the bracket 1012. The other end of one first folding rod 1019 has a first mounting portion 1015, and the other end of the other first folding rod 1019 has a second mounting portion 1016. One end of each of the two second folding rods 1020 is hinged to the bottom of the bracket 1012. The other end of one second folding rod 1020 has a third mounting portion 1017, and the other end of the other second folding rod 1020 has a fourth mounting portion 1018. At least one hydrophone can be fixedly mounted on the first mounting portion 1015, at least one hydrophone can be fixedly mounted on the second mounting portion 1016, and the third mounting portion 101... At least one hydrophone can be fixedly mounted on the 7th mounting part, and at least one hydrophone can be fixedly mounted on the fourth mounting part 1018; both first folding rods 1019 and two second folding rods 1020 can be folded to attach to the bracket 1012, and both first folding rods 1019 and two second folding rods 1020 can be unfolded and locked; after both first folding rods 1019 and two second folding rods 1020 are unfolded and locked, the first mounting part 1015, the second mounting part 1016, the third mounting part 1017 and the fourth mounting part 1018 can be placed on the four vertices of a regular tetrahedron; one end of the support rod is fixedly connected to the bracket 1012. The other end of the support rod is used to fix it to the unfolded first folding rod 1019 or the unfolded second folding rod 1020, which is convenient for folding for carrying or storage and for unfolding for use. It achieves a balance between portability and stability of the underwater target real-time monitoring system, and can flexibly switch between carrying and deployment, which improves the efficiency of the underwater target real-time monitoring system. In particular, it can meet the needs of frequently deploying the underwater target real-time monitoring system in different locations in marine environmental monitoring, improve the portability and deployment efficiency of the underwater target real-time monitoring system, and also help reduce the manufacturing and maintenance costs of the underwater target real-time monitoring system.

[0062] In a preferred embodiment of this invention, the support rod is made of stainless steel. The structure of the annular disk 1011 and the bracket 1012 is the same as that of the annular disk 1011 and the bracket 1012 of the simplified platform described above. The bracket 1012 is fixedly connected to the unfolded first folding rod 1019 or the unfolded second folding rod 1020 by the support rod. The support rod plays a role in reinforcement and stabilization to improve the overall structural stability. In addition, a support rod can also be provided between the two unfolded first folding rods 1019 to connect them as one unit, and a support rod can also be provided between the unfolded first folding rod 1019 and the unfolded second folding rod 1020 to connect them as one unit. A support rod can also be provided between the electronic instrument compartment 1013 and the bracket 1012 to connect them as one unit, thereby further improving the overall structural stability so as to form a stable tetrahedral structure, ensuring the stability and directivity of the acoustic array 102, thereby realizing real-time monitoring and automatic classification of underwater targets.

[0063] In a preferred embodiment of this invention, the platform 101 further includes a driving component, which is connected to the first folding rod 1019 and the second folding rod 1020 to provide power for the unfolding or folding of the first folding rod 1019 and the second folding rod 1020. In an optional embodiment, the driving component is a hydraulic system, which controls the operation of the hydraulic system to achieve precise control of the first folding rod 1019 and the second folding rod 1020, thereby realizing the automatic folding and unfolding function of the platform 101. In an optional embodiment, the driving component uses a motor to drive the first folding rod 1019 and the second folding rod 1020. Specifically, the driving component and the first folding rod 1019 and the second folding rod 1020 can use conventional transmission structures such as gears, chains, or belts to achieve power transmission, as long as the matching accuracy and transmission efficiency are ensured.

[0064] As an optional implementation, the annular disk 1011 and the support 1012 are manufactured by injection molding or machining. Injection molding has the advantages of high production efficiency and low cost; machining can ensure the precision and surface quality of the product. The first folding rod 1019, the second folding rod 1020 and the support rod are manufactured by extrusion molding or stretch molding. Extrusion molding and stretch molding can ensure that the folding rod and the support rod have good shape and dimensional accuracy, while meeting the requirements for material performance. The first folding rod 1019 and the second folding rod 1020 are made of aluminum alloy and high-density polyethylene (HDPE) and are connected to the support 1012 or the support rod by hinges, so that the frame 101 can be easily folded and unfolded.

[0065] In addition, the folding platform also includes a locking structure, which is installed on the support 1012 by means of bolts or pins. The locking structure is used to lock the platform 101 in a stable position after it is unfolded to prevent it from shaking or collapsing during operation. The locking structure should be designed to be simple, reliable, easy to operate and maintain. For example, a mechanical locking device or an electromagnetic locking device can be used to achieve the locking function.

[0066] Furthermore, the operation of the drive components and locking structure is controlled by receiving external commands or sensor signals through the microcontroller or embedded system, while ensuring that the microcontroller or embedded system can operate stably in complex underwater environments and has good anti-interference capabilities and data processing capabilities.

[0067] Power is supplied to the various components of the folding platform using battery packs or solar panels to meet the power requirements of the platform 101 in different working environments.

[0068] It should be noted that the connection between the annular disk 1011 and the bracket 1012 should be secure and reliable to prevent loosening or detachment during use. The coaxiality and perpendicularity of the annular disk 1011 and the bracket 1012 should also be ensured. During the assembly of the first folding rod 1019 and the second folding rod 1020, sufficient strength and stability should be ensured at the connection points between the first folding rod 1019 and the second folding rod 1020 and the bracket 1012. Furthermore, the connection points should be lubricated to reduce frictional resistance during folding and unfolding. The connection between the support rod and the bracket 1012 should also be carefully considered. During connection, support rods are installed at key locations such as the four corners or the middle of the bracket 1012, and are firmly fixed to the bracket 1012 using fasteners such as welding or bolts to enhance the structural rigidity and stability of the test bench 101. When installing the acoustic array 102 and the electronic instrument compartment 1013, the acoustic array 102 and the electronic instrument compartment 1013 are installed on the bracket 1012, and the positions of the acoustic array 102 and the electronic instrument compartment 1013 are ensured to be accurate, stable and reliable. At the same time, the acoustic array 102 and the electronic instrument compartment 1013 are waterproofed to ensure their normal operation in complex underwater environments.

[0069] After the assembly of test bench 101 is completed, overall debugging and testing should be carried out. Performance and stability tests should be conducted on test bench 101 by simulating the actual working environment to ensure that it meets design requirements and operates stably. During the testing process, relevant data should be recorded and analyzed to facilitate further optimization and improvement of test bench 101.

[0070] In this application example, the folding operation of the folding platform is as follows: The automatic folding mode of the hydraulic system is activated, and the opening and closing of the hydraulic valves are controlled to regulate the flow direction and volume of the hydraulic oil. Driven by the hydraulic system, the first folding rod 1019 and the second folding rod 1020 begin to fold automatically and press against the bracket 1012. After the first folding rod 1019 and the second folding rod 1020 are fully folded, it is checked whether the volume of the platform 101 has decreased to the expected value. If necessary, the parameter settings of the hydraulic system or the structure of the first folding rod 1019 and the second folding rod 1020 should be adjusted and optimized to improve folding efficiency.

[0071] The unfolding process of the folding test bench is as follows: The automatic unfolding mode of the hydraulic system is activated, and the opening and closing of the hydraulic valves are controlled to regulate the flow direction and volume of the hydraulic oil. Driven by the hydraulic system, the first folding rod 1019 and the second folding rod 1020 automatically unfold and form a stable support structure with the bracket 1012. After the first folding rod 1019 and the second folding rod 1020 are fully unfolded, support rods are installed, and the structural rigidity and stability of the test bench 101 are checked to ensure they meet the requirements. If necessary, the position or number of support rods should be adjusted and optimized to improve the stability of the test bench 101.

[0072] (III) Data Acquisition Module 103: The data acquisition module 103 is used to convert the underwater acoustic signals into digital signals and extract the audio features of the underwater acoustic signals. The data acquisition module 103 can realize simultaneous data acquisition from multiple hydrophones, ensuring data synchronization and consistency.

[0073] The data acquisition module 103 includes an analog-to-digital converter, a filter, an operational amplifier, and a first microcontroller connected in sequence. Furthermore, the data acquisition module 103 also includes a data storage unit and a communication interface. The data storage unit stores digital signals and audio characteristics, and the communication interface connects it to the fusion device 104.

[0074] (1) The analog-to-digital converter (ADC) is used to convert the underwater acoustic signal into a digital signal using the SAR (Successive Approximation Register) ADC algorithm. A high-precision ADC is employed to ensure that the converted digital signal accurately reflects the characteristics of the original analog signal. The SAR ADC algorithm features fast conversion speed, high accuracy, and low power consumption, making it suitable for real-time data acquisition scenarios. The SAR ADC algorithm performs the analog-to-digital conversion as follows:

[0075] ① Initialize the analog-to-digital converter module: Configure the sampling rate, resolution, and other parameters of the analog-to-digital converter. The sampling rate should be high enough to ensure that high-frequency components in the acoustic signal are captured; the resolution should be high enough to ensure that the converted digital signal has sufficient accuracy.

[0076] ② Initiate Analog-to-Digital Conversion: After the acoustic array 102 acquires the analog electrical signal, it initiates the analog-to-digital conversion process. The analog-to-digital conversion module converts the analog electrical signal into a digital signal.

[0077] ③ Read the analog-to-digital conversion result: After the conversion is completed, read the analog-to-digital conversion result and store it in a buffer for subsequent processing.

[0078] (2) The filter is used to filter the digital signal using the FIR (Finite Impulse Response) filtering algorithm to obtain a filtered signal. Filtering is the process of removing noise and interference components from a signal using a filter. This application uses a low-pass filter to remove high-frequency noise and improve signal quality. The design of the low-pass filter needs to be optimized according to the noise characteristics of the actual underwater environment to ensure that useful low-frequency signals are retained as much as possible while removing high-frequency noise.

[0079] The FIR filtering algorithm features linear phase and good stability, effectively removing high-frequency noise while preserving useful signal information. The specific filtering process is as follows:

[0080] ① Configure the FIR filter: Set the coefficients and order of the FIR filter. The coefficients and order of the filter should be reasonably selected according to the characteristics of the underwater acoustic signal and the noise situation to ensure the filtering effect.

[0081] ② Apply an FIR filter: The digital signal after analog-to-digital conversion is filtered by an FIR filter. The filter removes high-frequency noise and interference components from the signal, improving signal quality.

[0082] (3) An operational amplifier is used to amplify the filtered signal to obtain an amplified signal, ensuring the integrity and accuracy of the signal. The selection of the operational amplifier needs to consider its performance parameters such as gain, bandwidth, and noise to ensure that the amplified signal meets the requirements of subsequent processing. The specific amplification process is as follows:

[0083] ① Configure the operational amplifier: Set the gain and other parameters of the operational amplifier. The gain should be reasonably selected based on the strength of the underwater acoustic signal and the requirements of subsequent processing.

[0084] ② Operational Amplifier Application: The filtered signal is amplified by an operational amplifier. The operational amplifier amplifies the weak signal to an appropriate level for subsequent processing.

[0085] (4) The first microcontroller is used to extract the audio features of the amplified signal. The model of the first microcontroller is MCX-N947.

[0086] The audio features include MFCC (Mel-Frequency Cepstral Coefficient) and Mel spectrogram.

[0087] The process by which the first microcontroller extracts the Mel frequency cepstral coefficients is as follows: the amplified signal is sequentially processed by pre-emphasis, framing, windowing, fast Fourier transform, Mel filtering, logarithmic transform, and discrete cosine transform, wherein:

[0088] ① Pre-emphasis: The amplified signal is pre-emphasized to highlight the high-frequency resonance peak.

[0089] ② Framing: Divide the pre-emphasized signal into multiple frames.

[0090] ③ Windowing: Windowing is applied to each frame of the signal to reduce spectral leakage. Commonly used window functions include the Hamming window.

[0091] ④ Fast Fourier Transform (FFT): A Fast Fourier Transform is performed on the windowed signal in each frame to convert the signal from the time domain to the frequency domain, obtaining its spectrum. The Fast Fourier Transform is an efficient algorithm that can decompose a time-domain signal into a combination of different frequency components.

[0092] ⑤ Mel Filtering: A set of Mel filters is used to filter the spectrum, obtaining the sum of the energy of each Mel filter. Mel filters are designed to simulate the auditory characteristics of the human ear, enabling them to better capture the features of audio signals. A Mel filter is a set of triangular filters evenly distributed on the Mel scale; they simulate the auditory characteristics of the human ear and further process frequency domain signals.

[0093] ⑥ Logarithmic Transformation: Take the logarithm of the sum of the energy after Mel filtering to simulate the nonlinear perception of sound intensity by the human ear, obtaining the logarithmic energy representation of the Mel spectrum. The purpose of this step is to compress the spectral energy to a logarithmic scale in order to better reflect the acoustic characteristics.

[0094] ⑦ Discrete Cosine Transform (DAC) Processing: Perform a Discrete Cosine Transform on the logarithmic result, retaining the discrete cosine logarithm, as the MFCC feature. MFCC is an important audio feature that can well describe the spectral envelope and dynamic characteristics of audio signals.

[0095] The process of the first microcontroller extracting the Mel spectrum is as follows: the amplified signal is sequentially processed by frame segmentation, fast Fourier transform, Mel filtering, and logarithmic transform.

[0096] Mel spectrograms are an intuitive method of representing audio signals, clearly displaying the frequency distribution and temporal variations of an audio signal. They are generated by calculating the energy distribution of the signal at different frequencies. When extracting a Mel spectrogram, the signal is typically processed frame by frame. Each frame undergoes a Fast Fourier Transform (FFT) to obtain its frequency domain representation, followed by filtering and logarithmic transformation using a Mel filter bank. Finally, the Mel spectra of each frame are arranged temporally to form the Mel spectrogram.

[0097] In a specific application example, configuration and debugging steps are required before the data acquisition module 103 can be run.

[0098] Configuration steps: Set the sampling rate, resolution, and other parameters of the data acquisition module 103 according to actual needs. The sampling rate should be set according to the characteristics of the underwater acoustic signal and real-time requirements; the resolution should be set according to accuracy requirements. Configure the communication interface and protocol between the data acquisition module 103 and the fusion device 104. The communication interface and protocol should be selected and optimized according to actual needs and transmission distance.

[0099] Debugging steps: Use a standard signal source to send a test signal to the data acquisition module 103 and check whether the acquired signal is accurate. If there is an error, check and adjust the parameter settings or hardware connection of the data acquisition module 103. Connect the data acquisition module 103 to the fusion device 104 and test whether the communication is normal. If there is a communication failure, check and optimize the communication interface and protocol.

[0100] (iv) Fusion Device 104: The fusion device 104 is equipped with a lightweight deep learning model. The fusion device 104 is used to dynamically fuse the digital signal and the audio features based on domain-specific knowledge (underwater acoustics, oceanography, etc.) using the lightweight deep learning model to obtain fused features.

[0101] The fusion device 104 is a second microcontroller. Each of the second microcontrollers contains a lightweight deep learning model. The second microcontroller is either an STM32H757 or an STM32N657.

[0102] (v) Identification device 105: The identification device 105 is equipped with a lightweight deep learning model. The identification device 105 is used to identify targets based on the fused features using the lightweight deep learning model, so as to monitor underwater targets in real time.

[0103] In this application, the identification device 105 can identify a variety of underwater targets, including but not limited to common underwater targets such as underwater vehicles, fish, submarines, marine mammals, and underwater robots, as well as specific target features such as ship propeller parameters.

[0104] The identification device 105 includes a third microcontroller and a memory. Each third microcontroller is equipped with a lightweight deep learning model. The model of the third microcontroller is... i.MX 8M Mini or NXP i.MX 93. The memory is used to store the target recognition results.

[0105] In a specific application example, the third microcontroller uses a CNN (Convolutional Neural Networks) model based on TFLM (TensorFlow Lite Micro, a lightweight microcontroller version) for target recognition. The CNN model has powerful feature extraction and classification capabilities, enabling it to automatically learn feature representations of underwater targets and accurately identify unknown targets. The CNN model includes an input layer, multiple convolutional layers, pooling layers, fully connected layers, and an output layer.

[0106] The system consists of convolutional layers to extract features from underwater acoustic signals, pooling layers to reduce feature dimensionality and computational cost, and fully connected layers to map features to target categories. The output layer outputs the recognition results.

[0107] (1) Input layer: Receives fused features. During the training of the CNN model, the data in the input layer needs to undergo preprocessing operations such as normalization to ensure the stability and convergence of the CNN model training.

[0108] (2) Convolutional Layer: A series of learnable convolutional kernels (or filters) perform convolution operations on the fused features to extract local features such as edges and textures. A convolutional layer consists of multiple convolutional kernels (or filters). Each kernel slides across the fused features and performs convolution operations to obtain feature maps. These feature maps together constitute the output of the convolutional layer. The design of the convolutional layer needs to consider parameters such as the size, number, and stride of the convolutional kernels to ensure the effectiveness and efficiency of feature extraction.

[0109] Convolution can be expressed as the following mathematical formula:

[0110]

[0111] in, Let be the value of the output feature map of the l-th convolutional layer at position (i,j). Let be the weights of the l-th convolutional layer at position (m, n). Let b be the input value of the output feature map of the (l-1)th convolutional layer at position (i+m, j+n). l is the bias term of the l-th convolutional layer, f is the activation function, M is the lateral displacement, and N is the vertical displacement.

[0112] (3) Pooling layers: These are used to reduce feature dimensionality and computational cost while retaining important feature information. Pooling layers are usually placed after convolutional layers and perform downsampling operations on the feature maps, such as max pooling or average pooling. The design of pooling layers needs to consider parameters such as the size of the pooling window and the stride to ensure that useful feature information is retained as much as possible while reducing feature dimensionality.

[0113] (4) Fully Connected Layer: The fully connected layer flattens the feature map output by the pooling layer into a one-dimensional vector and maps it to the target category through a fully connected operation. A fully connected layer typically consists of multiple neurons, each connected to all neurons in the previous layer, acting as a classifier. The design of a fully connected layer requires consideration of parameters such as the number of neurons and the choice of activation function to ensure classification accuracy and efficiency.

[0114] (5) Output layer: Outputs the probability distribution corresponding to the target category. The output layer usually uses the Softmax function as the activation function to convert the score output by the fully connected layer into a probability distribution and selects the category with the highest probability as the recognition result.

[0115] The accuracy and efficiency of target recognition can be improved by training and optimizing the parameters of the CNN model. Therefore, the recognition device 105 also includes a deep learning accelerator to accelerate the process of training and optimizing the CNN model parameters. Specifically, a TFLM-based CNN model is trained on a server or PC using deep learning frameworks such as TensorFlow. The training process includes forward propagation, loss calculation, backpropagation, and parameter update. By adjusting the structure of the CNN model (such as kernel size, number of convolutional layers, number of fully connected layers, etc.), optimizing hyperparameters (such as learning rate, batch size, regularization parameters, etc.), and adopting advanced training strategies (such as transfer learning, data augmentation, early stopping, etc.), the recognition performance and generalization ability of the CNN model are improved.

[0116] The training process of a CNN model includes the following steps:

[0117] (1) Dataset Acquisition: Multiple publicly available underwater acoustic databases are used for model training, such as the ShipsEar dataset and the DeepShip dataset. These datasets contain acoustic signals of various underwater targets and provide corresponding label information (such as ship type, operating condition, etc.). In addition to publicly available datasets, data can also be collected and analyzed independently according to actual needs. For example, underwater acoustic monitoring equipment can be deployed in specific sea areas to collect long-term underwater acoustic signals, which can then be labeled and processed for model training.

[0118] (2) Training: Preprocessing and labeling the dataset. Preprocessing includes denoising and filtering; labeling involves classifying each signal in the dataset. A suitable deep learning algorithm and model architecture are selected, and the model is trained using the preprocessed dataset. During training, the model's parameters and structure are continuously adjusted to improve its recognition accuracy and generalization ability.

[0119] (3) Optimization: The trained model is validated and tested to evaluate its performance. If necessary, the model is further optimized and adjusted to improve its recognition accuracy and real-time performance. The optimized model is deployed to a third microcontroller and tested and validated using actual underwater acoustic signals to ensure the stability and reliability of the system.

[0120] This application uses an MCX-N947 microcontroller for data acquisition and an STM32H757 or STM32N657 microcontroller for feature fusion. The i.MX 8M Mini or NXP i.MX 93 microcontroller enables target identification, reducing the size and power consumption of the entire underwater target real-time monitoring system.

[0121] This application utilizes a collaborative design of hardware such as microcontrollers and microprocessors, and software such as deep learning frameworks. It establishes a three-layer computing architecture for multi-task and multi-modal fusion, namely a hierarchical functional division and tiered hardware configuration of "detection-fusion-recognition," reducing the performance impact between sensors, hardware / software platforms, and physical models within the underwater target real-time monitoring system, and improving the recognition accuracy under complex tasks. In the three-layer architecture, each layer achieves joint representation of multimodal features or heterogeneous hardware collaborative processing through joint representation. This application achieves a triple improvement in latency, energy efficiency, and accuracy in edge computing scenarios through a design concept of deep software and hardware adaptation. The core innovations of its hierarchical optimization framework of detection, processing, and recognition layers include: ensuring real-time performance in the detection layer, enabling heterogeneous computing in the processing layer, and achieving high-performance inference in the recognition layer. The impact of hardware characteristics on the performance of lightweight deep learning models in the key technologies of hardware collaborative optimization are mainly reflected in three key aspects: dynamic balance between clock speed and energy efficiency, collaborative optimization of hardware resources, and the adaptability of framework tools. The synergistic effect of these factors determines the model's real-time performance, energy efficiency, and computational efficiency on edge devices.

[0122] The underwater target real-time monitoring system features a lightweight deep learning model and mechanical collaborative control capabilities. Through a three-level architecture of "detection-fusion-recognition", it achieves multimodal data interaction and dynamic control collaboration, establishes a feedback loop between the mechanical system and the recognition model, and realizes closed-loop control of attitude adjustment, data fusion and noise recognition.

[0123] In another exemplary embodiment, the miniaturized underwater target real-time monitoring system based on hardware and software co-optimization and data interaction further includes a communication module 106. The communication module 106 is used to transmit target identification results to a remote monitoring center or cloud platform so that users can perform remote monitoring and data analysis.

[0124] The communication module 106 supports both wired and wireless communication methods, adapting to the needs of different application scenarios and improving the system's flexibility and scalability. For example, in marine monitoring missions, wireless communication methods (such as Wi-Fi or LoRa-WAN) can be selected for remote data transmission; in laboratory environments, wired communication methods (such as Ethernet or USB) can be selected for high-speed data transmission.

[0125] Wired communication methods include Ethernet, USB, RS-232 / 485 interfaces, which are characterized by high transmission speed and high stability. Wireless communication methods include Wi-Fi and LoRa-WAN for surface applications and underwater acoustic communication interfaces for underwater applications, which are characterized by high flexibility and ease of deployment.

[0126] The design of the communication module 106 needs to consider factors such as the compatibility of communication protocols and the security of data transmission to ensure the real-time performance and reliability of the data.

[0127] In another exemplary embodiment, the miniaturized underwater target real-time monitoring system based on hardware and software co-optimization and data interaction further includes a power management module 107.

[0128] The power management module 107 provides a stable and reliable power supply for the miniaturized underwater target real-time monitoring system based on hardware-software co-optimization and data interaction. It supports a wide voltage input range and features overvoltage protection, overcurrent protection, and short-circuit protection to ensure the safe and stable operation of the system. The design of the power management module 107 needs to consider factors such as power conversion efficiency and output voltage stability to ensure the power requirements of the miniaturized underwater target real-time monitoring system under different operating environments.

[0129] Specifically, the power management module 107 includes a lithium battery pack, a DC-DC converter, an overvoltage protection circuit, an overcurrent protection circuit, and a short-circuit protection circuit. The lithium battery pack serves as the main power source, boosting the voltage to the required level for the underwater target real-time monitoring system via the DC-DC converter. The overvoltage, overcurrent, and short-circuit protection circuits ensure that the underwater target real-time monitoring system will not be damaged under abnormal conditions, improving its safety and reliability. For example, when the battery voltage is too high or the current is too large, the overvoltage and overcurrent protection circuits will automatically cut off the power supply to prevent damage to the underwater target real-time monitoring system. The power management module 107 uses a lithium battery pack with a capacity of no less than 10,000mAh, supports battery replacement and USB charging, and has a battery life of no less than 7 days (continuous operation).

[0130] In another exemplary embodiment, the miniaturized underwater target real-time monitoring system based on hardware and software co-optimization and data interaction further includes a housing and encapsulation module 108. The test bench 101, data acquisition module 103, fusion device 104, identification device 105, communication module 106, and power management module 107 are all disposed inside the housing and encapsulation module 108.

[0131] The housing and encapsulation module 108 protects the test bench 101, data acquisition module 103, fusion device 104, identification device 105, communication module 106, and power management module 107 from external environmental influences, and ensures the waterproof performance and mechanical strength of the underwater target real-time monitoring system. The housing and encapsulation module 108 is made of high-strength ABS resin material, including the electronics compartment housing, frame housing, watertight connectors, and other components. It features corrosion resistance and a high waterproof rating, while also considering the module's sealing and heat dissipation performance to ensure the underwater target real-time monitoring system can operate stably for extended periods in harsh marine environments.

[0132] The outer shell is made of high-strength ABS resin, possessing excellent resistance to seawater corrosion, pressure, and impact. The internal layout and fixing structure of the shell ensure stable and reliable connections between modules. Simultaneously, the shell also features waterproof interfaces and sealing structures to ensure the waterproof performance of the underwater target real-time monitoring system. The instrument compartment is encapsulated using ABS resin injection molding, with dimensions of 210mm × 480mm (outer diameter × height) and a uniformly distributed 20mm wall thickness. Finite element analysis was used to optimize the reinforcing rib layout, improving the vibration resistance coefficient by 40%.

[0133] In summary, the miniaturized underwater target real-time monitoring system based on software and hardware co-optimization and data interaction provided in this application has advantages such as modular design, high-performance data acquisition and processing, intelligent target recognition, stable and reliable power management and communication, and good environmental adaptability. By adopting a hierarchical functional division of "detection-processing-recognition" and a tiered hardware configuration, combined with a dynamic resource management strategy, it can achieve a balance between efficiency and accuracy in scenarios such as industrial IoT, smart wearables, and smart homes, providing an efficient, accurate, and real-time solution for underwater target monitoring.

[0134] In one embodiment, the construction of a real-time underwater target monitoring system based on TinyDL includes the following two aspects:

[0135] (1) Lightweight model design and technology optimization.

[0136] First, select and optimize lightweight neural network architectures suitable for the TinyDL environment, such as MobileNet, SqueezeNet, and EfficientNet Lite. Based on underwater noise characteristics, adjust network structure parameters, such as depth, width, kernel size, and number of channels, employing strategies like depthwise separable convolution and grouped convolution to balance recognition accuracy and computational complexity. Simultaneously, consider lightweight recurrent neural networks or Transformer models (deep learning models based on self-attention mechanisms) suitable for time-series data.

[0137] Second, model compression and quantization techniques, including quantization, pruning, and knowledge distillation, are employed to reduce model size and improve inference speed. This involves minimizing the impact of different quantization methods on model accuracy by selecting an appropriate quantization bit width; developing pruning strategies specifically for propeller noise recognition tasks; and utilizing large, complex models as teacher models to enhance the generalization ability and recognition accuracy of smaller models through knowledge distillation.

[0138] Third, for the selected TinyDL hardware platform (such as the ARM Cortex-M series), use libraries such as TensorFlow to perform model conversion and optimization, and implement operations such as operator fusion, memory optimization, and code optimization to improve model running efficiency.

[0139] (2) Construction and performance evaluation of real-time monitoring system.

[0140] First, the optimized TinyDL model is deployed to a hardware platform and integrated with underwater acoustic sensors, data acquisition module 103, power management module 107, etc., to build a miniaturized real-time underwater acoustic target monitoring system. Low-power hardware circuitry is designed, and data transmission and processing flows are optimized to meet real-time requirements. The system adopts a modular architecture for easy maintenance and expansion, and a user-friendly interface is designed.

[0141] Secondly, the system was tested in both laboratory and actual underwater environments, and its performance was evaluated using metrics such as accuracy, inference speed, power consumption, and memory usage. The impact of environmental factors on system performance was analyzed, and the system's robustness and stability were assessed.

[0142] Third, based on the performance evaluation results, the applicability and limitations of the system in different application scenarios (such as marine environmental monitoring, underwater navigation safety, underwater target identification, and underwater structure health monitoring) are discussed. System improvement and optimization schemes are proposed to address the needs of specific application scenarios, promoting the practical application of the results.

[0143] The overall testing and verification process of the miniaturized underwater target real-time monitoring system provided in this application is as follows:

[0144] Test Procedure: Place the installed and debugged miniaturized underwater target real-time monitoring system in a simulated underwater environment for testing. During the test, simulate different types of underwater targets and underwater acoustic signal environments to comprehensively evaluate the system's performance. Use a remote monitoring center or cloud platform to receive and process the data transmitted by the miniaturized underwater target real-time monitoring system, and check whether the accuracy and real-time performance of the data meet the requirements. If errors or delays are found, adjust and optimize the communication module 106 or the data processing algorithm of the miniaturized underwater target real-time monitoring system.

[0145] Verification Steps: Based on the test results, evaluate and analyze the various performance indicators of the miniaturized underwater target real-time monitoring system to determine whether it meets the expected design requirements and usage needs. If necessary, further optimization and improvement work should be carried out on the miniaturized underwater target real-time monitoring system to improve its overall performance and reliability. Detailed technical documentation and user manuals should also be prepared for user reference and use.

[0146] The assembly and debugging process of the miniaturized underwater target real-time monitoring system provided in this application is as follows:

[0147] Hardware Assembly: Assemble the structural components of the simplified or folding platform, ensuring that the connections between the components are firm and reliable, meeting the requirements of mechanical strength and stability. Mount the acoustic array 102 on the first mounting section 1015, second mounting section 1016, third mounting section 1017, and fourth mounting section 1018, ensuring the hydrophones are arranged in a tetrahedral array. Install electronic equipment such as the data acquisition module 103, fusion device 104, identification device 105, communication module 106, and power management module 107 within the electronic instrument compartment 1013, and connect and secure them accordingly.

[0148] Software Debugging: Based on the MCUXpresso SDK development environment and MCUXpresso IDE debugging tool, C language was used for programming and debugging. First, the data acquisition module 103 was initialized and configured, including setting parameters for the analog-to-digital converter, filters, operational amplifiers, etc. Then, algorithms for data acquisition, preprocessing, and feature extraction were written, debugged, and optimized to ensure the data acquisition and preprocessing effects met requirements. Next, the fusion device 104 and recognition device 105 were initialized and the model was trained. Public datasets (such as the ShipsEar dataset, DeepShip dataset, etc.) or self-made datasets were used for model training, adjusting the model structure and parameters to improve recognition accuracy. After training, the model was exported and integrated into the fusion device 104 and recognition device 105. Finally, the communication module 106 and power management module 107 were initialized, configured, and tested. This ensured that the communication module 106 could transmit data normally to the remote monitoring center or cloud platform, and that the power management module 107 could provide a stable and reliable power supply.

[0149] The actual working process of the miniaturized underwater target real-time monitoring system provided in this application is as follows:

[0150] (1) System Deployment

[0151] Choose the appropriate rack type (simple rack or folding rack) based on the actual application scenario. For scenarios requiring rapid deployment, choose a simple rack; for scenarios requiring transportation and storage, choose a folding rack.

[0152] The acoustic array 102 is mounted on the stand 101 and connected to the data acquisition module 103. Ensure the orientation and position of the acoustic array 102 meet the monitoring requirements for accurate acquisition of underwater acoustic signals. Specifically, multiple high-performance broadband hydrophones are arranged in a tetrahedral array on the first mounting section 1015, second mounting section 1016, third mounting section 1017, and fourth mounting section 1018. During installation, ensure the relative positions of the hydrophones are accurate and use adhesives such as glue to firmly fix the hydrophones to the first mounting section 1015, second mounting section 1016, third mounting section 1017, and fourth mounting section 1018. When a simple stand is selected, the hydrophone frame with the installed hydrophones is fixed to the bracket 1012, and the bracket 1012 is placed on the annular disk 1011, connected to the annular disk 1011 using bolts or other fasteners. When connecting, ensure that the connection is secure and reliable, and check that the acoustic array 102 is in a horizontal position.

[0153] The system connects components such as the data acquisition module 103, fusion device 104, identification device 105, communication module 106, and power management module 107, ensuring correct electrical connections between each module. Specifically, a signal generator sends a test signal to the acoustic array 102, and the data acquisition module 103 acquires the received signal. The acquired signal is transmitted to the fusion device 104 for real-time analysis and processing. The accuracy of the identification result of the identification device 105 on the test signal is checked. If there is an error, the installation position of the acoustic array 102 or the signal processing algorithm should be adjusted and optimized to improve the identification accuracy.

[0154] The entire miniaturized underwater target real-time monitoring system is placed in the housing and encapsulation module 108, and the watertight connector and other components are sealed to ensure the waterproof performance of the miniaturized underwater target real-time monitoring system.

[0155] (2) Data Acquisition and Processing

[0156] The data acquisition module 103 is activated to begin acquiring underwater acoustic signals. The data acquisition module 103 converts analog electrical signals into digital signals and performs preprocessing operations such as filtering and amplification to improve signal quality and recognition accuracy.

[0157] During data acquisition, key audio features (such as MFCC and Mel spectrograms) are extracted and transmitted to the fusion device 104 for further processing. The feature extraction process employs efficient algorithms and mathematical formulas to ensure the accuracy and reliability of the features.

[0158] (3) Target recognition and data transmission

[0159] The fusion device 404 receives digital signals and audio features transmitted by the data acquisition module 103, and uses a TFLM-based CNN model to perform feature fusion and target recognition. The CNN model maps features to target categories and outputs the recognition results.

[0160] The communication module 106 transmits the identification results to a remote monitoring center or cloud platform for real-time monitoring and data analysis. Simultaneously, the communication module 106 can also transmit raw underwater acoustic signals or processed data to the remote monitoring center for storage and analysis.

[0161] (4) Power Management and Maintenance

[0162] The power management module 107 provides a stable and reliable power supply for the miniaturized underwater target real-time monitoring system and monitors parameters such as battery level and voltage. When the battery is low, the miniaturized underwater target real-time monitoring system can be charged or replaced via USB charging or battery replacement. The power management module 107 is designed to ensure sufficient endurance and safety performance in complex underwater environments.

[0163] Regular maintenance and upkeep of the miniaturized underwater target real-time monitoring system are essential. This includes checking the connections and operational status of all components to ensure long-term stable operation. Faulty components should be repaired or replaced promptly. Experiments show that the miniaturized underwater target real-time monitoring system provided in this application maintains an accuracy rate of over 90% even in sea state 5, demonstrating its practical application value.

[0164] The miniaturized underwater target real-time monitoring system provided in this application has broad application prospects. The following is a detailed analysis of its application prospects:

[0165] Marine environmental monitoring: This application can be used in the field of marine environmental monitoring to achieve real-time monitoring and automatic classification of underwater targets. Through the design of a foldable platform, the miniaturized real-time underwater target monitoring system can be more easily deployed in different sea areas, providing strong support for marine environmental protection and resource development.

[0166] Underwater Engineering Monitoring: In the field of underwater engineering, such as the laying of subsea pipelines and the construction of underwater tunnels, real-time monitoring of the underwater environment is required. The miniaturized real-time underwater target monitoring system provided in this application can meet the needs of these applications, providing strong support for the safe construction and quality control of underwater engineering projects.

[0167] Military reconnaissance and anti-submarine warfare: In the field of military reconnaissance and anti-submarine warfare, this application can be used for real-time monitoring and location of underwater targets. Through the design of a foldable platform, the miniaturized real-time underwater target monitoring system can be more easily deployed in complex underwater environments, providing strong support for military reconnaissance and anti-submarine operations.

[0168] Scientific Research and Education: In the field of marine science research, this application can be used for the observation and study of marine organisms, marine geology, and other phenomena. Simultaneously, in the field of education, it can also serve as experimental equipment for marine science teaching, helping students better understand and master marine science knowledge.

[0169] In summary, the beneficial effects of this application include at least the following:

[0170] 1) Miniaturized Design: The underwater target real-time monitoring system is compact in size, facilitating deployment and maintenance. This is particularly beneficial for the complex environment of marine monitoring, enhancing the system's applicability and flexibility. The compact structural design, with efficient interfaces connecting the modules, significantly reduces the system's size, simplifying deployment and maintenance.

[0171] 2) Low power consumption design: The low power consumption hardware and software design extends the endurance of the underwater target real-time monitoring system, reduces operating costs, and improves the economy and practicality of the underwater target real-time monitoring system.

[0172] 3) Strong real-time performance: Based on the high-performance processing and communication module 106, real-time data acquisition, processing and transmission are realized, ensuring the timeliness and accuracy of monitoring results.

[0173] 4) High recognition accuracy: The use of deep learning algorithms improves the accuracy of target recognition and can accurately identify a variety of underwater targets, including propeller cavitation noise.

[0174] 5) Strong environmental adaptability: It can operate stably in complex underwater environments and has good resistance to seawater corrosion, high pressure, and impact.

[0175] 6) High portability: Through the support and fixation of key components such as the acoustic array 102 by the foldable platform, the platform 101 has the characteristics of compact structure, easy folding and unfolding, and good stability, which greatly improves the portability and deployment efficiency of the underwater target real-time monitoring system.

[0176] It should be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with relevant regulations.

[0177] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0178] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A miniaturized real-time underwater target monitoring system based on software and hardware co-optimization and data interaction, characterized in that, The system includes: a detection device, a fusion device, and a recognition device; The detection device includes a stand, an acoustic array, and a data acquisition module; the stand supports the acoustic array, the data acquisition module, the fusion device, and the recognition device; the acoustic array is used to acquire underwater acoustic signals in real time; the data acquisition module converts the underwater acoustic signals into digital signals and extracts the audio features of the underwater acoustic signals. The fusion device is equipped with a lightweight deep learning model. The fusion device is used to dynamically fuse the digital signal and the audio features based on professional domain knowledge and the lightweight deep learning model to obtain fused features. The identification device is equipped with a lightweight deep learning model, which is used to identify targets based on the fused features, so as to monitor underwater targets in real time.

2. The miniaturized underwater target real-time monitoring system based on software and hardware co-optimization and data interaction according to claim 1, characterized in that, The acoustic array includes multiple hydrophones; the multiple hydrophones are arranged in a tetrahedral array on the platform to collect underwater acoustic signals from different directions.

3. The miniaturized underwater target real-time monitoring system based on software and hardware co-optimization and data interaction according to claim 2, characterized in that, The platform includes an annular disk, a support, an electronic instrument compartment, and a hydrophone frame. The bottom of the support is fixedly connected to the annular disk, and the support is fixedly sleeved outside the electronic instrument compartment. The annular disk is positioned below the electronic instrument compartment and can support the support and the electronic instrument compartment. The electronic instrument compartment can accommodate the data acquisition module, the fusion device, and the identification device. The hydrophone frame is fixedly mounted on the top of the support and has a first mounting part, a second mounting part, a third mounting part, and a fourth mounting part. At least one hydrophone can be fixedly mounted on the first mounting part, at least one hydrophone can be fixedly mounted on the second mounting part, at least one hydrophone can be fixedly mounted on the third mounting part, and at least one hydrophone can be fixedly mounted on the fourth mounting part. The first mounting part, the second mounting part, the third mounting part, and the fourth mounting part are respectively located at the four vertices of a regular tetrahedron.

4. The miniaturized underwater target real-time monitoring system based on software and hardware co-optimization and data interaction according to claim 2, characterized in that, The platform includes an annular disk, a support, an electronic instrument compartment, two first folding rods, two second folding rods, and several support rods. The bottom of the support is fixedly connected to the annular disk, and the support is fixedly sleeved outside the electronic instrument compartment. The annular disk is positioned below the electronic instrument compartment. The annular disk can support the support and the electronic instrument compartment. The electronic instrument compartment can accommodate the data acquisition module, the fusion device, and the identification device. One end of each of the two first folding rods is hinged to the top of the support. The other end of one first folding rod has a first mounting part, and the other end of the other first folding rod has a second mounting part. One end of each of the two second folding rods is hinged to the bottom of the support. The other end of one second folding rod has a third mounting part, and the other end of the other second folding rod has a fourth mounting part. At least one hydrophone can be fixedly mounted on the first mounting part, at least one hydrophone can be fixedly mounted on the second mounting part, at least one hydrophone can be fixedly mounted on the third mounting part, and at least one hydrophone can be fixedly mounted on the fourth mounting part; both first folding rods and both second folding rods can be folded to attach to the bracket, and both first folding rods and both second folding rods can be unfolded and locked; after both first folding rods and both second folding rods are unfolded and locked, the first mounting part, the second mounting part, the third mounting part, and the fourth mounting part can be respectively placed on the four vertices of a regular tetrahedron; one end of the support rod is fixedly connected to the bracket, and the other end of the support rod is used to fixally connect to the unfolded first folding rod or the unfolded second folding rod.

5. The miniaturized underwater target real-time monitoring system based on software and hardware co-optimization and data interaction according to claim 3 or 4, characterized in that, The electronic instrument compartment is equipped with microelectromechanical system acoustic sensors.

6. The miniaturized underwater target real-time monitoring system based on software and hardware co-optimization and data interaction according to claim 1, characterized in that, The data acquisition module includes an analog-to-digital converter, a filter, an operational amplifier, and a first microcontroller connected in sequence. The analog-to-digital converter is used to convert the underwater acoustic signal into a digital signal using a successive approximation analog-to-digital conversion algorithm. The filter is used to filter the digital signal using a finite impulse response filtering algorithm to obtain a filtered signal; The operational amplifier is used to amplify the filtered signal to obtain an amplified signal; The first microcontroller is used to extract the audio features of the amplified signal.

7. The miniaturized underwater target real-time monitoring system based on software and hardware co-optimization and data interaction according to claim 6, characterized in that, The first microcontroller is model MCX-N947.

8. The miniaturized underwater target real-time monitoring system based on software and hardware co-optimization and data interaction according to claim 6, characterized in that, The audio features include Mel frequency cepstral coefficients and Mel spectrograms; The process of the first microcontroller extracting the Mel frequency cepstral coefficients is as follows: the amplified signal is sequentially processed by pre-emphasis, framing, windowing, fast Fourier transform, Mel filtering, logarithmic transform, and discrete cosine transform. The process of the first microcontroller extracting the Mel spectrum is as follows: the amplified signal is sequentially processed by frame segmentation, fast Fourier transform, Mel filtering, and logarithmic transform.

9. The miniaturized underwater target real-time monitoring system based on software and hardware co-optimization and data interaction according to claim 1, characterized in that, The fusion device is a second microcontroller; the recognition device includes a third microcontroller and a memory; both the second microcontroller and the third microcontroller are equipped with lightweight deep learning models; the memory is used to store the target recognition results.

10. The miniaturized underwater target real-time monitoring system based on software and hardware co-optimization and data interaction according to claim 9, characterized in that, The second microcontroller is either an STM32H757 or an STM32N657, and the third microcontroller is of the following model: i.MX 8M Mini or NXP i.MX 93.

11. The miniaturized underwater target real-time monitoring system based on software and hardware co-optimization and data interaction according to claim 1, characterized in that, The system also includes a communication module; the communication module is used to transmit the target recognition results to a remote monitoring center or cloud platform so that users can perform remote monitoring.