A VHF and IP Voice Fusion Communication Method and Device for Special Vessels

By real-time environmental monitoring and dynamic optimization of radio frequency parameters, combined with complex convolutional attention networks and protocol conversion engines, deep integration of VHF and IP voice systems is achieved, solving the communication integration problem in polar ships and improving the environmental adaptability and reliability of the communication system.

CN122001863BActive Publication Date: 2026-06-30SHANGHAI MARITIME UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI MARITIME UNIVERSITY
Filing Date
2026-04-09
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

In polar vessels, the separate deployment of VHF and IP voice systems results in high communication system costs, poor environmental adaptability, and functional fragmentation, making it difficult to meet the needs of collaborative operations.

Method used

By employing real-time environmental monitoring and dynamically optimized RF front-end transmission parameters, combined with complex convolutional attention networks and protocol conversion engines, deep integration of VHF and IP voice is achieved. Through multimodal noise suppression and adaptive encoding and decoding, communication quality is guaranteed.

Benefits of technology

Significantly reduces hardware deployment and maintenance costs, improves equipment reliability and operational efficiency, ensures communication quality and reliability, enables rapid analysis and shipwide broadcasting of DSC distress signals, and provides highly reliable communication assurance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122001863B_ABST
    Figure CN122001863B_ABST
Patent Text Reader

Abstract

This invention provides a method and apparatus for VHF and IP voice fusion communication on special vessels. The method includes: real-time monitoring of ambient temperature and electromagnetic interference intensity; dynamically optimizing radio frequency front-end transmission parameters based on a pre-stored antenna parameter compensation library matched to ambient temperature; performing multimodal noise suppression on the received VHF audio signal, including: generating a time-frequency mask using a complex convolutional attention network based on electromagnetic interference intensity to suppress electromagnetic interference; converting the processed signal into IP voice data packets, and then inversely converting the IP voice data packets into VHF signals through a protocol conversion engine to achieve a bidirectional communication link. This invention improves the quality and reliability of voice communication in polar environments, ensuring navigation safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of ship communication technology, and in particular to a VHF and IP voice fusion communication method and apparatus for special ships. Background Technology

[0002] Currently, all special-purpose vessels in service in my country have established relatively complete computer network systems, providing various IP-based communication services within the vessels, such as VoIP, video surveillance, video conferencing, and eLTE digital trunking. Meanwhile, according to shipbuilding specifications, VHF communication systems are also an indispensable and important component of the communication and navigation systems of polar vessels. VHF is a traditional radio communication system widely used in aviation, maritime, and ground emergency communications. VoIP, on the other hand, is a communication method based on the Internet Protocol (IP) and carried over IP networks, achieving voice communication through packet switching technology. VHF and VoIP systems differ significantly in communication principles, transmission media, and communication protocols; currently, VHF and VoIP on in-service polar vessels operate and are used independently.

[0003] On commercial or other traditional vessels, VHF communication systems, as an important means of shipboard communication, are primarily used by crew members for onboard operations, and the need for integration and interoperability between the two systems is not particularly prominent. Specialty vessels, however, consist of crew members, onboard research team members, transit research team members, and expedition management personnel, typically numbering over a hundred. In addition to the shipboard VHF system, most personnel utilize IP-based communication services. Polar expeditions involve the unified scheduling and collaborative efforts of all parts of the vessel and all types of personnel. Close cooperation among all personnel is essential to ensure operational safety and the completion of tasks. In this scenario, the need for integration and interoperability of various communication methods within polar vessels is particularly prominent.

[0004] To address the aforementioned issues, existing technologies rely on combinations of general-purpose commercial equipment and lack integrated design tailored to the low-temperature, high-interference environment, multi-source noise characteristics, and essential collaborative operation requirements of polar vessels, thus limiting their deployment in polar scientific research scenarios. Therefore, there is an urgent need for a highly integrated, environmentally adaptive, and fully functional integrated communication device to achieve deep fusion of VHF and IP voice systems on polar vessels. Summary of the Invention

[0005] This invention proposes a VHF and IP voice fusion communication method and device for special vessels, which solves the problems of high cost, poor environmental adaptability and functional fragmentation of communication systems in polar environments.

[0006] The present invention specifically provides the following technical solution:

[0007] A VHF and IP voice fusion communication method for special vessels, the method comprising the following steps:

[0008] Real-time monitoring of ambient temperature and electromagnetic interference intensity;

[0009] Based on the pre-stored antenna parameter compensation library for ambient temperature matching, the RF front-end transmission parameters are dynamically optimized; the transmission parameters include antenna phase center offset, gain compensation value and impedance matching coefficient.

[0010] Multimodal noise suppression processing is performed on the received VHF audio signal, including:

[0011] Based on the electromagnetic interference intensity, a complex convolutional attention network is used to generate a time-frequency mask to suppress electromagnetic interference;

[0012] The processed signal is converted into IP voice data packets, and the IP voice data packets are then converted into VHF signals through a protocol conversion engine to achieve a two-way communication link.

[0013] Optionally, the matching process of the pre-stored antenna parameter compensation library specifically includes:

[0014] The root mean square error between the ambient temperature and the temperature range in the pre-stored parameter library is calculated in real time. If the root mean square error is less than a set threshold, the antenna phase center offset, gain compensation value and impedance matching coefficient of the corresponding temperature range are called.

[0015] When the ambient temperature is below -30℃, the deformation compensation coefficient of the titanium alloy phase change temperature control shell is additionally activated. The deformation compensation coefficient satisfies the following relationship:

[0016]

[0017] in, For phase change temperature control shell deformation, The coefficient of thermal expansion of titanium alloy is 8.6 × 10⁻⁶. -6 K -1 , This is the initial length of the antenna. For reference to ambient temperature, The wind load deformation coefficient is taken as 0.024 mm·m. 2 / N, For wind stress load, The angle between the wind direction and the antenna axis.

[0018] Optionally, a time-frequency mask is generated through the complex convolutional attention network. :

[0019]

[0020] in, For learnable convolutional kernels, For complex number convolution operators, For querying the matrix, Let be the transpose of the key matrix. This is the scaling factor;

[0021] The training loss function of the complex convolutional attention network for:

[0022]

[0023] in, The time-frequency matrix for pure speech. The time-frequency matrix is ​​the output of the network prediction. The regularization coefficient is . The noise distribution estimated for the network. For the noise prior distribution, For frequency variables, For the i-th order noise energy weight, Let i be the center frequency of the i-th order noise. For the i-th order noise bandwidth, It is the probability density function of a Gaussian distribution.

[0024] Optionally, the complex convolutional attention network adopts a joint architecture of complex conformer modules and a three-dimensional attention mechanism, specifically including:

[0025] The complex number conformer module employs a real-part dilated convolution path in the time dimension, with a kernel dilation factor of d. t The output features are:

[0026]

[0027] An imaginary part dilated convolution path is used in the frequency dimension, with a kernel dilation factor of d. f The output features are:

[0028]

[0029] Z r Z i These are the real and imaginary matrices of the input complex spectrum, respectively. t and d f Dynamically adjusted according to the intensity of electromagnetic interference;

[0030] The three-dimensional attention mechanism specifically involves performing a three-dimensional dynamic weighting of the features F=Fr+jFi output by the complex conformer module, based on channel, time, and frequency.

[0031]

[0032] in, , , These are channel attention weights, temporal attention weights, and frequency attention weights, respectively.

[0033] Weighted features Input a complex gated cyclic unit and output a time-frequency mask.

[0034] Optionally, the specific steps for the protocol conversion engine to implement a bidirectional communication link include a forward link from VHF to IP and a reverse link from IP to VHF.

[0035] The forward link specifically inputs the VHF audio signal after multimodal noise suppression into a dynamic codec. When the end-to-end transmission delay is less than 200 milliseconds, the ITU-T G.711 codec standard is used; when the delay is between 200 and 500 milliseconds, the ITU-T G.729A codec standard is switched. An environmentally aware data tag is encapsulated in the Real-Time Protocol (RTP) header, including a real-time temperature value tag X-Temp: {T} and an electromagnetic interference intensity tag X-EMI: {EMI}, where {T} and {EMI} are the ambient temperature and the electromagnetic interference intensity, respectively.

[0036] The reverse link specifically involves parsing the DSC distress identifier in the IP data packet. If the distress identifier indicates a distress state, the VHF signal reconstruction module is triggered to prioritize processing and is marked as the highest service level of IEEE 802.1p priority 7. The codec matching the sending end, including G.711 or G.729A, is invoked to decode the IP voice data, and electromagnetic interference is suppressed a second time through the complex convolutional attention network. The complex convolutional attention network is activated when the X-EMI tag value exceeds 20dB.

[0037] Optionally, the synchronization mechanism of the dynamic codec includes:

[0038] The sending end declares the current codec type in the Codec-Type extension field of the RTP header. The codec type includes 0x01 for G.711 and 0x02 for G.729A.

[0039] The receiving end parses the Codec-Type field and automatically switches to the same codec standard to ensure bidirectional voice codec consistency.

[0040] When the MOS score calculated by the voice quality closed-loop monitoring module is lower than 3.0, the bidirectional link synchronously degrades to the 6.3kbps low bit rate coding mode of the ITU-TG.723.1 coding standard.

[0041] Optionally, the guarantee mechanism for bidirectional transmission is as follows:

[0042] A primary and backup dual-queue buffer system is configured at the switch layer. The primary queue transmits the current voice frame in real time, while the backup queue stores the data for the next voice frame. When network congestion is detected, the queue switching is completed within 10 milliseconds.

[0043] Implement dynamic jitter buffer control, with a base buffer depth set at 50 milliseconds, and dynamically adjust the buffer depth according to the network latency change rate. Buffer depth = 50 + 20 × tanh (0.1 × latency change rate).

[0044] A forward error correction redundancy encapsulation mechanism is enabled. Redundant error correction packets are added for every 10 voice data packets transmitted. The number of redundant packets is dynamically configured according to the real-time packet loss rate: no redundant packets are added when the packet loss rate is less than 1%, 20% of redundant packets are added when the packet loss rate is between 1% and 5%, and 30% of redundant packets are added when the packet loss rate exceeds 5%.

[0045] Optionally, the IP voice transmission process includes closed-loop monitoring of voice quality:

[0046] The receiver periodically calculates the average opinion score (MOS) value: MOS = 4.5 - 0.008 × end-to-end delay - 0.032 × packet loss rate percentage;

[0047] When the MOS score is below 3.0, a codec downgrade command is sent to the transmitter to automatically switch to the 6.3kbps low bit rate coding mode of the ITU-T G.723.1 standard.

[0048] Establish an encrypted transmission quality log, link it to storage temperature, electromagnetic interference intensity, latency and MOS score data, and retain the log for no less than 3 months as required.

[0049] Optionally, the bidirectional communication link adopts a polar environment-enhanced design, including:

[0050] The network equipment meets the wide operating temperature range of -40℃ to +70℃ specified in IEC 60092-509 standard;

[0051] The backbone communication link uses armored optical cable with a tensile strength of not less than 2000N;

[0052] The switch is equipped with a phase change heat dissipation module, filled with paraffin-based composite material with a melting point of -35℃, and the latent heat of phase change is not less than 200J / g;

[0053] Key communication ports are equipped with π-type filters to provide an insertion loss of no less than 40 dB in the 30-300MHz frequency band;

[0054] The equipment cabinet meets the 90 dB shielding effectiveness requirement specified in the MIL-STD-461G standard.

[0055] The present invention also provides a VHF and IP voice converged communication device for special vessels, the device comprising:

[0056] The main control module uses a multi-core heterogeneous processor to process environmental perception data, control protocol conversion, and schedule resources.

[0057] The VHF signal processing module connects to the ship's VHF radio and has a built-in three-level electromagnetic shielding structure and DSC decoding unit. It is used to receive and preprocess VHF signals and output noise-reduced audio streams.

[0058] The speech enhancement module integrates an adaptive comb filter bank and a complex convolutional attention network hardware accelerator to suppress engine noise and electromagnetic interference and generate speech signals that conform to IP transmission standards.

[0059] The environmental perception unit includes a wide-temperature sensor, a triaxial magnetometer, and a combined navigation module. It collects temperature, electromagnetic interference intensity, and position data in real time, providing input for thermal deformation compensation and position correction.

[0060] The protocol conversion engine is configured to dynamically select codecs, encapsulation environment tags, and location metadata to convert VHF audio streams into IP voice data packets.

[0061] The network transmission module supports dual network port redundancy and Wi-Fi 6 protocol, and implements dual queue buffering, dynamic jitter control and forward error correction mechanism to ensure the reliability of voice transmission within the ship's local area network;

[0062] The polar-reinforced structure includes a phase-change temperature control housing and an armored connector;

[0063] The power management system integrates a lithium titanate backup battery to ensure battery life in extreme environments.

[0064] This invention offers the following beneficial technical effects: It provides a VHF and IP voice fusion communication method and device for special-purpose vessels. This invention integrates the traditionally separately deployed softswitch server, IP gateway, and noise suppression functions into a single device, significantly simplifying the system architecture and reducing hardware deployment and maintenance costs. Modular design supports rapid maintenance and replacement, effectively improving equipment reliability and operational efficiency. A phase-change temperature-controlled shell combined with a dynamic thermal deformation compensation mechanism ensures stable operation of the device in extremely low-temperature environments, effectively suppressing antenna parameter drift. Simultaneously, a multi-level electromagnetic shielding structure and electromagnetic interference suppression algorithm work synergistically to significantly improve communication quality in complex electromagnetic environments. It enables rapid analysis and ship-wide broadcasting of DSC distress signals, ensuring priority access to critical communication channels through service quality grading. Adaptive noise suppression technology dynamically tracks the characteristics of the ship's engine and electromagnetic interference, significantly improving voice clarity and intelligibility. Intelligent buffering and redundancy mechanisms effectively cope with network fluctuations, ensuring continuous voice transmission. Real-time voice quality monitoring triggers adaptive adjustments to the transmission strategy, establishing a closed loop for quality traceability and optimization. This invention systematically solves the problem of communication integration for polar ships through hardware-algorithm collaborative innovation, providing highly reliable communication support for polar scientific research operations. Attached Figure Description

[0065] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0066] Figure 1 A flowchart of a VHF and IP voice fusion communication method for special vessels is provided as an embodiment of the present invention.

[0067] Figure 2 The spectrum diagram of multi-source noise suppression effect provided for embodiments of the present invention.

[0068] Figure 3 A timing diagram for DSC distress call forwarding provided for embodiments of the present invention.

[0069] Figure 4 The diagram illustrates the performance analysis of protocol conversion latency comparison provided for embodiments of the present invention.

[0070] Figure 5 A schematic diagram of a VHF and IP voice fusion communication device for special vessels is provided as an embodiment of the present invention. Detailed Implementation

[0071] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0072] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. This application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0073] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and the structure and / or function of any markings described herein are merely illustrative. Based on this application, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number and aspects set forth herein can be used to implement the apparatus and / or practical methods.

[0074] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this application.

[0075] Additionally, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that practice can be carried out without these marked details.

[0076] The purpose of this invention is to provide a VHF and IP voice fusion communication method and device for special vessels, aiming to solve the problems of high cost, poor environmental adaptability and functional fragmentation of communication systems in polar environments.

[0077] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0078] In this invention, all the formula calculations are presented after parameter normalization and dimension removal.

[0079] Reference Figure 1 This application illustrates a VHF and IP voice fusion communication method for special vessels according to an embodiment of the present application. The method includes the following steps:

[0080] Real-time monitoring of ambient temperature and electromagnetic interference intensity. Specifically, temperature sensors are installed on the mast base atop the bridge to monitor actual temperature changes on the windward side; electromagnetic probes are deployed at the junction of the radar compartment and the communication and navigation compartment to capture the peak EMI levels across the entire ship; and the integrated navigation module is connected to the ship's INS system to share attitude data and improve positioning accuracy.

[0081] Based on the pre-stored antenna parameter compensation library for environmental temperature matching, the RF front-end transmission parameters are dynamically optimized; the transmission parameters include antenna phase center offset, gain compensation value and impedance matching coefficient.

[0082] Preferably, the matching process of the pre-stored antenna parameter compensation library specifically includes:

[0083] The root mean square error between the ambient temperature and the temperature range in the pre-stored parameter library is calculated in real time. If the root mean square error is less than a set threshold, the antenna phase center offset, gain compensation value and impedance matching coefficient of the corresponding temperature range are called.

[0084] When the ambient temperature is below -30℃, the deformation compensation coefficient of the titanium alloy phase change temperature control shell is additionally activated. The deformation compensation coefficient satisfies the following relationship:

[0085]

[0086] in, For phase change temperature control shell deformation, The coefficient of thermal expansion of titanium alloy is 8.6 × 10⁻⁶. -6 K -1 , This is the initial length of the antenna. For reference to ambient temperature, The wind load deformation coefficient is taken as 0.024 mm·m. 2 / N, For wind stress load, The angle between the wind direction and the antenna axis.

[0087] Since this invention is a pre-packaged product, dynamic correction using a thermal deformation compensation function is generally not employed. However, to improve adaptability, antenna transmission parameters can be dynamically corrected using a thermal deformation compensation function based on the ambient temperature. Exemplarily, the thermal deformation compensation function... Specifically:

[0088]

[0089] in, The coefficient of thermal expansion of titanium alloy is 8.6 × 10⁻⁶. -6 K -1 , The activation energy is set to 0.35 eV. Boltzmann's constant, For ambient temperature, The deformation of the phase change temperature control shell.

[0090] For example, antenna phase center offset ,in, For the operating wavelength, a compensation term is injected into the beamforming weight vector. Temperature-dependent gain attenuation model: Add compensation gain to the preamplifier stage of the power amplifier. The π-type matching network is dynamically adjusted based on the Smith chart, where the capacitance value... Inductance value .

[0091] Multimodal noise suppression processing is performed on the received VHF audio signal, including:

[0092] An adaptive comb filter is constructed based on the real-time rotational speed of the ship's main engine to suppress engine noise, and the center frequency is dynamically calculated according to the rotational speed.

[0093] Based on the intensity of the electromagnetic interference, a complex convolutional attention network is used to generate a time-frequency mask to suppress the electromagnetic interference.

[0094] For example, constructing the adaptive comb filter specifically includes:

[0095]

[0096] in, The fundamental frequency and harmonic frequencies of the noise generated by the operation of the ship's main engine. The harmonic order is given by RPM, which is the real-time rotational speed of the ship's main engine. The digital sampling rate of VHF audio signals. This is the notch depth control factor. This is the unit delay operator. Preferably, for diesel engines, k=1-3, corresponding to the fundamental frequency plus the second harmonic, while for gas turbines, k=1-5, because its high-frequency harmonics are significant. Satisfying Nyquist's theorem, the frequency band must be at least twice that of the 156-174 MHz band, preferably 48 kHz. Limiter =clamp(0.75, 0.95).

[0097] The specific method for dynamically adjusting the parameters of the adaptive comb filter is as follows:

[0098]

[0099] in, For input signal-to-noise ratio, This refers to the notch width. Preferably, it is a limiter. = clamp(0.75,0.95).

[0100] A time-frequency mask is generated using the complex convolutional attention network. :

[0101]

[0102] in, For learnable convolutional kernels, For complex number convolution operators, For querying the matrix, Let be the transpose of the key matrix. is the scaling factor. Q and K are obtained by projecting the VHF signal onto the time-frequency matrix generated by the STFT.

[0103] The training loss function of the complex convolutional attention network for:

[0104]

[0105] in, The time-frequency matrix for pure speech. The time-frequency matrix is ​​the output of the network prediction. The regularization coefficient is . The noise distribution estimated for the network. For the noise prior distribution, For frequency variables, For the i-th order noise energy weight, Let i be the center frequency of the i-th order noise. For the i-th order noise bandwidth, Let be the probability density function of a Gaussian distribution. This value is tied to the host machine's rotational speed; The bandwidth ratio remains constant; This corresponds to the energy decay model. Preferably, after cross-validation... The optimal value is 0.3.

[0106] Preferably, the complex convolutional attention network adopts a joint architecture of complex conformer modules and a three-dimensional attention mechanism, specifically including:

[0107] The complex number conformer module employs a real-part dilated convolution path in the time dimension, with a kernel dilation factor of d. t The output features are:

[0108]

[0109] An imaginary part dilated convolution path is used in the frequency dimension, with a kernel dilation factor of d. f The output features are:

[0110]

[0111] Z r Z i These are the real and imaginary matrices of the input complex spectrum, respectively. t and d f Dynamically adjusted according to the intensity of electromagnetic interference;

[0112] The three-dimensional attention mechanism specifically involves performing a three-dimensional dynamic weighting of the features F=Fr+jFi output by the complex conformer module, based on channel, time, and frequency.

[0113]

[0114] in, , , These are channel attention weights, temporal attention weights, and frequency attention weights, respectively.

[0115] Weighted features Input a complex gated cyclic unit and output a time-frequency mask.

[0116] like Figure 2 The image shown is a spectrum of the multi-source noise suppression effect, in which... Figure 2 Table (a) shows a spectrum comparison of the suppression effect, comprising four curves: the original noise (orange), the spectrum after G.711 processing (red), the spectrum after G.729A processing (blue), and the spectrum after G.723.1 processing (cyan). The frequency range indicator for the experiment is 20Hz-16kHz. Figure 2 As shown in (b), three key indicators are used for evaluation: noise suppression, speech distortion, and voice retention.

[0117] To adapt to polar environments, this invention employs proprietary positioning satellites to ensure high-precision positioning. However, to improve adaptability, when ordinary satellite positioning errors exceed a set threshold, ship inertial navigation data can be invoked to compensate for position deviations through motion equations. For example, when satellite positioning errors... At that time, the motion equation compensation position is activated:

[0118]

[0119] in, These are the compensated coordinates of the ship's position at the current moment. These are the coordinates of the position at the previous moment. For the ship's velocity vector, For the ship's acceleration vector, Update the sampling interval for location. Specifically, Provides east / north velocity components for the ship's Doppler log (DVL). This is the three-axis acceleration output from the IMU of the fiber optic gyroscope (FOG). Before compensation, the inertial navigation unit (ENU) coordinates need to be converted to the geodetic coordinate system (BLH).

[0120] When handling timing conflicts, comb filtering (i.e., time domain processing) is performed first, followed by CCAN time-frequency masking (i.e., frequency domain processing).

[0121] The processed signal is converted into IP voice data packets, and the IP voice data packets are then converted into VHF signals through a protocol conversion engine to achieve a two-way communication link.

[0122] Preferably, the specific steps for the protocol conversion engine to implement a bidirectional communication link include a forward link from VHF to IP and a reverse link from IP to VHF.

[0123] The forward link specifically inputs the VHF audio signal after multimodal noise suppression into a dynamic codec. When the end-to-end transmission delay is less than 200 milliseconds, the ITU-T G.711 codec standard is used; when the delay is between 200 and 500 milliseconds, the ITU-T G.729A codec standard is switched. An environmentally aware data tag is encapsulated in the Real-Time Protocol (RTP) header, including a real-time temperature value tag X-Temp: {T} and an electromagnetic interference intensity tag X-EMI: {EMI}, where {T} and {EMI} are the ambient temperature and the electromagnetic interference intensity, respectively.

[0124] The reverse link specifically involves parsing the DSC distress identifier in the IP data packet. If the distress identifier indicates a distress state, the VHF signal reconstruction module is triggered to prioritize processing and is marked as the highest service level of IEEE 802.1p priority 7. The codec matching the sending end, including G.711 or G.729A, is invoked to decode the IP voice data, and electromagnetic interference is suppressed a second time through the complex convolutional attention network. The complex convolutional attention network is activated when the X-EMI tag value exceeds 20dB.

[0125] The system dynamically selects the voice codec. When the end-to-end transmission latency is less than 200 milliseconds, it adopts the ITU-T G.711 codec standard. When the latency is between 200 and 500 milliseconds, it switches to the ITU-T G.729A codec standard. It also switches to a hysteresis range, for example, the latency must be continuously below 180ms before switching back to G.711 to avoid frequent switching that could cause voice interruptions.

[0126] An environmental awareness data tag is encapsulated in the real-time transmission protocol header, including a real-time temperature value tag X-Temp: {T} and an electromagnetic interference intensity tag X-EMI: {EMI}, where {T} and {EMI} are the ambient temperature and the electromagnetic interference intensity, respectively; the IETF draft standard RTP Header Extension is adopted, defining the fields TEMP, int16, in units of 0.1℃ and EMI, uint16, in units of 0.1dBm.

[0127] The synchronization mechanism of the dynamic codec includes: the transmitting end declares the current codec type in the Codec-Type extension field of the RTP header, wherein the codec type includes 0x01 for G.711 and 0x02 for G.729A; the receiving end parses the Codec-Type field and automatically switches to the same codec standard to ensure bidirectional voice codec consistency; when the MOS score calculated by the voice quality closed-loop monitoring module is lower than 3.0, the bidirectional link synchronization is downgraded to the 6.3kbps low bit rate coding mode of the ITU-T G.723.1 coding standard.

[0128] Embedded ship position metadata: GPS coordinates are directly recorded when the satellite positioning error is no more than 100 meters; when position compensation is enabled, the position source is marked as inertial navigation compensation data.

[0129] Based on the DSC distress flag, a Quality of Service (QoS) priority flag is triggered. Distress voice streams are flagged as IEEE 802.1p priority 7, the highest QoS level, while non-distress calls are flagged as priority 5, the standard voice level. The DSCP EF is bound to 802.1p priority 7 in the switch to ensure end-to-end consistency. For example... Figure 3 The DSC distress call forwarding sequence diagram shown below contains... Figure 3 (a) in the diagram is a distress timeline management diagram, which includes six key event nodes: the vessel initiates a DSC distress call (0.2 seconds), the shore station receives and confirms the call (0.8 seconds), the call is forwarded to the search and rescue center (1.5 seconds), the search and rescue center confirms the call (2.3 seconds), the rescue operation is deployed (3.0 seconds), and the rescue begins (4.2 seconds). Figure 3 (b) shows the experimental simulation data.

[0130] The following safeguards are implemented for bidirectional transmission:

[0131] Configure a primary and backup dual-queue buffer system at the switch layer. The primary queue transmits the current voice frame in real time, while the backup queue pre-stores the data for the next voice frame. When network congestion is detected, the queue switching is completed within 10 milliseconds. Add a frame boundary detector to trigger switching only at the boundary of a complete voice frame and pre-store the frame verification sequence.

[0132] Implement dynamic jitter buffer control, with a base buffer depth set at 50 milliseconds, and dynamically adjust the buffer depth according to the network latency change rate. Buffer depth = 50 + 20 × tanh (0.1 × latency change rate); the constraint is: 50ms ≤ buffer depth ≤ 200ms.

[0133] A forward error correction redundancy encapsulation mechanism is enabled. Redundant error correction packets are added for every 10 voice data packets transmitted. The number of redundant packets is dynamically configured according to the real-time packet loss rate: no redundant packets are added when the packet loss rate is less than 1%, 20% of redundant packets are added when the packet loss rate is between 1% and 5%, and 30% of redundant packets are added when the packet loss rate exceeds 5%.

[0134] The IP voice transmission process includes closed-loop voice quality monitoring:

[0135] The receiver periodically calculates the average opinion score (MOS) value: MOS = 4.5 - 0.008 × end-to-end delay - 0.032 × packet loss rate percentage;

[0136] When the MOS score falls below 3.0, a codec downgrade command is sent to the transmitter, automatically switching to the 6.3kbps low bit rate coding mode of the ITU-T G.723.1 standard. For example... Figure 4 As shown, the encoding and decoding performance analysis diagrams for conventional communication, network congestion, and polar environments are presented.

[0137] Establish an encrypted transmission quality log, link it to storage temperature, electromagnetic interference intensity, latency and MOS score data, and retain the log for no less than 3 months as required.

[0138] The bidirectional link communication employs a polar environment-enhanced design, including:

[0139] The network equipment meets the wide operating temperature range of -40℃ to +70℃ specified in IEC 60092-509 standard. The PCB material of the network equipment is made of high Tg FR-4 with a glass transition temperature ≥170℃, and the components have passed the -55℃ low temperature aging test.

[0140] The backbone communication link uses armored optical cable with a tensile strength of not less than 2000 Newtons;

[0141] The switch is equipped with a phase change heat dissipation module filled with paraffin-based composite material with a melting point of -35℃ and a latent heat of phase change of no less than 200 joules / gram. It adopts a dual phase change material, with paraffin as the main material, which has a melting point of -35℃, to adapt to low-temperature phase change; and fatty acids as the auxiliary material, which has a melting point of -50℃, to adapt to extremely low-temperature phase change.

[0142] Key communication ports are equipped with π-type filters to provide insertion loss of no less than 40 dB in the 30-300 MHz band. A common-mode choke is added to provide insertion loss of ≥20 dB in the 1-30 MHz band.

[0143] The equipment cabinet meets the 90 dB shielding effectiveness requirement specified in the MIL-STD-461G standard.

[0144] One embodiment of this application, such as Figure 5 As shown, a VHF and IP voice fusion communication device for special vessels is also provided, the device comprising:

[0145] The main control module uses a multi-core heterogeneous processor to process environmental perception data, control protocol conversion, and schedule resources.

[0146] The VHF signal processing module connects to the ship's VHF radio and has a built-in three-level electromagnetic shielding structure and DSC decoding unit. It is used to receive and preprocess VHF signals and output noise-reduced audio streams.

[0147] The speech enhancement module integrates an adaptive comb filter bank and a complex convolutional attention network hardware accelerator to suppress engine noise and electromagnetic interference and generate speech signals that conform to IP transmission standards.

[0148] The environmental perception unit includes a wide-temperature sensor, a triaxial magnetometer, and a combined navigation module. It collects temperature, electromagnetic interference intensity, and position data in real time, providing input for thermal deformation compensation and position correction.

[0149] The protocol conversion engine is configured to dynamically select codecs, encapsulation environment tags, and location metadata to convert VHF audio streams into IP voice data packets.

[0150] The network transmission module supports dual network port redundancy and Wi-Fi 6 protocol, and implements dual queue buffering, dynamic jitter control and forward error correction mechanism to ensure the reliability of voice transmission within the ship's local area network;

[0151] The polar-reinforced structure includes a phase-change temperature control housing and an armored connector;

[0152] The power management system integrates a lithium titanate backup battery to ensure battery life in extreme environments.

[0153] Preferably, the device of the present invention further includes some functional modules, such as a physical alarm button for distress, a dual-guard priority switching circuit, an MMSI tamper-proof storage unit, an independent power supply path for DSC, and noise / channel management hardware.

[0154] The above are exemplary embodiments disclosed in this invention. However, it should be noted that various changes and modifications can be made without departing from the scope of the embodiments of this invention as defined by the claims. The functions, steps, and / or actions of the methods according to the disclosed embodiments described herein do not need to be performed in any marked order. Furthermore, although the elements disclosed in the embodiments of this invention may be described or claimed individually, they may be understood as multiple unless explicitly limited to a singular.

[0155] In this specification, the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the descriptions of the embodiments described later are relatively simple, and relevant parts can be referred to the descriptions of the foregoing embodiments.

[0156] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A VHF and IP voice fusion communication method for special vessels, characterized in that, The method includes the following steps: Real-time monitoring of ambient temperature and electromagnetic interference intensity; Based on a pre-stored antenna parameter compensation library for environmental temperature matching, the RF front-end transmission parameters are dynamically optimized; these transmission parameters include antenna phase center offset, gain compensation value, and impedance matching coefficient; the matching process of the pre-stored antenna parameter compensation library specifically includes: The root mean square error between the ambient temperature and the temperature range in the pre-stored parameter library is calculated in real time. If the root mean square error is less than a set threshold, the antenna phase center offset, gain compensation value and impedance matching coefficient of the corresponding temperature range are called. When the ambient temperature is below -30℃, the deformation compensation coefficient of the titanium alloy phase change temperature control shell is additionally activated. The deformation compensation coefficient satisfies the following relationship: in, For phase change temperature control shell deformation, The coefficient of thermal expansion of titanium alloy is 8.6 × 10⁻⁶. -6 K -1 , This is the initial length of the antenna. For ambient temperature, For reference to ambient temperature, The wind load deformation coefficient is taken as 0.024 mm·m. 2 / N, For wind stress load, The angle between the wind direction and the antenna axis; Multimodal noise suppression processing is performed on the received VHF audio signal, including: Based on the electromagnetic interference intensity, a complex convolutional attention network is used to generate a time-frequency mask to suppress the electromagnetic interference; the time-frequency mask is generated through the complex convolutional attention network. : in, For learnable convolutional kernels, For complex number convolution operators, For querying the matrix, Let be the transpose of the key matrix. This is the scaling factor; The training loss function of the complex convolutional attention network for: in, The time-frequency matrix for pure speech. The time-frequency matrix is ​​the output of the network prediction. The regularization coefficient is . The noise distribution estimated for the network. For the noise prior distribution, For frequency variables, For the i-th order noise energy weight, Let i be the center frequency of the i-th order noise. For the i-th order noise bandwidth, The probability density function is a Gaussian distribution. The processed signal is converted into IP voice data packets, and the IP voice data packets are then converted into VHF signals through a protocol conversion engine to achieve a two-way communication link.

2. The VHF and IP voice fusion communication method for special vessels according to claim 1, characterized in that, The complex convolutional attention network adopts a joint architecture of complex conformer modules and a three-dimensional attention mechanism, specifically including: The complex number conformer module employs a real-part dilated convolution path in the time dimension, with a kernel dilation factor of d. t The output features are: An imaginary part dilated convolution path is used in the frequency dimension, with a kernel dilation factor of d. f The output features are: Z r Z i These are the real and imaginary matrices of the input complex spectrum, respectively. t and d f Dynamically adjusted according to the intensity of electromagnetic interference; The three-dimensional attention mechanism specifically involves performing a three-dimensional dynamic weighting of the features F=Fr+jFi output by the complex conformer module, based on channel, time, and frequency. in, , , These are channel attention weights, temporal attention weights, and frequency attention weights, respectively. Weighted features Input a complex gated cyclic unit and output a time-frequency mask.

3. The VHF and IP voice fusion communication method for special vessels according to claim 1, characterized in that, The specific steps for the protocol conversion engine to implement a bidirectional communication link include a forward link from VHF to IP and a reverse link from IP to VHF. The forward link specifically inputs the VHF audio signal after multimodal noise suppression into a dynamic codec. When the end-to-end transmission delay is less than 200 milliseconds, the ITU-T G.711 codec standard is used; when the delay is between 200 and 500 milliseconds, the ITU-T G.729A codec standard is switched. An environmentally aware data tag is encapsulated in the Real-Time Protocol (RTP) header, including a real-time temperature value tag X-Temp: {T} and an electromagnetic interference intensity tag X-EMI: {EMI}, where {T} and {EMI} are the ambient temperature and the electromagnetic interference intensity, respectively. The reverse link specifically involves parsing the DSC distress identifier in the IP data packet. If the distress identifier indicates a distress state, the VHF signal reconstruction module is triggered to prioritize processing and is marked as the highest service level of IEEE 802.1p priority 7. The codec matching the sending end, including G.711 or G.729A, is invoked to decode the IP voice data, and electromagnetic interference is suppressed a second time through the complex convolutional attention network. The complex convolutional attention network is activated when the X-EMI tag value exceeds 20dB.

4. The VHF and IP voice fusion communication method for special vessels according to claim 3, characterized in that, The synchronization mechanism of the dynamic codec includes: The sending end declares the current codec type in the Codec-Type extension field of the RTP header. The codec type includes 0x01 for G.711 and 0x02 for G.729A. The receiving end parses the Codec-Type field and automatically switches to the same codec standard to ensure bidirectional voice codec consistency. When the MOS score calculated by the voice quality closed-loop monitoring module is lower than 3.0, the bidirectional link synchronously degrades to the 6.3kbps low bit rate coding mode of the ITU-TG.723.1 coding standard.

5. A VHF and IP voice fusion communication method for special vessels according to claim 4, characterized in that, The specific mechanism for ensuring bidirectional transmission is as follows: A primary and backup dual-queue buffer system is configured at the switch layer. The primary queue transmits the current voice frame in real time, while the backup queue stores the data for the next voice frame. When network congestion is detected, the queue switching is completed within 10 milliseconds. Implement dynamic jitter buffer control, with a base buffer depth set at 50 milliseconds, and dynamically adjust the buffer depth according to the network latency change rate. Buffer depth = 50 + 20 × tanh (0.1 × latency change rate). A forward error correction redundancy encapsulation mechanism is enabled. Redundant error correction packets are added for every 10 voice data packets transmitted. The number of redundant packets is dynamically configured according to the real-time packet loss rate: no redundant packets are added when the packet loss rate is less than 1%, 20% of redundant packets are added when the packet loss rate is between 1% and 5%, and 30% of redundant packets are added when the packet loss rate exceeds 5%.

6. A VHF and IP voice fusion communication method for special vessels according to claim 5, characterized in that, The IP voice transmission process includes closed-loop voice quality monitoring: The receiver periodically calculates the average opinion score (MOS) value: MOS = 4.5 - 0.008 × end-to-end delay - 0.032 × packet loss rate percentage; When the MOS score is below 3.0, a codec downgrade command is sent to the transmitter to automatically switch to the 6.3kbps low bit rate coding mode of the ITU-T G.723.1 standard; Establish an encrypted transmission quality log, link it to storage temperature, electromagnetic interference intensity, latency and MOS score data, and retain the log for no less than 3 months as required.

7. A VHF and IP voice fusion communication method for special vessels according to claim 6, characterized in that, The bidirectional communication link adopts a polar environment-enhanced design, including: The network equipment meets the wide operating temperature range of -40℃ to +70℃ specified in IEC 60092-509 standard; The backbone communication link uses armored optical cable with a tensile strength of not less than 2000N; The switch is equipped with a phase change heat dissipation module, filled with paraffin-based composite material with a melting point of -35℃, and the latent heat of phase change is not less than 200J / g; Key communication ports are equipped with π-type filters to provide an insertion loss of no less than 40 dB in the 30-300MHz frequency band; The equipment cabinet meets the 90 dB shielding effectiveness requirement specified in the MIL-STD-461G standard.

8. A VHF and IP voice fusion communication device for special vessels, used to implement the VHF and IP voice fusion communication method for special vessels as described in any one of claims 1-7, characterized in that, The device includes: The main control module uses a multi-core heterogeneous processor to process environmental perception data, control protocol conversion, and schedule resources. The VHF signal processing module connects to the ship's VHF radio and has a built-in three-level electromagnetic shielding structure and DSC decoding unit. It is used to receive and preprocess VHF signals and output noise-reduced audio streams. The speech enhancement module integrates an adaptive comb filter bank and a complex convolutional attention network hardware accelerator to suppress engine noise and electromagnetic interference and generate speech signals that conform to IP transmission standards. The environmental perception unit includes a wide-temperature sensor, a triaxial magnetometer, and a combined navigation module. It collects temperature, electromagnetic interference intensity, and position data in real time, providing input for thermal deformation compensation and position correction. The protocol conversion engine is configured to dynamically select codecs, encapsulation environment tags, and location metadata to convert VHF audio streams into IP voice data packets. The network transmission module supports dual network port redundancy and Wi-Fi 6 protocol, and implements dual queue buffering, dynamic jitter control and forward error correction mechanism to ensure the reliability of voice transmission within the ship's local area network; The polar-reinforced structure includes a phase-change temperature control housing and an armored connector; The power management system integrates a lithium titanate backup battery to ensure battery life in extreme environments.

Citation Information

Patent Citations

  • Method and device for identifying radar target under non-uniform interference based on attention network

    CN121348252A

  • Baseband architecture for GNSS jamming mitigation

    US20250164646A1