Communication method and device of marine very high frequency equipment, electronic equipment and storage medium

By performing signal and data processing on the voice signals of marine VHF equipment, effective text data is extracted from the semantic text data, and effective text signals are generated and transmitted to the voice receiver, thus solving the problem of poor communication quality and achieving stable and efficient voice communication.

CN119559943BActive Publication Date: 2025-11-25SHANGHAI MERCHANT SHIP DESIGN & RES INST
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411680519.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-11-25
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

The existing shipboard VHF equipment has poor communication quality, mainly due to factors such as crew accents, environmental noise, low VHF equipment volume, and congested communication channels.

Method used

By performing signal processing on the original speech signal, semantic text data is determined, data processing is performed to extract effective text data, effective text signals are generated and transmitted to the speech receiver, thus realizing the target communication speech of the speech receiver.

Benefits of technology

It improves the understandability and effectiveness of voice receivers, avoids channel congestion when communication channel resources are limited, ensures stable transmission of voice signals, and enhances communication quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119559943B_ABST
    Figure CN119559943B_ABST
Patent Text Reader

Abstract

The application discloses a communication method and device of a marine very high frequency equipment, electronic equipment and a storage medium. The method comprises the following steps: determining an original voice signal, performing signal processing on the original voice signal, and determining semantic text data; performing data processing on the semantic text data, and determining effective text data in the semantic text data; generating an effective text signal corresponding to the effective text data, and transmitting the effective text signal to a voice receiving end through a target communication channel; and enabling the voice receiving end to determine target communication voice based on the effective text signal. According to the technical scheme, the communication quality of the marine very high frequency equipment can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer application, and in particular to a communication method and device of a marine very high frequency (VHF) equipment, an electronic device and a storage medium. BACKGROUND

[0002] The marine VHF equipment is an important communication channel between ships and ships, and between ships and land stations in navigation. At present, the VHF communication has the problem of poor communication quality caused by objective factors such as the accent of the crew, environmental noise, too low volume of the VHF equipment, and communication channel congestion.

[0003] In the related art, the transmission voice signal is usually converted into text, and the text data is further transmitted to realize end-to-end communication. That is, the text readability is used to improve the intelligibility of the communication information, and the low bandwidth characteristic of the text transmission is used to reduce the communication channel congestion to improve the communication quality, but the improvement effect of the communication quality is not obvious, and it is difficult to meet the high-quality communication demand of the user. SUMMARY

[0004] The present application provides a communication method and device of a marine VHF equipment, an electronic device and a storage medium to solve the technical problem of poor communication quality of the marine VHF equipment.

[0005] According to an aspect of the present application, a communication method of a marine VHF equipment is provided, which comprises:

[0006] determining an original voice signal, signal processing the original voice signal, and determining semantic text data;

[0007] data processing the semantic text data to determine effective text data in the semantic text data;

[0008] generating an effective text signal corresponding to the effective text data, and transmitting the effective text signal to a voice receiving end through a target communication channel;

[0009] so that the voice receiving end determines a target communication voice based on the effective text signal.

[0010] According to another aspect of the present application, a communication device of a marine VHF equipment is provided, which comprises:

[0011] a signal processing module for determining an original voice signal, signal processing the original voice signal, and determining semantic text data;

[0012] a data processing module for data processing the semantic text data to determine effective text data in the semantic text data;

[0013] a signal transmission module configured to generate an effective text signal corresponding to the effective text data, and transmit the effective text signal to a voice receiving end through a target communication channel;

[0014] a voice receiving module configured to determine a target communication voice based on the effective text signal at the voice receiving end.

[0015] According to another aspect of the present application, there is provided an electronic device, comprising:

[0016] at least one processor; and

[0017] a memory in communication with the at least one processor; wherein

[0018] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the communication method of the marine VHF device according to any one of the embodiments of the present application.

[0019] According to another aspect of the present application, there is provided a computer readable storage medium storing computer instructions for enabling a processor to perform the communication method of the marine VHF device according to any one of the embodiments of the present application when executed by the processor.

[0020] The technical solution of the embodiments of the present application determines an original voice signal, performs signal processing on the original voice signal to determine semantic text data, performs data processing on the semantic text data to determine effective text data in the semantic text data, generates an effective text signal corresponding to the effective text data, and transmits the effective text signal to a voice receiving end through a target communication channel, so that the voice receiving end determines a target communication voice based on the effective text signal. The present application extracts effective text data from semantic text data for voice communication, filters invalid data in text data corresponding to voice signals, improves the intelligibility and effectiveness of the received voice at the voice receiving end, and ensures the stability of voice signal transmission, thereby effectively improving the communication quality of the marine VHF device.

[0021] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiments description. Obviously, the drawings in the following description only show some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without any creative effort.

[0023] Figure 1 is a system architecture diagram of a marine VHF device according to an embodiment of the present application;

[0024] Figure 2 is a flow chart of a communication method of a marine VHF device according to an embodiment of the present application;

[0025] Figure 3 is a flow chart of a communication method of a marine VHF device according to an embodiment of the present application;

[0026] Figure 4 is a whole flow chart of data processing according to an embodiment of the present application;

[0027] Figure 5 is a whole flow chart of a communication method of a marine VHF device according to an embodiment of the present application;

[0028] Figure 6 is a structural schematic diagram of a communication device of a marine VHF device according to an embodiment of the present application;

[0029] Figure 7 is a structural schematic diagram of an electronic device for implementing a communication method of a marine VHF device according to an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to make the technical personnel in the art better understand the present application scheme, the following will combine the drawings in the embodiments of the present application, and clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort should belong to the scope of protection of the present application.

[0031] It should be noted that the terms "first", "second", and the like in the description and claims of the application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in other than the order illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0032] Before the embodiments of the application are described, the communication end and the marine VHF equipment involved in the communication method of the marine VHF equipment are described: the communication end can include a voice transmitting end and a voice receiving end, wherein the voice transmitting end can be a ship end or a land end, and the voice receiving end can be a ship end or a land end. The communication end can be installed with a marine VHF equipment, such as Figure 1 , Figure 1 is a system architecture diagram of a marine VHF equipment provided according to an embodiment of the application. The marine VHF equipment can include a transmitter for transmitting voice signals, a receiver for receiving voice signals, a display, a speaker, a microphone, a control unit, an intelligent analysis unit, and a panel unit, etc. The following embodiments are specifically described with the voice transmitting end as the execution subject.

[0033] Embodiment one

[0034] Figure 2 A flowchart of a communication method of a marine VHF equipment is provided for the first embodiment of the application. The embodiment can be applicable to the case of using a voice processing mode for marine VHF equipment communication. The method can be executed by a communication device of the marine VHF equipment, which can be realized in the form of hardware and / or software, and can be configured in a computer. As Figure 2 shown, the method includes:

[0035] S110, determining an original voice signal, performing signal processing on the original voice signal, and determining semantic text data.

[0036] The original voice signal can be understood as an unprocessed voice signal. In the embodiments of the application, the original voice signal can include environmental noise or invalid voice, etc. The invalid voice can be redundant voice or non-standard term voice, etc.

[0037] The semantic text data can be understood as text data related to the original voice signal. Optionally, the semantic text data can represent the specific semantics of the original voice signal.

[0038] S120, data processing is performed on the semantic text data to determine valid text data in the semantic text data.

[0039] The valid text data can be understood as valid text data in the semantic text data. Optionally, the valid text data can be determined by language model intent recognition and key data extraction on the semantic text data.

[0040] Optionally, the data processing on the semantic text data to determine the valid text data in the semantic text data comprises:

[0041] An intent recognition is performed on the semantic text data by a natural language model to determine at least one target semantic intent;

[0042] For each target semantic intent, subtext data corresponding to the target semantic intent is determined, and data extraction is performed on the subtext data by a language representation model to obtain key text data;

[0043] The valid text data in the semantic text data is determined according to the target semantic intent and the key text data.

[0044] The natural language model can be understood as a machine learning model with semantic intent recognition function. In the embodiments of the present application, the natural language model can be preset according to the scene requirement, which is not limited here.

[0045] The target semantic intent can be understood as the semantic intent represented by the semantic text data. Optionally, the target semantic intent can include navigation safety, port operation, emergency rescue, navigation coordination, information exchange, ship management or training and the like.

[0046] The subtext data can be understood as a text data segment in the semantic text data related to the target semantic intent. In the embodiments of the present application, the subtext data corresponding to different target semantic intents can be different.

[0047] The language representation model can be understood as a machine learning model with key data extraction function. In the embodiments of the present application, the language representation model can be preset according to the scene requirement, which is not limited here. Optionally, the language representation model can be a Bert model.

[0048] The key text data can be understood as key text data in the subtext data. Optionally, the key text data can include at least one of a receiver identifier, position-related data, time-related data, and event description data.

[0049] Optionally, the determining of the valid text data in the semantic text data according to the target semantic intention and the key text data comprises:

[0050] For each of the subtext data, in the case that the key text data is not empty and the key text data matches the target semantic intention, the target semantic intention and the key text data are encoded to obtain valid sub-data.

[0051] The valid text data in the semantic text data is determined according to at least one of the valid sub-data.

[0052] For example, for the target subtext data, the target semantic intention is navigation safety, and the event description data of the key text data is training description data. In this case, the key text data does not match the target semantic intention.

[0053] In the embodiment of the present application, the specific manner of encoding the target semantic intention and the key text data is not limited. Optionally, the encoding operation can be used to improve the data transmission security.

[0054] The valid sub-data can be understood as an effective data segment that can be communicated and transmitted.

[0055] S130, generating an effective text signal corresponding to the valid text data, and transmitting the effective text signal to a voice receiving end through a target communication channel.

[0056] The effective text signal can be understood as a communication signal related to the valid text data. The target communication channel can be a communication channel related to a marine very high frequency device.

[0057] The voice receiving end can be determined based on the effective text signal, that is, the effective text signal includes relevant communication information of the voice receiving end.

[0058] S140, so that the voice receiving end determines a target communication voice based on the effective text signal.

[0059] The target communication voice can be related to the valid text data.

[0060] The technical scheme of the embodiment of the present application determines an original voice signal, performs signal processing on the original voice signal to determine semantic text data, performs data processing on the semantic text data to determine effective text data in the semantic text data, generates effective text signals corresponding to the effective text data, and transmits the effective text signals to a voice receiving end through a target communication channel, so that the voice receiving end determines target communication voice based on the effective text signals. The present application extracts effective text data in semantic text data for voice communication, filters invalid data in text data corresponding to voice signals, improves the intelligibility and effectiveness of the received voice of the voice receiving end, and ensures the stability of voice signal transmission and effectively improves the communication quality of the marine VHF equipment.

[0061] Embodiment two

[0062] Figure 3 A flowchart of a communication method of a marine VHF equipment according to the second embodiment of the present application is provided, and the embodiment is a refinement of the signal processing on the original voice signal to determine semantic text data in the above-mentioned embodiment. As shown in the figure, Figure 3 the method comprises:

[0063] S210, determining an original voice signal.

[0064] S220, performing frequency spectrum denoising on the original voice signal through a Fourier transform algorithm to obtain a first voice signal after denoising.

[0065] The Fourier transform algorithm can be a short-time Fourier transform (STFT) algorithm. The short-time Fourier transform algorithm can divide the voice signal into short time segments.

[0066] The first voice signal can be understood as a voice signal after denoising.

[0067] Optionally, the method of performing frequency spectrum denoising on the original voice signal through the Fourier transform algorithm to obtain the first voice signal after denoising comprises:

[0068] performing slicing processing on the original voice signal through a preset signal slicing parameter to obtain a plurality of first signal segments, wherein the signal slicing parameter comprises a frame length and / or a frame shift;

[0069] determining at least one noise signal segment from the plurality of first signal segments according to a noise threshold, wherein the noise threshold is determined according to a root mean square energy value of the first signal segments;

[0070] performing short-time Fourier transform on the at least one noise signal segment to determine a noise signal spectrum, and performing short-time Fourier transform on the original speech signal to determine a speech signal spectrum;

[0071] determining a de-noised signal spectrum according to the speech signal spectrum and the noise signal spectrum, and performing inverse short-time Fourier transform on the de-noised signal spectrum to obtain a de-noised first speech signal.

[0072] The signal segment parameters can be understood as parameters for segmenting the speech signal. For example, the frame length can be 40 ms, and the frame shift can be 20 ms.

[0073] Specifically, the step of performing spectral de-noising on the original speech signal by the Fourier transform algorithm to obtain the de-noised first speech signal is calculated as follows:

[0074] 1. Segmenting the original speech signal by selecting signal segment parameters of a frame length of 40 ms and a frame shift of 20 ms to obtain a plurality of first signal segments with a length N of 40 ms.

[0075] 2. Determining a root mean square energy value of each first signal segment x(t, N) according to the following formula:

[0076]

[0077] wherein RMS(t, N) represents the root mean square energy value, N represents the frame length, and t represents the frame shift.

[0078] 3. Selecting the root mean square energy value at the 10th percentile as the noise threshold, and selecting the first signal segment with a root mean square energy value lower than the noise threshold as the noise signal segment.

[0079] 4. Performing short-time Fourier transform on the noise signal segment by STFT to determine a noise signal spectrum, and performing short-time Fourier transform on the original speech signal to determine a speech signal spectrum, according to the following formula:

[0080]

[0081] wherein x(t) represents an input signal (a signal segment or an original speech signal), ω(t) represents a preset Hamming window function, τ represents a time offset, ω represents an angular frequency, and X(τ, ω) represents an output spectrum (a noise signal spectrum or a speech signal spectrum).

[0082] 5. Calculate the average value N(ω) of the noise signal spectrum, determine the de-noised signal spectrum based on the average value N(ω), and realize the formula (spectral subtraction formula) as follows:

[0083] Y(τ,ω) = max(|X(τ,ω)| - aN(ω), 0)

[0084] wherein Y(τ,ω) represents the spectral subtraction output (de-noised signal spectrum), a represents the spectral subtraction coefficient, N(ω) represents the spectral subtraction input (average value of the noise signal spectrum), and X(τ,ω) represents the spectral subtraction input (speech signal spectrum). The value of a can be 1.

[0085] 6. Perform inverse transformation on the de-noised signal spectrum by an Inverse Short-Time Fourier Transform (ISTFT) algorithm to obtain the first speech signal after de-noising, and realize the formula as follows:

[0086]

[0087] wherein x(t) represents the output signal (first speech signal), X(τ,ω) represents the input signal (de-noised signal spectrum), ω(t) represents a preset Hamming window function, τ represents time offset, and ω represents angular frequency.

[0088] S230. Adjust the volume of the first speech signal by a first gain parameter to obtain a second speech signal after volume adjustment.

[0089] wherein the first gain parameter is determined according to the global root mean square energy value corresponding to the first speech signal.

[0090] Optionally, the volume adjustment of the first speech signal by the first gain parameter to obtain the second speech signal after volume adjustment comprises:

[0091] performing slicing processing on the first speech signal according to a preset signal slicing parameter to obtain a plurality of second signal segments;

[0092] for each second signal segment, performing volume analysis on the second signal segment according to an enhancement threshold, and in the case that the second signal segment is a low-volume segment, performing volume enhancement on the second signal segment according to a first gain parameter to obtain an enhanced signal segment;

[0093] determine the second speech signal after volume adjustment based on the enhanced signal segment.

[0094] Optionally, the determination of the second speech signal after volume adjustment based on the enhanced signal segment comprises:

[0095] In a case where the root mean square energy values of the plurality of second signal segments meet a volume compression condition, performing volume compression on at least one high-volume segment of the plurality of second signal segments whose root mean square energy value exceeds a compression threshold according to a second gain parameter to obtain a compressed signal segment, wherein the second gain parameter is determined according to a maximum value of the root mean square energy values of the plurality of second signal segments;

[0096] Determining a signal to be compensated according to the compressed signal segment;

[0097] Determining a target gain parameter according to the first gain parameter and the second gain parameter, and performing gain compensation on the signal to be compensated according to the target gain parameter to obtain the second speech signal after volume adjustment.

[0098] Specifically, the following formula is used to calculate the volume adjustment of the first speech signal by the first gain parameter to obtain the second speech signal after volume adjustment:

[0099] 1. Selecting a signal segment parameter of a frame length of 40 ms and a frame shift of 20 ms to segment the first speech signal to obtain a plurality of second signal segments.

[0100] 2. Determining a root mean square energy value RMS(t, N) of each second signal segment, setting a fixed energy value threshold, i.e., an enhancement threshold T1, and screening the second signal segment with RMS(t, N) < T1 as a low-volume segment.

[0101] 3. For each low-volume segment, calculating a global root mean square energy value RMS g of the first speech signal, determining an intermediate value RMS t according to RMS g , and determining a first gain parameter according to RMS t , which is realized by the following formula:

[0102]

[0103] wherein RMS t represents the intermediate value, and RMS g represents the global root mean square energy value.

[0104] θ = RMS t / RMS(t, N)

[0105] wherein θ represents the first gain parameter, RMS t represents the intermediate value, and RMS(t, N) represents the root mean square energy value of the second signal segment.

[0106] 4. Apply the first gain factor to the corresponding low-volume segment to obtain an enhanced signal segment of the low-volume segment.

[0107] 5. Determine the maximum and minimum values of RMS(t,N), i.e. RMS max and RMS min , if RMS max -RMS min > ΔRMS, the volume compression condition is met, i.e. the volume compression operation is needed, wherein ΔRMS is a preset threshold.

[0108] 6. Calculate the mean μ and standard deviation σ of each segment RMS(t,N), determine the dynamic energy threshold, i.e. the compression threshold T2, screen the second signal segment with RMS(t,N)>T2 as a high-volume segment, and realize the following formula:

[0109] T2 = μ + kσ

[0110] wherein T2 represents the compression threshold, μ represents the mean of RMS(t,N), σ represents the standard deviation of RMS(t,N), and k represents a preset constant value.

[0111] Alternatively, take the 75th percentile of the root mean square energy values of the second signal segments in the sequence as the compression threshold RMS thresh , i.e. when the amplitude of the second signal segment exceeds the compression threshold, the volume compression is performed, and the formula is as follows:

[0112]

[0113] wherein RMS thresh represents the compression threshold, RMS(t,N) represents the root mean square energy value, cr represents the second gain parameter, and g(t,N) represents the compressed signal segment.

[0114] 7. Set the basic compression ratio r b and the maximum compression ratio r m according to the scene requirement, calculate the target compression ratio cr according to r b and r m , take cr as the second gain parameter, and realize the following formula:

[0115]

[0116] wherein cr represents the target compression ratio, r b represents the basic compression ratio, r m represents the maximum compression ratio, RMS max represents the maximum value of the root mean square energy value of the second signal segment, and RMS min represents the minimum value of the root mean square energy value of the second signal segment.

[0117] 8. determining a target gain parameter, according to which the gain compensation is performed on the signal to be compensated, and realizing the formula as follows:

[0118]

[0119] Wherein g represents the target gain parameter, x(t) represents the output signal (second speech signal), x(t) represents the signal to be compensated. Wherein g can be the average of the first gain parameter and the second gain parameter.

[0120] S240, performing semantic recognition on the first speech signal through a speech recognition model, and determining the semantic text data corresponding to the second speech signal.

[0121] Wherein the speech recognition model (Automatic Speech Recognition, ASR) can have the function of recognizing the semantic text corresponding to the speech signal.

[0122] Specifically, between the application of the speech recognition model for semantic recognition, model quantization is performed, and the original model is quantized to an 8bit model, realizing the reduction of the consumption of hardware resources by the end side deployment; the model is exported to an open file format (Open Neural Network Exchange, ONNX) format, and the model is loaded in the embedded computing function module.

[0123] S250, performing data processing on the semantic text data, and determining the effective text data in the semantic text data.

[0124] Figure 4 According to an embodiment of the present application, a kind of data processing overall flow chart is provided. As shown in Figure 4 The overall flow of data processing can be as follows:

[0125] Specifically, 1, the semantic text data output by ASR is subjected to intent recognition. The target semantic intent identified can be as shown in Table 1 as follows:

[0126] Table 1

[0127]

[0128] 2, the key data of semantic text data is extracted by Bert model. The key text data can be extracted as shown in Table 2 as follows:

[0129] Table 2

[0130]

[0131] 3. If the extracted key text data is empty, or the key text data does not match the target semantic intent, the identified relevant text data (sub-text data) is considered invalid data and can be directly filtered out to obtain valid text data.

[0132] S260. Generate a valid text signal corresponding to the valid text data, and transmit the valid text signal to the voice receiver through the target communication channel.

[0133] S270, so that the voice receiver determines the target communication voice based on the valid text signal.

[0134] Specifically, for the voice receiver, the received valid text signal is decoded, and the valid text signal is converted into speech offline using a text-to-speech (TTS) model to obtain the target communication speech.

[0135] The technical solution of this invention involves performing spectral denoising on the original speech signal using a Fourier transform algorithm to obtain a denoised first speech signal; adjusting the volume of the first speech signal using a first gain parameter to obtain a volume-adjusted second speech signal, wherein the first gain parameter is determined based on the global root mean square energy value corresponding to the first speech signal; and performing semantic recognition on the first speech signal using a speech recognition model to determine the semantic text data corresponding to the second speech signal. This invention achieves high-quality speech signal denoising and highly accurate semantic recognition.

[0136] Optionally, Figure 5 This is an overall flowchart of a communication method for a marine VHF device according to an embodiment of the present invention. Figure 5 As shown, the overall process of the communication method of the marine VHF equipment may include: 1. Spectral denoising implemented by a denoising module (determining the first speech signal corresponding to the original speech signal); 2. Volume adjustment implemented by a speech enhancement module (determining the second speech signal corresponding to the first speech signal); 3. Semantic recognition implemented by a speech recognition module (determining the semantic text data corresponding to the second speech signal); 4. Effective text extraction implemented by a speech understanding module (determining the effective text data corresponding to the semantic text data); 5. Signal-to-speech conversion implemented by a speech generation module (determining the target communication speech corresponding to the effective text signal).

[0137] Based on the technical scheme of the embodiment, 1. The communication quality and efficiency of the VHF device are improved. The de-noising module and the speech enhancement module can ensure that the communication data is clearer and purer. Wireless transmission can improve the data transmission speed. The speech generation module can ensure the standardization and intelligibility of the voice output. 2. Shielding invalid information. The semantic understanding module can effectively filter out invalid information and release the communication pressure of the VHF channel. 3. The convenience is realized. The intelligent analysis module is added in the traditional VHF communication device, without changing the operation of the user. The high-quality VHF communication effect and efficiency can improve the user experience. 4. The safety and reliability are improved. The communication quality of the VHF is improved, thereby ensuring the safety of the ship.

[0138] Embodiment three

[0139] Figure 6 A structure diagram of a communication device of a marine very high frequency device is provided for the third embodiment of the application. As shown in the figure, Figure 6 The device comprises a signal processing module 310, a data processing module 320, a signal transmission module 330 and a voice receiving module 340.

[0140] The signal processing module 310 is configured to determine an original voice signal, perform signal processing on the original voice signal, and determine semantic text data. The data processing module 320 is configured to perform data processing on the semantic text data and determine valid text data in the semantic text data. The signal transmission module 330 is configured to generate an effective text signal corresponding to the valid text data, and transmit the effective text signal to a voice receiving end through a target communication channel. The voice receiving module 340 is configured to enable the voice receiving end to determine a target communication voice based on the effective text signal.

[0141] The technical scheme of the embodiment of the application determines an original voice signal, performs signal processing on the original voice signal, determines semantic text data, performs data processing on the semantic text data, determines valid text data in the semantic text data, generates an effective text signal corresponding to the valid text data, and transmits the effective text signal to a voice receiving end through a target communication channel. The voice receiving end determines a target communication voice based on the effective text signal. The application extracts valid text data from semantic text data for voice communication, filters invalid data in the text data corresponding to the voice signal, improves the intelligibility and effectiveness of the voice received by the voice receiving end, and effectively improves the communication quality of the marine very high frequency device.

[0142] Optionally, the signal processing module 310 comprises a spectrum denoising unit, a volume adjusting unit and a semantic recognition unit.

[0143] The spectrum denoising unit is configured to perform spectrum denoising on the original voice signal by a Fourier transform algorithm to obtain a first voice signal after denoising.

[0144] The volume adjusting unit is configured to perform volume adjustment on the first voice signal by a first gain parameter to obtain a second voice signal after volume adjustment, wherein the first gain parameter is determined according to a global root mean square energy value corresponding to the first voice signal.

[0145] The semantic recognition unit is configured to perform semantic recognition on the first voice signal by a voice recognition model to determine the semantic text data corresponding to the second voice signal.

[0146] Optionally, the spectrum denoising unit is specifically configured to:

[0147] perform slicing processing on the original voice signal by a preset signal slicing parameter to obtain a plurality of first signal segments, wherein the signal slicing parameter comprises a frame length and / or a frame shift;

[0148] determine at least one noise signal segment in the plurality of first signal segments according to a noise threshold, wherein the noise threshold is determined according to a root mean square energy value of the first signal segment;

[0149] perform short-time Fourier transform on at least one noise signal segment to determine a noise signal spectrum, and perform short-time Fourier transform on the original voice signal to obtain a voice signal spectrum;

[0150] determine a denoising signal spectrum according to the voice signal spectrum and the noise signal spectrum, and perform inverse short-time Fourier transform on the denoising signal spectrum to obtain the first voice signal after denoising.

[0151] Optionally, the volume adjusting unit comprises a slicing processing subunit, a volume analysis subunit and a signal determination subunit.

[0152] The slicing processing subunit is configured to perform slicing processing on the first voice signal according to a preset signal slicing parameter to obtain a plurality of second signal segments.

[0153] The volume analysis subunit is configured to, for each second signal segment, perform volume analysis on the second signal segment according to an enhancement threshold, and in a case where the second signal segment is a low-volume segment, perform volume enhancement on the second signal segment according to a first gain parameter to obtain an enhanced signal segment.

[0154] The signal determination subunit is configured to determine the second speech signal after volume adjustment based on the enhanced signal segment.

[0155] Optionally, the signal determination subunit is specifically configured to:

[0156] In a case where the root mean square energy values of the plurality of second signal segments meet a volume compression condition, performing volume compression on at least one high-volume segment with a root mean square energy value exceeding a compression threshold in the plurality of second signal segments according to a second gain parameter to obtain a compressed signal segment, wherein the second gain parameter is determined according to a maximum value in the root mean square energy values of the plurality of second signal segments;

[0157] determining a signal to be compensated according to the compressed signal segment;

[0158] determining a target gain parameter according to the first gain parameter and the second gain parameter, and performing gain compensation on the signal to be compensated according to the target gain parameter to obtain the second speech signal after volume adjustment.

[0159] Optionally, the signal transmission module 330 includes an intention recognition unit, a data extraction unit, and an effective text determination unit.

[0160] The intention recognition unit is configured to perform intention recognition on the semantic text data through a natural language model to determine at least one target semantic intention.

[0161] The data extraction unit is configured to determine, for each target semantic intention, subtext data corresponding to the target semantic intention, and perform data extraction on the subtext data through a language representation model to obtain key text data.

[0162] The effective text determination unit is configured to determine effective text data in the semantic text data according to the target semantic intention and the key text data.

[0163] Optionally, the effective text determination unit is specifically configured to:

[0164] For each subtext data, in a case where the key text data is not empty and the key text data matches the target semantic intention, encoding the target semantic intention and the key text data to obtain effective sub-data.

[0165] determining effective text data in the semantic text data according to at least one effective sub-data.

[0166] The communication device of the marine VHF equipment provided by the embodiment of the present application can execute the communication method of the marine VHF equipment provided by any embodiment of the present application, has the function module and the beneficial effect corresponding to the execution method.

[0167] Embodiment four

[0168] Figure 7 A structural schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.

[0169] As shown in Figure 7 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11, wherein the memory stores a computer program that can be executed by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0170] Various components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.

[0171] The processor 11 can be various general and / or special purpose processing components having processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the communication method of a marine VHF device.

[0172] In some embodiments, the communication method of a marine VHF device can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the communication method of a marine VHF device described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the communication method of a marine VHF device by any other appropriate means, such as by means of firmware.

[0173] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0174] Computer programs used to implement the methods of the application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, and partially on a machine or a remote machine or a server.

[0175] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0176] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0177] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0178] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.

[0179] It should be understood that the various forms of flow shown above can be used to reorder, add or delete steps. For example, each step described in the present application can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, which is not limited herein.

[0180] The above detailed description does not constitute a limitation on the scope of protection of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

Claims

1. A communication method for marine VHF equipment, characterized in that, include: The original speech signal is determined, and the original speech signal is denoised by Fourier transform algorithm to obtain the first denoised speech signal; The first speech signal is segmented according to preset signal segmentation parameters to obtain multiple second signal segments; For each of the second signal segments, a volume analysis is performed on the second signal segment according to an enhancement threshold. If the second signal segment is a low-volume segment, the volume of the second signal segment is enhanced according to a first gain parameter to obtain an enhanced signal segment. The first gain parameter is determined based on the global root mean square energy value corresponding to the first speech signal. When the root mean square energy values ​​of multiple second signal segments meet the volume compression condition, at least one high-volume segment whose root mean square energy value exceeds the compression threshold is compressed according to the second gain parameter to obtain a compressed signal segment. The second gain parameter is determined based on the maximum or minimum value among the root mean square energy values ​​of multiple second signal segments. The signal to be compensated is determined based on the compressed signal segment; A target gain parameter is determined based on the first gain parameter and the second gain parameter. Gain compensation is then performed on the signal to be compensated based on the target gain parameter to obtain a second voice signal with adjusted volume. The second speech signal is semantically recognized by a speech recognition model to determine the semantic text data corresponding to the second speech signal. The semantic text data is processed to determine the valid text data within it. Generate a valid text signal corresponding to the valid text data, and transmit the valid text signal to the voice receiver through the target communication channel; So that the voice receiver can determine the target communication voice based on the valid text signal.

2. The method according to claim 1, characterized in that, The step of performing spectral denoising on the original speech signal using a Fourier transform algorithm to obtain a denoised first speech signal includes: The original speech signal is segmented using preset signal segmentation parameters to obtain multiple first signal segments, wherein the signal segmentation parameters include frame length and / or frame shift. At least one noise signal segment among a plurality of the first signal segments is determined based on a noise threshold, wherein the noise threshold is determined based on the root mean square energy value of the first signal segment; Perform a short-time Fourier transform on at least one of the noise signal segments to determine the noise signal spectrum; and perform a short-time Fourier transform on the original speech signal to obtain the speech signal spectrum; The denoised signal spectrum is determined based on the speech signal spectrum and the noise signal spectrum, and a short-time inverse Fourier transform is performed on the denoised signal spectrum to obtain the first speech signal after denoising.

3. The method according to claim 1, characterized in that, The step of processing the semantic text data to determine the valid text data within the semantic text data includes: The semantic text data is subjected to intent recognition using a natural language model to determine at least one target semantic intent. For each target semantic intent, sub-text data corresponding to the target semantic intent is determined, and data extraction is performed on the sub-text data through a language representation model to obtain key text data; The effective text data in the semantic text data is determined based on the target semantic intent and the key text data.

4. The method according to claim 3, characterized in that, The step of determining the valid text data in the semantic text data based on the target semantic intent and the key text data includes: For each of the sub-text data, if the key text data is not empty and the key text data matches the target semantic intent, the target semantic intent and the key text data are encoded to obtain valid sub-data; Valid text data in the semantic text data is determined based on at least one of the valid sub-data.

5. A communication device for marine VHF equipment, characterized in that, include: The signal processing module is used to determine the original speech signal, perform signal processing on the original speech signal, and determine semantic text data. The data processing module is used to process the semantic text data and determine the valid text data in the semantic text data. The signal transmission module is used to generate a valid text signal corresponding to the valid text data, and transmit the valid text signal to the voice receiver through the target communication channel; A voice receiving module is configured to enable the voice receiving end to determine the target communication voice based on the valid text signal; The signal processing module includes: a spectrum denoising unit, a volume adjustment unit, and a semantic recognition unit; The spectrum denoising unit is used to perform spectrum denoising on the original speech signal using a Fourier transform algorithm to obtain a denoised first speech signal. The volume adjustment unit is used to adjust the volume of the first speech signal using a first gain parameter to obtain a second speech signal with adjusted volume, wherein the first gain parameter is determined based on the global root mean square energy value corresponding to the first speech signal. The semantic recognition unit is used to perform semantic recognition on the second speech signal through a speech recognition model to determine the semantic text data corresponding to the second speech signal. The volume adjustment unit includes: a segmentation processing subunit, a volume analysis subunit, and a signal determination subunit; The segmentation processing subunit is used to segment the first speech signal according to preset signal segmentation parameters to obtain multiple second signal segments. The volume analysis subunit is used to perform volume analysis on each of the second signal segments according to an enhancement threshold, and to enhance the volume of the second signal segment according to a first gain parameter when the second signal segment is a low volume segment, thereby obtaining an enhanced signal segment. The signal determination subunit is used to determine the volume-adjusted second speech signal based on the enhanced signal segment; The signal determination subunit is specifically used for: When the root mean square energy values ​​of multiple second signal segments meet the volume compression conditions, at least one high-volume segment whose root mean square energy value exceeds the compression threshold is compressed according to a second gain parameter to obtain a compressed signal segment. The second gain parameter is determined based on the maximum or minimum value among the root mean square energy values ​​of the multiple second signal segments. A signal to be compensated is determined based on the compressed signal segment. A target gain parameter is determined based on the first gain parameter and the second gain parameter. Gain compensation is performed on the signal to be compensated based on the target gain parameter to obtain the second speech signal with adjusted volume.

6. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the communication method of the marine VHF equipment according to any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the communication method of the marine VHF equipment according to any one of claims 1-4.

Citation Information

Patent Citations

  • Audio processing method and device, equipment, medium and computer program product

    CN114338623A

  • Offline speech recognition method and device, electronic equipment and readable storage medium

    CN115104151A

  • Voice information processing method and device and nonvolatile storage medium

    CN117636878A

  • Bandwidth efficient digital voice communication system and method

    US20060235692A1