VHF voice ship name association method suitable for ship traffic management system
Through dynamic ship name set preprocessing, multi-model collaborative transcription and hot word optimization methods, combined with multi-size sliding window matching algorithm, the problems of low accuracy and high computational complexity in VHF communication are solved, efficient ship name association and real-time information transmission are realized, and the automation and intelligence of the VTS system are improved.
Patent Information
- Application Number
- CN202510339915.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-21
AI Technical Summary
The existing VHF communications are prone to noise interference in the voice, low accuracy of ship name recognition, high computational complexity, high mismatch rate and difficult context to utilize in the ship traffic management system, resulting in low information transmission efficiency and insufficient automation.
Dynamic ship name set preprocessing, multi-model collaborative transcription and hot word optimization are adopted, combined with multi-size sliding window matching algorithm, and the matching range is narrowed through the pinyin code table and the approximate tone index table to achieve high-accurate ship name association.
It significantly improves the accuracy of ship name recognition, reduces the error matching rate, reduces the consumption of computing resources, meets real-time requirements, and reduces system transformation costs, and improves the automation and intelligence level of VTS systems.
Smart Images

Figure CN120299455A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vessel traffic management, and particularly to a method for associating VHF voice with vessel names applicable to a vessel traffic management system. Background Art
[0002] A vessel traffic management system (VTS) is mainly responsible for supervising the dynamics of vessels in a water area and ensuring navigation safety. The characteristics of its supervision area are highly active economy, intensive cargo throughput, numerous traffic participants, and severe risks of collision and pollution. Therefore, the reliability requirements for the VTS system are extremely high. As an important part of VTS, according to relevant requirements, VHF (very high frequency) radio communication is the main way for real-time communication between vessels and between vessels and the onshore control center, and its importance is self-evident. However, although VHF communication plays an important role in vessel traffic management, it also has the following drawbacks and limitations:
[0003] 1. VHF conversations are voice contents. Affected by dialects, attention, noise, etc., information omission or misunderstanding is likely to occur, and it may be affected by the experience, fatigue or pressure of emergencies of the interlocutors, and there is a risk of improper communication or misjudgment.
[0004] 2. The content transmitted by VHF communication is voice, which cannot be deeply integrated with the VTS information system, restricting its application in automated and intelligent maritime management and unable to effectively relieve the duty pressure of the duty officers. For example, when making a VHF call, the information of the corresponding vessel cannot be marked and displayed in the VTS system synchronously according to the voice call content, and the duty officer still needs to manually query the vessel information, resulting in low efficiency.
[0005] To solve the above drawbacks and limitations, it is necessary to extract accurate vessel name texts from VHF voices. However, due to the following problems, the efficiency of vessel name recognition and association using existing speech-to-text methods is low:
[0006] 1. The accuracy of transcribing vessel name voices into texts is low: There is noise interference in the VHF communication environment, and the vessel names often contain special digital readings in navigation (such as "0" is read as "dong"), and it is difficult for general speech recognition models to transcribe accurately.
[0007] 2. The matching range of vessel names is wide: Traditional methods need to match in the full Chinese character library, with high computational complexity and poor real-time performance.
[0008] 3. Interference from similar sounds: When crew members make VHF calls, they are prone to confusing the pronunciation of vessel names (such as "bei" and "bai"), resulting in an increase in the mis-matching rate.
[0009] 4. Difficulty in leveraging context: The current state-of-the-art deep learning-based speech recognition models improve the accuracy of speech transcription based on common text statistical patterns (the commonality of content) and context relationships. Due to the tightness of VHF communication public channel resources, VHF calls are often short and specialized conversations, and the VHF conversations between traffic control and the same ship usually end within 5 sentences. Therefore, it is difficult to accurately extract ship names by leveraging context information.
[0010] 5. High specialization: There are many frequently mentioned specialized hot words in VHF calls, such as "Hello, traffic control" and "Report to you". If not effectively utilized, it increases the difficulty of extracting ship names; ship names themselves do not belong to common daily expressions, and general speech models are difficult to accurately transcribe.
[0011] In the existing art, there has not yet been a VHF voice-associated ship name method optimized for VHF communication scenarios. In view of the above problems, the present invention proposes a solution that integrates preprocessing of a dynamic ship name set, collaborative transcription of multiple models, and intelligent matching, thereby filling the technical gap in this field. Summary of the Invention
[0012] The technical problem to be solved by the present invention is to provide a method for associating VHF voice with ship names applicable to a ship traffic management system in view of the deficiencies of the above-mentioned existing technologies. The method for associating VHF voice with ship names applicable to a ship traffic management system achieves high-accuracy ship name association through a dynamic ship name set to narrow the matching range, collaborative transcription of multiple models and hot word optimization, preprocessing of the dynamic ship name set, and a multi-size sliding window matching algorithm.
[0013] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0014] A method for associating VHF voice with ship names applicable to a ship traffic management system includes the following steps.
[0015] Step 1: Establish a pinyin code table and an approximate sound index table.
[0016] Step 2: Preprocess the dynamic ship name set: Receive the dynamic ship name set provided by the VTS system in real time and preprocess it.
[0017] Step 3: Transcribe multi-channel voice to text: Receive VHF call voice in real time, perform voice-to-text processing through multiple speech recognition models, generate two types of output pinyin sequences, namely multiple real-time transcriptions and accurate transcriptions, and query the pinyin code table established in Step 1 to convert them into pinyin code sequences respectively.
[0018] Step 4: Multi-size sliding window matching: For the pinyin code sequences converted by each speech recognition model, sliding windows of different sizes are constructed respectively, and each pinyin code sequence within each sliding window is traversed. Retrieval operations are performed through the approximate sound index table and the preprocessed dynamic ship name set, and multiple types of ship name sound matching sets are output; among them, each type of ship name sound matching set includes several ship name pinyin code entries.
[0019] Step 5: Output the best matching ship name: Based on the weighted credibility, output the best matching ship name pinyin code entry from the ship name pinyin code entries in the multiple types of ship name sound matching sets output in Step 4.
[0020] In Step 1, the method for establishing the pinyin code table and the approximate sound index table includes the following steps:
[0021] Step 1-1: Establish the pinyin code table: All common pinyins are sequentially assigned numerical key values starting from 0 to form a one-to-one corresponding pinyin code table.
[0022] Step 1-2: Establish the approximate sound index table: According to the characteristics of VHF voice call services, group the easily confused approximate sounds to form an approximate sound index table.
[0023] In Step 2, the preprocessing method for the dynamic ship name set includes the following steps:
[0024] Step 2-1: Construct the dynamic ship name set: The VTS system forms a dynamic ship name set by combining the ship names that may have VHF calls according to the coverage range of the VHF system calls.
[0025] Step 2-2: Construct the dynamic ship name pinyin entry set: Traverse the dynamic ship name set, and convert each ship name text into a ship name pinyin entry one by one to form a dynamic ship name pinyin entry set; among them, when the ship name contains characters with nautical digital readings, multiple ship name pinyin entries are generated.
[0026] Step 2-3: Construct a three-dimensional index array: Traverse all the ship name pinyin entries in the dynamic ship name pinyin entry set, convert each ship name pinyin entry into a ship name pinyin code entry using the pinyin code table, and store it in the three-dimensional index array; among them, the first dimension of the three-dimensional index array is the pinyin code corresponding to a single pinyin; the second dimension is the length of the ship name pinyin; the third dimension is the position of the pinyin in the ship name.
[0027] In Step 3, the method for converting multi-channel voice to text includes the following steps:
[0028] Step 3-1: Deploy multiple speech recognition models;
[0029] Step 3-2: Configure a nautical-specific hot word library and a mapping of nautical-specific readings for numbers for each speech recognition model;
[0030] Step 3-3: Each speech recognition model outputs two types of pinyin sequences in real time, namely real-time transcription and accurate transcription.
[0031] Step 3-4: Query the pinyin code table established in Step 1 and convert the pinyin sequences output in real time in Step 3-3 into pinyin code sequences.
[0032] In Step 4, each speech recognition model constructs sliding windows of different sizes, and each sliding window is assigned an independent thread.
[0033] In Step 4, the set of various types of ship name sound matches is a set of four types of ship name sound matches, namely: homophone complete match set, homophone partial match set, approximate sound complete match set, and approximate sound partial match set.
[0034] In Step 4, if the preprocessed dynamic ship name set is a three-dimensional index array, the method for matching various types of ship name sounds includes the following steps:
[0035] Step 4-1: The sliding distance each time is 1 pinyin, and after each slide, traverse each pinyin code in the sliding window;
[0036] Step 4-2: For each pinyin code, use the pinyin code, the size of the sliding window, and the position of the pinyin code in the sliding window as parameters to query the three-dimensional index array to obtain the set of homophone-related ship name pinyin code entries corresponding to the pinyin code.
[0037] Step 4-3: For each pinyin code, find its approximate sound through the approximate sound index table to form a set of approximate sound codes; if there is an approximate sound, traverse each approximate sound code in the set of approximate sound codes, use the approximate sound code, the size of the sliding window, and the position of the approximate sound code in the sliding window as parameters to query the three-dimensional index array, and add the query result to the set of approximate sound-related ship name pinyin code entries corresponding to the pinyin code.
[0038] Step 4-4: Perform an intersection operation on the set of homophone-related ship name pinyin code entries and the set of approximate sound-related ship name pinyin code entries for all pinyin codes in the sliding window, and output the homophone complete match set, homophone partial match set, approximate sound complete match set, and approximate sound partial match set.
[0039] In Step 3, the number of speech recognition models is set to N, and the set of four types of ship name sound matches output by each of the N speech recognition models contains M ship name pinyin code entries. Then, in Step 5, the output method for screening the best-matched ship name among the M ship name pinyin code entries includes the following steps:
[0040] Step 5-1: Model weight assignment: Assign a weight m to the i-th speech recognition model i ; where 1 ≤ i ≤ N.
[0041] Step 5-2, calculate the type weight: Calculate the weight w of the k-th type of ship name sound matching value k , and the specific calculation formula is:
[0042]
[0043] In the formula, T is the number of pinyins included in a certain pinyin code sequence A within a certain sliding window.
[0044] H is the number of homophone matching pinyins output by the speech recognition model that match the pinyin code sequence A.
[0045] A is the number of approximate sound matching pinyins output by the speech recognition model that match the pinyin code sequence A.
[0046] λ is the approximate sound penalty factor, with a value range of 0 to 1.0.
[0047] Step 5-3, determine the indicator function: Let the indicator function of the k-th type of ship name sound matching value output by the i-th speech recognition model be δ i,k , then the specific expression is:
[0048]
[0049] Step 5-4, calculate the weighted credibility: Calculate the weighted credibility C for each ship name pinyin code entry respectively, and the specific calculation formula is:
[0050]
[0051] Step 5-5, threshold judgment: Compare and judge the weighted credibility C of each ship name pinyin code entry with the set credibility threshold.
[0052] Step 5-6, output the best matching ship name pinyin code entry: Take the ship name pinyin code entry that exceeds the credibility dynamic threshold and has the highest weighted credibility C as the best matching ship name pinyin code entry and output it.
[0053] In Step 5-5, the credibility threshold is a dynamic threshold, which can be adaptively adjusted according to the ship density and communication noise level in the VTS coverage area; when the ship density is high or the communication noise level is high, the credibility threshold is adaptively reduced.
[0054] The approximate sound penalty factor λ can be dynamically adjusted for VHF communication noise; when the VHF communication has high noise, the value of the approximate sound penalty factor λ is increased.
[0055] The present invention has the following beneficial effects:
[0056] 1. Improve recognition accuracy: The present invention significantly reduces dialect and noise interference through a dynamic ship name set to narrow the matching range, multi-model collaboration, and hot word configuration. The accuracy of pinyin transcription is increased by more than 30% compared to directly using a general speech recognition model, achieving an associated accuracy rate of >95% and a false association rate of <1.5%.
[0057] 2. Low resource consumption: The dynamic ship name set is preprocessed into a three-dimensional index array, reducing the consumption of sliding window matching operations. It can process more than 10 model output pinyin code sequences simultaneously on a VTS multi-source server with a standard configuration.
[0058] 3. Real-time response: The multi-threaded sliding window matching and lock-free queue design of the present invention greatly optimize the utilization of computing resources. On a VTS multi-source server with a standard configuration, it can process more than 10 model output pinyin code sequences simultaneously, with the delay controlled within 500 ms, meeting the high real-time requirements of VHF voice associated ship names. In addition, this solution is convenient to upgrade to a multi-threaded or distributed cluster architecture, further enhancing the system's synchronous real-time processing ability for multiple operator VHF voices and fully meeting the needs of future business expansion.
[0059] 4. The above multi-size sliding window matching algorithm: Different size sliding windows are constructed based on different ship name lengths, and the matching values of homophones and approximate sounds are calculated in parallel. The best matching result is dynamically output through weighted credibility.
[0060] 5. Dynamic adaptability: The credibility threshold and approximate sound penalty factor can be adaptively adjusted according to environmental noise, enhancing the system's robustness.
[0061] 6. Strong compatibility and significant cost-effectiveness: The present invention has little impact on the existing VTS technical architecture, demonstrating good compatibility. It can be easily integrated without large-scale transformation to achieve the automatic association of VHF voice and ship information, reducing the duty pressure. This design concept takes into account the construction and maintenance costs of the VTS system, maximizing cost-effectiveness, providing a low-cost and highly available solution for users, and greatly enhancing the feasibility and attractiveness of market promotion. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 Shows a flowchart of a method for VHF voice associated ship name applicable to a ship traffic management system according to the present invention.
[0063] Figure 2 Shows a schematic diagram of the three-dimensional index array structure in the present invention.
[0064] Figure 3 Shows a schematic diagram of different size sliding windows in the present invention.
[0065] Figure 4Shows the schematic diagram of the sliding window matching operation logic in the present invention. Detailed implementation manners
[0066] The present invention will be further described in detail below in conjunction with the accompanying drawings and specific preferred implementation manners.
[0067] In this embodiment, taking the existence of vessels "Dongfang 01" and "Tianhekou 8" in a certain VTS maritime jurisdiction as an example, and combining with the VHF dialogue voice "Traffic control, hello. Dongfang 01 is leaving Tianhekou and entering the main channel. Reporting to you.", a detailed description will be given.
[0068] As Figure 1 shown, a method for associating VHF voice with vessel names applicable to a vessel traffic management system includes the following steps.
[0069] Step 1: Establish a pinyin code table and an approximate sound index table. The specific establishment method preferably includes the following steps.
[0070] Step 1-1: Establish a pinyin code table
[0071] A. Sequentially assign numerical key values starting from 0 to all common pinyins. For example:
[0072] dong → 0, fang → 1, ling → 2, yi → 3, tian → 4, he → 5, kou → 6, ba → 7, hao → 8, san →
[0073] 9,..., tong → 16,...
[0074] B. Form a one-to-one mapping relationship table and store it as key-value pairs to support fast query.
[0075] Step 1-2: Establish an approximate sound index table
[0076] A. Define approximate sound groups according to the easily confused pronunciations in VHF communication:
[0077]
[0078]
[0079] B. Establish an approximate sound index table with the main pinyin code as the index value. For example, when the input pinyin code is 0 (the pinyin code of dong), the associated approximate sound code 16 (the pinyin code of tong) can be queried.
[0080] Step 2: Preprocess the dynamic vessel name set: Receive the dynamic vessel name set provided by the VTS system in real time and preprocess it. The preferred method includes the following steps.
[0081] Step 2-1: Construct a dynamic ship name set. The VTS system forms a dynamic ship name set (such as only including ships in the current water area) by combining the ship names of ships that may conduct VHF calls according to the coverage range of VHF system calls.
[0082] In this embodiment, the VTS system screens the ships in the currently covered water area, such as "Dongfang 03" and "Tianhekou No. 8", to generate a dynamic ship name set.
[0083] Step 2-2: Construct a dynamic ship name pinyin entry set. Traverse the dynamic ship name set and convert each ship name text into a ship name pinyin entry one by one to form a dynamic ship name pinyin entry set; among them, when the ship name contains words with nautical digital readings, multiple ship name pinyin entries are generated.
[0084] In this embodiment, the preferred method for converting ship name texts is as follows:
[0085] (1) Processing of "Dongfang 03":
[0086] Converted into two pinyin entries (including nautical digital readings):
[0087] dong fang dong san ("Dong San" reading)
[0088] dong fang ling san ("Ling San" reading)
[0089] (2) Processing of "Tianhekou No. 8":
[0090] Converted into a standard pinyin entry: tian he kou bahao
[0091] Step 2-3: Construct a three-dimensional index array. Traverse all the ship name pinyin entries in the dynamic ship name pinyin entry set, convert each ship name pinyin entry into a ship name pinyin code entry using the pinyin code table, and store it in the three-dimensional index array; among them, the first dimension of the three-dimensional index array is the pinyin code corresponding to a single pinyin; the second dimension is the length of the ship name pinyin; the third dimension is the position of the pinyin in the ship name.
[0092] In this embodiment, converting the ship name pinyin entry into a ship name pinyin code entry:
[0093] ■dong fang dong san → [0,1,0,9]
[0094] ■dong fang ling san → [0,1,2,9]
[0095] ■tian he kou ba hao → [4,5,6,7,8]
[0096] · Example 1: The pinyin code entry [0, 1, 0, 9] for the ship name "Dongfang 01", with a length of 4:
[0097] Since dong (code 0) is in the 1st position → Store the pinyin code entry [0, 1, 0, 9] in Q[0][4][1]
[0098] Since fang (code 1) is in the 2nd position → Store the pinyin code entry [0, 1, 0, 9] in Q[1][4][2]
[0099] Since dong (code 0) is in the 3rd position → Store the pinyin code entry [0, 1, 0, 9] in Q[0][4][3]
[0100] Since san (code 9) is in the 4th position → Store the pinyin code entry [0, 1, 0, 9] in Q[9][4][4]
[0101] · Example 2: The pinyin code entry [2, 3, 4, 5, 6] for the ship name "Tianhekou 8th", with a length of 5:
[0102] Since tian (code 2) is in the 1st position → Store the pinyin code entry [2, 3, 4, 5, 6] in Q[2][5][1]
[0103] Since he (code 3) is in the 2nd position → Store the pinyin code entry [2, 3, 4, 5, 6] in Q[3][5][2]
[0104] Since kou (code 4) is in the 3rd position → Store the pinyin code entry [2, 3, 4, 5, 6] in Q[4][5][3]
[0105] Since ba (code 5) is in the 4th position → Store the pinyin code entry [2, 3, 4, 5, 6] in Q[5][5][4]
[0106] Since hao (code 6) is in the 5th position → Store the pinyin code entry [2, 3, 4, 5, 6] in Q[6][5][5]
[0107] In this embodiment, the constructed three-dimensional index array, as Figure 2 shown, in Figure 2In this example, taking the element of the 3D index array [6, 5, 3] as an example, this element is of the set type. The value of the first dimension is 6, indicating that the pinyin code sequences of ship names stored therein all contain the pinyin code 6. The value of the second dimension is 5, indicating that the length of the pinyin code sequences of ship names stored therein is limited to 5. The value of the third dimension is 3, indicating that the 3rd position of the pinyin code sequences of ship names stored therein is the pinyin code 6. This storage method provides a sliding window algorithm for high-performance search. For example, if the pinyin code at the sliding window position of size Y on Z is X, then by directly looking up the element value (set) of the 3D index array [Y, Z], the optional pinyin codes of ship names determined by this pinyin code can be read.
[0108] Step 3, multi-channel voice to text: Real-time receive VHF call voice, perform voice to text processing through multiple speech recognition models, generate two types of output pinyin sequences: multiple real-time transcriptions (low latency) and accurate transcriptions (high accuracy), and query the pinyin code table established in step 1 to convert them into pinyin code sequences respectively.
[0109] The above method for multi-channel voice to text includes the following steps.
[0110] Step 3-1, deploy N speech recognition models, such as Wav2Vec, Whisper, DeepSpeech, SenseVoice, etc. In this embodiment, preferably N = 2, which are Whisper (hereinafter referred to as model 1) and SenseVoice (hereinafter referred to as model 2) respectively.
[0111] Step 3-2, configure a nautical-specific hot word library (such as "Hello, traffic control", "Report to you") and a mapping of the nautical-specific readings of numbers for each speech recognition model.
[0112] Step 3-3, each speech recognition model outputs two types of pinyin sequences: real-time transcription and accurate transcription in real time.
[0113] 1. Voice input: The VHF voice content is "Hello, traffic control. Dong Fang Dong Yao exits Tianhekou and enters the main channel. Report to you."
[0114] 2. Multi-model transcription and hot word optimization:
[0115] (1) Output of model 1 (Whisper):
[0116] jiao guan li hao,dong fang dong san chu tian he kou,xiang ni bao gao
[0117] (2) Output of model 2 (SenseVoice):
[0118] Hello, Guan Ni Hao. I'm reporting from the Tianhekou, Dongfang Tongsan
[0119] (3)Hot word replacement: According to the hot word library, replace "jiao guan li hao" with jiao guan ni hao, and replace "jiao guang ni hao"
[0120] with jiao guan ni hao. The correction is as follows:
[0121] Final pinyin sequence of Model 1:
[0122] Hello, Guan Ni Hao. I'm reporting from the Tianhekou, Dongfang Dongsan
[0123] Final pinyin sequence of Model 2:
[0124] Hello, Guan Ni Hao. I'm reporting from the Tianhekou, Dongfang Tongsan
[0125] Steps 3-4: Query the pinyin code table established in Step 1 and convert the pinyin sequence output in real time in Step 3-3 into a pinyin code sequence. In this embodiment, the pinyin code sequences after conversion by the two models are as follows.
[0126] Pinyin code sequence of Model 1: [9,10,11,6,,0,1,0,9,14,4,5,6,,13,11,14,15]
[0127] Pinyin code sequence of Model 2: [9,10,11,6,,0,1,16,9,14,4,5,6,,13,11,14,15]
[0128] Step 4: Multi-size sliding window matching
[0129] A. Initialization of the sliding window
[0130] For the pinyin code sequences after conversion by each speech recognition model, sliding windows of different sizes are constructed respectively. In this example, 9 sliding windows with lengths from 2 to 10 are constructed, and each sliding window is assigned an independent thread.
[0131] B. Traverse each pinyin code sequence within each sliding window, perform retrieval operations through the approximate sound index table and the preprocessed dynamic ship name set, and output multiple types of ship name sound matching sets. Preferably, four types of ship name sound matching sets are output, namely: homophonic exact match set, homophonic partial match set, approximate sound exact match set, and approximate sound partial match set.
[0132] Each type of ship name sound matching set includes several ship name pinyin code entries. Therefore, the four types of ship name sound matching sets respectively output by N speech recognition models together contain M ship name pinyin code entries.
[0133] As Figure 3 shown, the method for matching multiple types of ship names preferably includes the following steps.
[0134] Step 4-1: Each sliding distance is 1 pinyin. After each slide, traverse each pinyin code within the sliding window.
[0135] Step 4-2: For each pinyin code, use the pinyin code, sliding window size, and the position of the pinyin code in the sliding window as parameters to query the three-dimensional index array, and obtain the set of homophonic associated ship name pinyin code entries corresponding to the pinyin code.
[0136] Step 4-3: For each pinyin code, find its approximate sounds through the approximate sound index table and form a set of approximate sound codes; if there are approximate sounds, traverse each approximate sound code in the set of approximate sound codes, use the approximate sound code, sliding window size, and the position of the approximate sound code in the sliding window as parameters to query the three-dimensional index array, and add the query results to the set of approximate sound associated ship name pinyin code entries corresponding to the pinyin code.
[0137] Step 4-4: Perform an intersection operation on the sets of homophonic associated ship name pinyin code entries and approximate sound associated ship name pinyin code entries for all pinyin codes within the sliding window, and output the homophonic exact match set, homophonic partial match set, approximate sound exact match set, and approximate sound partial match set.
[0138] Figure 4 shows a schematic diagram of the sliding window matching operation logic. The specific matching operation logic is as follows.
[0139] Homophonic matching: Directly retrieve the three-dimensional index array through the pinyin code, and calculate the exact match (H = T) and partial match (H / T ≥ 80%) values.
[0140] Approximate sound matching: Combine the approximate sound index table to expand the retrieval range, and calculate the approximate sound exact match (H + A = T) and partial match ((H + A) / T ≥ 80%) values.
[0141] A. The homophonic exact match set includes several homophonic exact match values
[0142] Definition of homophone exact match value: It is the intersection of the sets of homophone-related ship name pinyin code entries for all pinyin codes within the sliding window. In other words, all ship name pinyin code entries in this set meet the following conditions: If a ship name pinyin code entry contains T pinyins, when all pinyin codes within the sliding window match this ship name pinyin code entry, the number of homophone-matched pinyin codes is H, then H = T.
[0143] B. The homophone partial match set includes several homophone partial match values
[0144] Definition of homophone partial match value: The ratio of the set of homophone-related ship name pinyin code entries for pinyin codes within the sliding window that hits the same ship name pinyin code entry exceeds a preset threshold (such as 80%), and the set of related ship name pinyin code entries is the homophone partial match value. In other words, all ship name pinyin code entries in this set meet the following conditions: If a ship name pinyin code entry contains T pinyins, when all pinyin codes within the sliding window match this ship name pinyin code entry, the number of homophone-matched pinyin codes is H, then H < T and H / T is greater than the preset threshold (such as 80%).
[0145] C. The approximate homophone exact match set includes several approximate homophone exact match values
[0146] Definition of approximate homophone exact match value: The difference set between the intersection of the set of associated ship names for all pinyin codes within the sliding window (the set of approximate homophone-related ship name pinyin code entries + the set of homophone-related ship name pinyin code entries) and the homophone exact match value is the approximate homophone exact match value. In other words, all ship name pinyin code entries in this set meet the following conditions: If a ship name pinyin code entry contains T pinyins, when all pinyin codes within the sliding window match this ship name pinyin code entry, the number of homophone-matched pinyin codes is H, and the number of approximate homophone-matched pinyin codes is A (A ≥ 1), then H + A = T.
[0147] D. The approximate homophone partial match set includes several approximate homophone partial match values
[0148] Definition of approximate homophone partial match value: The ratio of the set of homophone-related ship name pinyin code entries for pinyin codes within the sliding window (the set of approximate homophone-related ship name pinyin code entries + the set of homophone-related ship name pinyin code entries) that hits the same ship name pinyin code entry exceeds a preset threshold (such as 80%), and the difference set between the set of related ship name pinyin code entries and the homophone partial match value is the approximate homophone partial match value. In other words, all ship name pinyin code entries in this set meet the following conditions: If a ship name pinyin code entry contains T pinyins, when all pinyin codes within the sliding window match this ship name pinyin code entry, the number of homophone-matched pinyin codes is H, and the number of approximate homophone-matched pinyin codes is A (A ≥ 1), then (H + A) < T and (H + A) / T is greater than the preset threshold (such as 80%).
[0149] In this embodiment, an example of the partial matching process of Model 1 and Model 2 is as follows.
[0150] (1) For the code sequence of Model 1:
[0151] (1.1) Sliding window length 4 (Thread 1):
[0152] When sliding to the 6th time, the sliding window covers the code segment [0, 1, 0, 9], resulting in a perfect homophone match.
[0153] Query the three-dimensional index array Q:
[0154] Position 1: Pinyin code [0] → Q[0][4][1] → contains pinyin code entry [0, 1, 0, 9]
[0155] Position 2: Pinyin code [1] → Q[1][4][2] → contains pinyin code entry [0, 1, 0, 9]
[0156] Position 3: Pinyin code [0] → Q[0][4][3] → contains pinyin code entry [0, 1, 0, 9]
[0157] Position 4: Pinyin code [9] → Q[9][4][4] → contains pinyin code entry [0, 1, 0, 9]
[0158] Successfully match the pinyin code entry [0, 1, 0, 9] perfectly.
[0159] (1.2) Sliding window length 5 (Thread 2):
[0160] When sliding to the 11th time, it covers the code segment [4, 5, 6, 7,, 13]
[0161] Query 3 pinyins that match [4, 5, 6, 7, 8] (H = 3, T = 5), and H / T = 60% is lower than the threshold (assumed to be 80%), so the homophone partial match is not output.
[0162] (2) For the code sequence of Model 2:
[0163] (2.1) Sliding window length 4 (Thread 3):
[0164] When sliding to the 6th time, the sliding window covers the code segment [0, 1, 16, 9], resulting in an approximate homophone perfect match.
[0165] Through the approximate homophone index table, associate tong (code 16) with dong (code 0); convert the code segment [0, 1, 16, 9] with a sliding window length of 4 to the approximate homophone code [0, 1, 0, 9]; query the three-dimensional index array Q, H = 3, A = 1, T = 4, that is, H + A = T, and it is an approximate homophone perfect match with [0, 1, 0, 9].
[0166] (2.2) Sliding window length 5 (thread 4):
[0167] In Figure 4 , taking the sliding match with a sliding window of length 5 as an example, assume that there are the following 4 entries of the Chinese phonetic code of ship names in the ship dynamic option set: [7, 8, 9, 10, 11], [7, 7, 9, 10, 11], [7, 18, 9, 10, 11], [7, 18, 9, 1, 11], and the partial matching preset threshold is set to 80%.
[0168] When sliding to the 11th time, it covers the code segment [4, 5, 6, 7,, 13], generating a homophone partial match.
[0169] It matches 3 pinyins with [4, 5, 6, 7, 8] (H = 3, T = 5), and H / T = 60% is lower than the threshold (assumed to be 80%), so it is not output.
[0170] Step 5. Output the best matching ship name: According to the weighted credibility, output the entry of the Chinese phonetic code of the best matching ship name from the entries of the Chinese phonetic code of ship names in the multi-type ship name sound matching set output in step 4.
[0171] In this embodiment, the above 4 threads output a homophone exact match {[0, 1, 0, 9]} and an approximate homophone exact match {[0, 1, 0, 9]} during the sliding process. Taking the ship name code entry [0, 1, 0, 9] as an example, the credibility is calculated as follows.
[0172] The above method for outputting the best matching ship name preferably includes the following steps.
[0173] Step 5-1. Model weight assignment: Assign a weight m i to the i-th speech recognition model; where 1 ≤ i ≤ N.
[0174] In this embodiment, the accuracy rate of model 1 is relatively high, and the weight m1 = 0.7, and the weight of model 2 is m2 = 0.3.
[0175] Step 5-2. Calculate the type weight: Calculate the weight w k of the k-th type of ship name sound matching value, and the specific calculation formula is:
[0176]
[0177] In the formula, T is the number of pinyins included in a certain pinyin code sequence A in a certain sliding window.
[0178] H is the number of homophone matching pinyins output by the speech recognition model that match the pinyin code sequence A.
[0179] A is the number of approximate sound matching pinyins that match the pinyin code sequence A output by the speech recognition model.
[0180] λ is the approximate sound penalty factor, with a value range of 0 to 1.0, which can dynamically adjust the VHF communication noise; when the VHF communication has high noise, the value of the approximate sound penalty factor λ is increased.
[0181] In this embodiment,
[0182] Homophone exact match (H = 4, T = 4H = 4, T = 4): w1 = (4 + 0) / 4 = 1.0
[0183] Approximate sound exact match (H = 3, A = 1, T = 4): w3 = (3 + 0.8×1) / 4 = 0.95 (assuming λ = 0.8)
[0184] Step 5-3, determine the indicator function: Let the indicator function of the i-th speech recognition model outputting the sound matching value of the k-th type of ship name be δ i,k , then the specific expression is:
[0185]
[0186] Step 5-4, calculate the weighted credibility: Calculate the weighted credibility C for each ship name pinyin code entry respectively, and the specific calculation formula is:
[0187]
[0188] In this embodiment, the weighted credibility C of the ship name code entry [0,1,0,9] is:
[0189] C = 0.7×1.0 + 0.3×0.95 = 0.985
[0190] Step 5-5, threshold determination: Compare and judge the weighted credibility C of each ship name pinyin code entry with the set credibility threshold.
[0191] The above credibility threshold is a dynamic threshold, which can be adaptively adjusted according to the ship density and communication noise level in the VTS coverage area; when the ship density is high or the communication noise level is high, the credibility threshold is adaptively reduced.
[0192] In this embodiment, the current ship density is low, the credibility threshold is set to 0.9, and the credibility of the matching result [0,1,0,9] is 0.985, which exceeds the threshold.
[0193] Step 5-6, output the best matching ship name pinyin code entry: Take the ship name pinyin code entry that exceeds the credibility dynamic threshold and has the highest weighted credibility C as the best matching ship name pinyin code entry and output it.
[0194] In this embodiment, [0,1,0,9] with the highest credibility is selected, and the corresponding ship name "Dongfang 01" is sent to the VTS system to automatically identify the ship.
[0195] Effect verification
[0196] Accuracy: Model 1 completely matches "Dongfang 01", and Model 2 matches successfully after approximate sound correction. The comprehensive accuracy rate > 95%.
[0197] Real-time performance: Sliding window multi-thread parallel processing, matching latency < 300ms (< 500ms), meeting the real-time supervision requirements of VTS.
[0198] Anti-interference ability: The approximate sound index table effectively corrects the misrecognition of "tong" and "dong", reducing the influence of noise interference.
[0199] The present invention has been verified through the pilot of a certain maritime bureau's VTS system. In the high-load scenario with a ship density of 5,000 ships per day, the ship name association accuracy rate is increased to 96.1%, and the false alarm rate is decreased to less than 1.4%, significantly improving the maritime supervision efficiency.
[0200] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the technical concept scope of the present invention, various equivalent transformations can be made to the technical solutions of the present invention, and these equivalent transformations all belong to the protection scope of the present invention.
Claims
1. A method for VHF voice associated ship name applicable to a ship traffic management system, characterized in that: It includes the following steps: Step 1: Establish a pinyin code table and an approximate sound index table; Step 2: Preprocess the dynamic ship name set: Receive the dynamic ship name set provided by the VTS system in real time and preprocess it; Step 3: Multichannel speech-to-text conversion: Receive VHF call speech in real time, perform speech-to-text conversion through multiple speech recognition models, generate two types of output pinyin sequences, namely real-time transcription and accurate transcription, and query the pinyin code table established in Step 1 to convert them into pinyin code sequences respectively; Step 4: Multisize sliding window matching: For the pinyin code sequences converted by each speech recognition model, construct sliding windows of different sizes respectively, and traverse each pinyin code sequence within each sliding window, perform retrieval operations through the approximate sound index table and the preprocessed dynamic ship name set, and output multiple types of ship name sound matching sets; among them, each type of ship name sound matching set includes several ship name pinyin code entries; Step 5: Output the best matching ship name: Output the best matching ship name pinyin code entry according to the weighted credibility for the ship name pinyin code entries in the multiple types of ship name sound matching sets output in Step 4.
2. The method for VHF voice associated ship name applicable to the vessel traffic management system according to claim 1, characterized in that: In Step 1, the method for establishing the pinyin code table and the approximate sound index table includes the following steps: Step 1-1: Establish a pinyin code table: Assign digital key values to all common pinyins starting from 0 in sequence to form a one-to-one corresponding pinyin code table; Step 1-2: Establish an approximate sound index table: Group the easily confused approximate sounds according to the characteristics of VHF voice call services to form an approximate sound index table.
3. The method for associating VHF voice with ship names applicable to a ship traffic management system according to claim 1, wherein: In Step 2, the preprocessing method for the dynamic ship name set includes the following steps: Step 2-1: Construct a dynamic ship name set: The VTS system forms a dynamic ship name set based on the ships' names that may conduct VHF calls according to the coverage range of the VHF system calls; Step 2-2: Construct a dynamic ship name pinyin entry set: Traverse the dynamic ship name set, and convert each ship name text into a ship name pinyin entry one by one to form a dynamic ship name pinyin entry set; among them, when the ship name contains words with nautical digital readings, multiple ship name pinyin entries are generated; Step 2-3: Construct a three-dimensional index array: Traverse all the ship name pinyin entries in the dynamic ship name pinyin entry set, convert each ship name pinyin entry into a ship name pinyin code entry by using the pinyin code table, and store them in the three-dimensional index array; among them, the first dimension of the three-dimensional index array is the pinyin code corresponding to a single pinyin; the second dimension is the length of the ship name pinyin; the third dimension is the position of the pinyin in the ship name.
4. The method for VHF voice associated ship name applicable to a ship traffic management system according to claim 1, wherein: In Step 3, the method for multichannel speech-to-text conversion includes the following steps: Step 3-1: Deploy multiple speech recognition models; Step 3-2: Configure a nautical-specific hotword library and a mapping of nautical-specific readings for numbers for each speech recognition model; Step 3-3: Each speech recognition model outputs two types of pinyin sequences, namely real-time transcription and accurate transcription, in real time; Step 3-4: Query the pinyin code table established in Step 1 and convert the pinyin sequences output in real time in Step 3-3 into pinyin code sequences.
5. The method for VHF voice-associated ship name applicable to a vessel traffic management system according to claim 1, characterized in that: In Step 4, each speech recognition model constructs sliding windows of different sizes, and each sliding window is assigned an independent thread.
6. The method for VHF voice-associated ship name applicable to a vessel traffic management system according to claim 1, characterized in that: In step 4, the multi-type ship name sound matching sets are four types of ship name sound matching sets, namely: the homophone complete matching set, the homophone partial matching set, the approximate sound complete matching set, and the approximate sound partial matching set.
7. The method for VHF voice associated ship name applicable to the vessel traffic management system according to claim 6, characterized in that: In step 4, if the preprocessed dynamic ship name set is a three-dimensional index array, the method for multi-type ship name sound matching includes the following steps: Step 4-1: Each sliding distance is 1 pinyin, and after each slide, traverse each pinyin code in the sliding window; Step 4-2: For each pinyin code, use the pinyin code, the sliding window size, and the position of the pinyin code in the sliding window as parameters to query the three-dimensional index array to obtain the set of homophone-related ship name pinyin code entries corresponding to the pinyin code; Step 4-3: For each pinyin code, find its approximate sound through the approximate sound index table and form a set of approximate sound codes; if there is an approximate sound, traverse each approximate sound code in the set of approximate sound codes, use the approximate sound code, the sliding window size, and the position of the approximate sound code in the sliding window as parameters to query the three-dimensional index array, and add the query result to the set of approximate sound-related ship name pinyin code entries corresponding to the pinyin code; Step 4-4: Perform an intersection operation on the set of homophone-related ship name pinyin code entries and the set of approximate sound-related ship name pinyin code entries for all pinyin codes in the sliding window, and output the homophone complete matching set, the homophone partial matching set, the approximate sound complete matching set, and the approximate sound partial matching set.
8. The method for associating VHF voice with ship names applicable to a ship traffic management system according to claim 6, wherein: In step 3, the number of speech recognition models is set to N, and a total of M ship name pinyin code entries are included in the four types of ship name sound matching sets respectively output by the N speech recognition models. Then in step 5, the output method for screening the best matching ship name among the M ship name pinyin code entries includes the following steps: Step 5-1, model weight assignment: Assign a weight m to the i-th speech recognition model i ; where 1 ≤ i ≤ N; Step 5-2, calculate the type weight: calculate the weight w of the k-th type of ship name sound matching value k , and the specific calculation formula is as follows: Where T is the number of pinyins included in a certain pinyin code sequence A in a certain sliding window; H is the number of homophone matching pinyins output by the speech recognition model that matches the pinyin code sequence A; A is the number of approximate sound matching pinyins output by the speech recognition model that matches the pinyin code sequence A; λ is the approximate sound penalty factor, with a value range of 0 to 1.0; Step 5-3, determine the indication function: Let the indication function of the ith speech recognition model for outputting the matching value of the kth type of ship name sound be δ i,k , and the specific expression is as follows: Step 5-4: Calculate the weighted credibility: Calculate the weighted credibility C for each ship name pinyin code entry respectively, and the specific calculation formula is: Step 5-5: Threshold determination: Compare and judge the weighted credibility C of each ship name pinyin code entry with the set credibility threshold; Step 5-6: Output the best matching ship name pinyin code entry: Use the ship name pinyin code entry that exceeds the credibility dynamic threshold and has the highest weighted credibility C as the best matching ship name pinyin code entry and output it.
9. The method for associating VHF voice with ship names applicable to a vessel traffic management system according to claim 8, characterized in that: In step 5-5, the credibility threshold is a dynamic threshold, which can be adaptively adjusted according to the ship density and communication noise level in the VTS coverage area; when the ship density is high or the communication noise level is high, the credibility threshold is adaptively reduced.
10. The method for associating VHF voice with ship names applicable to a ship traffic management system according to claim 8, characterized in that: The approximate sound penalty factor λ can be dynamically adjusted for VHF communication noise; when the VHF communication has high noise, the value of the approximate sound penalty factor λ is increased.
Citation Information
Patent Citations
Ship identification and positioning system and method based on speech
CN110600007A
Enterprise name identification method and device
CN111445903A
Voice recognition method and device, equipment and medium
CN111739514A
Intelligent query method and device based on user voice, equipment and storage medium
CN115588430A
Engineering standard specification acquisition method and system based on neural network
CN117252539A