A real-time voice call translation device and a call translation method

By real-time identification of breaking sentence marks in voice data in real-time in real-time voice call translation device, the problem of inaccurate recognition of breaking sentence marks in the prior art is solved, and efficient and accurate real-time voice call translation is achieved to meet the needs of cross-language communication.

CN119808795BActive Publication Date: 2025-08-01SHENZHEN FENGLIYUAN ENERGY SAVING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411848481.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-16
Publication Date
2025-08-01
Estimated Expiration
2044-12-16

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify sentence breaking signs in voice data in continuous conversations, resulting in insufficient accuracy and timeliness of real-time voice call translation.

Method used

Through the sentence breaking mark recognition module, including text translation submodule, sentence breaking mark recognition submodule, curve fitting subunit, interval recognition subunit, endpoint time determination subunit and pause mark determination subunit, we can identify sentence breaking marks in the speech data in real time, and combine word participle model and sentence breaking mark recognition and judgment unit to determine whether sentence breaking marks exist in the speech text data.

Benefits of technology

It realizes timely and accurate translation of the latest received voice data, improves the accuracy and timeliness of voice processing, and ensures the completeness and translation quality of input statements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119808795B_ABST
    Figure CN119808795B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of call translation, and specifically discloses a real-time voice call translation device and a call translation method. The device includes: a sentence-breaking flag recognition module for determining in real time whether there is a sentence-breaking flag in the latest received voice data as a real-time sentence-breaking flag recognition result; an input sentence acquisition module for, when the real-time sentence-breaking flag recognition result indicates that there is a sentence-breaking flag in the latest received voice data, acquiring the input sentence text of the latest received input language based on the latest recognized sentence-breaking flag in the latest received voice data; a sentence text translation module for translating the input sentence text of the latest received input language based on the target translation language to obtain a real-time voice call translation result, so as to achieve timely and accurate translation of the latest received voice data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of call translation, and in particular to a real-time voice call translation device and a call translation method. Background Art

[0002] With the development of society, global communication is becoming increasingly frequent across all sectors. In cross-border economic and trade, political exchanges, and cultural dissemination, the need for communication between people of different language backgrounds is growing. Language barriers have become a major barrier to communication, leading to an increasingly urgent need for real-time voice call translation. Advances in network communication technology have enabled the rapid transmission of voice data, providing a communication guarantee for real-time voice call translation. Real-time voice data transmission is now possible via mobile networks, Wi-Fi, and other methods. Continuous advancements in speech recognition technology, which can convert speech into text, provide a foundation for voice call translation. For example, neural networks with speech recognition capabilities can generate speech recognition text for the speech to be translated. Some translation software already exists on the market, providing translation assistance in daily communication. These software provide a reference and foundation for the development of real-time voice call translation technology.

[0003] In voice call translation applications, accuracy and readability are typically achieved only after a complete sentence is captured. Therefore, identifying punctuation markers is crucial for accurate and timely voice call translation. However, identifying punctuation markers in the interlocutor's voice data during continuous conversations is a significant challenge.

[0004] Therefore, the present invention provides a real-time voice call translation device and a call translation method. Summary of the Invention

[0005] The present invention provides a real-time voice call translation device and a call translation method, which are used to accurately identify and determine the sentence segmentation marks in the latest received voice data, thereby achieving timely and accurate translation of the latest received voice data.

[0006] The present invention provides a real-time voice call translation device, comprising:

[0007] A sentence segmentation mark recognition module is used to determine in real time whether there is a sentence segmentation mark in the latest received voice data as a real-time sentence segmentation mark recognition result;

[0008] An input sentence acquisition module, configured to, when the real-time segmentation mark recognition result indicates that the segmentation mark exists in the latest received voice data, acquire the input sentence text of the latest received input language based on the segmentation mark newly recognized in the latest received voice data;

[0009] A statement text translation module, which is used to determine the target translation language according to the international area code of the incoming / outgoing call, and translate the input statement text in the latest received input language based on the target translation language to obtain a real-time voice call translation result.

[0010] Preferably, the sentence segmentation flag recognition module includes:

[0011] A text translation sub-module, which is used to convert the real-time received voice data into text data form containing pause flags to obtain the real-time voice text data in the input language;

[0012] A sentence segmentation flag recognition sub-module, which is used to recognize whether there is a sentence segmentation flag in the real-time voice text data in the input language as the real-time sentence segmentation flag recognition result.

[0013] Preferably, the text translation sub-module includes:

[0014] A text translation unit, which is used to convert the real-time received voice data into text data form to obtain the initial voice text data in the input language;

[0015] A pause flag recognition unit, which is used to recognize all pause flags in the real-time received voice data in real time;

[0016] A flag marking unit, which is used to mark all pause flags in the initial voice text data in the input language in real time to obtain the real-time voice text data in the input language.

[0017] Preferably, the pause flag recognition unit includes:

[0018] A curve fitting sub-unit, which is used to fit the short-time energy curve of the input voice signal in real time based on the short-time energy values of the voice signal of the real-time received voice data in each preset short time period;

[0019] An interval recognition sub-unit, which is used to recognize all short-time energy abnormal curve intervals in the short-time energy curve of the input voice signal in real time;

[0020] An endpoint time determination sub-unit, which is used to determine the start time and end time of the pause flag of all short-time energy abnormal curve intervals in the short-time energy curve of the input voice signal;

[0021] A pause flag determination sub-unit, which is used to determine all pause flags in the real-time received voice data based on the start time and end time of the pause flag of all short-time energy abnormal curve intervals in the short-time energy curve of the input voice signal;

[0022] Among them, all abnormal short time periods of each short-time energy abnormal curve interval include the left abnormal short time period and the right abnormal short time period of each short-time energy abnormal curve interval.

[0023] Preferably, the interval recognition subunit includes:

[0024] The first mean calculation end is used to determine all the energy sudden drop curve segments in the short-time energy curve of the input speech signal, and calculate the average short-time energy value of each energy sudden drop curve segment;

[0025] The second mean calculation end is used to calculate the average short-time energy value of all parts of the short-time energy curve of the input speech signal except all the energy sudden drop curve segments in the short-time energy curve of the input speech signal, as the reference average short-time energy value;

[0026] The sudden drop ratio calculation end is used to calculate the ratio of the average short-time energy value of each energy sudden drop curve segment to the reference average short-time energy value, as the average energy sudden drop ratio of each energy sudden drop curve segment;

[0027] The curve segment screening end is used to regard all the energy sudden drop curve segments with an average energy sudden drop ratio less than the preset threshold among all the energy sudden drop curve segments as all the short-time energy abnormal curve intervals in the short-time energy curve of the input speech signal.

[0028] Preferably, the endpoint time determination subunit includes:

[0029] The endpoint time period determination end is used to regard the preset short time period corresponding to the short-time energy value at the left endpoint of each short-time energy abnormal curve interval as the left abnormal short time period, and at the same time, regard the preset short time period corresponding to the short-time energy value at the right endpoint of the short-time energy abnormal curve interval as the right abnormal short time period;

[0030] The endpoint time determination end is used to identify the start time of the pause flag in the left abnormal short time period of each short-time energy abnormal curve interval based on the left neighborhood speech signal segment and the right neighborhood speech signal segment of all the abnormal short time periods of each short-time energy abnormal curve interval, and at the same time, identify the end time of the pause flag in the right abnormal short time period of each short-time energy abnormal curve interval.

[0031] Preferably, the endpoint time determination end includes:

[0032] The time period merging sub-end is used to merge each abnormal short time period of the short-time energy abnormal curve interval with the time periods corresponding to the left neighborhood speech signal segment and the right neighborhood speech signal segment corresponding thereto, to obtain the reference time period of each short-time energy abnormal curve interval;

[0033] A signal segment division sub - end, which is used to determine multiple sampling points at preset intervals in each abnormal short - time period of each short - time energy abnormal curve interval, and based on each sampling point, divide the voice signal of the corresponding part of the reference period of the short - time energy abnormal curve interval into a left reference voice signal segment and a right reference voice signal segment corresponding to the sampling point;

[0034] A zero - crossing rate difference sub - end, which is used to take the difference between the zero - crossing rate of the left reference voice signal segment and the zero - crossing rate of the right reference voice signal segment of each sampling point as the left - right zero - crossing rate difference of each sampling point;

[0035] A moment determination sub - end, which is used to take the sampling point with the largest left - right zero - crossing rate difference among all sampling points in the left abnormal short - time period of each short - time energy abnormal curve interval as the start moment of the pause flag. At the same time, take the sampling point with the smallest left - right zero - crossing rate difference among all sampling points in the right abnormal short - time period of each short - time energy abnormal curve interval as the start moment of the pause flag.

[0036] Preferably, the sentence - breaking flag recognition sub - module includes:

[0037] A word - segmentation unit, which is used to input the real - time voice text data of the input language into the word - segmentation model to obtain the word - segmentation result of the real - time voice text data of the input language;

[0038] A duration determination unit, which is used to determine the duration of all pause flags in the real - time voice text data of the input language;

[0039] A sentence - breaking flag recognition and judgment unit, which is used to judge whether there is a sentence - breaking flag in the real - time voice text data of the input language as the real - time sentence - breaking flag recognition result based on the part - of - speech and meaning of each word included in the word - segmentation result and the duration of all pause flags.

[0040] Preferably, the sentence - breaking flag recognition and judgment unit includes:

[0041] A model acquisition sub - unit, which is used to acquire the sentence - breaking flag recognition model;

[0042] A sentence - breaking flag judgment sub - unit, which is used to input the part - of - speech and meaning of all words included in the word - segmentation result and the duration of all pause flags into the sentence - breaking flag recognition model to judge whether there is a sentence - breaking flag in the real - time voice text data of the input language as the real - time sentence - breaking flag recognition result.

[0043] The present invention provides a real - time voice call translation method, which is applied to any one of the above - mentioned real - time voice call translation devices, and includes:

[0044] S1: Real - time judge whether there is a sentence - breaking flag in the latest received voice data as the real - time sentence - breaking flag recognition result;

[0045] S2: When the recognition result of the real-time sentence-breaking flag indicates that there is a sentence-breaking flag in the latest received voice data, based on the latest recognized sentence-breaking flag in the latest received voice data, obtain the input sentence text of the latest received input language;

[0046] S3: Determine the target translation language according to the international area code of the incoming / outgoing call, and translate the input sentence text of the latest received input language based on the target translation language to obtain the real-time voice call translation result.

[0047] The beneficial effects of the present invention compared with the prior art are as follows: By accurately identifying and judging the sentence-breaking flag in the latest received voice data, the timely and accurate translation of the latest received voice data is realized.

[0048] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in this application document.

[0049] The technical solutions of the present invention will be further described in detail below through the accompanying drawings and embodiments. Description of the Drawings

[0050] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0051] Figure 1 It is a schematic diagram of the internal functional modules of the real-time voice call translation device in the embodiment of the present invention;

[0052] Figure 2 It is a schematic diagram of the internal functional sub-modules of the sentence-breaking flag recognition module in the embodiment of the present invention;

[0053] Figure 3 It is a flowchart of the real-time voice call translation method in the embodiment of the present invention. Detailed Embodiments

[0054] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0055] Embodiment 1:

[0056] The present invention provides a real-time voice call translation device, referring to Figure 1 , including:

[0057] The sentence break flag recognition module is used to determine in real time whether there is a sentence break flag in the latest received voice data as the real-time sentence break flag recognition result;

[0058] The input sentence acquisition module is used to, when the real-time sentence break flag recognition result indicates that there is a sentence break flag in the latest received voice data, obtain the input sentence text of the latest received input language based on the latest recognized sentence break flag in the latest received voice data;

[0059] The sentence text translation module is used to determine the target translation language according to the international area code of the incoming / outgoing call, and translate the input sentence text of the latest received input language based on the target translation language to obtain the real-time voice call translation result.

[0060] In this embodiment, the voice data is the voice data of the party to be translated obtained during the real-time voice call.

[0061] In this embodiment, the sentence break flag is the flag between adjacent sentences in the voice data, and its manifestation forms include: the pause duration of the voice signal and the interval duration between adjacent sentences determined according to the meaning of the voice text.

[0062] In this embodiment, the real-time sentence break flag recognition result includes two cases: there is a sentence break flag in the latest received voice data and there is no sentence break flag in the latest received voice data.

[0063] In this embodiment, obtaining the input sentence text of the latest received input language based on the latest recognized sentence break flag in the latest received voice data means:

[0064] Regarding the content between the last recognized sentence break flag and the latest recognized sentence break flag in the latest received voice data as the input sentence text of the latest received input language.

[0065] In this embodiment, translating the input sentence text of the latest received input language based on the target translation language to obtain the real-time voice call translation result means:

[0066] Translating the input sentence text of the latest received input language into the target translation language to obtain the real-time voice call translation result.

[0067] In this embodiment, the real-time voice call translation device includes an embedded computer module, a Bluetooth module, a WiFi module, a power supply and storage module, and text-to-speech software (which includes the above-mentioned sentence break flag recognition module, input sentence acquisition module, and sentence text translation module);

[0068] The implementation principle and process of the device of the present invention:

[0069] 1. Call Preparation:

[0070] It needs to be used in conjunction with the mobile phone software.

[0071] The supporting mobile phone software needs to enable the dialing permission of the mobile phone system, call record permission and microphone permission.

[0072] When the user makes a call, it should be ensured that this device is in a Bluetooth connection state with the mobile phone software.

[0073] When the user makes a call in a different language, one party starts the software, pre-selects the language of the other party, and then dials or answers the call.

[0074] 2. During the cross-language voice call:

[0075] The software simultaneously monitors the voice input by the microphone and the voice information of the other party in the voice call network, packs them separately, and sends them to the translation device via Bluetooth.

[0076] This device receives the voice file transmitted by the software via Bluetooth, performs voice-to-text conversion through the built-in software, and automatically calls the corresponding offline / online translation software API according to the WiFi connection status.

[0077] Through the offline / online text translation software, the text file is translated between the languages of the two parties in the call.

[0078] The translated text file is converted into a voice file through the text-to-voice software and sent back to the mobile phone software via Bluetooth.

[0079] The mobile phone software plays the corresponding file through the microphone or uploads it to the voice call network.

[0080] The beneficial effects of the above technologies are as follows: The sentence break flag recognition module can judge the sentence break flags in the voice data in real time, improving the accuracy and timeliness of voice processing. The input sentence acquisition module can obtain the input sentence text based on the latest recognized sentence break flags, ensuring the integrity and accuracy of the input sentence. The sentence text translation module can translate the input sentence text into the target language for the latest input sentence text, quickly obtaining the real-time voice call translation result, improving the efficiency and practicality of translation. The overall solution helps to achieve smooth and accurate real-time voice call translation, meeting the needs of cross-language communication. It can effectively avoid translation errors caused by improper sentence breaking or inaccurate sentence acquisition, improving the translation quality.

[0081] Embodiment 2:

[0082] On the basis of Embodiment 1, the sentence break flag recognition module refers to Figure 2 , including:

[0083] A text translation sub-module, which is used to convert the real-time received voice data into text data form including pause flags, and obtain the real-time voice text data of the input language;

[0084] A sentence-breaking flag recognition sub-module, which is used to identify whether there is a sentence-breaking flag in the real-time voice text data of the input language as the real-time sentence-breaking flag recognition result.

[0085] In this embodiment, the input language is the language in the voice data that needs to be translated currently.

[0086] In this embodiment, the pause flag refers to the pause duration that appears in the voice signal.

[0087] The beneficial effects of the above technologies are as follows: The text translation sub-module can convert voice data into text data form including pause flags, providing a clearer and more accurate basis for subsequent processing. It helps to improve the quality and readability of converting voice data into text. The sentence-breaking flag recognition sub-module can accurately judge the sentence-breaking flags in the real-time voice text data, improving the accuracy of subsequent processing. It provides strong support for realizing efficient and accurate voice processing and related applications. It makes the entire voice processing process more standardized and intelligent.

[0088] Embodiment 3:

[0089] Based on Embodiment 2, the text translation sub-module includes:

[0090] A text translation unit, which is used to convert the real-time received voice data into text data form, and obtain the initial voice text data of the input language;

[0091] A pause flag recognition unit, which is used to recognize all pause flags in the real-time received voice data in real time;

[0092] A flag marking unit, which is used to mark all pause flags in the initial voice text data of the input language in real time, and obtain the real-time voice text data of the input language.

[0093] The beneficial effects of the above technologies are as follows: The text translation unit can convert real-time voice data into text data, laying a foundation for subsequent processing and improving the convenience of data processing. The pause flag recognition unit can recognize the pause flags in the voice in real time, which helps to more accurately grasp the rhythm and semantics of the voice. The flag marking unit marks the pause flags in the text data in real time, making the text more conform to the characteristics of the original voice, enhancing the expressiveness and accuracy of the text. The overall process can improve the quality and accuracy of voice-to-text conversion, providing more reliable data support for related applications. It helps to achieve a more natural and practical voice text processing effect for communication.

[0094] Embodiment 4:

[0095] Based on Embodiment 3, the pause flag recognition unit includes:

[0096] A curve fitting subunit, configured to fit a short-time energy curve of the input speech signal in real time based on the short-time energy values of the speech signal of the real-time received speech data in each preset short time period;

[0097] An interval recognition subunit, configured to recognize all short-time energy abnormal curve intervals in the short-time energy curve of the input speech signal in real time;

[0098] An endpoint moment determination subunit, configured to determine the start moment and the end moment of the pause flag for all short-time energy abnormal curve intervals in the short-time energy curve of the input speech signal;

[0099] A pause flag determination subunit, configured to determine all pause flags in the real-time received speech data based on the start moment and the end moment of the pause flag for all short-time energy abnormal curve intervals in the short-time energy curve of the input speech signal;

[0100] Wherein, all abnormal short time periods of each short-time energy abnormal curve interval include the left abnormal short time period and the right abnormal short time period of each short-time energy abnormal curve interval.

[0101] In this embodiment, the preset short time period is a time period with a preset shorter duration, such as twenty milliseconds.

[0102] In this embodiment, the short-time energy value is a parameter used to describe the energy magnitude of the speech signal in a short time.

[0103] In this embodiment, the short-time energy curve of the input speech signal is a curve obtained by connecting and fitting the short-time energy values of all preset short time periods in sequence.

[0104] In this embodiment, the short-time energy abnormal curve interval is the speech signal of part of the speech data during the continuous period when the short-time energy of the short-time energy curve of the input speech signal starts to drop sharply and the short-time energy stabilizes at the value after the sharp drop.

[0105] In this embodiment, the start moment of the pause flag is the start moment of the pause flag.

[0106] In this embodiment, the end moment of the pause flag is the end moment of the pause flag.

[0107] In this embodiment, determining all pause flags in the real-time received speech data based on the start moment and the end moment of the pause flag for all short-time energy abnormal curve intervals in the short-time energy curve of the input speech signal means:

[0108] From the start moment of the pause flag to the end moment of the pause flag in all short-time energy abnormal curve intervals in the short-time energy curve of the input speech signal, regard this period as all pause flags in the real-time received speech data.

[0109] The beneficial effects of the above technology are as follows: The curve fitting subunit can more intuitively display the energy change trend of the speech signal by fitting the short-time energy curve of the input speech signal. The interval recognition subunit can identify the short-time energy abnormal curve intervals in real time, which helps to accurately locate the areas where pauses may exist. The endpoint moment determination subunit can determine the start and end moments of the abnormal curve intervals, providing key time points for accurately judging pauses. The pause flag determination subunit determines the pause flags based on these moments, improving the accuracy and reliability of pause judgment. The overall solution can effectively identify the pause flags accurately from complex speech data, providing strong support for subsequent operations in speech processing.

[0110] Embodiment 5:

[0111] On the basis of Embodiment 4, the interval recognition subunit includes:

[0112] The first mean calculation end is used to determine all energy sudden drop curve segments in the short-time energy curve of the input speech signal and calculate the average short-time energy value of each energy sudden drop curve segment;

[0113] The second mean calculation end is used to calculate the average short-time energy value of all parts of the short-time energy curve of the input speech signal except all energy sudden drop curve segments as the reference average short-time energy value;

[0114] The sudden drop ratio calculation end is used to calculate the ratio of the average short-time energy value of each energy sudden drop curve segment to the reference average short-time energy value as the average energy sudden drop ratio of each energy sudden drop curve segment;

[0115] The curve segment screening end is used to regard all energy sudden drop curve segments with an average energy sudden drop ratio less than the preset threshold in all energy sudden drop curve segments as all short-time energy abnormal curve intervals in the short-time energy curve of the input speech signal.

[0116] In this embodiment, determining all energy sudden drop curve segments in the short-time energy curve of the input speech signal includes:

[0117] Identify all energy sudden drop points in the short-time energy curve of the input speech signal;

[0118] And in the short-time energy curve of the input voice signal, determine the longest time sequence period starting from each energy sudden drop point, where the ratio of the difference between the maximum short-time energy value and the minimum short-time energy value included to the absolute value of the difference between the short-time energy value at the corresponding energy sudden drop point and the short-time energy value at the previous adjacent sampling point of the energy sudden drop point is less than a preset ratio (such as one-tenth), and use it as the energy sudden drop curve segment.

[0119] In this embodiment, the average short-time energy value of the energy sudden drop curve segment is the average value of the short-time energy values of all sampling points within the energy sudden drop curve segment.

[0120] In this embodiment, calculating the average short-time energy value of all parts of the short-time energy curve of the input voice signal except all energy sudden drop curve segments is the average value of the short-time energy values of all sampling points within all parts of the short-time energy curve of the input voice signal except all energy sudden drop curve segments.

[0121] In this embodiment, the preset threshold is a threshold of the average energy sudden drop ratio prepared in advance for screening the abnormal curve interval of short-time energy.

[0122] The beneficial effects of the above technology are as follows: The first mean calculation end can determine the energy sudden drop curve segment and calculate its average short-time energy value, providing basic data for subsequent analysis. The second mean calculation end calculates the average short-time energy value except the energy sudden drop curve segment as a reference, making the analysis more comprehensive and comparable. The sudden drop ratio calculation end obtains the average energy sudden drop ratio of each energy sudden drop curve segment, which helps to quantify the degree of energy change. The curve segment screening end screens out the abnormal curve interval of short-time energy according to the preset threshold, improving the accuracy and scientificity of the judgment of the abnormal interval. The overall solution can accurately identify the abnormal interval from the short-time energy curve of the input voice signal, providing strong support for the optimization and improvement of voice processing.

[0123] Embodiment 6:

[0124] Based on Embodiment 4, the endpoint time determination sub-unit includes:

[0125] The endpoint time period determination end is used to regard the preset short time period corresponding to the short-time energy value at the left endpoint of each short-time energy abnormal curve interval as the left abnormal short time period, and at the same time, regard the preset short time period corresponding to the short-time energy value at the right endpoint of the short-time energy abnormal curve interval as the right abnormal short time period;

[0126] The endpoint moment determination end is used to identify the start moment of the pause flag in the left abnormal short time period of each short-time energy abnormal curve interval and, at the same time, identify the end moment of the pause flag in the right abnormal short time period of each short-time energy abnormal curve interval based on the left neighborhood speech signal segment and the right neighborhood speech signal segment of all abnormal short time periods in each short-time energy abnormal curve interval.

[0127] In this embodiment, the preset short time period corresponding to the short-time energy value at the left endpoint of the short-time energy abnormal curve interval is the preset short time period corresponding to the short-time energy value of the sampling point located at the leftmost side in the short-time energy abnormal curve interval.

[0128] In this embodiment, the preset short time period corresponding to the short-time energy value at the right endpoint of the short-time energy abnormal curve interval is the preset short time period corresponding to the short-time energy value of the sampling point located at the rightmost side in the short-time energy abnormal curve interval.

[0129] In this embodiment, the left neighborhood speech signal segment and the right neighborhood speech signal segment of the abnormal short time period are the speech signal segments with a preset duration that are adjacent to the abnormal short time period and are located on the left and right sides of the abnormal short time period respectively in the speech signal of the real-time received speech data.

[0130] The beneficial effects of the above technology are as follows: The endpoint time period determination end can accurately define the short time periods corresponding to the short-time energy values at the left and right endpoints of the short-time energy abnormal curve interval, providing a clear range for subsequent determination of the pause moment. It helps to more accurately locate the abnormal short time period and improve the accuracy of pause flag recognition. The endpoint moment determination end identifies the start and end moments of the pause flag by analyzing the neighborhood speech signal segments of the abnormal short time period, making the pause judgment more reliable. It can effectively and accurately find the start and end moments of the pause from the complex short-time energy abnormal curve, improving the accuracy and quality of speech processing. It provides more accurate pause information for speech-related applications, such as speech recognition, translation, etc., improving the user experience and application effect.

[0131] Embodiment 7:

[0132] Based on Embodiment 6, the endpoint moment determination end includes:

[0133] The time period merging sub-end is used to merge each abnormal short time period of the short-time energy abnormal curve interval with the time period corresponding to the left neighborhood speech signal segment and the time period corresponding to the right neighborhood speech signal segment to obtain the reference time period of each short-time energy abnormal curve interval;

[0134] The signal segment division sub - end is used to determine multiple sampling points in each abnormal short - time segment of each short - time energy abnormal curve interval at a preset interval, and based on each sampling point, divide the speech signal of the corresponding part of the speech data in the reference period of the corresponding short - time energy abnormal curve interval into a left - reference speech signal segment and a right - reference speech signal segment corresponding to the sampling point;

[0135] The zero - crossing rate difference sub - end is used to take the difference between the zero - crossing rate of the left - reference speech signal segment and the zero - crossing rate of the right - reference speech signal segment of each sampling point as the left - right zero - crossing rate difference of each sampling point;

[0136] The moment determination sub - end is used to take the sampling point with the largest left - right zero - crossing rate difference among all sampling points in the left abnormal short - time segment of each short - time energy abnormal curve interval as the start moment of the pause flag. At the same time, take the sampling point with the smallest left - right zero - crossing rate difference among all sampling points in the right abnormal short - time segment of each short - time energy abnormal curve interval as the start moment of the pause flag.

[0137] In this embodiment, merging each abnormal short - time segment of the short - time energy abnormal curve interval with the corresponding time periods of the left - neighborhood speech signal segment and the right - neighborhood speech signal segment means:

[0138] Connect the time periods of the left - neighborhood speech signal segment corresponding to each abnormal short - time segment of the short - time energy abnormal curve interval, this abnormal short - time segment, and the time periods of the right - neighborhood speech signal segment corresponding to it in sequence.

[0139] In this embodiment, the preset interval is a pre - prepared time interval for determining sampling points, for example, 20 milliseconds.

[0140] In this embodiment, dividing the speech signal of the corresponding part of the speech data in the reference period of the corresponding short - time energy abnormal curve interval into a left - reference speech signal segment and a right - reference speech signal segment corresponding to the sampling point means taking the speech signal segment on the left side of the sampling point in the speech signal of the corresponding part of the speech data in the reference period of the corresponding short - time energy abnormal curve interval as the left - reference speech signal segment, and the speech signal segment on the right side of the sampling point as the right - reference speech signal segment.

[0141] In this embodiment, the zero - crossing rate of the left (right) reference speech signal segment refers to the number of times the speech signal crosses the zero value per unit time, and its value is the ratio of the number of times the left (right) reference speech signal segment crosses the zero value to the duration of the left (right) reference speech signal segment.

[0142] The beneficial effects of the above technology are as follows: The time period merging sub-terminal provides a richer data basis for a more comprehensive analysis of pauses by merging abnormally short time periods and their adjacent voice signal segments. The signal segment division sub-terminal realizes the fine processing of voice data by determining sampling points and dividing voice signal segments, improving the accuracy of analysis. The zero-crossing rate difference sub-terminal calculates the left and right zero-crossing rate differences, providing an effective quantitative index for judging pause moments. The moment determination sub-terminal can accurately find the start and end moments of pause markers, greatly improving the accuracy and reliability of pause judgment. The overall solution can accurately identify the moments of pause markers from complex voice signals, helping to improve the quality and effect of voice processing.

[0143] Embodiment 8:

[0144] Based on Embodiment 2, the sentence segmentation marker recognition sub-module includes:

[0145] The word segmentation unit is used to input the real-time voice text data of the input language into the word segmentation model to obtain the word segmentation result of the real-time voice text data of the input language;

[0146] The duration determination unit is used to determine the duration of all pause markers in the real-time voice text data of the input language;

[0147] The sentence segmentation marker recognition and judgment unit is used to judge whether there is a sentence segmentation marker in the real-time voice text data of the input language as the real-time sentence segmentation marker recognition result based on the part of speech and meaning of each word included in the word segmentation result and the duration of all pause markers.

[0148] In this embodiment, the word segmentation model is trained in advance using a large amount of text data of the input language and the analysis results obtained after manually performing word segmentation operations on the foregoing text data as training samples. During the training process, the text data is used as the model input quantity, and the word segmentation result is used as the model output quantity.

[0149] In this embodiment, the word segmentation result includes the result of dividing the real-time voice text data into multiple words.

[0150] The beneficial effects of the above technology are as follows: The word segmentation unit uses the word segmentation model to perform word segmentation on the real-time voice text data, providing a more refined and accurate basis for subsequent analysis and processing. The duration determination unit can determine the duration of pause markers, providing important time dimension information for judging whether it is a sentence segmentation marker. The sentence segmentation marker recognition and judgment unit comprehensively considers part of speech, meaning, and pause duration, improving the accuracy and rationality of sentence segmentation marker recognition. It helps to more accurately understand and process the input voice text, improving the quality and effect of voice text processing. It provides more accurate sentence segmentation judgment for related voice applications, such as voice translation and speech-to-text, enhancing the practicality and reliability of the applications.

[0151] Example 9:

[0152] Based on Example 8, the sentence segmentation marker recognition and judgment unit includes:

[0153] A model acquisition subunit for acquiring a sentence segmentation marker recognition model;

[0154] A sentence segmentation marker judgment subunit for inputting the part-of-speech and meaning of all words included in the word segmentation result and the duration of all pause markers into the sentence segmentation marker recognition model to judge whether there is a sentence segmentation marker in the real-time speech text data of the input language as the real-time sentence segmentation marker recognition result.

[0155] In this embodiment, the sentence segmentation marker recognition model is pre-trained by using the part-of-speech and meaning of all words included in the word segmentation result of a large amount of real-time speech text data of the input language, and the duration of all pause markers, as well as the judgment result of whether there is a sentence segmentation marker in the real-time speech text data of the input language judged manually based on the foregoing features as training samples;

[0156] During the training process, the part-of-speech and meaning of all words included in the word segmentation result of the real-time speech text data of the input language and the duration of all pause markers are the model input quantities, and the judgment result of whether there is a sentence segmentation marker in the real-time speech text data of the input language judged manually based on the foregoing features is the model output quantity.

[0157] The beneficial effects of the above technologies are as follows: The model acquisition subunit acquires the sentence segmentation marker recognition model, providing an effective tool and basis for accurate judgment. The sentence segmentation marker judgment subunit improves the scientificity and accuracy of judgment by inputting multi-dimensional information such as part-of-speech, meaning, and pause duration into the model for judgment. With the help of a special model, the subjectivity and error of manual judgment can be reduced, and the reliability of sentence segmentation marker recognition can be improved. It helps to process a large amount of speech text data more efficiently, improving the processing speed and efficiency. It can adapt to the characteristics of different languages and contexts, enhancing the generality and applicability of sentence segmentation marker recognition.

[0158] Example 10:

[0159] The present invention provides a real-time voice call translation method, which is applied to the real-time voice call translation device described in any one of Examples 1 to 9, refer to Figure 3 , including:

[0160] S1: Real-time judge whether there is a sentence segmentation marker in the latest received voice data as the real-time sentence segmentation marker recognition result;

[0161] S2: When the real-time sentence segmentation flag recognition result indicates that there is a sentence segmentation flag in the latest received voice data, based on the latest recognized sentence segmentation flag in the latest received voice data, obtain the input sentence text of the latest received input language;

[0162] S3: Determine the target translation language according to the international area code of the incoming / outgoing call, and translate the input sentence text of the latest received input language based on the target translation language to obtain the real-time voice call translation result.

[0163] The beneficial effects of the above technology are as follows: S1 can judge the sentence segmentation flag in the voice data in real time, improving the accuracy and timeliness of voice processing. S2 can obtain the input sentence text based on the latest recognized sentence segmentation flag, ensuring the integrity and accuracy of the input sentence. S3 can translate the target language for the latest input sentence text to quickly obtain the real-time voice call translation result, improving the efficiency and practicality of translation. The overall solution helps to achieve smooth and accurate real-time voice call translation, meeting the needs of cross-language communication. It can effectively avoid translation errors caused by improper sentence segmentation or inaccurate sentence acquisition, improving the translation quality.

[0164] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.

Claims

1. A real-time voice call translation device, characterized in that, Including: The sentence segmentation flag recognition module is used to determine in real time whether there is a sentence segmentation flag in the latest received voice data as the real-time sentence segmentation flag recognition result. It also includes: The text translation sub-module, including: The pause flag recognition unit is used to recognize in real time all the pause flags in the received voice data, including: The curve fitting sub-unit is used to fit in real time the short-time energy curve of the input voice signal based on the short-time energy values of the voice signal of the received voice data in each preset short time period; The interval recognition sub-unit is used to recognize in real time all the short-time energy abnormal curve intervals in the short-time energy curve of the input voice signal; The end point time determination sub-unit is used to determine the start time and end time of the pause flags of all the short-time energy abnormal curve intervals in the short-time energy curve of the input voice signal; The pause flag determination sub-unit is used to determine all the pause flags in the received voice data in real time based on the start time and end time of the pause flags of all the short-time energy abnormal curve intervals in the short-time energy curve of the input voice signal; Among them, all the abnormal short time periods of each short-time energy abnormal curve interval include the left abnormal short time period and the right abnormal short time period of each short-time energy abnormal curve interval; The input sentence acquisition module is used to, when the real-time sentence segmentation flag recognition result is that there is a sentence segmentation flag in the latest received voice data, obtain the input sentence text of the latest received input language based on the latest recognized sentence segmentation flag in the latest received voice data; The sentence text translation module is used to determine the target translation language according to the international area code of the incoming / outgoing call, and translate the input sentence text of the latest received input language based on the target translation language to obtain the real-time voice call translation result.

2. The real-time voice call translation device according to claim 1, characterized in that, The sentence segmentation flag recognition module, including: The text translation sub-module is used to convert the received voice data in real time into a text data form containing pause flags to obtain the real-time voice text data of the input language; The sentence segmentation flag recognition sub-module is used to recognize whether there is a sentence segmentation flag in the real-time voice text data of the input language as the real-time sentence segmentation flag recognition result.

3. The real-time voice call translation device according to claim 2, wherein The text translation sub-module, including: The text translation unit is used to convert the received voice data in real time into a text data form to obtain the initial voice text data of the input language; The pause flag recognition unit is used to recognize in real time all the pause flags in the received voice data; The flag marking unit is used to mark all the pause flags in the initial voice text data of the input language in real time to obtain the real-time voice text data of the input language.

4. The real-time voice call translation device according to claim 1, wherein, The interval recognition sub-unit, including: The first mean calculation end is used to determine all the energy sudden drop curve segments in the short-time energy curve of the input voice signal and calculate the average short-time energy value of each energy sudden drop curve segment; The second mean calculation end is used to calculate the average short-time energy value of all the parts of the short-time energy curve of the input voice signal except all the energy sudden drop curve segments in the short-time energy curve of the input voice signal as the reference average short-time energy value; A sudden drop ratio calculation end, which is used to calculate the ratio of the average short-time energy value of each energy sudden drop curve segment to the reference average short-time energy value as the average energy sudden drop ratio of each energy sudden drop curve segment; A curve segment screening end, which is used to regard all energy sudden drop curve segments with an average energy sudden drop ratio less than a preset threshold among all energy sudden drop curve segments as all short-time energy abnormal curve intervals in the short-time energy curve of the input voice signal.

5. The real-time voice call translation device according to claim 1, wherein An endpoint moment determination subunit, including: An endpoint time period determination end, which is used to regard the preset short time period corresponding to the short-time energy value at the left endpoint of each short-time energy abnormal curve interval as the left abnormal short time period, and at the same time, regard the preset short time period corresponding to the short-time energy value at the right endpoint of the short-time energy abnormal curve interval as the right abnormal short time period; An endpoint moment determination end, which is used to identify the starting moment of the pause flag in the left abnormal short time period of each short-time energy abnormal curve interval based on the left neighborhood voice signal segment and the right neighborhood voice signal segment of all abnormal short time periods of each short-time energy abnormal curve interval, and at the same time, identify the ending moment of the pause flag in the right abnormal short time period of each short-time energy abnormal curve interval.

6. The real-time voice call translation device according to claim 5, wherein An endpoint moment determination end, including: A time period merging sub-end, which is used to merge each abnormal short time period of the short-time energy abnormal curve interval with the time periods corresponding to the left neighborhood voice signal segment and the right neighborhood voice signal segment corresponding thereto to obtain a reference time period for each short-time energy abnormal curve interval; A signal segment division sub-end, which is used to determine a plurality of sampling points at a preset interval in each abnormal short time period of each short-time energy abnormal curve interval, and based on each sampling point, divide the voice signal of the partial voice data corresponding to the reference time period of the corresponding short-time energy abnormal curve interval into a left reference voice signal segment and a right reference voice signal segment corresponding to the sampling point; A zero-crossing rate difference calculation sub-end, which is used to regard the difference between the zero-crossing rate of the left reference voice signal segment and the zero-crossing rate of the right reference voice signal segment of each sampling point as the left-right zero-crossing rate difference of each sampling point; A moment determination sub-end, which is used to regard the sampling point with the largest left-right zero-crossing rate difference among all sampling points in the left abnormal short time period of each short-time energy abnormal curve interval as the starting moment of the pause flag, and at the same time, regard the sampling point with the smallest left-right zero-crossing rate difference among all sampling points in the right abnormal short time period of each short-time energy abnormal curve interval as the starting moment of the pause flag.

7. The real-time voice call translation device according to claim 2, wherein, A sentence-breaking flag recognition sub-module, including: A word segmentation unit, which is used to input the real-time voice text data of the input language into a word segmentation model to obtain a word segmentation result of the real-time voice text data of the input language; A duration determination unit, which is used to determine the duration of all pause flags in the real-time voice text data of the input language; A sentence-breaking flag recognition judgment unit, which is used to judge whether there is a sentence-breaking flag in the real-time voice text data of the input language based on the part-of-speech and meaning of each word included in the word segmentation result and the duration of all pause flags as a real-time sentence-breaking flag recognition result.

8. The real-time voice call translation device according to claim 7, wherein A sentence-breaking flag recognition judgment unit, including: A model acquisition sub-unit, which is used to acquire a sentence-breaking flag recognition model; A sentence-breaking marker judgment sub-unit, which is used to input the part-of-speech and meaning of all words included in the word segmentation result and the duration of all pause markers into a sentence-breaking marker recognition model, and judge whether there is a sentence-breaking marker in the real-time speech text data of the input language as the real-time sentence-breaking marker recognition result.

9. A real-time voice call translation method, characterized in that, Applied to the real-time voice call translation device described in any one of claims 1 to 8, comprising: S1: Real-time judge whether there is a sentence-breaking marker in the latest received voice data as the real-time sentence-breaking marker recognition result; S2: When the real-time sentence-breaking marker recognition result is that there is a sentence-breaking marker in the latest received voice data, based on the latest recognized sentence-breaking marker in the latest received voice data, obtain the input sentence text of the input language of the latest received; S3: Determine the target translation language according to the international area code of the incoming / outgoing call, and translate the input sentence text of the input language of the latest received based on the target translation language to obtain the real-time voice call translation result.

Citation Information

Patent Citations

  • Automatic voice sentence segmentation algorithm based on cyclic neural network

    CN109887499A

  • Translation method, device and equipment and storage medium thereof

    CN119132305A