An intelligent real-time language synchronous translation system and its terminal

Through an intelligent real-time language synchronous translation system, the use of knowledge graph model and speech signal processing technology for speech segmentation and semantic optimization is solved, and the problem of insufficient accuracy and fluency in traditional translation technology is achieved, and higher quality cross-language communication is achieved.

CN119920244BActive Publication Date: 2025-06-24CHANGCHUN VOCATIONAL INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510405742.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-06-24
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

When traditional real-time language translation technology faces complex language structures, accent changes or unclear contexts, the accuracy of the translation is poor and the flexibility of adaptability to different speech scenarios is lacking, resulting in poor real-time and fluency of the translation.

Method used

An intelligent real-time language synchronous translation system is adopted, which includes a speech segmentation module, a scene recognition module, a syntax rhythm analysis module, a semantic optimization module and a speech translation module. Phrasic segmentation and semantic analysis are performed through the knowledge graph model, combined with scene recognition and syntactic rhythm analysis, semantic information is optimized and processed, and finally translated through the language translation model.

Benefits of technology

It significantly improves the accuracy and fluency of translation, can better adapt to different speech scenarios, solves semantic ambiguity problems, improves the accuracy and depth of semantic understanding, and provides higher quality translation services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119920244B_ABST
    Figure CN119920244B_ABST
Patent Text Reader

Abstract

This application relates to an intelligent real-time language synchronous translation system and its terminal. The system includes: a speech segmentation module for performing semantic analysis on the target speech to be translated, segmenting the target speech with the obtained semantic information to obtain speech units; a scene recognition module for analyzing the speech units to obtain scene information; a syntactic rhythm analysis module for converting the speech units into text and then performing syntactic analysis, and analyzing the speech units in combination with this information and the semantic information to obtain speech rhythm information and sentence structure information; a semantic optimization module for correcting the semantic information based on the sentence structure information, scene information, and speech rhythm information to obtain optimized semantic information; a speech translation module for translating the optimized semantic information based on the target language to obtain a target language translation. Through the cooperation of multiple modules, the system can accurately grasp the speech semantics, fit different scenarios, and effectively improve the accuracy, fluency, and practicality of translation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer and communication technologies, and particularly relates to an intelligent real-time language synchronous translation system and its terminal. Background Art

[0002] With the continuous advancement of the globalization process, communication and cooperation among countries have become increasingly frequent, and the demand for cross-language and cross-cultural exchanges has been growing. Language translation plays a crucial role in multiple fields such as international trade, diplomacy, tourism, education, and healthcare. It is not only a tool for building a bridge between different languages and cultures but also the key to achieving barrier-free information circulation and promoting understanding and cooperation.

[0003] However, traditional real-time language translation technologies usually rely on simple speech recognition and translation models. These methods generally first convert audio signals into text through speech recognition technology and then use translation models to convert the text into the target language. Although this method can achieve basic translation functions, when facing complex language structures, accent variations, or ambiguous contexts, the accuracy of translation is often affected, and it lacks the ability to flexibly adapt to different speech scenarios and cannot perform intelligent dynamic adjustments, resulting in poor real-time performance and fluency of translation. Summary of the Invention

[0004] Based on this, in view of the above technical problems, the object of the present invention is to provide an efficient and intelligent real-time language synchronous translation system to improve the accuracy and fluency of translation and meet the language translation needs in different scenarios.

[0005] In a first aspect, the present application provides an intelligent real-time language synchronous translation system, which includes:

[0006] A speech segmentation module, configured to perform semantic analysis on the target speech to be translated through a knowledge graph model to obtain semantic information, and segment the target speech based on the semantic information to obtain speech units;

[0007] A scene recognition module, configured to analyze the speech units to determine the speech scene where the current conversation is located and obtain scene information;

[0008] A syntactic rhythm analysis module, configured to convert the speech units into speech text, perform syntactic analysis on the speech text to obtain syntactic structure information, and analyze the speech units based on speech signal processing technology in combination with the syntactic structure information and semantic information to obtain speech rhythm information;

[0009] A semantic optimization module, configured to perform a correction operation on the semantic information based on the syntactic structure information, scene information, and speech rhythm information to obtain optimized semantic information, and the correction operation includes semantic adjustment, semantic screening, and semantic enhancement;

[0010] A speech translation module, which is used to input the optimized semantic information into a language translation model for translation based on the target language to obtain a target language translation.

[0011] In one embodiment, the speech segmentation module includes:

[0012] A speech segmentation subunit, which is used for:

[0013] Perform frame segmentation on the target speech to obtain speech frames;

[0014] Extract features based on each speech frame to obtain acoustic features, where the acoustic features include MFCC coefficients;

[0015] Perform phoneme recognition on the MFCC coefficients to obtain a phoneme sequence, and perform semantic constraint on the phoneme sequence based on semantic information to obtain an optimized phoneme sequence;

[0016] Construct a cost matrix according to the dynamic programming algorithm, and segment the optimized phoneme sequence to obtain speech units.

[0017] In one embodiment, the calculation formula of the cost matrix is as follows:

[0018] ;

[0019] Where, is the cost matrix, is the phoneme recognition accuracy function, is the semantic score function, is the acoustic feature continuity function of the target speech, , and are the weight coefficients of the corresponding functions respectively, represents the optimized phoneme sequence.

[0020] In one embodiment, the calculation formula of the phoneme recognition accuracy function is:

[0021] ;

[0022] Where, is the phoneme recognition accuracy function, represents the possibility that the kth factor is correctly recognized.

[0023] In one embodiment, the syntactic rhythm analysis module includes:

[0024] An information extraction subunit, which is used for:

[0025] Analyze the pause intervals in the speech units based on the endpoint detection algorithm of the energy threshold and the zero-crossing rate to obtain the pause duration;

[0026] Analyze the pitch change of the speech unit according to the autocorrelation function method to obtain the pitch change trajectory;

[0027] Calculate the speech rate information of the speech unit based on the syllable segmentation algorithm;

[0028] The dynamic analysis subunit is used to analyze the pause duration, pitch change trajectory, speech rate information, sentence structure information, and semantic information based on the dynamic programming algorithm to obtain the speech rhythm information.

[0029] In one embodiment, the speech segmentation module further includes:

[0030] The semantic extraction subunit is used for:

[0031] Based on the named entity recognition technology, perform entity recognition on the speech text to obtain key entities;

[0032] Identify and match the key entities through the knowledge graph model, and extract the relationships between the key entities to obtain relationship information;

[0033] Perform semantic reasoning on the key entities and relationship information through the knowledge graph model to obtain extended semantic information;

[0034] Combine the extended semantic information with the attention mechanism for importance evaluation to obtain optimized extended semantic information;

[0035] Fuse the optimized extended semantic information, key entities, and relationship information to obtain semantic information.

[0036] In one embodiment, the system further includes a translation evaluation module for:

[0037] Evaluate the target language translation according to the source language text of the target speech to obtain an evaluation result;

[0038] When the evaluation result is uncertain, send a correction instruction to the user, and the correction instruction includes displaying the problematic translation and providing correction options.

[0039] In a second aspect, the present application also provides an intelligent real-time language synchronous translation method, which includes:

[0040] Perform semantic analysis on the target speech to be translated through the knowledge graph model to obtain semantic information, and segment the target speech based on the semantic information to obtain speech units;

[0041] Analyze the speech unit to determine the speech scenario of the current conversation to obtain scenario information;

[0042] Convert the speech unit into speech text, perform syntactic analysis on the speech text to obtain sentence structure information, and analyze the speech unit based on speech signal processing technology, combining the sentence structure information and semantic information to obtain speech rhythm information;

[0043] Based on the sentence structure information, scene information, and speech rhythm information, perform a correction operation on the semantic information to obtain optimized semantic information. The correction operation includes semantic adjustment, semantic screening, and semantic enhancement;

[0044] Input the optimized semantic information into a language translation model based on the target language for translation to obtain the target language translation.

[0045] In a third aspect, the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the system according to any one of the first aspects.

[0046] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the system according to any one of the first aspects.

[0047] For the above intelligent real-time language synchronous translation system, method, computer device, and storage medium, the speech segmentation module uses the knowledge graph model to perform semantic analysis on the target speech to be translated, extracts semantic information, and segments the target speech based on this semantic information to obtain speech units. This module can ensure that the speech signal is scientifically decomposed and provides clear speech unit data for subsequent processing. Based on the speech units, through the scene recognition module, combined with keywords in the speech and background noise, etc., context analysis is performed to determine the specific speech scene where the current conversation is located, and scene information is generated, providing important context support for the accurate understanding of semantics. Secondly, the syntactic rhythm analysis module can convert the speech unit into speech text. On the one hand, syntactic analysis can be performed to clarify the sentence structure information. On the other hand, based on speech signal processing technology, combined with the sentence structure information and semantic information, the speech unit can be further analyzed to extract speech rhythm information. These information help the translation system better understand the pauses, accents, and tones in the speech signal, ensuring that the translation result conforms to the grammar and rhythm of the target speech.

[0048] The semantic optimization module can perform correction operations such as semantic adjustment, semantic screening, and semantic enhancement on semantic information based on sentence structure information, scene information, and speech rhythm information, which can ensure that the optimized semantic information is more in line with the actual context and expression requirements, thereby improving the accuracy and fluency of translation. Finally, the system uses the speech translation module to input the optimized semantic information into the language translation model for translation with the target language as the guide, and can generate high-quality target language translations to meet the needs of users for real-time language synchronous translation.

[0049] Compared with traditional speech translation systems, this system uses a knowledge graph model for speech segmentation, achieving more accurate division of speech units from a semantic perspective and improving the accuracy of basic processing. And through the collaborative work of multiple modules, it comprehensively considers multi-dimensional information such as scenes, syntax, and speech rhythm to optimize semantics, effectively solving the problem of semantic ambiguity, improving the accuracy and depth of semantic understanding, and then significantly improving the quality and fluency of translation, providing more reliable and efficient translation services for users' real-time language communication in different scenarios, and enhancing the intelligence level and practicality of the entire system. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0051] Figure 1 Schematic structural diagram of an intelligent real-time language synchronous translation system provided for an exemplary embodiment of the present invention;

[0052] Figure 2 Flowchart of an intelligent real-time language synchronous translation method provided for an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] In order to make the purpose, technical solutions, and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0054] In one embodiment, as Figure 1 shown, an intelligent real-time language synchronous translation system 100 is provided. In this embodiment, it is exemplified that the system is applied to a terminal. It can be understood that the system can also be applied to a server, or to a system including a terminal and a server, and is realized through the interaction between the terminal and the server. In this embodiment, the system includes:

[0055] The speech segmentation module 101 is used to perform semantic analysis on the target speech to be translated through the knowledge graph model to obtain semantic information, and to segment the target speech based on the semantic information to obtain speech units.

[0056] Specifically, the knowledge graph model contains a large amount of language knowledge, entity relationships, semantic associations and other information. By matching and reasoning the content in the target speech with this knowledge, the semantic information contained in the target speech can be obtained, such as determining the exact reference of a word in a sentence, understanding the logical association between words, etc., providing sufficient semantic basis for subsequent speech segmentation operations. The target speech can be segmented based on semantic information to obtain speech units. A speech unit refers to a relatively independent and complete speech segment in terms of semantics, which can be a complete sentence, phrase, or a phrase with specific semantics. Compared with traditional speech segmentation methods that often rely on acoustic features, such as physical properties such as energy and frequency of speech to divide speech segments, dividing speech units based on semantic information can better follow the logic and expression habits of the speech itself, and can avoid the problem of cutting off semantically coherent content or mixing different semantic contents that should be separated, which may occur when segmenting based solely on acoustic features, thereby improving the accuracy and effectiveness of the entire translation system and laying a good foundation for subsequent accurate translation.

[0057] The scene recognition module 102 is used to analyze the speech unit, determine the speech scene of the current dialogue, and obtain scene information.

[0058] In the scene recognition module 102, the speech unit can be analyzed from multiple angles, such as extracting keywords for recognition. There are often specific high-frequency words in different scenes, such as "menu", "order", "waiter" and other words may appear in the restaurant scene. And the background noise characteristics of the speech are recognized. For example, the office scene may be accompanied by keyboard tapping, meeting discussion, etc., and the outdoor scene will have background sounds such as vehicle driving and wind. In addition, the scene recognition module 102 can also be combined with other information, such as visual scene information recognized by the camera image, or location information determined by the positioning system, etc. for fusion recognition, and finally determine the specific speech scene of the current speech, and output the scene information. This scene information is very important for the subsequent accurate understanding of semantics, the selection of appropriate translation expressions, etc., because the same sentence may have different meanings and appropriate expressions in different scenes. This information helps the subsequent semantic optimization module to perform more accurate semantic screening and adjustment of ambiguous words and sentences, so that the translation results are more in line with the actual dialogue situation, improve the accuracy and practicality of the entire translation, make the translation more in line with the language expression habits in the corresponding scene, and better assist users in cross-language communication.

[0059] The syntactic rhythm analysis module 103 is used to convert a speech unit into speech text, perform syntactic analysis on the speech text to obtain sentence structure information, and analyze the speech unit based on speech signal processing technology in combination with the sentence structure information and semantic information to obtain speech rhythm information.

[0060] Specifically, the syntactic rhythm analysis module 103 can analyze and process the acoustic features (such as information on the frequency, energy, duration, etc. of speech) in the speech unit through speech recognition technology, identify and convert it into corresponding text content, so that more convenient and in-depth analysis can be carried out in text form subsequently. By performing syntactic analysis on the speech text according to grammar rules, the clear hierarchical relationship and grammatical structure between the parts of the sentence can be sorted out, thereby obtaining sentence structure information. This sentence structure information plays a key role in accurately understanding the sentence semantics, grasping the semantic focus, and subsequent processing in combination with other information. And based on speech signal processing technology in combination with the sentence structure information, the speech unit is analyzed. For example, the speech rhythm information can be obtained by analyzing acoustic features such as the pause duration, pitch change, and speech rate change in the speech unit signal. And by combining semantic information in this process, the analysis of the speech rhythm is made more in line with the actual meaning expression and grammatical logic of the language. For example, after determining key components such as the subject, predicate, and object based on the sentence structure, it can be analyzed how the rhythm characteristics in the corresponding speech parts of these components reflect semantic contents such as the sentence focus and modification relationship, and finally the speech rhythm information of the entire speech unit is accurately refined. These rhythm information have important impacts on subsequent semantic optimization and the shaping of the natural fluency of the translation, and can help the system generate a translation result that is more in line with the language expression habits and can better reflect the speaker's intention.

[0061] The semantic optimization module 104 is used to perform a correction operation on the semantic information based on the sentence structure information, scene information, and speech rhythm information to obtain optimized semantic information. The correction operation includes semantic adjustment, semantic screening, and semantic enhancement.

[0062] Specifically, the semantic optimization module 104 can comprehensively and specifically correct the original semantic information based on the sentence structure information, scene information, and speech rhythm information obtained by the above modules. Schematically, the semantic information can be adjusted by considering the grammatical relationships of the components in the sentence and the influence of the collocation rules between words on the semantics through the sentence structure information, excluding other semantic possibilities that do not conform to grammatical logic, so that the semantics of each word in the sentence is accurate and error-free, thus ensuring the accuracy of the semantic understanding of the entire sentence. Since the same word often has different meanings in different scenarios, in order to make the semantics more in line with the actual dialogue situation, the semantic information can be screened by combining the scene information to avoid the semantic ambiguity problem caused by polysemy, making the entire semantic information closely fit the specific dialogue scene, and then enhancing the rationality and accuracy of semantic understanding. The speech rhythm information can be used to comprehensively consider that in the process of expression, the speaker will emphasize certain content through some characteristic changes in speech (such as pitch elevation, speech rate slowdown, etc.), and these emphasized parts often have a more important position semantically, so as to strengthen the semantics, making the core semantic focus of the entire sentence clearer, and enabling the optimized semantic information to better reflect the true expression intention of the speaker.

[0063] The speech translation module 105 is used to input the optimized semantic information into a language translation model for translation based on the target language to obtain the target language translation.

[0064] The speech translation module 105 can input the optimized semantic information obtained after being processed layer by layer by multiple previous modules into the corresponding language translation model for translation operations according to the target language set by the user, so as to generate the final target language translation. The language translation model is trained based on a large number of bilingual or multilingual parallel corpora, such as corresponding sentences, phrases, etc. between the source language and the target language, learning the language patterns, vocabulary correspondence relationships, grammar structure conversions, etc. in these, and then building the mapping ability from one language to another. When the optimized semantic information is input into the language translation model, the model can go through a series of complex operations and conversion processes and finally output the target language translation. This process realizes the important function of helping users communicate across languages in real time and accurately.

[0065] The system first performs semantic analysis on the target speech to be translated through the speech segmentation module 101 to obtain semantic information, and based on this information, segments the target speech to obtain speech units. This process makes the division of speech units more in line with the internal logic and semantic structure of the language, effectively avoiding the semantic fragmentation problem that may be brought about by traditional speech segmentation methods, and providing a reliable guarantee for the efficient operation of subsequent modules. Through the scene recognition module 102, it can analyze the keyword information in the speech units, keenly capture features such as background noise, and judge the speech scene where the current conversation is located to obtain scene information. This step enables the system to flexibly adjust the semantic understanding strategy according to different scene environments, effectively cope with various complex and changeable communication scenarios, ensure that the translation result can highly match the actual context, avoid semantic deviation or translation errors caused by unclear scenes, and further improve the intelligence level and user experience of the entire translation system.

[0066] Secondly, in the syntactic rhythm analysis module 103, the system can convert the speech units into speech texts, perform syntactic analysis on them, extract sentence structure information, and clarify the grammatical relationships and hierarchical structures of each component in the sentence. Based on speech signal processing technology, this module can analyze the speech units by combining the sentence structure information and semantic information to obtain speech rhythm information. These information help to accurately grasp the semantic focus, logical level, and the emotional intention of the speaker of the sentence, etc., enabling the system to more subtly perceive the internal emotions and expression intentions of the language, further optimize the interpretation of semantics and the quality of translation, and improve the natural fluency and readability of the translation. In the semantic optimization module 104, the system can perform correction and optimization such as semantic adjustment, semantic screening, and semantic strengthening on the semantic information based on the sentence structure information, scene information, and speech rhythm information, effectively solving common problems such as semantic ambiguity and fuzziness in the language, so as to obtain optimized semantic information that is more in line with the expression habits of the target speech. Finally, through the speech translation module 105, the optimized semantic information can be input into the language translation model for translation according to the target language to obtain the target language translation, meeting the user's real-time and accurate language synchronous translation needs. This module combines all the foregoing processing results and realizes high-quality and real-time speech translation through a deep learning model.

[0067] Through the collaborative work of the above five modules, the system can accurately parse speech information in different scenarios, efficiently optimize semantic expressions, and quickly and accurately convert them into the target language, providing users with smooth, natural, and accurate translation services. Compared with traditional language translation systems, this system significantly improves the intelligence level of speech processing and the accuracy of translation through speech segmentation, scene judgment, grammar analysis, semantic optimization, and language translation. It effectively reduces translation errors caused by inaccurate semantic understanding, inappropriate scene adaptation, etc., enhances the practicality and reliability of the system in different language communication scenarios, and provides strong technical support for promoting cross-language communication and cooperation globally.

[0068] In an exemplary embodiment, the speech segmentation module includes:

[0069] A speech segmentation subunit for:

[0070] Performing frame segmentation on the target speech to obtain speech frames;

[0071] Extracting features based on each speech frame to obtain acoustic features, where the acoustic features include MFCC coefficients;

[0072] Performing phoneme recognition on the MFCC coefficients to obtain a phoneme sequence, and performing semantic constraint on the phoneme sequence based on semantic information to obtain an optimized phoneme sequence;

[0073] Constructing a cost matrix according to the dynamic programming algorithm, and segmenting the optimized phoneme sequence to obtain speech units.

[0074] Specifically, in the process of speech processing, in order to facilitate subsequent analysis and processing, it is first necessary to perform frame segmentation on the continuous target speech. By dividing it into shorter time segments, speech frames are obtained. Schematically, the target speech can be framed by determining the frame length and frame shift. The frame length can be selected to be around 20 - 30 milliseconds. The frame shift is generally less than the frame length, such as 10 - 15 milliseconds, which can represent the overlapping part between two adjacent frames. By determining an appropriate frame shift, it can be ensured that too much information of the target speech signal will not be lost during the frame segmentation process, enabling subsequent processing to capture speech features more comprehensively.

[0075] Since different speech contents have differences in acoustic performance, these differences can be quantified by extracting acoustic features in speech frames, such as MFCC coefficients (Mel-Frequency Cepstral Coefficients). MFCC coefficients can simulate the perceptual characteristics of the human ear for speech and characterize the acoustic properties of speech frames from multiple dimensions, such as the energy ratio of different frequency bands, spectral shape and other information. After obtaining the MFCC coefficients, they can be used as inputs to identify the phoneme sequence through a pre-trained phoneme recognition model. Schematically, a phoneme is the smallest speech unit that can distinguish meanings in speech. The phoneme recognition model can be constructed based on machine learning, for example, in a way that combines the hidden Markov model with a deep neural network. Through training, it can learn the relationship between the MFCC coefficients and the corresponding phonemes in a large number of speech samples. When the MFCC coefficients of a new speech frame are input, it can then output the corresponding phoneme sequence, that is, determine which phonemes this speech is composed of in sequence.

[0076] Since the phoneme sequence obtained relying on acoustic features may not be completely accurate or conform to the semantic logic of the language, semantic information can be introduced to adjust the initially recognized phoneme sequence, exclude unreasonable phoneme combinations, and select a phoneme arrangement that is more in line with the semantics, so as to obtain an optimized phoneme sequence, making the information at the phoneme level more consistent with the actual content to be expressed by the language, and preparing for more accurate segmentation of speech units in the future. Based on the obtained optimized phoneme sequence, the cost of splitting the optimized phoneme sequence into different speech units can be measured by comprehensively considering multiple factors, such as the accuracy of phoneme recognition, semantic rationality, and the continuity of the acoustic features of the speech signal, etc. A suitable quantization function is set for each factor, and weight coefficients are assigned according to the importance of different factors in the segmentation decision, so as to construct a cost matrix. Based on the dynamic programming algorithm, the path with the minimum cumulative cost, that is, the optimal speech unit segmentation method, can be found through this cost matrix. Splitting the optimized phoneme sequence according to this path can obtain speech units that conform to the actual semantics and acoustic features of the speech. These speech units lay a solid foundation for realizing high-quality intelligent real-time language synchronous translation services.

[0077] In an exemplary embodiment, the calculation formula of the cost matrix is as follows:

[0078] ;

[0079] where, is the cost matrix, is the phoneme recognition accuracy function, is the semantic score function, is the acoustic feature continuity function of the target speech, , and are the weight coefficients of the corresponding functions, indicating the optimized phoneme sequence.

[0080] The above cost matrix formula is constructed by comprehensively considering factors such as phoneme recognition accuracy, semantic rationality, and the continuity of acoustic features of speech signals. For each factor, a suitable quantization function is set, and weight coefficients are assigned according to their importance in the segmentation decision. This process helps to find the path with the minimum cumulative cost based on the dynamic programming algorithm, that is, the optimal speech unit segmentation method.

[0081] In an exemplary embodiment, the calculation formula of the phoneme recognition accuracy function is:

[0082] ;

[0083] where, is the phoneme recognition accuracy function, represents the possibility that the k-th factor is correctly recognized.

[0084] The above phoneme recognition accuracy function formula is determined by calculating the average of the possibilities that the factors from the i-th to the j-th are correctly recognized. Specifically, first sum up the possibilities of correct recognition for each factor from i to j, and then divide by the number of factors to obtain the average possibility of correct recognition.

[0085] In an exemplary embodiment, the syntactic rhythm analysis module includes:

[0086] An information extraction subunit for:

[0087] Analyzing the pause intervals in the speech unit based on the endpoint detection algorithm of energy threshold and zero-crossing rate to obtain the pause duration;

[0088] Analyzing the pitch change of the speech unit according to the autocorrelation function method to obtain the pitch change trajectory;

[0089] Calculating the speech rate information of the speech unit based on the syllable segmentation algorithm;

[0090] A dynamic analysis subunit for analyzing the pause duration, pitch change trajectory, speech rate information together with the sentence structure information and semantic information based on the dynamic programming algorithm to obtain the speech rhythm information.

[0091] Specifically, a voice signal has a certain amount of energy when it is being uttered, and the energy is lower during pauses. By setting an energy threshold, when the energy of the voice signal is below this threshold, it is considered to be in a pause state. And the zero-crossing rate can represent the number of times the voice signal crosses the zero axis per unit time. In a voice signal, voiced sounds usually have a lower zero-crossing rate, while unvoiced sounds and pauses have a higher zero-crossing rate. In the information extraction sub-unit, an endpoint detection algorithm that combines the energy threshold and the zero-crossing rate can be used to more accurately detect the endpoints of the voice, that is, to determine the start and end positions of the voice, and then find the pause intervals. By calculating the time length of the pause intervals, the pause duration can be obtained. The autocorrelation function can measure the similarity of a signal with itself at different time delays. And in voice signal analysis, pitch is related to the fundamental frequency of the voice signal. By using the autocorrelation function, the periodic characteristics of the voice signal can be found, and then the pitch can be determined. In the information extraction sub-unit, the autocorrelation function method can be applied frame by frame to the voice units to obtain the pitch values at each time point or time period, and these pitch values can be connected in chronological order to obtain the pitch change trajectory.

[0092] Schematically, a syllable is one of the basic units of speech, and the speech rate can be measured by the number of syllables per unit time. The syllable segmentation algorithm can determine the boundaries of syllables based on the acoustic features of the speech, such as energy, frequency, etc., and then segment the syllables. After the segmentation is completed, the speech rate can be calculated by counting the number of syllables within a certain period of time and combining the length of this period. By using the dynamic programming algorithm, the pause duration, pitch change trajectory, speech rate information, sentence structure information, and semantic information can be comprehensively considered to construct a state space containing these information, and the optimal combination method can be found through iterative calculation to determine the speech rhythm and obtain the speech rhythm information. This speech rhythm information can include the stress distribution, rhythm pattern, etc. in the sentence, which helps to understand the semantic focus and expression style of the sentence and makes the generated or recognized speech more natural and fluent.

[0093] In an exemplary embodiment, the speech segmentation module further includes:

[0094] A semantic extraction sub-unit, for:

[0095] Based on named entity recognition technology, perform entity recognition on the speech text to obtain key entities;

[0096] Through the knowledge graph model, identify and match the key entities, and extract the relationships between the key entities to obtain relationship information;

[0097] Perform semantic reasoning on the key entities and relationship information through the knowledge graph model to obtain extended semantic information;

[0098] Combine the extended semantic information with the attention mechanism for importance evaluation to obtain optimized extended semantic information;

[0099] Fuse the optimized extended semantic information, key entity and relationship information to obtain semantic information.

[0100] Specifically, named entity recognition can identify entities with specific meanings from speech texts, such as person names, place names, organization names, time, dates, currencies, etc., providing a basis for further semantic analysis and knowledge extraction. A knowledge graph is a graph-based data structure that can use entities as nodes and the relationships between entities as edges. Thus, it can identify and match the identified key entities, determine whether there are corresponding nodes for these entities in the pre-constructed knowledge graph, and analyze the connections between these key entities to obtain relationship information. This relationship information can reveal the hidden semantic structure in the text, helping to understand the logic and meaning of the entire text, such as causal relationships, ownership relationships, etc. By using the knowledge graph model and performing semantic reasoning based on the existing key entity and relationship information, the implicit semantic content that is not directly expressed but implied in the speech text can be mined to obtain extended semantic information. This extended semantic information can help fill in the possible semantic gaps in the text, enabling the system to process and analyze speech content from a more comprehensive perspective and providing richer semantic materials for subsequent language processing tasks. And through the attention mechanism, the importance of each element in the extended semantic information can be evaluated, removing some extended semantic content that may be less relevant or unimportant, making the subsequent semantic fusion and processing more efficient and accurate to obtain the extended semantic information. Finally, by fusing the optimized extended semantic information with the original key entity and relationship information, a complete semantic information with rich connotations can be formed.

[0101] In an exemplary embodiment, the system further includes a translation evaluation module for:

[0102] Evaluate the target language translation according to the source language text of the target speech to obtain an evaluation result;

[0103] When the evaluation result indicates uncertainty, send a correction instruction to the user, and the correction instruction includes displaying the problematic translation and providing correction options.

[0104] Schematically, the translation evaluation module can compare the source language text of the target speech with the target language translation through various methods. For example, through a statistics-based evaluation method, a large-scale bilingual parallel corpus can be used to calculate the similarity between the translation and a high-quality reference translation, that is, to evaluate the translation quality by calculating statistical indicators such as the matching probability of vocabulary and the coherence of phrases. When the evaluation result is uncertain, that is, there may be problems such as semantic inaccuracy, grammar errors, or inappropriate vocabulary in the translation, the system will send a correction instruction to the user. This correction instruction will first display the problematic translation, allowing the user to intuitively see where the problem lies, and can provide multiple correction options based on the system's own language knowledge and possible correction directions to assist the user in modifying the translation. This translation evaluation module can evaluate the translation based on the source language text and feedback the result, and send an instruction containing the problematic translation and correction options to the user when the result is uncertain, which helps to ensure the translation quality, assist the user in modification, and promote system optimization, improving the user experience and trust.

[0105] Based on the same inventive concept, as Figure 2 shown, an embodiment of the present application further provides an intelligent real-time language synchronous translation method, which includes:

[0106] S201: Perform semantic analysis on the target speech to be translated through a knowledge graph model to obtain semantic information, and segment the target speech based on the semantic information to obtain speech units;

[0107] S202: Analyze the speech units to determine the speech scenario of the current conversation to obtain scenario information;

[0108] S203: Convert the speech units into speech text, perform syntactic analysis on the speech text to obtain sentence structure information, and analyze the speech units based on speech signal processing technology, combining the sentence structure information and semantic information to obtain speech rhythm information;

[0109] S204: Based on the sentence structure information, scenario information, and speech rhythm information, perform a correction operation on the semantic information to obtain optimized semantic information, and the correction operation includes semantic adjustment, semantic screening, and semantic enhancement;

[0110] S205: Input the optimized semantic information into a language translation model for translation based on the target language to obtain a target language translation.

[0111] In the above intelligent real-time language synchronous translation method, first, the knowledge graph model is used to perform semantic analysis on the target speech to be translated, obtain semantic information, and segment the target speech based on the semantic information to obtain speech units. This process can effectively extract key information in the speech and convert it into speech units, providing a semantic basis for subsequent translation processing. By comprehensively considering various key elements in the speech units, such as keywords, background noise characteristics, etc., further analysis is carried out to judge the speech scenario of the current conversation and obtain scenario information. This information helps to determine the context behind the speech, thereby improving the accuracy and context adaptability of translation. Secondly, the speech units will be converted into speech texts for syntactic analysis to obtain sentence structure information. Based on speech signal processing technology, the system can analyze the speech units through the sentence structure information and semantic information, and further extract speech rhythm information. This speech rhythm information can reflect characteristics such as the speaking speed and pauses of the speaker, providing auxiliary support for the tone and expression mode of translation. By integrating these information, it is possible to better grasp the rhythm and semantics of the speech, more subtly perceive the inner emotions and expression intentions of the language, and further optimize the semantic interpretation and translation quality.

[0112] Finally, the semantic information can be corrected by combining the sentence structure information, scenario information, and speech rhythm information to obtain optimized semantic information. This process makes the semantics more accurate, clear, coherent, and logical, can more accurately reflect the true intention of the speaker, effectively solves common problems such as semantic ambiguity and fuzziness in language, provides a solid and reliable semantic basis for subsequent high-quality translation, and thus ensures the accuracy and coherence of the translation content. Based on the target language, the optimized semantic information is input into the language translation model for translation to obtain the translation of the target language. Through this series of steps, this method can achieve efficient and accurate real-time language synchronous translation, greatly improving the fluency and practicality of translation, and ensuring that the translation result conforms to the context.

[0113] Furthermore, segmenting the target speech based on the semantic information to obtain speech units includes:

[0114] Performing a framing operation on the target speech to obtain speech frames;

[0115] Performing feature extraction based on each speech frame to obtain acoustic features, and the acoustic features include MFCC coefficients;

[0116] Performing phoneme recognition on the MFCC coefficients to obtain a phoneme sequence, and performing semantic constraint on the phoneme sequence based on the semantic information to obtain an optimized phoneme sequence;

[0117] Constructing a cost matrix according to the dynamic programming algorithm, segmenting the optimized phoneme sequence, and obtaining speech units.

[0118] Furthermore, the calculation formula of the cost matrix is as follows:

[0119] ;

[0120] Wherein, is the cost matrix, is the phoneme recognition accuracy function, is the semantic score function, is the acoustic feature continuity function of the target speech, , and are the weight coefficients of the corresponding functions respectively, represents the optimized phoneme sequence.

[0121] Furthermore, the calculation formula of the phoneme recognition accuracy function is:

[0122] ;

[0123] Wherein, is the phoneme recognition accuracy function, represents the possibility that the k-th factor is correctly recognized.

[0124] Furthermore, according to the speech signal processing technology, the speech unit is analyzed by combining the sentence structure information and semantic information to obtain the speech rhythm information, including:

[0125] Analyze the pause interval in the speech unit based on the endpoint detection algorithm of energy threshold and zero-crossing rate to obtain the pause duration;

[0126] Analyze the pitch change of the speech unit according to the autocorrelation function method to obtain the pitch change trajectory;

[0127] Calculate the speech rate information of the speech unit based on the syllable segmentation algorithm;

[0128] Analyze the pause duration, pitch change trajectory, speech rate information with the sentence structure information and semantic information based on the dynamic programming algorithm to obtain the speech rhythm information.

[0129] Furthermore, the target speech to be translated is semantically analyzed through the knowledge graph model to obtain semantic information, including:

[0130] Based on the named entity recognition technology, entity recognition is performed on the speech text to obtain key entities;

[0131] The key entities are identified and matched through the knowledge graph model, and the relationships between the key entities are extracted to obtain relationship information;

[0132] The key entities and relationship information are semantically inferred through the knowledge graph model to obtain extended semantic information;

[0133] Combine the extended semantic information with the attention mechanism for importance evaluation to obtain optimized extended semantic information;

[0134] Fuse the optimized extended semantic information, key entity and relationship information to obtain semantic information.

[0135] Furthermore, the method further includes:

[0136] Evaluate the target language translation according to the source language text of the target voice to obtain an evaluation result;

[0137] When the evaluation result indicates uncertainty, send a correction instruction to the user. The correction instruction includes displaying the problematic translation and providing correction options.

[0138] In an exemplary embodiment, the present invention further provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of an intelligent real-time language synchronous translation system of the present application. A multi-core processor is preferred to improve the parallel processing ability of the system. Memory: Provide sufficient temporary storage space to support the operation of the program and the processing of data. The memory capacity should be large enough to accommodate a large amount of supply information and computing tasks.

[0139] In an exemplary embodiment, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of an intelligent real-time language synchronous translation system of the present application. The computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), solid state drive (SSD, Solid State Drives) or optical disc, etc. Among them, the random access memory may include resistive random access memory (ReRAM, Resistance Random Access Memory) and dynamic random access memory (DRAM, Dynamic Random Access Memory).

[0140] The above embodiments only represent several implementation manners of the embodiments of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the embodiments of the application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the embodiments of the present application.

Claims

1. An intelligent real-time language synchronous translation system, characterized in that: The system comprises: A speech segmentation module is used to perform semantic analysis on the target speech to be translated through a knowledge graph model to obtain semantic information, and to segment the target speech based on the semantic information to obtain speech units; A scene recognition module is used to analyze the speech unit, determine the speech scene of the current conversation, and obtain scene information; A syntactic rhythm analysis module, used for converting the speech unit into speech text, performing syntactic analysis on the speech text to obtain sentence structure information, and analyzing the speech unit in combination with the sentence structure information and the semantic information according to speech signal processing technology to obtain speech rhythm information; A semantic optimization module, used for performing a correction operation on the semantic information based on the sentence structure information, the scene information and the speech rhythm information to obtain optimized semantic information, wherein the correction operation includes semantic adjustment, semantic screening and semantic reinforcement; The speech translation module is used to input the optimized semantic information into a language translation model for translation based on the target language to obtain a target language translation.

2. The system according to claim 1, characterized in that The speech segmentation module comprises: The speech segmentation subunit is used to: Performing a frame operation on the target speech to obtain speech frames; Performing feature extraction based on each of the speech frames to obtain acoustic features, wherein the acoustic features include MFCC coefficients; Performing phoneme recognition on the MFCC coefficients to obtain a phoneme sequence, and performing semantic constraints on the phoneme sequence based on the semantic information to obtain an optimized phoneme sequence; A cost matrix is ​​constructed according to a dynamic programming algorithm, and the optimized phoneme sequence is segmented to obtain the speech unit.

3. The system according to claim 2, characterized in that The calculation formula of the cost matrix is ​​as follows: ; in, is the cost matrix, is the phoneme recognition accuracy function, is the semantic score function, is the acoustic feature continuity function of the target speech, , and are the weight coefficients of the corresponding functions, represents the optimized phoneme sequence.

4. The system according to claim 3, characterized in that The calculation formula of the phoneme recognition accuracy function is: ; in, is the phoneme recognition accuracy function, represents the probability that the kth factor is correctly identified.

5. The system according to claim 1, characterized in that The syntactic rhythm analysis module includes: The information extraction subunit is used to: Analyze the pause interval in the speech unit based on the endpoint detection algorithm of energy threshold and zero-crossing rate to obtain the pause duration; Analyzing the pitch change of the speech unit according to the autocorrelation function method to obtain a pitch change trajectory; Calculating the speech rate information of the speech unit based on the syllable segmentation algorithm; The dynamic analysis subunit is used to analyze the pause duration, the pitch change trajectory, the speech rate information, the sentence structure information and the semantic information based on a dynamic programming algorithm to obtain the speech rhythm information.

6. The system according to claim 1, characterized in that The speech segmentation module also includes: Semantic extraction subunit, used to: Based on named entity recognition technology, entity recognition is performed on the speech text to obtain key entities; The key entities are identified and matched by the knowledge graph model, and the relationships between the key entities are extracted to obtain relationship information; Perform semantic reasoning on the key entities and the relationship information through the knowledge graph model to obtain extended semantic information; The extended semantic information is combined with an attention mechanism to perform importance evaluation to obtain optimized extended semantic information; The optimized and extended semantic information, the key entity and the relationship information are integrated to obtain the semantic information.

7. The system according to claim 1, characterized in that The system further comprises a translation evaluation module, which is used to: Evaluate the target language translation according to the source language text of the target speech to obtain an evaluation result; When the evaluation result indicates that there is uncertainty, a correction instruction is sent to the user, wherein the correction instruction includes displaying the problematic translation and providing correction options.

8. An intelligent real-time language synchronous translation method, characterized in that: The method comprises: Performing semantic analysis on the target speech to be translated by using a knowledge graph model to obtain semantic information, and segmenting the target speech based on the semantic information to obtain speech units; Analyze the speech unit, determine the speech scene of the current dialogue, and obtain scene information; Converting the speech unit into speech text, performing syntactic analysis on the speech text to obtain sentence structure information, and analyzing the speech unit in combination with the sentence structure information and the semantic information according to speech signal processing technology to obtain speech rhythm information; Based on the sentence structure information, the scene information and the speech rhythm information, a correction operation is performed on the semantic information to obtain optimized semantic information, wherein the correction operation includes semantic adjustment, semantic screening and semantic reinforcement; The optimized semantic information is input into a language translation model for translation based on the target language to obtain a target language translation.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the system according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the system according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Speech recognition method for multiple languages

    CN118553231A

  • Method and device for translating object information and acquiring derivative information

    US20170293611A1