A cross-system speech real-time acquisition and bidirectional simultaneous interpretation method and device
Patent Information
- Application Number
- CN202610970939.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-25
AI Technical Summary
由于参与通信的双方通常使用不同通信平台进行语音交互,语音数据之间存在连续对话、多轮交互以及上下文关联特征,而现有技术往往将各语音片段独立处理,缺少对语音交互关系和会话演化过程的关联分析,导致不同语音内容之间的上下文联系难以有效保持
[0072]首先,本发明通过提取双向语音序列中的语音交互事件并构建会话交互图,利用改进型DyRep模型对会话交互关系进行动态关系演化分析,能够有效建立不同语音交互事件之间的关联关系,相较于现有技术对语音内容独立处理的方式,能够更好地保持会话连续性,提高后续语义理解的准确性。
Smart Images

Figure CN122821985A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech processing technology, and in particular to a method and apparatus for real-time cross-system speech acquisition and bidirectional simultaneous interpretation. Background Technology
[0002] With the increasing prevalence of cross-regional collaboration, international business exchanges, and remote conferencing applications, the demand for real-time voice communication between speakers of different languages is growing. To eliminate communication barriers caused by language differences, existing technologies typically employ speech recognition, machine translation, and speech synthesis to achieve real-time conversion between different languages, thereby enabling cross-language voice interaction.
[0003] In existing technologies, most cross-system speech translation solutions follow a process of speech acquisition, speech recognition, text translation, and speech synthesis. Since the parties involved in the communication typically use different communication platforms for voice interaction, there are continuous dialogue, multi-turn interactions, and contextual relationships among the speech data. However, existing technologies often process each speech segment independently, lacking correlation analysis of the speech interaction relationships and conversation evolution process, making it difficult to effectively maintain the contextual connections between different speech contents.
[0004] Meanwhile, existing technologies lack the ability to model the dynamic changes in session states, making it difficult to combine historical interaction content with current semantics for inference. When contextual references, referential delivery, and continuous question-and-answer scenarios occur during communication, semantic comprehension biases, inconsistencies in translation results, and poor continuity in two-way communication can easily arise, thus affecting the quality of cross-system voice communication and simultaneous interpretation.
[0005] Therefore, how to provide a method and apparatus for real-time cross-system speech acquisition and two-way simultaneous interpretation is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0006] One objective of this invention is to propose a method and apparatus for real-time cross-system speech acquisition and bidirectional simultaneous interpretation. This invention fully utilizes dynamic graph relation evolution analysis, speech recognition, state path search, cross-language semantic mapping, and speech synthesis technology. By constructing a conversation interaction graph and combining it with an improved DyRep model, the dynamic correlation relationships in the speech interaction process are analyzed. The Viterbi algorithm is used to achieve optimal semantic state path search, and the SeamlessM4T model is combined to complete cross-language semantic mapping and bidirectional semantic reconstruction. This enables real-time acquisition of cross-system speech data, continuous semantic understanding, and bidirectional simultaneous interpretation, and has the advantages of strong contextual continuity, high translation accuracy, smooth bidirectional interaction, and strong cross-platform adaptability.
[0007] A method for real-time cross-system speech acquisition and bidirectional simultaneous interpretation according to an embodiment of the present invention includes the following steps:
[0008] S1. Collect voice data from the communication platform participating in two-way voice communication at a preset frequency, and preprocess it to construct a two-way voice sequence;
[0009] S2. Perform voice interaction event extraction on the bidirectional voice sequence, calculate the interaction relationship between each voice interaction event, construct a conversation interaction graph, and use the improved DyRep model to perform dynamic relationship evolution analysis on the conversation interaction graph to obtain the conversation association sequence.
[0010] S3. Perform speech recognition processing on the bidirectional speech sequence, and construct a state transition path by combining the conversation association sequence. Use the Viterbi algorithm to perform optimal path search processing on the state transition path to construct a semantic text sequence.
[0011] S4. Perform context association analysis on the semantic text sequence to determine the context dependencies between each semantic text. Perform cross-language semantic mapping on each semantic text using the SeamlessM4T model, and perform bidirectional semantic reorganization based on the context dependencies to generate a bidirectional translation sequence.
[0012] S5. Perform speech synthesis processing on the bidirectional translation sequence to generate the target speech sequence, and send the target speech sequence to the corresponding communication platform.
[0013] Optionally, the communication platform refers to a communication carrier that participates in two-way voice communication and generates voice data, the two-way voice communication refers to the communication process of two-way transmission of voice information between two communication platforms, and the preprocessing includes noise reduction, silence filtering and voice framing processing.
[0014] Optionally, S2 specifically includes:
[0015] S21. Perform speech behavior detection processing on the two-way speech sequence, determine the speech time period, speech communication platform and response object corresponding to each speech behavior, and construct a speech interaction event set;
[0016] S22. Calculate the temporal and response relationships among the voice interaction events in the set of voice interaction events, and construct an interaction relationship set;
[0017] S23. Construct a conversation interaction graph based on the set of voice interaction events and the set of interaction relationships, where voice interaction events are used as graph nodes and interaction relationships are used as connecting edges.
[0018] S24. Input the conversation interaction graph into the improved DyRep model;
[0019] The improved DyRep model includes an event state encoding unit and a response chain evolution unit;
[0020] S25. In the event state encoding unit, event state encoding processing is performed on each graph node of the session interaction graph to obtain the event state sequence;
[0021] S26. Input the event state sequence into the response chain evolution unit, construct the session response chain according to the response association, and perform dynamic relationship evolution analysis processing based on the session response chain to generate the association evolution sequence.
[0022] S27. Calculate the association state corresponding to each graph node based on the association evolution sequence, and sort the association states according to the time order to construct the session association sequence.
[0023] Optionally, S25 specifically includes:
[0024] S251. In the event state encoding unit, extract the voice interaction events corresponding to each graph node in the conversation interaction graph;
[0025] S252. Construct an event state feature set based on the time correlation and response correlation of each voice interaction event;
[0026] S253. Use the EM algorithm to perform iterative estimation processing on the event state feature set and calculate the event state probability corresponding to each voice interaction event.
[0027] S254. Perform parameter update processing on the event state feature set based on the event state probability, and recalculate the event state probability corresponding to each voice interaction event.
[0028] S255. Determine whether the difference between the event state probabilities obtained from two adjacent iterations meets the preset condition:
[0029] If the conditions are not met, continue with iterative estimation and parameter update processing;
[0030] If satisfied, output the probability of the current event state;
[0031] S256. Based on the current event state probability, perform state encoding processing on each event state feature to generate the corresponding event state vector and construct the event state sequence.
[0032] Optionally, S26 specifically includes:
[0033] S261. In the response chain evolution unit, each event state vector in the event state sequence is read sequentially as the current state vector, and the state difference value and response correlation value between the current state vector and each non-current state vector are calculated.
[0034] S262. Construct a matching cost matrix based on the difference values of each state and the corresponding response correlation values;
[0035] S263. The Hungarian algorithm is used to perform optimal matching solution processing on the matching cost matrix to determine the optimal matching relationship between each event state vector.
[0036] S264. Establish response connection relationships between the event state vectors corresponding to each optimal matching relationship, and construct a session response chain;
[0037] S265. Perform correlation propagation analysis along the session response chain and calculate the propagation impact value corresponding to each event state vector.
[0038] S266. Based on the propagation impact value, perform dynamic association update processing on the state vectors of each event in the session response chain to generate an association evolution sequence.
[0039] Optionally, S3 specifically includes:
[0040] S31. Perform speech segmentation processing on the bidirectional speech sequence;
[0041] S32. Based on the preset semantic state library, perform speech recognition processing on each speech segment to determine the corresponding candidate semantic state set;
[0042] The semantic state library represents a pre-built set of semantic states, used to match the corresponding candidate semantic state set based on speech segments;
[0043] S33. Extract each associated state from the session association sequence, calculate the state association relationship between each associated state, and construct a state transition relationship set;
[0044] S34. Based on the candidate semantic state set and the state transition relationship set, establish the state transition path between each candidate semantic state;
[0045] S35. Use the Viterbi algorithm to perform optimal path search processing on each state transition path and calculate the path score of each state transition path.
[0046] S36. Determine the target path based on the path score, and perform semantic decoding processing on each candidate semantic state in the target path to generate the corresponding semantic text fragments.
[0047] S37. Perform association and recombination processing on each semantic text fragment in chronological order to construct a semantic text sequence.
[0048] Optionally, S4 specifically includes:
[0049] S41. Perform context association analysis on the semantic text sequence to calculate the context dependencies between each semantic text, specifically including:
[0050] Extract each semantic text from the semantic text sequence, analyze the related content between each semantic text, and establish the association relationship between semantic texts with related content;
[0051] Identify the contextual association paths between semantic texts along the association relationships, and determine the corresponding contextual dependencies;
[0052] S42. Using the SeamlessM4T model, cross-language semantic mapping and context consistency correction are performed on the semantic text sequence according to the context dependency to obtain the corrected semantic fragment.
[0053] S43. Perform bidirectional semantic reorganization processing on each corrected semantic segment according to the context dependency to construct a bidirectional translation sequence.
[0054] Optionally, S42 specifically includes:
[0055] S421. Perform cross-language semantic mapping processing on the semantic text sequence using the SeamlessM4T model to generate the target semantic fragment;
[0056] S422. Extract the semantic relationships between target semantic segments and construct a semantic rule set by combining contextual dependencies;
[0057] S423. Sequentially use each target semantic segment as the current semantic segment;
[0058] S424. Based on the semantic rule set, the RETE algorithm is used to perform rule matching processing on the current semantic segment;
[0059] S425. Identify target semantic fragments that do not meet the semantic rule set based on the rule matching results, and perform context consistency correction processing in combination with the corresponding context dependency relationship to generate corrected semantic fragments.
[0060] Optionally, S5 specifically includes:
[0061] S51. Extract the voice interaction events corresponding to the bidirectional translation sequence, determine the corresponding communication platform based on the voice interaction events, and obtain the voice transmission parameters of the corresponding communication platform.
[0062] The voice transmission parameters represent the voice transmission configuration parameters corresponding to the communication platform, including voice encoding format, bit rate, and transmission protocol parameters;
[0063] S52. Perform speech synthesis processing on the bidirectional translation sequence to generate the target speech sequence;
[0064] S53. Perform communication adaptation processing on the target speech sequence according to the speech transmission parameters, and send the adapted target speech sequence to the corresponding communication platform.
[0065] A cross-system real-time speech acquisition and two-way simultaneous interpretation device according to an embodiment of the present invention includes:
[0066] The data acquisition module is used to acquire voice data from the communication platform participating in two-way voice communication at a preset frequency, and to preprocess and construct two-way voice sequences.
[0067] The evolutionary analysis module is used to extract voice interaction events from bidirectional speech sequences, calculate the interaction relationships between each voice interaction event, construct a conversation interaction graph, and use an improved DyRep model to perform dynamic relationship evolution analysis on the conversation interaction graph to obtain the conversation association sequence.
[0068] The speech recognition module is used to perform speech recognition processing on bidirectional speech sequences, and construct state transition paths by combining them with conversation association sequences. The Viterbi algorithm is used to perform optimal path search processing on the state transition paths to construct semantic text sequences.
[0069] The semantic processing module is used to perform context association analysis on the semantic text sequence, determine the context dependencies between each semantic text, perform cross-language semantic mapping on each semantic text through the SeamlessM4T model, and perform bidirectional semantic reorganization based on the context dependencies to generate a bidirectional translation sequence.
[0070] The translation synthesis module is used to perform speech synthesis processing on the bidirectional translation sequence, generate the target speech sequence, and send the target speech sequence to the corresponding communication platform.
[0071] The beneficial effects of this invention are:
[0072] First, this invention extracts voice interaction events from bidirectional speech sequences and constructs a conversation interaction graph. It then uses an improved DyRep model to perform dynamic relationship evolution analysis on conversation interaction relationships, which can effectively establish the correlation between different voice interaction events. Compared with the existing technology that processes voice content independently, this invention can better maintain conversation continuity and improve the accuracy of subsequent semantic understanding.
[0073] Secondly, this invention constructs state transition paths by combining conversation association sequences and uses the Viterbi algorithm to search for the optimal path, enabling the speech recognition process to perform association analysis based on historical conversation states. At the same time, it combines contextual dependencies to perform cross-language semantic mapping and bidirectional semantic reorganization processing, which can improve the semantic coherence and contextual consistency of the translation results, thereby improving the accuracy of bidirectional simultaneous interpretation.
[0074] Finally, this invention performs communication adaptation processing on the target speech sequence according to the speech transmission parameters corresponding to the communication platform, so as to realize the accurate transmission of translation results back to the corresponding communication platform, improve the real-time performance and adaptation capability in the cross-system speech communication process, and thus improve the overall communication efficiency of cross-system two-way simultaneous interpretation. Attached Figure Description
[0075] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0076] Figure 1 This is a flowchart of a cross-system real-time speech acquisition and bidirectional simultaneous interpretation method proposed in this invention;
[0077] Figure 2 This is a flowchart of the conversation association sequence generation process for a cross-system real-time speech acquisition and bidirectional simultaneous interpretation method proposed in this invention.
[0078] Figure 3 This is a module structure diagram of a cross-system real-time voice acquisition and two-way simultaneous interpretation device proposed in this invention. Detailed Implementation
[0079] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0080] refer to Figures 1-2 A method for real-time cross-system speech acquisition and bidirectional simultaneous interpretation includes the following steps:
[0081] S1. Collect voice data from the communication platform participating in two-way voice communication at a preset frequency, and preprocess it to construct a two-way voice sequence;
[0082] S2. Perform voice interaction event extraction on the bidirectional voice sequence, calculate the interaction relationship between each voice interaction event, construct a conversation interaction graph, and use the improved DyRep model to perform dynamic relationship evolution analysis on the conversation interaction graph to obtain the conversation association sequence.
[0083] S3. Perform speech recognition processing on the bidirectional speech sequence, and construct a state transition path by combining the conversation association sequence. Use the Viterbi algorithm to perform optimal path search processing on the state transition path to construct a semantic text sequence.
[0084] S4. Perform context association analysis on the semantic text sequence to determine the context dependencies between each semantic text. Perform cross-language semantic mapping on each semantic text using the SeamlessM4T model, and perform bidirectional semantic reorganization based on the context dependencies to generate a bidirectional translation sequence.
[0085] S5. Perform speech synthesis processing on the bidirectional translation sequence to generate the target speech sequence, and send the target speech sequence to the corresponding communication platform.
[0086] In this embodiment, the communication platform refers to the communication carrier that participates in two-way voice communication and generates voice data. The communication platform can be different types of voice communication carriers, used to carry the real-time voice interaction process between the two parties.
[0087] Two-way voice communication refers to the communication process of transmitting voice information in both directions between two communication platforms. In actual communication, both communication platforms can act as the voice sender and voice receiver in the current voice interaction event, and continuously switch roles according to the content of the dialogue.
[0088] Preprocessing includes noise reduction, silence filtering, and speech framing.
[0089] After acquiring the voice data, the voice data is first subjected to noise reduction processing to suppress the impact of environmental noise, equipment background noise and transmission noise on the voice content.
[0090] Then, a silent segment filtering process is performed to filter out silent data segments that last for more than 300ms and have an energy value lower than a preset threshold.
[0091] After completing the silence segment filtering, the speech data is processed by speech framing in a 20ms time window, and continuous speech frames are generated using a 10ms frame shift length, thereby constructing a two-way speech sequence.
[0092] In this embodiment, S2 specifically includes:
[0093] S21. Perform speech behavior detection processing on the two-way speech sequence, determine the speech time period, speech communication platform and response object corresponding to each speech behavior, and construct a speech interaction event set; specifically, perform speech activity detection on continuous speech frames in the two-way speech sequence, when the duration of continuous effective speech exceeds 500ms, identify the corresponding speech segment as a speech behavior, and record the corresponding speech start time, speech end time, speech communication platform and corresponding response object, thereby forming a speech interaction event set;
[0094] S22. Calculate the time interval between adjacent voice interaction events based on the occurrence time sequence of each voice interaction event, and determine the response association relationship by combining the correspondence between the response object and the speaking communication platform. When the time interval between two voice interaction events is less than the preset time interval and there is a response correspondence relationship, establish the corresponding time association relationship and response association relationship, and construct an interaction association relationship set.
[0095] S23. Construct a conversation interaction graph based on the set of voice interaction events and the set of interaction relationships. Each graph node corresponds to a voice interaction event, and each connecting edge corresponds to the time relationship or response relationship between two graph nodes, thus forming a conversation interaction graph that can characterize the conversation evolution process.
[0096] S24. Input the conversation interaction graph into the improved DyRep model;
[0097] The improved DyRep model includes an event state encoding unit and a response chain evolution unit;
[0098] S25. In the event state encoding unit, event state encoding processing is performed on each graph node of the session interaction graph to obtain the event state sequence;
[0099] S26. Input the event state sequence into the response chain evolution unit, construct the session response chain according to the response association, and perform dynamic relationship evolution analysis processing based on the session response chain to generate the association evolution sequence.
[0100] S27. Calculate the association state corresponding to each graph node based on the association evolution sequence, and sort the association states according to the time order to construct the session association sequence.
[0101] In this embodiment, S25 specifically includes:
[0102] S251. In the event state encoding unit, each graph node in the session interaction graph is read sequentially, and the voice interaction event corresponding to each graph node is extracted. At the same time, the time association relationship and response association relationship corresponding to the voice interaction event are obtained.
[0103] S252. Associate and integrate the time correlation and response correlation of each voice interaction event, and construct the corresponding event state feature set;
[0104] S253. Use the EM algorithm to perform iterative estimation processing on the event state feature set and calculate the event state probability corresponding to each voice interaction event.
[0105] S254. Perform parameter update processing on the event state feature set according to the currently calculated event state probability, and recalculate the event state probability corresponding to each voice interaction event based on the updated event state feature set.
[0106] S255. Determine whether the difference between the event state probabilities obtained from two adjacent iterations meets the preset condition:
[0107] If the conditions are not met, continue with iterative estimation and parameter update processing;
[0108] If satisfied, output the probability of the current event state;
[0109] The preset condition is that the difference between the event state probabilities obtained from two adjacent iterations is less than a preset value.
[0110] S256. Perform state encoding processing on the corresponding event state features according to the current event state probability to generate the corresponding event state vector, and construct the event state sequence according to the time order of the voice interaction events corresponding to each graph node in the conversation interaction graph.
[0111] In this embodiment, S26 specifically includes:
[0112] S261. In the response chain evolution unit, each event state vector in the event state sequence is read sequentially as the current state vector, and the current state vector is compared with the other event state vectors in the event state sequence. The corresponding state difference value and response correlation value are calculated based on the temporal correlation and response correlation between the voice interaction events corresponding to each event state vector.
[0113] S262. Use the state difference value and the response correlation value as matrix elements in the matching cost matrix. The smaller the state difference value and the higher the response correlation value, the smaller the corresponding matching cost value, thereby constructing the matching cost matrix.
[0114] S263. The Hungarian algorithm is used to perform optimal matching solution processing on the matching cost matrix to determine the event state vector combination with the minimum matching cost, which is the corresponding optimal matching relationship.
[0115] S264. Based on the optimal matching relationship, establish response connection relationships between the corresponding event state vectors, and connect each response connection relationship in the order of the occurrence of the voice interaction events to construct a session response chain.
[0116] S265. Using the event state vectors in the session response chain as propagation nodes, perform correlation propagation analysis and processing step by step along the response connection relationship, and calculate the corresponding propagation impact value based on the correlation changes between the event state vectors during the propagation process.
[0117] S266. Update the associated states corresponding to each event state vector in the session response chain according to the propagation impact value, and record the change process of the updated associated states according to the propagation order in the session response chain, thereby generating an associated evolution sequence.
[0118] In this embodiment, S3 specifically includes:
[0119] S31. The bidirectional speech sequence is continuously divided according to the speech pause position. When a silent segment with a duration of more than 300ms occurs between adjacent speech contents, the corresponding position is used as the speech segment division position to obtain multiple speech segments.
[0120] S32. Input each speech segment into the semantic state database for matching analysis, and determine the corresponding candidate semantic state set based on the matching results. Each speech segment corresponds to multiple candidate semantic states.
[0121] The semantic state library represents a pre-built set of semantic states, which is used to match the corresponding candidate semantic state set based on speech segments;
[0122] S33. Read each associated state in the session association sequence in sequence, and calculate the corresponding state association relationship according to the previous and subsequent association relationship of each associated state in the session association sequence, thereby constructing a state transition relationship set;
[0123] S34. Take each candidate semantic state in the candidate semantic state set as a path node and the state association relationship in the state transition relationship set as a path connection relationship to establish the corresponding state transition path.
[0124] S35. Use the Viterbi algorithm to perform optimal path search processing on each state transition path and calculate the path score of each state transition path.
[0125] S36. Select the state transition path with the highest path score as the target path, and perform semantic decoding processing on each candidate semantic state in the target path to generate the corresponding semantic text fragments.
[0126] S37. Based on the time position of each speech segment in the bidirectional speech sequence, sort and reassemble the corresponding semantic text segments to construct a semantic text sequence.
[0127] In this embodiment, S4 specifically includes:
[0128] S41. Perform context association analysis on the semantic text sequence to calculate the context dependencies between each semantic text, specifically including:
[0129] Read each semantic text in the semantic text sequence in sequence and determine whether there is any common related content between different semantic texts. When two semantic texts correspond to the same related content, establish a relationship between the two semantic texts.
[0130] Starting from the relationships between each semantic text, the connections between the relationships are traced level by level, and corresponding contextual association paths are established. The contextual dependencies between each semantic text are determined based on the contextual association paths.
[0131] S42. Input the semantic text sequence and the corresponding context dependency into the SeamlessM4T model. Perform cross-language semantic mapping processing on each semantic text through the SeamlessM4T model, and perform context consistency correction processing on the cross-language semantic mapping result according to the context dependency to obtain the corrected semantic fragment.
[0132] S43. Determine the association order between segments based on the contextual dependencies corresponding to each corrected semantic segment, and perform bidirectional semantic reorganization processing on each corrected semantic segment according to the contextual association path to construct a bidirectional translation sequence.
[0133] In this embodiment, S42 specifically includes:
[0134] S421. Input the semantic text sequence into the SeamlessM4T model, identify the language type corresponding to each semantic text, determine the corresponding semantic mapping direction based on the identified language type, and perform cross-language semantic mapping processing to generate the target semantic fragment corresponding to each semantic text.
[0135] S422. Analyze the related content between each target semantic fragment, establish the corresponding semantic relationship, and then integrate the semantic relationship with the context dependency relationship to construct a semantic rule set;
[0136] S423. Read each target semantic segment sequentially as the current semantic segment according to the corresponding order of the target semantic segments in the semantic text sequence;
[0137] S424. Based on the semantic rule set, the RETE algorithm is used to perform rule matching processing on the current semantic segment to determine whether the current semantic segment satisfies the corresponding semantic rule;
[0138] S425. Filter target semantic segments that do not meet the semantic rule set according to the rule matching results, and perform context consistency correction processing on the target semantic segments according to the corresponding context dependency relationship to generate corrected semantic segments; for target semantic segments that meet the semantic rule set, they are directly output as corrected semantic segments.
[0139] In this embodiment, S5 specifically includes:
[0140] S51. Extract the voice interaction events corresponding to the bidirectional translation sequence, determine the corresponding communication platform based on the voice interaction events, and obtain the voice transmission parameters of the corresponding communication platform.
[0141] Voice transmission parameters represent the voice transmission configuration parameters corresponding to the communication platform, including voice encoding format, bit rate, and transmission protocol parameters;
[0142] S52. Perform speech synthesis processing on the bidirectional translation sequence to generate the target speech sequence;
[0143] S53. Perform communication adaptation processing on the target speech sequence according to the speech transmission parameters, perform encoding conversion processing on the target speech sequence according to the speech encoding format of the corresponding communication platform, perform bitrate adaptation processing on the target speech sequence according to the bitrate, and perform protocol encapsulation processing on the target speech sequence according to the transmission protocol parameters, thereby generating the adapted target speech sequence, and sending the adapted target speech sequence to the corresponding communication platform.
[0144] refer to Figure 3 A cross-system real-time speech acquisition and two-way simultaneous interpretation device, comprising:
[0145] The data acquisition module is used to acquire voice data from the communication platform participating in two-way voice communication at a preset frequency, and to preprocess and construct two-way voice sequences.
[0146] The evolutionary analysis module is used to extract voice interaction events from bidirectional speech sequences, calculate the interaction relationships between each voice interaction event, construct a conversation interaction graph, and use an improved DyRep model to perform dynamic relationship evolution analysis on the conversation interaction graph to obtain the conversation association sequence.
[0147] The speech recognition module is used to perform speech recognition processing on bidirectional speech sequences, and construct state transition paths by combining them with conversation association sequences. The Viterbi algorithm is used to perform optimal path search processing on the state transition paths to construct semantic text sequences.
[0148] The semantic processing module is used to perform context association analysis on the semantic text sequence, determine the context dependencies between each semantic text, perform cross-language semantic mapping on each semantic text through the SeamlessM4T model, and perform bidirectional semantic reorganization based on the context dependencies to generate a bidirectional translation sequence.
[0149] The translation synthesis module is used to perform speech synthesis processing on the bidirectional translation sequence, generate the target speech sequence, and send the target speech sequence to the corresponding communication platform.
[0150] Example 1: To verify the feasibility of this invention in practice, it was applied to a cross-language remote collaborative communication scenario. In this scenario, the two parties use different communication platforms to conduct real-time voice communication. One party participates in the communication through an office communication platform, while the other party participates through a remote collaborative communication platform. During their daily communication, both parties need to conduct continuous communication on topics such as technical issues, business communication, solution confirmation, and problem feedback. Because they use different languages, they need to utilize two-way simultaneous interpretation to achieve real-time communication. In actual communication, situations such as continuous question-and-answer sessions, contextual references, pronoun references, and topic switching frequently occur, placing high demands on translation accuracy, contextual continuity, and real-time response capabilities.
[0151] Traditional translation methods typically employ a processing model of speech recognition, text translation, and speech synthesis, with each speech segment processed independently, lacking dynamic correlation analysis of the conversation process. When referential expressions such as "this solution," "the above content," and "the current problem" appear during the dialogue, the system is prone to semantic comprehension bias because it cannot effectively identify the relationship between historical conversation content and the current statement. Furthermore, different communication platforms use different speech encoding formats and transmission mechanisms, which can easily lead to increased latency, discontinuous speech transitions, and inconsistencies in the translation results during transmission, impacting the efficiency of cross-language communication.
[0152] In this embodiment, the system first collects voice data from the communication platform participating in two-way voice communication at a preset frequency, and performs noise reduction processing, silence segment filtering processing, and voice framing processing to construct a two-way voice sequence. Subsequently, the system detects the speaking behavior in the two-way voice sequence, identifies the corresponding speaking communication platform, responding object, and speaking time information, and constructs a set of voice interaction events. Based on the temporal and response correlations between each voice interaction event, a session interaction graph is constructed. An improved DyRep model is used to perform evolutionary analysis on the dynamic interaction relationships in the session interaction graph to obtain a session association sequence, enabling the system to identify the correlations between different session contents and the session evolution process.
[0153] After obtaining the conversation association sequence, the system performs speech recognition processing on the bidirectional speech sequence and constructs a state transition path based on the conversation association sequence. It then uses the Viterbi algorithm to perform optimal path search processing to determine the optimal semantic state path and generate the corresponding semantic text sequence. When continuous question-and-answer sessions, contextual references, and omitted expressions occur during communication, the system can combine historical conversation states with the current semantic state to complete semantic recognition, improving the accuracy of semantic understanding.
[0154] Subsequently, the system performs contextual association analysis on the semantic text sequence, establishing contextual dependencies between the semantic texts, and performs cross-language semantic mapping through the SeamlessM4T model. Simultaneously, it combines contextual dependencies to complete bidirectional semantic reorganization and contextual consistency correction, ensuring the translation maintains semantic logical consistency. After translation, the system determines the corresponding communication platform based on the voice interaction events, obtains the voice transmission parameters of the corresponding communication platform, performs communication adaptation processing on the target voice sequence, and finally sends the target voice sequence to the corresponding communication platform in real time, achieving cross-system bidirectional simultaneous interpretation.
[0155] During the testing process, two communication platforms were connected, with 48 participants. The system ran continuously for 30 days, generating approximately 320 hours of voice data, collecting 523,600 voice segments, generating 12,680 voice interaction events, and constructing a conversation interaction graph with 12,680 nodes and 38,750 interaction-related edges. The system processed real-time voice data at a 20ms voice frame interval, generating an average of 1,630 voice frames per minute. A modified DyRep model was used to perform dynamic relationship evolution analysis on the conversation interaction graph, generating 84,200 conversation-related states; the Viterbi algorithm was used to construct 59,300 state transition paths; and the SeamlessM4T model was used to complete cross-language semantic mapping processing, generating 185,400 bidirectional translated sentences. The testing included 3,250 continuous question-and-answer scenarios, 2,860 contextual reference scenarios, 2,140 pronoun reference scenarios, and 1,780 cross-topic switching scenarios. To verify the actual effect of the present invention, it was compared with traditional real-time translation schemes, and the statistical results are shown in the table below.
[0156] Table 1. Comparison of Cross-System Two-Way Simultaneous Interpretation Effectiveness
[0157] Speech recognition accuracy 92.4% 97.8% Semantic understanding accuracy 88.7% 96.9% Contextuality retention rate 81.5% 95.8% Multi-turn question answering accuracy 83.2% 95.1% Refers to content recognition accuracy 79.6% 94.3% Translation accuracy 90.1% 97.2% Translation consistency 84.8% 96.4% Two-way conversation continuity score 82.7 points 96.1 points Average translation delay 1.62 seconds 0.71 seconds Communication platform adaptation success rate 91.3% 99.1% Translation result return success rate 94.6% 99.4% Accuracy of voice interaction event correlation recognition 80.9% 96.7% Session state recognition accuracy 84.1% 95.9% User satisfaction rating 86.5 points 97.3 points
[0158] As shown in Table 1, by constructing a conversation interaction graph and combining it with an improved DyRep model to perform dynamic relationship evolution analysis on voice interaction relationships, this invention can effectively identify the correlation between different voice interaction events during the conversation, thereby increasing the voice recognition accuracy from 92.4% to 97.8% and the semantic understanding accuracy from 88.7% to 96.9%.
[0159] Meanwhile, by constructing state transition paths through the Viterbi algorithm and combining them with contextual dependencies to perform cross-language semantic mapping and bidirectional semantic reorganization, the contextual association retention rate reached 95.8% and the multi-turn question answering recognition accuracy reached 95.1%, significantly improving the semantic understanding ability in continuous conversation scenarios.
[0160] Furthermore, by performing communication adaptation processing for different communication platforms, the communication platform adaptation success rate reached 99.1%, the translation result return success rate reached 99.4%, and the average translation latency was reduced to 0.71 seconds.
[0161] This demonstrates that the present invention can effectively solve the problems of insufficient contextual continuity, poor consistency of translation results, and insufficient cross-system adaptability in the prior art, and has high practical value and application promotion value.
[0162] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for real-time cross-system speech acquisition and bidirectional simultaneous interpretation, characterized in that, Includes the following steps: S1. Collect voice data from the communication platform participating in two-way voice communication at a preset frequency, and preprocess it to construct a two-way voice sequence; S2. Perform voice interaction event extraction on the bidirectional voice sequence, calculate the interaction relationship between each voice interaction event, construct a conversation interaction graph, and use the improved DyRep model to perform dynamic relationship evolution analysis on the conversation interaction graph to obtain the conversation association sequence. S3. Perform speech recognition processing on the bidirectional speech sequence, and construct a state transition path by combining the conversation association sequence. Use the Viterbi algorithm to perform optimal path search processing on the state transition path to construct a semantic text sequence. S4. Perform context association analysis on the semantic text sequence to determine the context dependencies between each semantic text. Perform cross-language semantic mapping on each semantic text using the SeamlessM4T model, and perform bidirectional semantic reorganization based on the context dependencies to generate a bidirectional translation sequence. S5. Perform speech synthesis processing on the bidirectional translation sequence to generate the target speech sequence, and send the target speech sequence to the corresponding communication platform.
2. The method for real-time cross-system speech acquisition and bidirectional simultaneous interpretation according to claim 1, characterized in that, The communication platform refers to the communication carrier that participates in two-way voice communication and generates voice data. The two-way voice communication refers to the communication process of transmitting voice information bidirectionally between two communication platforms. The preprocessing includes noise reduction, silence filtering and voice framing processing.
3. The method for real-time cross-system speech acquisition and bidirectional simultaneous interpretation according to claim 1, characterized in that, S2 specifically includes: S21. Perform speech behavior detection processing on the two-way speech sequence, determine the speech time period, speech communication platform and response object corresponding to each speech behavior, and construct a speech interaction event set; S22. Calculate the temporal and response relationships among the voice interaction events in the voice interaction event set, and construct an interaction relationship set; S23. Construct a conversation interaction graph based on the set of voice interaction events and the set of interaction relationships, where voice interaction events are used as graph nodes and interaction relationships are used as connecting edges. S24. Input the conversation interaction graph into the improved DyRep model; The improved DyRep model includes an event state encoding unit and a response chain evolution unit; S25. In the event state encoding unit, event state encoding processing is performed on each graph node of the session interaction graph to obtain the event state sequence; S26. Input the event state sequence into the response chain evolution unit, construct the session response chain according to the response association, and perform dynamic relationship evolution analysis processing based on the session response chain to generate the association evolution sequence. S27. Calculate the association state corresponding to each graph node based on the association evolution sequence, and sort the association states according to the time order to construct the session association sequence.
4. The method for real-time cross-system speech acquisition and bidirectional simultaneous interpretation according to claim 3, characterized in that, Specifically, S25 includes: S251. In the event state encoding unit, extract the voice interaction events corresponding to each graph node in the conversation interaction graph; S252. Construct an event state feature set based on the time correlation and response correlation of each voice interaction event; S253. Use the EM algorithm to perform iterative estimation processing on the event state feature set and calculate the event state probability corresponding to each voice interaction event. S254. Perform parameter update processing on the event state feature set based on the event state probability, and recalculate the event state probability corresponding to each voice interaction event. S255. Determine whether the difference between the event state probabilities obtained from two adjacent iterations meets the preset condition: If the conditions are not met, continue with iterative estimation and parameter update processing; If satisfied, output the probability of the current event state; S256. Based on the current event state probability, perform state encoding processing on each event state feature to generate the corresponding event state vector and construct the event state sequence.
5. The method for real-time cross-system speech acquisition and bidirectional simultaneous interpretation according to claim 3, characterized in that, S26 specifically includes: S261. In the response chain evolution unit, each event state vector in the event state sequence is read sequentially as the current state vector, and the state difference value and response correlation value between the current state vector and each non-current state vector are calculated. S262. Construct a matching cost matrix based on the difference values of each state and the corresponding response correlation values; S263. The Hungarian algorithm is used to perform optimal matching solution processing on the matching cost matrix to determine the optimal matching relationship between each event state vector. S264. Establish response connection relationships between the event state vectors corresponding to each optimal matching relationship, and construct a session response chain; S265. Perform correlation propagation analysis along the session response chain and calculate the propagation impact value corresponding to each event state vector. S266. Based on the propagation impact value, perform dynamic association update processing on the state vectors of each event in the session response chain to generate an association evolution sequence.
6. The method for real-time cross-system speech acquisition and bidirectional simultaneous interpretation according to claim 1, characterized in that, S3 specifically includes: S31. Perform speech segmentation processing on the bidirectional speech sequence; S32. Based on the preset semantic state library, perform speech recognition processing on each speech segment to determine the corresponding candidate semantic state set; The semantic state library represents a pre-built set of semantic states, used to match the corresponding candidate semantic state set based on speech segments; S33. Extract each associated state from the session association sequence, calculate the state association relationship between each associated state, and construct a state transition relationship set; S34. Based on the candidate semantic state set and the state transition relationship set, establish the state transition path between each candidate semantic state; S35. Use the Viterbi algorithm to perform optimal path search processing on each state transition path and calculate the path score of each state transition path. S36. Determine the target path based on the path score, and perform semantic decoding processing on each candidate semantic state in the target path to generate the corresponding semantic text fragments. S37. Perform association and recombination processing on each semantic text fragment in chronological order to construct a semantic text sequence.
7. The method for real-time cross-system speech acquisition and bidirectional simultaneous interpretation according to claim 1, characterized in that, S4 specifically includes: S41. Perform context association analysis on the semantic text sequence to calculate the context dependencies between each semantic text, specifically including: Extract each semantic text from the semantic text sequence, analyze the related content between each semantic text, and establish the association relationship between semantic texts with related content; Identify the contextual association paths between semantic texts along the association relationships, and determine the corresponding contextual dependencies; S42. Using the SeamlessM4T model, cross-language semantic mapping and context consistency correction are performed on the semantic text sequence according to the context dependency to obtain the corrected semantic fragment. S43. Perform bidirectional semantic reorganization processing on each corrected semantic segment according to the context dependency to construct a bidirectional translation sequence.
8. The method for real-time cross-system speech acquisition and bidirectional simultaneous interpretation according to claim 7, characterized in that, S42 specifically includes: S421. Perform cross-language semantic mapping processing on the semantic text sequence using the SeamlessM4T model to generate the target semantic fragment; S422. Extract the semantic relationships between target semantic segments and construct a set of semantic rules by combining them with contextual dependencies; S423. Sequentially use each target semantic segment as the current semantic segment; S424. Based on the semantic rule set, the RETE algorithm is used to perform rule matching processing on the current semantic segment; S425. Identify target semantic fragments that do not meet the semantic rule set based on the rule matching results, and perform context consistency correction processing in combination with the corresponding context dependency relationship to generate corrected semantic fragments.
9. The method for real-time cross-system speech acquisition and bidirectional simultaneous interpretation according to claim 1, characterized in that, S5 specifically includes: S51. Extract the voice interaction events corresponding to the bidirectional translation sequence, determine the corresponding communication platform based on the voice interaction events, and obtain the voice transmission parameters of the corresponding communication platform. The voice transmission parameters represent the voice transmission configuration parameters corresponding to the communication platform, including voice encoding format, bit rate, and transmission protocol parameters; S52. Perform speech synthesis processing on the bidirectional translation sequence to generate the target speech sequence; S53. Perform communication adaptation processing on the target speech sequence according to the speech transmission parameters, and send the adapted target speech sequence to the corresponding communication platform.
10. A cross-system real-time speech acquisition and two-way simultaneous interpretation device, executing the cross-system real-time speech acquisition and two-way simultaneous interpretation method according to any one of claims 1 to 9, characterized in that, include: The data acquisition module is used to acquire voice data from the communication platform participating in two-way voice communication at a preset frequency, and to preprocess and construct two-way voice sequences. The evolutionary analysis module is used to extract voice interaction events from bidirectional speech sequences, calculate the interaction relationships between each voice interaction event, construct a conversation interaction graph, and use an improved DyRep model to perform dynamic relationship evolution analysis on the conversation interaction graph to obtain the conversation association sequence. The speech recognition module is used to perform speech recognition processing on bidirectional speech sequences, and construct state transition paths by combining them with conversation association sequences. The Viterbi algorithm is used to perform optimal path search processing on the state transition paths to construct semantic text sequences. The semantic processing module is used to perform context association analysis on the semantic text sequence, determine the context dependencies between each semantic text, perform cross-language semantic mapping on each semantic text through the SeamlessM4T model, and perform bidirectional semantic reorganization based on the context dependencies to generate a bidirectional translation sequence. The translation synthesis module is used to perform speech synthesis processing on the bidirectional translation sequence, generate the target speech sequence, and send the target speech sequence to the corresponding communication platform.