Multi-party conversation translation method and device, electronic equipment and storage medium

By setting overlapping regions between time windows, speaker identification and speech feature fusion are performed, solving the problem of inconsistent translation results in existing technologies and achieving real-time and semantic coherence in dialogue translation.

CN121393468BActive Publication Date: 2026-04-17BEIJING SUPERHEXA CENTURY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING SUPERHEXA CENTURY TECH CO LTD
Filing Date
2025-10-30
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing dialogue translation methods are insufficient in terms of the coherence of translation results, especially when switching time windows, which can easily lead to semantic fragmentation.

Method used

By setting overlapping areas between adjacent time windows, speaker identification is performed, and the speaker is selected based on the duration of the voice information. Voice features are then filtered and integrated to ensure the coherence of the translation results.

Benefits of technology

It achieves real-time and semantic coherence in dialogue translation, avoids abrupt changes in speech features when switching time windows, and improves the coherence of translation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121393468B_ABST
    Figure CN121393468B_ABST
Patent Text Reader

Abstract

The application provides a multi-party conversation translation method and device, electronic equipment and storage medium, and belongs to the technical field of voice processing. The method comprises the following steps: acquiring voice conversation information in a current time window; performing speaker identity recognition on the voice conversation information, and selecting a first main speaker in the current time window based on the time length of the voice information; dividing the first voice feature of the first main speaker into a first voice feature of a first time period and a first voice feature of a second time period; screening the second voice feature of the first time period from the second voice feature of the first main speaker in the previous time window, performing feature fusion on the first voice feature and the second voice feature of the first time period, and obtaining a first fused voice feature; performing voice recognition on the first voice feature of the second time period based on the first fused voice feature, and performing conversation translation. The multi-party conversation translation method and device, electronic equipment and storage medium provided by the application can improve the coherence of the translation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of speech processing technology, and more specifically, relates to a multi-party dialogue translation method and apparatus, electronic device, and storage medium. Background Technology

[0002] With continuous technological advancements, dialogue translation capabilities can be integrated into wearable devices, bringing immense convenience to people. For example, in multi-person meetings, smart glasses accurately capture the speeches of participants speaking different languages ​​through a microphone array and translate them in real time into the user's specified language, eliminating the need for human interpreters or post-meeting transcription. This allows users to keep up with the meeting's pace and reduces information loss due to translation delays.

[0003] The coherence of existing dialogue translation methods needs further improvement. Summary of the Invention

[0004] The purpose of this application is to provide a multi-party dialogue translation method, apparatus, electronic device, and storage medium to improve the coherence of translation results.

[0005] A first aspect of this application provides a multi-party dialogue translation method, including:

[0006] Obtain voice dialogue information within the current time window; wherein, the current time window overlaps with the previous time window, and the previous time window is the time window preceding the current time window;

[0007] Speaker identification is performed on the voice dialogue information to obtain voice information corresponding to multiple speakers. Feature extraction is performed on the voice information corresponding to multiple speakers to obtain the first voice features corresponding to multiple speakers. The duration of the voice information of each speaker is counted. Based on the duration of the voice information, the first speaker in the current time window is selected from multiple speakers.

[0008] Based on the timestamp information corresponding to the first speech feature of the first speaker, the first speech feature of the first speaker is divided into the first speech feature corresponding to the first time period and the first speech feature corresponding to the second time period; wherein, the first time period is the time period in which there is a time overlap between the current time window and the previous time window, and the second time period is the time period in the current time window excluding the first time period;

[0009] From the second speech features of the first speaker in the previous time window, the second speech features corresponding to the first time period are selected. The first speech features corresponding to the first time period and the second speech features corresponding to the first time period are fused to obtain the first fused speech features. Speech recognition is performed on the first speech features corresponding to the second time period based on the first fused speech features.

[0010] The dialogue is translated and output based on the results of speech recognition.

[0011] A second aspect of this application provides a multi-party dialogue translation apparatus, comprising:

[0012] The data acquisition module is used to acquire voice dialogue information within the current time window; wherein, there is a time overlap between the current time window and the previous time window, and the previous time window is the time window preceding the current time window;

[0013] The speech separation module is used to identify the speaker in the speech dialogue information, obtain speech information corresponding to multiple speakers, extract features from the speech information corresponding to multiple speakers, obtain first speech features corresponding to multiple speakers, count the duration of each speaker's speech information, and select the first speaker in the current time window from multiple speakers based on the duration of the speech information.

[0014] The feature extraction module is used to divide the first speaker's first voice feature into a first voice feature corresponding to a first time period and a second voice feature corresponding to a second time period based on the timestamp information corresponding to the first voice feature of the first speaker; wherein, the first time period is the time period in which there is a time overlap between the current time window and the previous time window, and the second time period is the time period in the current time window excluding the first time period;

[0015] The feature fusion module is used to filter the second speech features corresponding to the first time period from the second speech features of the first speaker in the previous time window, fuse the first speech features corresponding to the first time period and the second speech features corresponding to the first time period to obtain the first fused speech features, and perform speech recognition on the first speech features corresponding to the second time period based on the first fused speech features.

[0016] The dialogue translation module is used to translate and output dialogues based on the results of speech recognition.

[0017] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the multi-party dialogue translation method described above.

[0018] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the multi-party dialogue translation method described above.

[0019] The beneficial effects of the multi-party dialogue translation method, apparatus, electronic device, and storage medium provided in this application are as follows:

[0020] This application embodiment acquires voice dialogue information through a sliding time window and translates and outputs the voice dialogue information within each time window in real time, enabling real-time dialogue translation. Furthermore, to avoid semantic fragmentation caused by the same sentence being divided into two time windows, an overlapping region is set between adjacent time windows. Speaker identification is performed on the voice dialogue information within the current time window to obtain voice information corresponding to multiple speakers. The first speaker is selected based on the duration of each speaker's voice information. On this basis, the second voice features of the first speaker in the overlapping period (first time period) of the previous time window can be filtered and fused with the first voice features of the first speaker in the overlapping period (first time period) of the current time window. This achieves a smooth transition between the two time windows, avoids abrupt changes in voice features due to window switching, and ensures the semantic coherence of the speaker.

[0021] Based on the first fused speech features obtained after fusion, speech recognition is performed on the first speech features of the first speaker in the non-overlapping time period (second time period) within the current time window, and dialogue translation is performed based on the speech recognition results, which can improve the coherence of the translation results. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 A flowchart illustrating a multi-party dialogue translation method provided in an embodiment of this application;

[0024] Figure 2 A structural block diagram of a multi-party dialogue translation device provided in an embodiment of this application;

[0025] Figure 3 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0026] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0027] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.

[0028] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a multi-party dialogue translation method provided in an embodiment of this application. The method can be executed by an electronic device and may include the following steps: S101-S105:

[0029] S101: Obtain voice dialogue information within the current time window; wherein, there is a time overlap between the current time window and the previous time window, and the previous time window is the time window preceding the current time window.

[0030] In this embodiment, voice dialogue information can be obtained through a sliding time window, and the voice dialogue information within each time window can be translated and output in real time, thereby realizing real-time translation of the dialogue and matching the immediacy requirements of the dialogue.

[0031] Furthermore, to avoid the same sentence being divided into two time windows, which would lead to semantic fragmentation, an overlapping area can be set between two adjacent time windows. For example, the previous time window is 0 to 2 seconds, the current time window is 1 to 3 seconds, and the two time windows overlap by 1 second.

[0032] S102: Perform speaker identification on the voice dialogue information to obtain the voice information corresponding to multiple speakers, extract features from the voice information corresponding to multiple speakers to obtain the first voice features corresponding to multiple speakers, count the duration of the voice information of each speaker, and select the first speaker in the current time window from multiple speakers based on the duration of the voice information.

[0033] In this embodiment, considering that each person's acoustic characteristics (including voiceprint, timbre and speaking rhythm) are different, speaker identification can be performed based on acoustic characteristics, and the voice dialogue information in the current time window can be divided into voice information corresponding to multiple speakers respectively.

[0034] Based on this, feature extraction is performed on the speech information corresponding to multiple speakers to obtain the first speech features corresponding to each speaker. Meanwhile, considering that translating all speakers' statements indiscriminately would lead to chaotic translation output (for example, if speaker A's long sentence is translated halfway through and speaker B's short sentence is suddenly inserted, it would cause users to lose comprehension), this embodiment counts the duration of each speaker's speech information, prioritizes speakers according to the duration of their speech information, and selects the speaker with the longest corresponding speech information as the first speaker within the current time window.

[0035] In this embodiment, considering that the first speaker's voice information can represent the core dialogue content within the current time window, only the first speaker's voice information is translated within the current time window to improve the real-time performance of the dialogue translation. In actual use, translation results from other speakers can also be inserted during the translation gaps in the speaker's speech as needed. The processing procedure for the other speakers' voice information is the same as that for the first speaker's voice information, and will not be elaborated here.

[0036] S103: Based on the timestamp information corresponding to the first speech feature of the first speaker, the first speech feature of the first speaker is divided into the first speech feature corresponding to the first time period and the first speech feature corresponding to the second time period; wherein, the first time period is the time period in which there is a time overlap between the current time window and the previous time window, and the second time period is the time period in the current time window excluding the first time period.

[0037] In this embodiment, the first speech feature of the first speaker can be divided into the first speech feature corresponding to the first time period (overlapping time period) and the first speech feature corresponding to the second time period (non-overlapping time period) based on the timestamp information. For example, if the current time window is 1 to 3 seconds, the first time period can be 1 to 2 seconds and the second time period can be 2 to 3 seconds.

[0038] S104: Select the second speech feature corresponding to the first time period from the second speech feature of the first speaker in the previous time window, fuse the first speech feature corresponding to the first time period and the second speech feature corresponding to the first time period to obtain the first fused speech feature; perform speech recognition on the first speech feature corresponding to the second time period based on the first fused speech feature.

[0039] In this embodiment, the same method as described above can be used to divide the second speech features of the first speaker within the previous time window into second speech features corresponding to overlapping time periods (i.e., the first time period) and second speech features corresponding to non-overlapping time periods. Then, feature fusion is performed on the first speech features corresponding to the first time period and the second speech features corresponding to the first time period. This can achieve a smooth transition between the two time windows, avoid abrupt changes in speech features due to window switching, and ensure the semantic coherence of the speaker.

[0040] Based on this, the fused speech features and the first speech features corresponding to the second time period can be concatenated in chronological order to form the speech feature sequence of the first speaker in the current time window. Automatic speech recognition (ASR) is then performed based on the speech feature sequence of the first speaker in the current time window, converting the speech feature sequence into text, and selecting the speech feature sequence in the second time period as the speech recognition result of the first speaker.

[0041] It should be noted that if the current time window is the first time window, the above steps S103~S104 will not be executed, and speech recognition will be performed directly based on the first speech feature of the first speaker within the first time window.

[0042] S105: Translate the dialogue based on the results of speech recognition and output the result.

[0043] In this embodiment, the text content identified in the above steps is input into the translation model, which can be converted into the target language (such as Chinese-English translation) and output the final translation result.

[0044] As can be seen from the above, this embodiment acquires voice dialogue information through a sliding time window and translates and outputs the voice dialogue information within each time window in real time, thus achieving real-time dialogue translation. Furthermore, to avoid semantic fragmentation caused by the same sentence being divided into two time windows, an overlapping region is set between adjacent time windows. Speaker identification is performed on the voice dialogue information within the current time window to obtain voice information corresponding to multiple speakers. The first speaker is selected based on the duration of each speaker's voice information. On this basis, the second voice feature of the first speaker in the overlapping period (first period) of the previous time window is filtered and fused with the first voice feature of the first speaker in the overlapping period (first period) of the current time window. This achieves a smooth transition between the two time windows, avoids abrupt changes in voice features due to window switching, and ensures the semantic coherence of the speaker.

[0045] Based on the first fused speech features obtained after fusion, speech recognition is performed on the first speech features of the first speaker in the non-overlapping time period (second time period) within the current time window, and dialogue translation is performed based on the speech recognition results, which can improve the coherence of the translation results.

[0046] In one embodiment of this application, if the first speaker in the current time window is not the same speaker as the second speaker in the previous time window, and the second speaker is included among the multiple speakers in the current time window, the multi-party dialogue translation method further includes:

[0047] Filter the third speech features corresponding to the first time period from the third speech features of the second speaker in the previous time window;

[0048] From the fourth speech features of the second speaker within the current time window, select the fourth speech features corresponding to the first time period and the fourth speech features corresponding to the second time period; fuse the third speech features corresponding to the first time period and the fourth speech features corresponding to the first time period to obtain the second fused speech features; perform speech recognition on the fourth speech features corresponding to the second time period based on the second fused speech features.

[0049] In this embodiment, the first speaker is the speaker within the time window, and the second speaker is the speaker within the previous time window. The identity of the second speaker within the previous time window and the third speech feature of the second speaker can be cached. If the first speaker in the current time window and the second speaker in the previous time window are not the same speaker, and the second speaker is included among the multiple speakers in the current time window, then the third speech feature corresponding to the first time period (overlapping time period) of the current time window is selected from the third speech feature of the second speaker in the previous time window, and the fourth speech feature corresponding to the first time period (overlapping time period) of the previous time window is selected from the fourth speech feature of the second speaker in the current window. The third speech feature corresponding to the first time period and the fourth speech feature corresponding to the first time period are fused to obtain the second fused speech feature, which can achieve a smooth transition of the second speaker's speech feature between the two windows.

[0050] Based on this, the second fused speech features and the fourth speech features corresponding to the second time period are concatenated in chronological order to form the speech feature sequence of the second speaker within the current time window. Speech recognition is performed based on the speech feature sequence of the second speaker within the current time window, and the speech feature sequence of the second time period is selected as the speech recognition result of the second speaker.

[0051] Finally, dialogue translation is performed based on the speech recognition results of the first speaker and the second speaker respectively, which can preserve contextual relevance and ensure overall semantic coherence.

[0052] As can be seen from the above, this embodiment can achieve a smooth transition of the second speaker's voice features between the two windows when the speaker switches from the second speaker to the first speaker, and can perform dialogue translation based on the voice recognition results of the first speaker and the second speaker respectively, thereby preserving contextual association and ensuring overall semantic coherence.

[0053] In one embodiment of this application, a first speech feature corresponding to a first time period and a second speech feature corresponding to the first time period are fused to obtain a first fused speech feature, including:

[0054] The first weight and the second weight are determined based on the signal-to-noise ratio of the voice dialogue information in the current time window and the signal-to-noise ratio of the voice dialogue information in the previous time window; wherein, the sum of the first weight and the second weight is 1;

[0055] The first weight is used as the weight of the first speech feature corresponding to the first time period, and the second weight is used as the weight of the second speech feature corresponding to the first time period. The first speech feature and the second speech feature corresponding to the first time period are weighted and summed to obtain the first fused speech feature.

[0056] In this embodiment, a weighted summation method can be used to fuse the first speech feature corresponding to the first time period and the second speech feature corresponding to the first time period. Specifically, the signal-to-noise ratio of the speech dialogue information in the current time window and the previous time window can be calculated respectively, and the first weight and the second weight can be determined based on the signal-to-noise ratio of the two windows.

[0057] For example, the first weight and the second weight can be calculated using the following formula:

[0058] ;

[0059] in, Indicates the first weight. Indicates the second weight. This indicates the signal-to-noise ratio of the voice dialogue information within the current time window. This indicates the signal-to-noise ratio of the voice dialogue information within the previous time window.

[0060] Based on the first and second weights determined by the above formula, feature fusion is performed on the first speech feature and the second speech feature corresponding to the first time period. This can fully utilize the features of the time window with a high signal-to-noise ratio. For example, if the signal-to-noise ratio of the current time window is high, its speech features are clearer and have a higher weight, which can reduce the impact of noise from the previous time window on the fusion result. Conversely, if the speech features of the previous time window are clearer, then the speech features of the previous time window will be the main focus, avoiding the noise of the current time window from contaminating the effective information.

[0061] In one embodiment of this application, speaker identification is performed on the voice dialogue information to obtain voice information corresponding to multiple speakers, including:

[0062] Acoustic features are extracted from the speech dialogue information, and speech is separated based on the acoustic features to obtain multiple speech segments; each speech segment contains the speech stream of a single speaker.

[0063] Cluster analysis is performed on multiple speech segments based on the acoustic features corresponding to each speech segment to obtain the speech information corresponding to multiple speakers.

[0064] In this embodiment, acoustic features may include Mel-frequency cepstral coefficients or log-Mel spectra. These acoustic features effectively characterize the spectral properties and temporal variations of speech and can be used to distinguish different speakers. Based on these acoustic features, a speech separation algorithm is used to separate mixed speech, resulting in multiple speech segments. Each speech segment contains only a continuous speech stream from a single speaker. For example, from mixed speech involving multiple speakers, the speech segments of speaker A and speaker B can be separated, achieving the decomposition from mixed speech to single-speaker speech segments. The speech separation algorithm can be implemented using existing nonnegative matrix factorization (NMF) or fully convolutional temporal audio separation networks (Conv-TasNet).

[0065] Based on multiple speech segments, cluster analysis can be performed on these segments according to their corresponding acoustic features. For example, methods such as Euclidean distance or cosine similarity can be used to calculate the feature similarity between speech segments. Speech segments with similar acoustic features can be grouped into one category (same speaker), while those with large feature differences can be placed into different categories (different speakers). All speech segments belonging to the same speaker can then be concatenated to obtain the complete speech information of that speaker. Other methods can also be used for cluster analysis of multiple speech segments, as detailed in the following examples.

[0066] As can be seen from the above, this embodiment first performs speech separation based on the acoustic features of the speech dialogue information to obtain multiple speech segments. Then, it performs cluster analysis on the multiple speech segments to achieve speech separation of multiple speakers and obtain the speech information corresponding to each speaker, thus providing a data foundation for subsequent accurate translation.

[0067] In one embodiment of this application, cluster analysis is performed on multiple speech segments based on the acoustic features corresponding to each speech segment to obtain speech information corresponding to multiple speakers, including:

[0068] Retrieve multiple historical cluster centers prior to the current time window; each historical cluster center corresponds to a different speaker;

[0069] For each speech segment, calculate the distance between the speech segment and each historical cluster center, and determine the historical cluster center with the smallest corresponding distance and the corresponding distance less than the distance threshold as the target cluster center for the speech segment, and assign the speech segment to the speaker corresponding to the target cluster center;

[0070] If there are speech segments without speaker identification, hierarchical clustering is performed on these segments to obtain new cluster centers; based on the new cluster centers, the speech segments without speaker identification are assigned to new speakers.

[0071] Multiple speech segments corresponding to each speaker are spliced ​​together to obtain speech information corresponding to each speaker.

[0072] In this embodiment, during the calculation of the first time window, multiple speech segments within the first time window can be hierarchically clustered to obtain multiple cluster centers. Each cluster center corresponds to a unique speaker (e.g., the historical centers of speakers A and B are C_A and C_B, respectively), forming a historical speaker feature database.

[0073] During subsequent time window calculations, the historical speaker feature database can be continuously updated. Specifically, for each speech segment within the current time window, the distance (such as Euclidean distance or cosine distance) between its acoustic features and all historical cluster centers is calculated. If a historical cluster center has the smallest distance to the speech segment, and this distance is less than a preset distance threshold (e.g., 5.0), then the speech segment is assigned to the speaker corresponding to that historical cluster center. For example, if a speech segment has the smallest distance to historical cluster center C_A, and this distance is less than the distance threshold, then the speech segment is assigned to speaker A.

[0074] If, after matching with the historical cluster centers, there are still speech segments that have not been assigned to any speaker, then hierarchical clustering is performed on these unassigned speech segments to generate new cluster centers (such as C_D, C_E). Each new cluster center corresponds to a new speaker, and based on each new cluster center, the unassigned speech segments are assigned to the new speaker. For example, if a speech segment corresponds to cluster center C_D, then that speech segment is assigned to speaker D.

[0075] After dividing multiple speech segments into speaker segments, the multiple speech segments corresponding to each speaker are spliced ​​together to obtain the speech information corresponding to each speaker.

[0076] As can be seen from the above, this embodiment prioritizes matching known speakers through historical cluster centers, which avoids redundant clustering calculations and ensures the stability of the speaker's identity across time windows. For speech segments that cannot be matched with historical cluster centers, hierarchical clustering is used to continuously add new cluster centers, enabling accurate identification of speakers who join midway through the speech.

[0077] In one embodiment of this application, the multi-party dialogue translation method further includes:

[0078] For each historical cluster center, if there are no speech segments assigned to that historical cluster center within M consecutive time windows, then that historical cluster center is deleted.

[0079] In this embodiment, considering that some people may leave midway through a multi-party dialogue, in order to further reduce the amount of computation, this embodiment can pre-set the number of comparisons M. If, within M consecutive time windows (e.g., M=20), a certain historical cluster center is not matched with any speech segment, it indicates that the speaker corresponding to the historical cluster center is likely no longer participating in the current dialogue. The historical cluster center can be deleted, and the matching of speech segments with the historical cluster center will no longer be performed, thereby reducing the amount of computation for matching.

[0080] In one embodiment of this application, the multi-party dialogue translation method further includes, for each historical cluster center:

[0081] The threshold range determined by the distance threshold is divided into segments to obtain multiple segment ranges;

[0082] For each segment range, a historical acoustic feature corresponding to a speech segment is selected, and the distance between the historical acoustic feature and the historical cluster center is within the range of that segment.

[0083] For each speech segment, calculate the distance between that speech segment and each historical cluster center, including:

[0084] Calculate the baseline distance between the speech segment and each historical cluster center;

[0085] Calculate the distance between the speech segment and each historical acoustic feature corresponding to each historical cluster center to obtain multiple derived distances between the speech segment and each historical cluster center;

[0086] The average of the baseline distance and multiple derived distances is used as the distance between the speech segment and the historical cluster center.

[0087] In this embodiment, considering that changes in a speaker's emotions can lead to differences in tone, or the impact of sudden noise interference, causing fluctuations in the speaker's acoustic characteristics—for example, a speaker's acoustic characteristics may be close to the cluster center when speaking normally, but deviate from the center when excited—then a single historical cluster center comparison method may lead to misjudgment of the cluster center, resulting in misidentification of the speaker's identity. To avoid this problem, this embodiment selects multiple historical acoustic features related to the historical cluster centers, and uses the historical cluster centers and multiple historical acoustic features together for speaker identification.

[0088] Specifically, taking Mel-frequency cepstral coefficients as acoustic features and Euclidean distance as the method to calculate the distance between a speech segment and each historical cluster center, the distance threshold can be set to 5.0, corresponding to a threshold range of 0~5.0. This threshold range is then segmented, for example, into five segments: 0~1.0, 1.0~2.0, 2.0~3.0, 3.0~4.0, and 4.0~5.0. Within each segment, a historical acoustic feature corresponding to a speech segment is selected as the typical feature of that segment. For example, if the distance between the historical acoustic feature corresponding to a certain acoustic segment and the corresponding historical cluster center C_A is 1.3, falling within the 1.0~2.0 segment, then the historical acoustic feature corresponding to this acoustic segment is taken as the typical feature of the 1.0~2.0 segment. Using the same method, the historical acoustic features corresponding to the historical cluster center C_A within each of the five segments can be obtained.

[0089] Based on this, when calculating the distance between each speech segment and each historical cluster center, we can first calculate the distance between the acoustic features corresponding to the speech segment and the historical cluster center, and use this distance as the baseline distance; then calculate the distance between the acoustic features corresponding to the speech segment and the historical acoustic features of each segment, and use this distance as the derived distance; finally, we use the average of the baseline distance and multiple derived distances as the distance between the speech segment and the historical cluster center.

[0090] As can be seen from the above, this embodiment divides the threshold range determined by the distance threshold into segments, selects typical historical acoustic features within each segment, and uses historical cluster centers and multiple historical acoustic features together for speaker identification, which helps to improve the accuracy of speaker identification.

[0091] In one embodiment of this application, the multi-party dialogue translation method further includes:

[0092] If the speaker is the same in N consecutive time windows before the current time window, then the length of the current time window is increased by the first step length.

[0093] In this embodiment, the initial length of the current time window can be preset, for example, 2 seconds. When the speaker remains consistent for N consecutive time windows (e.g., 10), it indicates that the current dialogue is in a state of continuous output by a single subject (e.g., product introduction, training, etc.). At this time, a first step length can be added to the initial length to increase the length of the current time window, thereby adapting to the natural rhythm of long, continuous speech and avoiding information fragmentation caused by fixed short time windows. The first step length is a preset constant, and those skilled in the art can set a specific value for the first step length according to actual needs, such as 2 seconds, 5 seconds, etc.

[0094] Corresponding to the multi-party dialogue translation method in the above embodiments, Figure 2 This is a structural block diagram of a multi-party dialogue translation apparatus provided according to an embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. References Figure 2 The multi-party dialogue translation device 20 includes: a data acquisition module 21, a speech separation module 22, a feature extraction module 23, a feature fusion module 24, and a dialogue translation module 25.

[0095] The data acquisition module 21 is used to acquire voice dialogue information within the current time window; wherein, there is a time overlap between the current time window and the previous time window, and the previous time window is the time window preceding the current time window;

[0096] The speech separation module 22 is used to identify the speaker in the speech dialogue information, obtain the speech information corresponding to multiple speakers, extract features from the speech information corresponding to multiple speakers, obtain the first speech features corresponding to multiple speakers, count the duration of the speech information of each speaker, and select the first speaker in the current time window from multiple speakers based on the duration of the speech information.

[0097] The feature extraction module 23 is used to divide the first speech feature of the first speaker into the first speech feature corresponding to the first time period and the first speech feature corresponding to the second time period based on the timestamp information corresponding to the first speech feature of the first speaker; wherein, the first time period is the time period in which there is a time overlap between the current time window and the previous time window, and the second time period is the time period in the current time window excluding the first time period;

[0098] Feature fusion module 24 is used to filter the second speech features corresponding to the first time period from the second speech features of the first speaker in the previous time window, fuse the first speech features corresponding to the first time period and the second speech features corresponding to the first time period to obtain the first fused speech features; and perform speech recognition on the first speech features corresponding to the second time period based on the first fused speech features.

[0099] The dialogue translation module 25 is used to translate and output dialogues based on the results of speech recognition.

[0100] In one embodiment of this application, if the first speaker in the current time window is not the same speaker as the second speaker in the previous time window, and the second speaker is included among the multiple speakers in the current time window, the feature fusion module 24 is further configured to:

[0101] Filter the third speech features corresponding to the first time period from the third speech features of the second speaker in the previous time window;

[0102] Filter the fourth speech features corresponding to the first time period and the fourth speech features of the second speaker within the current time window;

[0103] The third speech feature corresponding to the first time period and the fourth speech feature corresponding to the first time period are fused to obtain the second fused speech feature.

[0104] Speech recognition is performed based on the second fused speech features and the fourth speech features corresponding to the second time period.

[0105] In one embodiment of this application, the feature fusion module 24 is specifically used for:

[0106] The first speech feature corresponding to the first time period and the second speech feature corresponding to the first time period are fused to obtain the first fused speech feature, including:

[0107] The first weight and the second weight are determined based on the signal-to-noise ratio of the voice dialogue information in the current time window and the signal-to-noise ratio of the voice dialogue information in the previous time window; wherein, the sum of the first weight and the second weight is 1;

[0108] The first weight is used as the weight of the first speech feature corresponding to the first time period, and the second weight is used as the weight of the second speech feature corresponding to the first time period. The first speech feature and the second speech feature corresponding to the first time period are weighted and summed to obtain the first fused speech feature.

[0109] In one embodiment of this application, the speech separation module 22 is specifically used for:

[0110] Acoustic features are extracted from the speech dialogue information, and speech is separated based on the acoustic features to obtain multiple speech segments; each speech segment contains the speech stream of a single speaker.

[0111] Cluster analysis is performed on multiple speech segments based on the acoustic features corresponding to each speech segment to obtain the speech information corresponding to multiple speakers.

[0112] In one embodiment of this application, the speech separation module 22 is further configured to:

[0113] Retrieve multiple historical cluster centers prior to the current time window; each historical cluster center corresponds to a different speaker;

[0114] For each speech segment, calculate the distance between the speech segment and each historical cluster center, and determine the historical cluster center with the smallest corresponding distance and the corresponding distance less than the distance threshold as the target cluster center for the speech segment, and assign the speech segment to the speaker corresponding to the target cluster center;

[0115] If there are speech segments without speaker identification, hierarchical clustering is performed on these segments to obtain new cluster centers; based on the new cluster centers, the speech segments without speaker identification are assigned to new speakers.

[0116] Multiple speech segments corresponding to each speaker are spliced ​​together to obtain speech information corresponding to each speaker.

[0117] In one embodiment of this application, for each historical cluster center, the speech separation module 22 is further configured to:

[0118] The threshold range determined by the distance threshold is divided into segments to obtain multiple segment ranges;

[0119] For each segment range, a historical acoustic feature corresponding to a speech segment is selected, and the distance between the historical acoustic feature and the historical cluster center is within the range of that segment.

[0120] When calculating the distance between each speech segment and each historical cluster center, the speech separation module 22 is specifically used for:

[0121] Calculate the baseline distance between the speech segment and each historical cluster center;

[0122] Calculate the distance between the speech segment and each historical acoustic feature corresponding to each historical cluster center to obtain multiple derived distances between the speech segment and each historical cluster center;

[0123] The average of the baseline distance and multiple derived distances is used as the distance between the speech segment and the historical cluster center.

[0124] In one embodiment of this application, the data acquisition module 21 is specifically used for:

[0125] If the speaker is the same in N consecutive time windows before the current time window, then the length of the current time window is increased by the first step length.

[0126] See Figure 3 , Figure 3 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 3 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of each module / unit in the above-described device embodiments, for example... Figure 2 The functions of the data acquisition module 21, speech separation module 22, feature extraction module 23, feature fusion module 24, and dialogue translation module 25 are shown.

[0127] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0128] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.

[0129] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store preset constants such as distance thresholds and step lengths.

[0130] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation methods described in the multi-party dialogue translation method provided in the embodiments of this application, or they can execute the implementation methods of the electronic devices described in the embodiments of this application, which will not be elaborated here.

[0131] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0132] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0133] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0134] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0135] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connections shown or discussed may be indirect coupling or communication connections through some interfaces or units, or they may be electrical, mechanical, or other forms of connection.

[0136] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.

[0137] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0138] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A multi-party dialogue translation method, characterized in that, include: Obtain voice dialogue information within the current time window; wherein, the current time window overlaps with the previous time window, and the previous time window is the time window preceding the current time window; Speaker identification is performed on the voice dialogue information to obtain voice information corresponding to multiple speakers. Feature extraction is performed on the voice information corresponding to multiple speakers to obtain the first voice features corresponding to multiple speakers. The duration of the voice information of each speaker is counted. Based on the duration of the voice information, the first speaker in the current time window is selected from multiple speakers. Based on the timestamp information corresponding to the first speech feature of the first speaker, the first speech feature of the first speaker is divided into the first speech feature corresponding to the first time period and the first speech feature corresponding to the second time period; wherein, the first time period is the time period in which there is a time overlap between the current time window and the previous time window, and the second time period is the time period in the current time window excluding the first time period; From the second speech features of the first speaker in the previous time window, the second speech features corresponding to the first time period are selected. The first speech features corresponding to the first time period and the second speech features corresponding to the first time period are fused to obtain the first fused speech features. Speech recognition is performed on the first speech features corresponding to the second time period based on the first fused speech features. The dialogue is translated and output based on the results of speech recognition.

2. The multi-party dialogue translation method as described in claim 1, characterized in that, If the first speaker in the current time window is not the same speaker as the second speaker in the previous time window, and the second speaker is included among the multiple speakers in the current time window, the multi-party dialogue translation method further includes: Filter the third speech features corresponding to the first time period from the third speech features of the second speaker in the previous time window; Filter the fourth speech features corresponding to the first time period and the fourth speech features of the second speaker within the current time window; The third speech feature corresponding to the first time period and the fourth speech feature corresponding to the first time period are fused to obtain the second fused speech feature. Speech recognition is performed on the fourth speech feature corresponding to the second time period based on the second fused speech feature.

3. The multi-party dialogue translation method as described in claim 1, characterized in that, The first speech feature corresponding to the first time period and the second speech feature corresponding to the first time period are fused to obtain the first fused speech feature, including: A first weight and a second weight are determined based on the signal-to-noise ratio (SNR) of the voice dialogue information in the current time window and the SNR of the voice dialogue information in the previous time window; wherein the sum of the first weight and the second weight is 1. The first weight is used as the weight of the first speech feature corresponding to the first time period, and the second weight is used as the weight of the second speech feature corresponding to the first time period. The first speech feature and the second speech feature corresponding to the first time period are weighted and summed to obtain the first fused speech feature.

4. The multi-party dialogue translation method as described in claim 1, characterized in that, The step of identifying the speaker's identity from the voice dialogue information to obtain voice information corresponding to multiple speakers includes: Acoustic features are extracted from the speech dialogue information, and speech separation is performed on the speech dialogue information based on the acoustic features to obtain multiple speech segments; each speech segment contains the speech stream of a single speaker; Cluster analysis is performed on the multiple speech segments based on the acoustic features corresponding to each speech segment to obtain the speech information corresponding to multiple speakers.

5. The multi-party dialogue translation method as described in claim 4, characterized in that, The clustering analysis of the multiple speech segments based on the acoustic features corresponding to each speech segment yields speech information corresponding to multiple speakers, including: Retrieve multiple historical cluster centers prior to the current time window; each historical cluster center corresponds to a different speaker; For each speech segment, calculate the distance between the speech segment and each historical cluster center, and determine the historical cluster center with the smallest corresponding distance and the corresponding distance less than the distance threshold as the target cluster center for the speech segment, and assign the speech segment to the speaker corresponding to the target cluster center; If there are speech segments without speaker identification, hierarchical clustering is performed on these segments to obtain new cluster centers; based on these new cluster centers, the speech segments without speaker identification are assigned to new speakers. Multiple speech segments corresponding to each speaker are spliced ​​together to obtain speech information corresponding to each speaker.

6. The multi-party dialogue translation method as described in claim 5, characterized in that, For each historical cluster center, the multi-party dialogue translation method further includes: The threshold range determined by the distance threshold is divided into segments to obtain multiple segment ranges; For each segment range, a historical acoustic feature corresponding to a speech segment is selected, and the distance between the historical acoustic feature and the historical cluster center is within the segment range; The step of calculating the distance between each speech segment and each historical cluster center includes: Calculate the baseline distance between the speech segment and each historical cluster center; Calculate the distance between the speech segment and each historical acoustic feature corresponding to each historical cluster center to obtain multiple derived distances between the speech segment and each historical cluster center; The average of the baseline distance and the plurality of derived distances is used as the distance between the speech segment and the historical cluster center.

7. The multi-party dialogue translation method as described in claim 1, characterized in that, Also includes: If the speaker is the same in N consecutive time windows before the current time window, then the length of the current time window is increased by the first step length.

8. A multi-party dialogue translation device, characterized in that, include: The data acquisition module is used to acquire voice dialogue information within the current time window; wherein, there is a time overlap between the current time window and the previous time window, and the previous time window is the time window preceding the current time window; The speech separation module is used to identify the speaker in the speech dialogue information, obtain speech information corresponding to multiple speakers, extract features from the speech information corresponding to multiple speakers, obtain first speech features corresponding to multiple speakers, count the duration of each speaker's speech information, and select the first speaker in the current time window from multiple speakers based on the duration of the speech information. The feature extraction module is used to divide the first speaker's first voice feature into a first voice feature corresponding to a first time period and a second voice feature corresponding to a second time period based on the timestamp information corresponding to the first voice feature of the first speaker; wherein, the first time period is the time period in which there is a time overlap between the current time window and the previous time window, and the second time period is the time period in the current time window excluding the first time period; The feature fusion module is used to filter the second speech features corresponding to the first time period from the second speech features of the first speaker in the previous time window, fuse the first speech features corresponding to the first time period and the second speech features corresponding to the first time period to obtain the first fused speech features, and perform speech recognition based on the first fused speech features and the first speech features corresponding to the second time period. The dialogue translation module is used for dialogue translation based on the results of speech recognition.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Speech translation method and device and storage medium

    CN114765024A

  • Device and method for voice translation

    US20170286407A1