A multi-language translation method and terminal based on voiceprint recognition and speech separation
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN MAXVISION TECH
- Filing Date
- 2026-04-21
- Publication Date
- 2026-08-07
AI Technical Summary
[0004]本申请实施例的目的在于提供一种基于声纹识别与语音分离的多语言翻译方法及终端,以解决现有技术面对多人动态场景下的实时翻译过程中存在的门控不准确的技术问题
[0050]本申请提供的基于声纹识别与语音分离的多语言翻译方法及终端的有益效果在于:与现有技术相比,通过引入声源定位轨迹并与声纹分离音轨进行时域相关性匹配,将空间维度信息作为身份验证的强约束,在多人、动态环境中显著提高了目标音轨的锁定和跟踪能力,解决了多人动态场景下的实时翻译过程中存在的门控不准确的技术问题。
Smart Images

Figure CN122090862B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of translation machine technology, and more specifically, relates to a multilingual translation method and terminal based on voiceprint recognition and speech separation. Background Technology
[0002] Homophonic translation uses voiceprint separation to obtain the pure audio track of the target speaker and uses it as a timbre reference for speech synthesis, which effectively improves the coherence and immersiveness of the translation broadcast.
[0003] However, existing homophonic translation schemes still face challenges in complex real-world applications. First, pure voiceprint recognition (VPR) suffers a significant drop in the quality of the separated audio tracks in noisy, low signal-to-noise ratio, or public settings with multiple speakers, leading to voiceprint matching failures and the incorrect filtering of key content from the target speaker, resulting in a higher rejection rate. Second, existing schemes cannot effectively handle changes in the speaker's location. When the target speaker moves or turns their head, the voiceprint features may fluctuate due to changes in the impulse response received by the microphone array, causing misjudgments by the gating system. Summary of the Invention
[0004] The purpose of this application is to provide a multilingual translation method and terminal based on voiceprint recognition and speech separation, so as to solve the technical problem of inaccurate gating in the real-time translation process of multiple people in dynamic scenarios in the prior art.
[0005] To achieve the above objectives, the technical solution adopted in this application is: to provide a multilingual translation method based on voiceprint recognition and speech separation, comprising the following steps:
[0006] Step S1: Acquire the initial mixed audio in real time, perform a first separation operation based on voiceprint recognition on the initial mixed audio to obtain several independent audio tracks; at the same time, perform a second separation operation based on sound source localization on the initial mixed audio to obtain several localization trajectories related to the time domain signal;
[0007] Step S2: Associate and match the several independent audio tracks with the several positioning trajectories to establish audio track-position track association pairs;
[0008] Step S3: Select the independent audio track of the target speaker from the audio track-bit track association pair as the target audio track;
[0009] Step S4: Using the spatial information of the positioning trajectory associated with the target audio track, perform spatial compensation on the speaker gating determination process of the target audio track;
[0010] Step S5: Based on the target audio track confirmed after spatial compensation as a timbre reference, generate real-time translated speech.
[0011] In a preferred embodiment, the spatial compensation in step S4 specifically includes:
[0012] When the similarity between the current audio block and the target speaker's reference voiceprint is lower than a preset threshold, if the source direction estimation result of the audio block deviates from the historical average azimuth angle of the target speaker by less than a first angle threshold, it is determined to be a spatial match, and the audio block is temporarily allowed to enter the speech recognition process.
[0013] In a preferred embodiment, the spatial compensation in step S4 specifically includes:
[0014] If the voiceprint similarity of the current audio block matches, but the estimated source direction of the audio block changes by more than the second angle threshold relative to the historical average azimuth angle of the target speaker, the output of the audio block is temporarily suppressed and the voiceprint re-registration process is triggered.
[0015] In a preferred embodiment, the association matching method in step S2 is as follows:
[0016] Extract the time-domain signal sequence features of the effective speech segments in the several independent audio tracks, perform correlation matching between the time-domain signal sequence features and the time-domain fluctuation features of the several positioning trajectories in the same time period, and determine the positioning trajectory with the highest correlation as the associated positioning trajectory of the independent audio track.
[0017] In a preferred embodiment, the method for correlation matching of time-domain fluctuation features in step S2 includes:
[0018] Each independent audio track is framed, and the short-time energy or logarithmic energy of each frame is calculated to form an energy envelope sequence E. track [n], and normalize it;
[0019] Differential operations are performed on each positioning trajectory to obtain the instantaneous angular velocity sequence ω[t], which is then smoothed and normalized to form the angular velocity fluctuation sequence Ω. traj [n];
[0020] For the energy envelope sequence E track [n] Perform speech activity detection, mark the start and end boundaries of speech segments for each independent audio track, and obtain the event set {V1,V2,…,V...} K};Regarding the angular velocity fluctuation sequence Ω traj [n] Perform zero-crossing and peak detection to mark the starting moments of significant fluctuations in angular velocity, obtaining the event set {W1, W2, ..., W...}. L}; Calculate the time overlap rate and event sequence consistency of two event sets to filter out candidate matching pairs;
[0021] Within the time window of the candidate matching pair, slide the window with a fixed step size and calculate the energy envelope sequence E track [n] and the angular velocity fluctuation sequence Ω traj [n] for the Pearson correlation coefficient r, and record the maximum correlation coefficient r max and its corresponding time delay τ opt ;
[0022] If any of the following conditions is met, it is determined that the independent audio track and the positioning trajectory are successfully matched:
[0023] Condition A: r max ≥T1 and ∣τ opt ∣≤T2, where T1 is the correlation coefficient threshold and T2 is the time delay threshold;
[0024] Condition B: r max <T1, but the overlap rate of the two event sets ≥T3 and the envelope peak phase difference of the two sequences ≤T4, where T3 is the overlap rate threshold and T4 is the phase difference threshold.
[0025] In a preferred embodiment, the method for selecting the target audio track in step S3 includes:
[0026] Receive the voice input of the questioner and extract the question keywords therefrom;
[0027] Predict the answer keywords according to the question keywords in the preset industry word library;
[0028] Perform real-time speech transcription on each independent audio track to obtain independent texts;
[0029] Select the independent audio track corresponding to the independent text that contains the answer keywords and has the highest matching degree as the target audio track.
[0030] In a preferred embodiment, the method further includes:
[0031] When obtaining the initial mixed audio, synchronously obtain the video stream and perform human target tracking and detection to obtain several human target tracking trajectories;
[0032] Perform spatial position association matching on several of the human target tracking trajectories and several of the positioning trajectories to establish a human - sound source - timbre correspondence;
[0033] Use the human target tracking trajectory to update the sound source positioning trajectory of the target speaker in real time to compensate for the missing positioning information during the period when the target speaker moves but does not speak.
[0034] In a preferred embodiment, the method for spatial position association matching includes:
[0035] The human target tracking trajectory and the positioning trajectory are projected onto the same spatial coordinate system, and the azimuth sequence and distance sequence of each trajectory are extracted respectively.
[0036] For each human body trajectory and each positioning trajectory, calculate the direction matching score S. dir And distance matching score S dist :
[0037]
[0038] Where θ diff θ is the average azimuth deviation. max The preset directional tolerance threshold;
[0039]
[0040] Where d diff d represents the average distance deviation. max The preset distance tolerance threshold;
[0041] Calculate the total matching score S:
[0042]
[0043] Where w dir With w dist w is the weighting coefficient. dir +w dist =1, and satisfies:
[0044]
[0045] in Estimating the variance of the direction for sound source localization. To estimate the variance of the distance;
[0046] When the total matching score S is greater than the preset matching threshold T5, it is determined that the human body trajectory is associated with the positioning trajectory, and a human body-sound source correspondence is established.
[0047] The aforementioned correspondence is then fused with the established audio track-position track association to form a human body-sound source-timbre triplet association.
[0048] In a preferred embodiment, the acquisition origin of the video stream is the same as the acquisition origin of the initial mixed audio, and the video stream is an RGBD video stream.
[0049] This application also provides a multilingual translation terminal based on voiceprint recognition and speech separation, including a main body, which is provided with a microphone array, a camera, a processing unit and a display screen, wherein the microphone array is used to collect initial mixed audio, the camera is used to synchronously acquire video streams, and the processing unit is used to execute the multilingual translation method based on voiceprint recognition and speech separation as described above.
[0050] The beneficial effects of the multilingual translation method and terminal based on voiceprint recognition and speech separation provided in this application are as follows: Compared with the prior art, by introducing the sound source localization trajectory and performing temporal correlation matching with the voiceprint-separated audio track, and using spatial dimension information as a strong constraint for identity verification, the locking and tracking capability of the target audio track is significantly improved in multi-person, dynamic environments, and the technical problem of inaccurate gating in real-time translation in multi-person dynamic scenarios is solved. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0052] Figure 1 This is a schematic diagram illustrating the application scenarios of the first and second embodiments of this application.
[0053] Figure 2 A flowchart of a multilingual translation method based on voiceprint recognition and speech separation provided in the first embodiment of this application.
[0054] Figure 3 This is a flowchart of a method for correlation matching of time-domain fluctuation features provided in the first embodiment of this application.
[0055] Figure 4 The flowchart shows the human body-sound source association and localization compensation method based on video stream provided in the second embodiment of this application.
[0056] Figure 5 A flowchart of a spatial location association matching method provided in the second embodiment of this application.
[0057] Figure 6 This is a three-dimensional structural diagram of a multilingual translation terminal based on voiceprint recognition and speech separation provided in the third embodiment of this application. Detailed Implementation
[0058] To make the technical problems, technical solutions, and beneficial effects to be solved by this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and are not intended to limit the scope of this application.
[0059] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0060] The multilingual translation method based on voiceprint recognition and speech separation provided in the first embodiment of this application will now be described. It is particularly suitable for multi-person dynamic scenarios in complex environments. For example, please refer to... Figure 1 In public places with multiple windows, not only does the tone, direction, and position of the target speaker change as they move from position a to position b, but there are also people nearby interfering with the conversation from the windows on both sides and the passage behind them.
[0061] Please refer to the following: Figure 2 and Figure 3 The multilingual translation method based on voiceprint recognition and speech separation includes:
[0062] Step S1: Acquire the initial mixed audio in real time, perform a first separation operation based on voiceprint recognition on the initial mixed audio to obtain several independent audio tracks; at the same time, perform a second separation operation based on sound source localization on the initial mixed audio to obtain several localization trajectories related to the time domain signal.
[0063] Understandably, the system acquires the initial mixed audio (e.g., 16kHz sampling rate, 16bit PCM format) in real time via a microphone array.
[0064] The first separation operation can employ a deep learning-based speech separation model, such as MossFormer2, to process the mixed audio and separate it into several independent audio streams, i.e., independent audio tracks. Each audio track theoretically corresponds to a primary acoustic source. Simultaneously, speaker embedding vectors are extracted for each audio track (e.g., 192-dimensional embeddings are extracted using ECAPA-TDNN).
[0065] The second separation operation can utilize the geometry of the microphone array to perform source localization (DOA) algorithms on the mixed audio, such as Controlled Response Power-Phase (SRP-PHAT) or Multiple Signal Classification (MUSIC) algorithms. This algorithm generates direction estimates (e.g., azimuth θ and elevation φ) for one or more sound sources at each time point, forming time-varying localization trajectories. Each trajectory is a sequence of time-domain signals recording the positional changes of potential sound sources.
[0066] Step S2: Associate and match the several independent audio tracks with the several positioning trajectories to establish audio track-position track association pairs.
[0067] Understandably, step S2 is to correlate the results of the voiceprint domain and spatial domain obtained in step S1.
[0068] For details, please also refer to Figure 3 The method for correlation matching of time-domain fluctuation features in step S2 includes:
[0069] Each independent audio track is framed, and the short-time energy or logarithmic energy of each frame is calculated to form an energy envelope sequence E. track [n], and normalize it;
[0070] Differential operations are performed on each positioning trajectory to obtain the instantaneous angular velocity sequence ω[t], which is then smoothed and normalized to form the angular velocity fluctuation sequence Ω. traj [n];
[0071] For the energy envelope sequence E track [n] Perform speech activity detection, mark the start and end boundaries of speech segments for each independent audio track, and obtain the event set {V1,V2,…,V...} K};Regarding the angular velocity fluctuation sequence Ω traj [n] Perform zero-crossing and peak detection to mark the starting moments of significant fluctuations in angular velocity, obtaining the event set {W1, W2, ..., W...}. L}; Calculate the time overlap rate and event sequence consistency of two event sets to filter out candidate matching pairs;
[0072] Within the time window of the candidate matching pairs, the energy envelope sequence E is calculated using a sliding window with a fixed step size. track [n] and the angular velocity fluctuation sequence Ω traj The Pearson correlation coefficient r of [n] is recorded, along with the maximum correlation coefficient r. max and its corresponding time delay τ opt ;
[0073] If any of the following conditions are met, the independent audio track is considered to have successfully matched the positioning trajectory:
[0074] Condition A: r max ≥ T1 and |τ opt | ≤ T2, where T1 is the correlation coefficient threshold and T2 is the time delay threshold;
[0075] Condition B: r max < T1, but the overlap rate of the two event sets ≥ T3 and the envelope peak phase difference of the two sequences ≤ T4, where T3 is the overlap rate threshold and T4 is the phase difference threshold.
[0076] It can be understood that for each independent audio track, the frame length of each frame obtained by frame division can be 20 - 40 ms and the frame shift is 10 ms. Marking the starting moment of significant angular velocity fluctuations corresponds to the slight movement of the speaker's head or the sudden change of orientation. The window length can be 256 - 512 points.
[0077] On the one hand, traditional methods usually only compare the amplitude or direction itself. The present invention first proposes to perform a correlation analysis on the energy envelope of the independent audio track and the instantaneous angular velocity of the positioning trajectory. Because when a person speaks, there is a non - linear but statistically significant correlation between the change in the amplitude of vocal cord vibration and the change in the relative angle between the head / microphone. Especially in natural conversations, there are often slight head rotations when emphasizing stress. On the other hand, by separately detecting speech activity events and angular velocity fluctuation events, first performing coarse - grained alignment, and then calculating the cross - correlation in a fine - grained manner, the computational amount is greatly reduced, and the anti - interference ability in a low - signal - to - noise ratio environment is improved. In addition, two matching conditions are set: direct matching with a high correlation coefficient; still determining a match when the correlation coefficient is low but the event overlap rate is high, which covers scenarios where reverberation, multipath, etc. cause waveform distortion but event synchrony is still retained, enhancing the robustness.
[0078] Step S3: From the audio - track - position - track association pairs, select the independent audio track corresponding to the target speaker as the target audio track.
[0079] It can be understood that the traditional method usually requires the user to pre - register the voiceprint, compare the embedding vectors of each audio track with the registered embedding vector, and the most - matched audio track is the target audio track. For example, after the user swipes the certificate, the identity information of the certificate is read, and the voiceprint is extracted according to the identity information.
[0080] However, this method is obviously not applicable to the one - to - many question - answering language translation scenario in a complex environment. However, the method for selecting the target audio track in step S3 provided in the first embodiment of the present application includes:
[0081] Receiving the voice input of the questioner and extracting the question keywords from it;
[0082] Predicting the answer keywords in a preset industry word library according to the question keywords;
[0083] Performing real - time speech transcription on each independent audio track to obtain independent texts;
[0084] The independent audio track corresponding to the independent text containing the keywords of the answer and with the highest matching degree is selected as the target audio track.
[0085] Understandably, the system predicts answer keywords based on question keywords and matches them with the real-time transcribed text of each audio track to select the target audio track. It does not rely on voiceprint registration, but utilizes dialogue semantics to help locate the speaker, adapting to the randomness of complex environments.
[0086] Step S4: Using the spatial information of the positioning trajectory associated with the target audio track, perform spatial compensation for the speaker gating determination process of the target audio track.
[0087] Understandably, in continuous audio stream processing, this method employs pre-embedding gating, for example, an audio block every 320ms.
[0088] Under normal circumstances, the cosine similarity (sim) is calculated between the embedded vector of the extracted audio block and the reference embedded vector of the target audio track. If the similarity (sim) is greater than a preset threshold, the audio block is allowed to enter automatic speech recognition. However, in complex scenarios, multiple people may be speaking simultaneously, and the speakers' timbres may be similar, or even the same person's timbres may differ depending on their tone of voice. Therefore, after establishing the audio track-position track association pair, the spatial information of sound source localization becomes another dimension, which is unaffected by timbre interference and can compensate for the deficiencies in the speaker gating process of the target audio track.
[0089] Specifically, when the voiceprint is weak but spatial matching is successful, the spatial compensation method in step S4 provided in the first embodiment of this application specifically includes:
[0090] When the similarity between the current audio block and the target speaker's reference voiceprint is lower than a preset threshold, if the source direction estimation result of the audio block deviates from the historical average azimuth angle of the target speaker by less than a first angle threshold, it is determined to be a spatial match, and the audio block is temporarily allowed to enter the speech recognition process.
[0091] Understandably, if the similarity (sim) is less than a preset threshold, but the deviation of the estimated sound source direction of the audio block from the historical average azimuth angle of the location trajectory associated with the target audio track is less than a first angle threshold (e.g., 15°), spatial compensation is triggered for release. The system determines that the voiceprint mismatch may be caused by noise or a brief change in timbre, but the speaker is still in the original position, so the audio block is temporarily released, and the original audio is output to the automatic speech recognition system.
[0092] Specifically, when voiceprint matching is achieved but spatial anomalies occur, the spatial compensation method in step S4 provided in the first embodiment of this application specifically includes:
[0093] If the voiceprint similarity of the current audio block matches, but the estimated source direction of the audio block changes by more than the second angle threshold relative to the historical average azimuth angle of the target speaker, the output of the audio block is temporarily suppressed and the voiceprint re-registration process is triggered.
[0094] Understandably, if the similarity (sim) is greater than a preset threshold, but the deviation between the estimated sound source direction of the audio block and the historical average azimuth angle of the location trajectory associated with the target audio track is greater than a second angle threshold (e.g., 30°), spatial anomaly suppression is triggered. The system determines that it may be that a non-target speaker has moved to the target location, or that their voiceprint has matched by chance. In this case, the audio block is temporarily replaced with silence and sent to automatic speech recognition, triggering a post-verification process. If the verification fails, a prompt is made to re-register the voiceprint.
[0095] Step S5: Based on the target audio track confirmed after spatial compensation as a timbre reference, generate real-time translated speech.
[0096] Understandably, in step S5, once the entire target audio track, confirmed by spatial compensation gating, is formed, it is directly reused as a timbre reference. The system performs machine translation to obtain the target language text, then uses a zero-shot TTS model (such as YourTTS) with the target audio track as the timbre condition to synthesize translated speech, which is finally broadcast to the client for playback via the data channel, achieving homophonic translation.
[0097] Compared with the prior art, the multilingual translation method based on voiceprint recognition and speech separation provided in the first embodiment of this application introduces the sound source localization trajectory and performs temporal correlation matching with the voiceprint-separated audio track, and uses spatial dimension information as a strong constraint for identity verification. This significantly improves the locking and tracking capability of the target audio track in multi-person, dynamic environments, and solves the technical problem of inaccurate gating in real-time translation in multi-person dynamic scenarios.
[0098] The specific solution in step S4 of the first embodiment mentions that when voiceprint matching is achieved but spatial anomalies occur, the output of the audio block will be temporarily suppressed. However, in multi-person scenarios with mobility, a "blind" effect will occur: for example, in a background check scenario involving a family unit, please refer to [link to relevant documentation]. Figure 1 ,by Figure 1 As shown in the verification channel in the right window, multiple people can be present in the verification area at the same time, and there is no fixed position. If people move quickly or move before speaking (such as the target speaker moving from position c to position d and then to position e to bypass family members in order to swipe their ID), the output of the audio block will be frequently temporarily suppressed.
[0099] To adapt to more complex scenarios involving multiple people with mobility, this application also provides a second embodiment of a multilingual translation method based on voiceprint recognition and speech separation.
[0100] Please see Figure 4 The second embodiment of this application, based on the first embodiment, further includes the following:
[0101] During the acquisition of the initial mixed audio, a video stream is simultaneously acquired and human target tracking and detection are performed to obtain several human target tracking trajectories. Specifically, for example, the system simultaneously acquires an RGBD video stream, whose image origin is pre-calibrated and aligned with the center of the microphone array. Human detection and tracking algorithms (such as deep learning-based multi-target tracking algorithms) are used to detect and track the bodies of all personnel present, generating a positional trajectory for each person that changes over time. This process does not rely on mouth movement detection, so the positional trajectory continues to update even when the speaker stops speaking and moves. However, the biggest drawback of human target tracking trajectories obtained based on human target tracking and detection is that it cannot identify the target speaker, essentially rendering it "deaf."
[0102] The system spatially correlates and matches several human target tracking trajectories with several positioning trajectories to establish a correspondence between human body, sound source, and timbre. Specifically, at the initial stage of system startup or periodically, the spatial coordinates of the person's position trajectory in the video are correlated with the positioning trajectory obtained from sound source localization. For example, if the sound source localization trajectory shows activity at an azimuth angle of 30°, and the video tracking also shows a person at that angle, then a correspondence of "person A - sound source trajectory α - audio track A" is established. This allows "deaf" users to identify the speaker without looking at their mouth.
[0103] The system utilizes human target tracking to update the speaker's sound source localization trajectory in real time, compensating for the loss of localization information caused by the speaker's body movement. Specifically, when the speaker stops speaking and moves to a new location (e.g., from 30° to 60°) before speaking again, there is no effective output for sound source localization during the movement, but video tracking continuously provides their position trajectory. When the person speaks again, the system prioritizes using the 60° position information provided by the video to quickly lock onto or associate the sound source localization result, avoiding the need for complex re-scanning of voiceprints or angles, thus adapting to more complex scenarios with multiple people in motion.
[0104] In addition, it can also solve the problem of ambiguous sound sources. If the sound source is located at multiple angles (such as 50° and 70°) and produces ambiguous results, while video tracking shows that the target person is located at 60°, the system uses video information to assist decision-making and prioritizes sound sources near 60°.
[0105] Through the complementarity of the video and sound sources, the system enables even a "blind" person to see the real-time location of the target speaker in multi-person scenarios with frequent personnel movement, while still maintaining stable tracking of the target speaker and high-quality homophonic translation.
[0106] Preferably, the origin of the video stream acquisition is the same as the origin of the initial mixed audio acquisition, and the video stream is an RGBD video stream. It is understood that using the same origin eliminates the need for audio and video coordinate calibration; RGBD depth information improves the accuracy and robustness of sound source localization, and solves the problems of personnel overlap and occlusion, as well as tracking during movement.
[0107] However, while the second embodiment of this application relies on spatial location association matching, and sound source localization (such as microphone array DOA estimation) is relatively accurate in judging azimuth, it was found in actual testing to be ambiguous in judging distance, especially when the distance exceeds one meter. Thus, the resulting positioning trajectories are also actually ambiguous. In this case, it is very difficult to achieve accurate spatial location association matching between the human target tracking trajectories and the positioning trajectories.
[0108] For this purpose, please refer to Figure 5 The second embodiment of this application provides a method for spatial location association matching, specifically including:
[0109] The human target tracking trajectory and the positioning trajectory are projected onto the same spatial coordinate system, and the azimuth sequence and distance sequence of each trajectory are extracted respectively.
[0110] For each human body trajectory and each positioning trajectory, calculate the direction matching score S. dir And distance matching score S dist :
[0111]
[0112] Where θ diff θ is the average azimuth deviation. max The preset directional tolerance threshold;
[0113]
[0114] Where d diff d represents the average distance deviation. max The preset distance tolerance threshold;
[0115] Calculate the total matching score S:
[0116]
[0117] Where w dir With w dist w is the weighting coefficient. dir+w dist =1, and satisfies:
[0118]
[0119] in Estimating the variance of the direction for sound source localization. To estimate the variance of the distance;
[0120] When the total matching score S is greater than the preset matching threshold T5, it is determined that the human body trajectory is associated with the positioning trajectory, and a human body-sound source correspondence is established.
[0121] The aforementioned correspondence is then fused with the established audio track-position track association to form a human body-sound source-timbre triplet association.
[0122] Understandably, the human target tracking trajectory and the sound source localization trajectory are projected onto the same coordinate system, and the direction matching score and distance matching score are calculated separately. Since sound source localization is sensitive to direction but ambiguous about distance, a weighted summation method is used to calculate the total matching score. Based on the characteristics of the sound source localization trajectory, the smaller the variance of the direction estimation or the larger the variance of the distance estimation, the higher the direction weight. This matching strategy cleverly adapts to the real-time confidence of sound source localization, overcoming the inherent defect of inaccurate distance estimation in sound source localization.
[0123] In order to implement the multilingual translation method based on voiceprint recognition and speech separation in the first or second embodiment described above, the third embodiment of this application also provides a multilingual translation terminal based on voiceprint recognition and speech separation.
[0124] For details, please refer to Figure 6 The multilingual translation terminal based on voiceprint recognition and speech separation includes a main body 100, which is equipped with a microphone array 10, a camera 20, a processing unit (not shown) and a display screen 30. The microphone array 10 is used to collect initial mixed audio, the camera 20 is used to synchronously acquire video streams, and the processing unit is used to execute the multilingual translation method based on voiceprint recognition and speech separation as described above.
[0125] It is understood that the processing unit processes the audio stream captured by the microphone array 10 and the video stream acquired by the camera 20. Since the first and second embodiments have already been described, the specific processing procedures will not be repeated here.
[0126] The display screen 30 is used to show the translated text or other relevant information for easy viewing by the user. Through the collaborative work of these components, the entire terminal enables one-to-many question-and-answer language translation in complex environments, providing users with efficient and accurate translation services.
[0127] Furthermore, the multilingual translation terminal based on voiceprint recognition and speech separation is configured in two units, which are interconnected. For example, one unit acts as the host, facing the questioner, and is connected to a PC. The PC is connected to an identity verification terminal to obtain the respondent's identity information. The other unit acts as the slave, facing the respondent, thus enabling seamless verification between both parties.
[0128] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A multilingual translation method based on voiceprint recognition and speech separation, characterized in that, include: Step S1: Acquire the initial mixed audio in real time, perform a first separation operation based on voiceprint recognition on the initial mixed audio to obtain several independent audio tracks; at the same time, perform a second separation operation based on sound source localization on the initial mixed audio to obtain several localization trajectories related to the time domain signal; Step S2: Associate and match the time-domain signal sequence features of the several independent audio tracks with the time-domain fluctuation features of the several positioning trajectories to establish audio track-position track association pairs; The association matching method in step S2 is as follows: Extract the time-domain signal sequence features of the effective speech segments in the several independent audio tracks, perform correlation matching between the time-domain signal sequence features and the time-domain fluctuation features of the several positioning trajectories in the same time period, and determine the positioning trajectory with the highest correlation as the associated positioning trajectory of the independent audio track. The method for correlation matching of time-domain fluctuation features in step S2 includes: Each independent audio track is divided into frames, and the short-time energy or logarithmic energy of each frame is calculated to form an energy envelope sequence E. track [n], and normalize it; Differential operations are performed on each positioning trajectory to obtain the instantaneous angular velocity sequence ω[t], which is then smoothed and normalized to form the angular velocity fluctuation sequence Ω. traj [n]; For the energy envelope sequence E track [n] Perform speech activity detection, mark the start and end boundaries of speech segments for each independent audio track, and obtain the event set {V1,V2,…,V...} K };Regarding the angular velocity fluctuation sequence Ω traj [n] Perform zero-crossing and peak detection to mark the starting moments of significant fluctuations in angular velocity, obtaining the event set {W1, W2, ..., W...}. L }; Calculate the time overlap rate and event sequence consistency of two event sets to filter out candidate matching pairs; Within the time window of the candidate matching pairs, the energy envelope sequence E is calculated using a sliding window with a fixed step size. track [n] and the angular velocity fluctuation sequence Ω traj The Pearson correlation coefficient r of [n] is recorded, along with the maximum correlation coefficient r. max and its corresponding time delay τ opt ; If any of the following conditions are met, the independent audio track is considered to have successfully matched the positioning trajectory: Condition A: r max ≥T1 and |τ opt |≤T2, where T1 is the correlation coefficient threshold and T2 is the time delay threshold; Condition B: r max <T1, but the overlap rate of the two event sets ≥ T3 and the envelope peak phase difference of the two sequences ≤ T4, where T3 is the overlap rate threshold and T4 is the phase difference threshold; Step S3: From the audio track-bit track association pairs, select the independent audio track of the corresponding target speaker as the target audio track by semantic matching between the question and the answer; Step S4: Using the spatial information of the positioning trajectory associated with the target audio track, perform spatial compensation on the speaker gating determination process of the target audio track; Step S5: Based on the target audio track confirmed after spatial compensation as a timbre reference, generate real-time translated speech.
2. The method according to claim 1, characterized in that, The spatial compensation in step S4 specifically includes: When the similarity between the current audio block and the target speaker's reference voiceprint is lower than a preset threshold, if the source direction estimation result of the audio block deviates from the historical average azimuth angle of the target speaker by less than a first angle threshold, it is determined to be a spatial match, and the audio block is temporarily allowed to enter the speech recognition process.
3. The method according to claim 1, characterized in that, The spatial compensation in step S4 specifically includes: If the voiceprint similarity of the current audio block matches, but the estimated source direction of the audio block changes by more than the second angle threshold relative to the historical average azimuth angle of the target speaker, the output of the audio block is temporarily suppressed and the voiceprint re-registration process is triggered.
4. The method according to claim 1, characterized in that, The method for selecting the target audio track in step S3 includes: It receives voice input from the questioner and extracts keywords from the question. The answer keywords are predicted from a preset industry thesaurus based on the question keywords. Each independent audio track is transcribed into real-time speech to obtain independent text; The independent audio track corresponding to the independent text containing the keywords of the answer and with the highest matching degree is selected as the target audio track.
5. The method according to claim 1, characterized in that, The method further includes: While acquiring the initial mixed audio, the video stream is acquired simultaneously and human target tracking and detection are performed to acquire several human target tracking trajectories; The spatial location association and matching of several human target tracking trajectories and several positioning trajectories are performed to establish a human-sound source-timbre correspondence relationship; The sound source localization trajectory of the target speaker is updated in real time using the human target tracking trajectory to compensate for the loss of localization information during the period when the target speaker moves but does not speak.
6. The method according to claim 5, characterized in that, The method for spatial location association and matching includes: The human target tracking trajectory and the positioning trajectory are projected onto the same spatial coordinate system, and the azimuth sequence and distance sequence of each trajectory are extracted respectively. For each human body trajectory and each positioning trajectory, calculate the direction matching score S. dir And distance matching score S dist : , Where θ diff θ is the average azimuth deviation. max The preset directional tolerance threshold; , Where d diff d represents the average distance deviation. max The preset distance tolerance threshold; Calculate the total matching score S: , Where w dir With w dist w is the weighting coefficient. dir +w dist =1, and satisfies: , in Estimating the variance of the direction for sound source localization. To estimate the variance of the distance; When the total matching score S is greater than the preset matching threshold T5, it is determined that the human body trajectory is associated with the positioning trajectory, and a human body-sound source correspondence is established. The aforementioned correspondence is then fused with the established audio track-position track association to form a human body-sound source-timbre triplet association.
7. The method according to claim 6, characterized in that, The acquisition origin of the video stream is the same as the acquisition origin of the initial mixed audio, and the video stream is an RGBD video stream.
8. A multilingual translation terminal based on voiceprint recognition and speech separation, characterized in that, The device includes a main body, which is equipped with a microphone array, a camera, a processing unit, and a display screen. The microphone array is used to acquire initial mixed audio, the camera is used to synchronously acquire video streams, and the processing unit is used to execute the multilingual translation method based on voiceprint recognition and speech separation as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Video translation method and system based on artificial intelligence
CN121012949A
Real-time duplex translation method based on multi-channel parallel processing and corresponding product
CN121237095A