The invention relates to the technical field of simultaneous interpretation, and discloses a simultaneous interpretation
data processing method and
system based on a POE
microphone array. The method comprises the following steps: synchronously acquiring multi-language original audio streams and
meeting place environment
noise spectrum features through a distributed
microphone array powered by the
Ethernet; after
time domain framing is carried out on the audio
stream, adaptive filtering is carried out by using a
dynamic noise reduction
weight coefficient to obtain a primary pure voice segment; dividing the multi-
language speech endpoint detection model into independent speech units with language labels through a pre-trained multi-
language speech endpoint detection model, and matching a corresponding
acoustic model to generate a phoneme-level
time alignment sequence; comparing and outputting a term replacement
instruction stream in real time in combination with a simultaneous transfer term
library, and generating an intermediate
semantic representation vector after fusion; and the low-
delay encoder converts the voice parameter sequence into a target language voice parameter sequence, and drives the waveform synthesizer to generate final simultaneous transmission audio. The method optimizes the whole process
processing, gives consideration to the simultaneous transmission accuracy and real-time performance, and is suitable for a multilingual
meeting place scene.