The purpose of this disclosure is to provide a
speech training noise enhancement
system and method based on a
hybrid noise generation model, comprising: an input module, a
noise environment enhancement module, a speech
noise enhancement module, and an output module; wherein, the input module is used to acquire simplified description information of the noise environment and clean speech data to be enhanced; the noise environment enhancement module converts the simplified description information of the noise environment into a structured noise
event sequence with temporal features; the speech
noise enhancement module generates multi-source
hybrid noise according to the noise
event sequence and adds the multi-source
hybrid noise to the clean speech data according to preset rules to obtain noisy speech data; the output module is used to output the noisy speech data for use in
speech model training. This disclosure fully integrates the interactive features of superposition, cancellation, and interference of different noise sources in the time dimension, solving the technical
bottleneck that traditional mixing methods can only achieve linear superposition.