The application provides a
speech recognition method based on bimodal mixed
contrast enhancement, an electronic device, a
chip, a storage medium and a program product; the method comprises the following steps: acquiring multi-
modal speech training data; a bimodal mixed contrast learning model is constructed and trained, wherein samples with the same emotional
label under the same mode are taken as
positive sample pairs, samples with different emotional labels are taken as
negative sample pairs, intra-
modal contrast loss is calculated to enhance the ability of the bimodal mixed contrast learning model to distinguish intra-
modal fine-grained emotional features; samples with the same emotional
label under different
modes are taken as
positive sample pairs, samples with different emotional labels under different
modes are taken as
negative sample pairs, inter-modal contrast loss is calculated to realize the alignment and complementarity of different modal feature spaces; the intra-modal contrast loss and the inter-modal contrast loss are fused to obtain multi-modal mixed contrast loss, and the
model parameters of the bimodal mixed contrast learning model are optimized; based on the features extracted by the bimodal mixed contrast learning model, a downstream
speech recognition task is trained.