The invention discloses a multi-
source data-based staphylographic
large model pre-training method and
system, and the method comprises the steps: collecting multi-source staphylographic texts of news, social contact, government affairs, academic and the like, and extracting the vocabulary coverage and language style features, thereby achieving the layering of corpus quality and the recognition of variants; analyzing a flection change structure of the
staphylococcus by adopting a word form reduction and
affix separation technology, and generating a training sample unit carrying a form
label; high-value
mask positions are identified based on morphological complexity, word-level,
phrase-level and
sentence-level
mask tasks are configured, and learning of grammatical relationships such as subject-model consistency and name-form consistency is enhanced in combination with a morphological consistency constraint enhancement model; a course learning strategy is adopted to progressively execute training tasks according to difficulty, finally, optimal parameters are screened through multi-dimensional evaluation to output a staphylography pre-training model, and a high-quality
language representation basis can be provided for downstream tasks such as staphylography text classification,
named entity recognition and
machine translation.