The application discloses a text language automatic detection method and
system based on
syllable and
affix features, relates to the field of
natural language processing, and comprises the following steps: text preprocessing and
basic language unit extraction are performed on a text to be detected to obtain a text
syllable sequence and an
affix list; based on the
affix list, a multidimensional pragmatic scene distribution vector in the context of the text
syllable sequence is extracted, and core
connectivity in a pre-constructed morphological paradigm
knowledge graph is acquired; the function load value of the affix is matched from an affix function load
database, the affix function load value is normalized, and a weighted affix
feature vector is generated;
syntax dependency
relationship analysis and preliminary structure construction are performed on the text to be detected, and a dependency
syntax tree of the text to be detected is generated. The application improves the detection accuracy and result
interpretability of highly similar languages and mixed code texts by explicitly modeling and quantitatively evaluating deep morphological-syntactic rule features of languages.