The application provides a cross-language text clustering method and
system based on high-dimensional vector space manifold alignment, and relates to the technical field of
natural language processing.The method comprises the following steps: constructing a source language and target language
feature matrix through
feature extraction, generating a pseudo-
anchor point matrix by unsupervised topological matching, performing space centering
processing, and executing orthogonal Procrustes analysis; based on the
orthogonal transformation matrix, scaling factor and translation vector, performing
rigid body transformation on the source language
feature matrix, fusing the target language
feature matrix, constructing a unified manifold space, performing clustering
processing, and obtaining the semantic cluster division result of the cross-language text. Through the application, the technical problem of the prior art that the local topological structure of the source language
semantic space is destroyed due to the dependence on large-scale parallel corpus for forced
space mapping, which affects the cross-language text clustering accuracy and causes clustering drift can be solved, the topological
coincidence of different language texts in a unified space is realized, and the cross-language text clustering accuracy is improved.