The application belongs to the technical field of
database benchmark test, and discloses a data generation method and
system for multi-model
database query performance evaluation. A
graph sampling algorithm based on probability propagation is used to extract a seed
data set that retains topological features from a large-scale academic network to construct a probability
distribution model. In the
generation process, a unified logical entity object is constructed based on statistical characteristics, and four kinds of
modal data, i.e. relational, document, graph structure and vector, are synchronously analyzed and mapped to ensure strong consistency of cross-model
semantics. A specific domain statistical
language model and deep semantic coding are introduced to realize
semantic alignment of unstructured text and high-dimensional vectors. Through a closed-loop
time sequence evolution mechanism, incremental features are fed back to the
probability model in real time to drive the next
time step generation. The application effectively solves the problems of existing generated data, such as loss of real data features, weak cross-model consistency and lack of
time sequence causality, and provides a high-fidelity and scalable test benchmark for multi-model
database performance evaluation.