The embodiment of the application provides a
data deduplication method, device, equipment, storage medium and program product, and relates to the field of
big data. The method comprises the following steps: obtaining a preset marked historical
data set, a preset unmarked historical
sample pool and to-be-deduplicated
data records, generating training candidate pairs and obtaining text sequences through
serialization, inputting the text sequences into a model to obtain a repetition probability, performing uncertainty sampling based on the probability and supplementing the marked samples to the marked historical
data set, iteratively updating the
model parameters to obtain a target model, inputting the to-be-deduplicated data into the target model to obtain a target probability, and performing retention or merging
processing on the data. The embodiment increases the mechanism of iteratively expanding the marked samples based on the unmarked
sample pool and continuously updating the
model parameters, improves the learning ability and generalization ability of the model for the repetition characteristics of multi-source heterogeneous data, improves the accuracy and
processing efficiency of multi-
source data deduplication, and reduces the dependence on initial artificial marking samples.