The invention provides a two-stage
data deduplication method and
system without data timeliness, and the method comprises the steps: obtaining a source end real-
time data stream, dividing the source end real-
time data stream according to a time window, a time identifier and an
object identifier, and enabling the data of the same time window to be subjected to
data deduplication; dividing the data with the same
object identifier and the same space identifier in the same time window into the same group to obtain multiple groups of data to be subjected to duplicate removal; performing pre-de-duplication
processing on multiple groups of data to be de-duplicated through two de-duplication paths, sending the data to a service calculation module for service calculation, and
synchronizing the data subjected to pre-de-duplication to a real-
time data stream at a target end; and performing post-merging
processing on the multiple groups of data after the business calculation, and writing a result into a target end data table. According to the method, the
data deduplication is divided into two stages, namely the pre-deduplication stage and the post-merging stage, and the pre-deduplication is carried out through two paths, so that the size of a deduplication time window is not limited on the premise of lossless data timeliness, and streaming computing resources and
database storage resources can be effectively reduced.