This invention discloses a
data annotation method, apparatus, storage medium, and program product, relating to the field of
artificial intelligence. The method includes: inputting a target sample into a target
language model to obtain N candidate labels, wherein the target sample includes data samples to be labeled; calculating the similarity between each candidate
label and S standard labels to obtain N similarity sets, wherein each similarity set includes the similarity between one candidate
label and the S standard labels, and the standard labels include labels defined based on
domain knowledge and
business requirements of the target sample's domain; and determining the target
label for the target sample from the S standard labels based on the N similarity sets. This invention solves the technical problem in related technologies where labels directly annotated by models are difficult to utilize directly, leading to low utilization rates of sample labels.