The invention discloses a member reasoning
attack method and device oriented to category fairness. According to the method, under the
black box access condition, through joint modeling of a shadow model and a
reference model, difficulty calibration and category fairness enhancement are realized. Firstly, a shadow model is used for fitting behaviors of a target model and generating data with member tags; secondly, training a plurality of reference models, and calculating difficulty calibration scores for input samples to correct member
score deviations; then, target member scores, calibration scores and category labels are fused to construct
attack features, a supervised comparative learning (MSCL)
mechanism based on member identities is introduced, and sample pairs which are the same members or non-members are used as
positive sample pairs for feature alignment; furthermore, adaptive weights are distributed for different categories according to the estimated value of the category memory degree, so that the low-memory degree category obtains higher attention in training, and the problem of
vulnerability imbalance among the categories is relieved. And finally, sample member identity judgment is realized through the member probability output by the
attack model. According to the method, the inter-category
vulnerability difference can be remarkably reduced while the overall attack performance is maintained, and more fair privacy
risk assessment is realized; and meanwhile, the number of required reference models is small, the calculation overhead is low, and good expandability and practical value are achieved.