A sparse positive sample risk discrimination method and system based on cluster analysis

By employing a sparse positive sample risk discrimination method based on cluster analysis, we can accurately identify heterogeneous structures within healthy individuals, construct a dedicated anomaly detection model, solve the problem of insufficient detection accuracy in early breast cancer screening, and achieve accurate identification of high-risk samples and interpretability of results.

CN122091258BActive Publication Date: 2026-07-07SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG UNIV
Filing Date
2026-04-27
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

Existing methods for early breast cancer screening and risk assessment struggle to accurately capture local substructure boundaries in the context of multimodal distribution in healthy individuals, resulting in insufficient detection accuracy and an inability to effectively distinguish between high-risk and normal samples.

Method used

A sparse positive sample risk discrimination method based on cluster analysis is adopted. Historical clinical datasets are preprocessed and feature evaluated, projected onto the feature space for cluster analysis, data subgroups are determined and a dedicated anomaly detection model is constructed. The local probability distribution and anomaly score of the samples are calculated by combining Mahalanobis distance and the anomaly detection model, and the decision threshold is optimized to achieve risk discrimination.

Benefits of technology

It significantly improved the identification accuracy of high-risk samples, reduced the risk of model misjudgment, improved operational efficiency, and enhanced the clinical transparency and credibility of the model through the interpretation of key risk features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122091258B_ABST
    Figure CN122091258B_ABST
Patent Text Reader

Abstract

The application discloses a sparse positive sample risk discrimination method and system based on cluster analysis, relates to the technical field of data processing, and comprises the following steps: obtaining positive samples and negative samples of a historical clinical data set, performing double evaluation on the historical clinical data through a feature evaluation model to obtain a target feature subset; projecting the negative samples to the target feature subset to obtain a feature space, performing cluster analysis to determine a plurality of data subgroups and a plurality of cluster centers, determining a corresponding anomaly detection model for each data subgroup, determining the local anomaly scores of the corresponding data subgroups through the anomaly detection model; calculating the Mahalanobis distance between the to-be-tested sample and each cluster center, determining the main subgroups and adjacent subgroups corresponding to the to-be-tested sample; determining the anomaly scores corresponding to the main subgroups and the adjacent subgroups; determining a verification set according to the positive samples, performing index maximization processing according to the verification set to determine a decision threshold, comparing the anomaly scores with the decision threshold, and obtaining a risk discrimination result.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • CN114443338A

  • CN116304712A