Intelligent recognition and classification method for biological samples
By using intelligent identification and classification methods for biological samples, a biological category map is obtained, classification and quantitative labels are screened, and an intelligent classification model is constructed. This solves the problems of accuracy and efficiency in biological sample identification and classification in existing technologies, and realizes intelligent and automated processing of biological samples.
Patent Information
- Application Number
- CN202510075950.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-01-17
AI Technical Summary
Existing biological sample identification and classification methods struggle to extract accurate classification features from complex and diverse data structures, resulting in low classification accuracy and reliability, and inefficient data processing.
A biological sample intelligent identification and classification method is adopted. By acquiring a biological category map, screening and quantifying classification labels, reconstructing the classification label map, determining the strong correlation features of the labels and performing feature screening and validity processing, an intelligent classification model is constructed, which includes a sample identification unit and a classification decision unit. Feature extraction and classification are performed, and the results are visualized by combining confidence verification.
It has achieved intelligent and automated identification and classification of biological samples, improving the accuracy and efficiency of classification, and ensuring the reliability and visualization of classification results.
Smart Images

Figure CN119889462B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field related to data identification, and specifically to a method for intelligent identification and classification of biological samples. Background Art
[0002] In today's era of rapid advancement in biotechnology, the accurate identification and efficient classification of biological samples have become indispensable key technologies in many fields such as life sciences, medical diagnosis, and drug development. With the rapid development of high-throughput technologies such as genomics, proteomics, and metabolomics, biological sample data has shown explosive growth. These data not only come from a variety of different technical platforms (multi-source), but also have complex structures (heterogeneous), posing huge challenges to traditional identification and classification methods. Traditional biological sample identification and classification often rely on manual observation and empirical judgment, which is not only inefficient but also easily affected by subjective factors. It cannot ensure the accuracy and effectiveness of classification labels, and it is difficult to cope with large-scale, highly complex sample data, resulting in sample classification errors due to improper label selection. At the same time, it is difficult to extract features useful for classification from complex biological sample data and perform effective processing.
[0003] Therefore, in the current technologies related to biological sample identification and classification, there is a technical problem that it is difficult to extract accurate classification features from complex and diverse data structures and process them efficiently, which leads to low accuracy and reliability of data classification and poor data processing efficiency. Summary of the Invention
[0004] This application provides a method for intelligent identification and classification of biological samples. By adopting technical means such as optimizing classification labels, feature extraction, building intelligent classification models, and confidence verification, it solves the technical problems existing in existing biological sample identification and classification, such as the difficulty in extracting accurate classification features from complex and diverse data structures and performing efficient processing, which leads to low accuracy and reliability of data classification and poor data processing efficiency. It realizes the intelligence and automation of sample identification and classification, and achieves the technical effect of improving the accuracy and efficiency of biological sample classification.
[0005] The present application provides a method for intelligent identification and classification of biological samples, the method comprising: obtaining a biological category map, screening classification quantitative labels, and reconstructing and determining a classification label map, wherein the classification quantitative labels are screened based on distinguishability, stability, and uniqueness; traversing the classification label map, determining label strong correlation features and performing feature screening and validity processing, determining effective correlation features, and measuring feature validity based on classification decision requirements; associating the classification label map with the effective correlation features to construct an intelligent classification model, the intelligent classification model comprising a sample identification unit and a classification decision unit, the sample identification unit comprising a feature dimensionality reduction block, a feature granularity classification block, and a feature contrast enhancement block; reading biological sample data, the biological sample data comprising multi-source heterogeneous data; transmitting the biological sample data to the intelligent classification model, performing feature extraction processing and classification based on the data structure, and determining a sample classification result; performing confidence verification on the sample classification result, and visualizing the result on a display terminal.
[0006] In a possible implementation, the strong association feature of the tag is determined and the feature screening is performed, and the following processing is performed: if it is less than the first association threshold, the first tag association feature is left vacant; if it is greater than the first association threshold and less than the second association threshold, the first classification weight is configured for the second tag association feature, wherein the second tag association feature is in the initialized feature state; if it is greater than the second association threshold, the third tag association feature is feature converted and a second classification weight is configured, wherein the second classification weight is greater than the first classification weight and the second association threshold is greater than the first association threshold.
[0007] In a possible implementation, the effective association feature is determined by performing the following processing: traversing the third label association feature to determine the classification feature element; based on the classification feature element, converting the mapped third label association feature to determine the conversion feature; based on the second label association feature and the conversion feature, determining the effective association feature.
[0008] In a possible implementation, the determination of the sample classification result further performs the following processing: extracting biological label features in combination with the sample identification unit; performing block rotation processing on the biological label features to determine effective biological label features; transferring the effective biological label features to the classification decision unit, performing bottom-up classification and division, and determining the sample classification result.
[0009] In a possible implementation, the sample identification unit includes a feature dimensionality reduction block, a feature granularity grading block and a feature contrast enhancement block, and also performs the following processing: the ports of the feature dimensionality reduction block, the feature granularity grading block and the feature contrast enhancement block are provided with feature state checkpoints based on the effective associated features, and block idle rounds can be executed.
[0010] In a possible implementation, the biological label feature is subjected to block rotation processing, and the following processing is also performed: the biological label feature is a fuzzy extraction feature; based on the feature state level of the feature dimensionality reduction block, the biological label feature is segmented to determine the pre-processed label feature and the idle wheel label feature; the pre-processed label feature is traversed to determine the reduced dimension label feature by direct dimensionality reduction and lower-level decomposition dimensionality reduction; the reduced dimension label feature and the idle wheel label feature are integrated to perform block flow processing.
[0011] In a possible implementation, after determining the valid biological tag features, the following processing is further performed: traversing the valid associated features, determining the classification risk features based on the misclassification probability, wherein the misclassification probability is based on historical classification record mining; and performing feature passivation processing on the classification risk features based on the probability ratio.
[0012] In a possible implementation, the method for intelligent identification and classification of biological samples further performs the following processing: the intelligent classification model is expandable, including update extensions and new extensions; sample classification data of a preset period is interactively obtained, and classification defects are located based on the abnormal classification data traceability; and based on the classification defects, the intelligent classification model is updated and learned.
[0013] The present application proposes a method for intelligent identification and classification of biological samples to obtain a biological category map, screen classification quantitative labels, and reconstruct and determine the classification label map, wherein the classification quantitative labels are screened based on distinguishability, stability, and uniqueness; the classification label map is traversed to determine the strong correlation features of the labels and perform feature screening and validity processing, determine the effective correlation features, and measure the feature validity based on the classification decision requirements; associate the classification label map with the effective correlation features to construct an intelligent classification model, wherein the intelligent classification model includes a sample identification unit and a classification decision unit, wherein the sample identification unit includes a feature dimensionality reduction block, a feature granularity grading block, and a feature contrast enhancement block; read biological sample data, wherein the biological sample data includes multi-source heterogeneous data; transmit the biological sample data to the intelligent classification model, perform feature extraction processing and classification based on the data structure, and determine the sample classification result; perform confidence verification on the sample classification result, and visualize the result on a display terminal. It solves the technical problem of existing biological sample identification and classification that it is difficult to extract accurate classification features from complex and diverse data structures and process them efficiently, which leads to low accuracy and reliability of data classification and poor data processing efficiency. It realizes the intelligence and automation of sample identification and classification, and achieves the technical effect of improving the accuracy and efficiency of biological sample classification. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments of the present disclosure are briefly introduced below. Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in precise order. Instead, various steps may be processed in reverse order or simultaneously as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0015] Figure 1 A schematic diagram of a process flow of a method for intelligent identification and classification of biological samples provided in an embodiment of the present application;
[0016] Figure 2 A schematic diagram of a process for determining effective correlation features in a method for intelligent identification and classification of biological samples provided in an embodiment of the present application. DETAILED DESCRIPTION
[0017] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below.
[0018] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0019] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict, and the terms “first\second” involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. The terms “including” and “having” and any variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products, or devices. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used herein are for the purpose of describing the embodiments of this application only.
[0020] The present application embodiment provides a method for intelligent identification and classification of biological samples, such as Figure 1 As shown, the method includes:
[0021] Step S100, obtain a biological category map, screen classification quantitative labels, and reconstruct and determine the classification label map, wherein the classification quantitative labels are screened based on distinguishability, stability, and uniqueness. Obtain a biological category map from a biological classification database. The biological category map is a structured representation map of biological classification, showing the hierarchical structure and classification relationship of biological classification. The nodes of the biological category map represent broad biological classification units. For example, the plant kingdom includes angiosperms, angiosperms include dicots, which are further divided into roses, chrysanthemums, etc.; the animal kingdom includes chordates, which include mammals, which are further divided into primates, carnivores, etc. The edges of the map represent the hierarchical relationship between classification units; classification quantitative labels are labels used to quantify and describe biological characteristics, and are used for biological classification and identification. Specifically, feature data of different biological categories are extracted from the biological category map, including morphological characteristics (such as leaf shape, inflorescence), gene sequence characteristics, ecological characteristics (such as habitat), etc., using information gain, Gini coefficient, etc. Calculate the distinguishability of each feature and evaluate the distinguishing ability of the feature; use analysis of variance (ANOVA) or stability index to evaluate the stability of each feature and analyze the degree of variation of the feature under different conditions; use uniqueness score or uniqueness index to evaluate the uniqueness of each feature and determine which features are unique in a specific biological category; then screen out important quantitative labels based on the distinguishability, stability and uniqueness criteria, that is, screen the classification quantitative labels; finally, adjust and optimize the original biological map based on the new classification quantitative labels, that is, use the screened quantitative labels to reconstruct the biological classification map, update the node and edge information, ensure that the new map can more accurately reflect the biological classification relationship, determine the specific label of each classification unit, and map it to the reconstructed map to ensure that the new map has a good hierarchical structure and classification accuracy.
[0022] Step S200, traverse the classification label map, determine the strong correlation features of the labels and perform feature screening and validity processing, determine the effective correlation features, and measure the feature validity based on the classification decision requirements. Traversing the classification label map refers to accessing each node (representing the classification label) and edge (representing the relationship between labels) in the map to fully understand the structure and content of the map, obtain detailed information of each classification label, including its attributes, associated labels, and relationships with other biological entities, etc., determine the strong correlation features of the labels and perform feature screening and validity processing, wherein the strong correlation features of the labels refer to features that are highly correlated with a certain classification label and can significantly distinguish the label from other labels. Specifically, the statistical indicators such as the frequency of occurrence and co-occurrence rate of features between different labels are calculated to determine which features have a strong correlation with specific labels, and then screen them from many features through feature selection algorithms (such as filtering, encapsulation, embedded, etc.). Select features related to the classification task, exclude those features that are irrelevant to the classification or have weak correlation, perform feature standardization and normalization on the filtered classification features, and further process the filtered features to reduce redundancy and noise between features, improve their effectiveness, and obtain effective associated features. For example, for a certain feature, its refined lower-dimensional decomposition elements can meet the classification requirements, etc. Guided by classification requirements, measuring feature effectiveness means further measuring and selecting the features that best meet the classification requirements based on specific classification decision requirements. For example, use feature importance analysis (such as random forest feature importance, SHAP value, etc.) to measure the contribution of each feature to the classification decision, and effectively evaluate the importance of features in the classification task.
[0023] In one possible implementation, step S200 further includes step S210, where the first tag-associated feature is left blank if the correlation is less than a first correlation threshold. When the correlation (or relevance, importance, etc.) between a feature and a tag is below a preset threshold (i.e., the first correlation threshold), the feature is considered weakly associated with the tag. In subsequent processing or analysis, the feature is left blank or ignored to reduce the impact of noise on the results.
[0024] In step S220, if the correlation is greater than the first correlation threshold and less than the second correlation threshold, a first classification weight is assigned to the second tag-associated feature, where the second tag-associated feature is in an initialized feature state. If the correlation of a tag-associated feature is between the first and second correlation thresholds, it is considered a second tag-associated feature, and it is considered to have a certain correlation with the tag, but it is not strong enough. The feature is retained and assigned a relatively low classification weight (i.e., the first classification weight). This can preserve the feature's information to a certain extent while reducing its influence in the final decision or analysis. The second tag-associated feature is in an initialized feature state, which refers to the original state of these features in the dataset before the weights are assigned.
[0025] In step S230, if the correlation is greater than the second correlation threshold, the third tag-associated feature is converted and a second classification weight is assigned, wherein the second classification weight is greater than the first classification weight, and the second correlation threshold is greater than the first correlation threshold. If the correlation of a tag-associated feature is greater than the second correlation threshold, it is considered a third tag-associated feature, and the correlation between the feature and the tag is considered to be very strong. The feature is converted (possibly to better fit the model or improve the expressiveness of the feature) and a relatively high classification weight (i.e., the second classification weight) is assigned to it to fully utilize this strong correlation. It is also noted that the second classification weight is greater than the first classification weight, and the second correlation threshold is greater than the first correlation threshold. In general, the correlation threshold is used to assess the strength of the relationship between the feature and the tag, helping to distinguish which features are useful to the model and which may be redundant or noisy. The first and second correlation thresholds divide the feature correlation into three intervals: below the first threshold, between the first and second thresholds, and above the second threshold.
[0026] In one possible implementation, Figure 2 As shown, step S230 further includes step S231, traversing the third tag-associated features to determine classification feature elements. All feature items contained in the third tag-associated features are traversed to determine which feature elements (or feature attributes, dimensions, perspectives, etc.) can meet the classification requirements. Specifically, this includes in-depth analysis and understanding of the features to identify which features or parts of the features are effective for distinguishing different categories (or labels). For example, if the feature is a set of points in a three-dimensional space, it may be necessary to select a two-dimensional plane as a classification feature element because this two-dimensional plane may contain sufficient information to distinguish different categories.
[0027] Step S232: Based on the classification feature elements, the mapped third label-associated features are converted to determine converted features. After the classification feature elements are determined, the third label-associated features mapped to these elements are converted. Specifically, the original features are converted into a form more suitable for the classification task using methods such as dimensionality reduction (such as PCA and LDA), feature selection (such as selecting the most useful features based on information gain and chi-square tests), feature encoding (such as label encoding), and feature construction (such as constructing new features based on the original features). The converted features are called converted features, which generally have better classification performance and contain the most important information in the original features for the classification task.
[0028] Step S233: Determine the effective association features based on the second label association features and the converted features. Finally, determine the effective association features by combining the second label association features (i.e., features that have been processed or weighted) with the converted features. This is done by obtaining effective association features based on feature importance assessment or feature relevance assessment. Effective association features are defined as a set of features that are both important and effective for the classification task, encompassing both directly useful portions of the original features and new features derived through conversion. These features are used as input features for model training.
[0029] Step S300, associating the classification label map with the effective association features to construct an intelligent classification model, wherein the intelligent classification model includes a sample identification unit and a classification decision unit, wherein the sample identification unit includes a feature dimensionality reduction block, a feature granularity classification block, and a feature contrast enhancement block. Match and associate each label in the classification label map with its corresponding valid associated features. Through association, the intelligent classification model can use these features to identify and distinguish samples of different categories, thereby improving the accuracy and efficiency of classification. The intelligent classification model includes a sample recognition unit and a classification decision unit. Specifically, the sample recognition unit is the core part of the intelligent classification model, which is responsible for extracting key features from the original data and processing these features to enhance its classification ability. The sample recognition unit includes a feature dimensionality reduction block, a feature granularity classification block and a feature contrast enhancement block. Among them, the feature dimensionality reduction block uses dimensionality reduction techniques such as principal component analysis (PCA) and linear discriminant analysis (LDA) to map high-dimensional feature space to low-dimensional space to reduce computational complexity and avoid overfitting, reduce the number of features, and try to retain important information in the original data; the feature granularity classification block is based on the complexity and importance of the features. , dividing it into different granularity levels, so that different strategies can be adopted according to the features of different levels in subsequent processing to improve the precision and accuracy of classification; the feature contrast enhancement block uses histogram equalization, contrast stretching and other technologies to adjust the distribution range of features to make them more suitable for processing by the classification algorithm. By transforming or enhancing the contrast between features, the distinction between samples of different categories in the feature space is made more obvious; the classification decision unit is the final decision-making part of the intelligent classification model. It classifies the input samples according to the features extracted and processed by the sample recognition unit. The classification decision unit is constructed and trained based on decision trees, random forests, support vector machines or neural networks. For example, for large-scale data sets and high-dimensional features, models such as random forests or neural networks that can handle complex data and nonlinear relationships can be used. The parameters of the selected model are tuned through historical data to improve the accuracy and generalization ability of classification.
[0030] In one possible implementation, step S300 further includes step S310, and the intelligent classification model is extensible, including update extensions and new extensions. The scalability of the intelligent classification model refers to the ability of the model to continuously improve its performance and scope of application by adding new functions, optimizing algorithms or adjusting parameters over time and as business needs change, specifically including update extensions and new extensions. Update extensions generally refer to improvements and optimizations to the existing functions of the model, including re-labeling of model training data, fine-tuning of model parameters or retraining of the entire model. For example, according to new business rules or data characteristics, the model's classification algorithm is adjusted, feature selection is optimized or the model structure is improved to improve classification accuracy and efficiency; new extensions refer to adding new functions to the model or supporting new classification categories. For example, as the business scope expands, new data categories may need to be included in the classification scope. At this time, the model needs to be newly extended.
[0031] Step S320: Interact with the preset period of sample classification data and locate classification defects based on abnormal classification data traceability. The preset period of sample classification data refers to sample data collected from actual business scenarios at a predetermined period (such as daily, weekly, or monthly), which is then classified and processed. Some misclassified samples, i.e., abnormal classification data, may be found in the preset period of sample classification data. Based on this abnormal data, traceability analysis can be used to locate the root cause of the classification defect. These defects may arise from data quality issues (such as noise, missing values), improper feature selection, or model overfitting or underfitting.
[0032] Step S330: Based on the classification defect, the intelligent classification model is updated and learned. After determining the root cause of the classification defect, the intelligent classification model is updated and learned. The purpose of the updated learning is to improve the model's classification performance and generalization capabilities by adjusting model parameters, optimizing the algorithm structure, or adding new training data. For example, the model can be retrained using new or adjusted training data, and model parameters can be fine-tuned based on the existing model to improve its performance.
[0033] Step S400 reads biological sample data, which includes multi-source heterogeneous data. Biological sample data refers to data obtained from organisms (such as humans, animals, plants, and microorganisms) for purposes such as scientific research, medical diagnosis, and drug development. It typically contains rich biological information and comes from multiple different sources. Biological sample data is heterogeneous. For example, multi-source data may include genomic data (such as DNA sequences and RNA sequences), proteomic data (such as protein sequences and mass spectrometry data), metabolomic data (such as mass spectrometry data of metabolites), phenotypic data such as morphological characteristics (such as leaf shape and inflorescence), physiological indicators (such as weight and height), and environmental data (such as habitat information, climate conditions, and geographic distribution). Data heterogeneity refers to data structure (structured tabular data, unstructured text, and image data), data format (such as text files and images), and data attributes (such as quantitative gene expression levels and qualitative classification labels).
[0034] Step S500: The biological sample data is transmitted to the intelligent classification model, where feature extraction, processing, and classification are performed based on the data structure to determine a sample classification result. The collected biological sample data (including genomic data, epigenomic data, transcriptomic data, proteomic data, metabolomic data, and clinical information data) is transmitted to the intelligent classification model for in-depth analysis of the data structure. Specifically, biological sample data is typically highly complex and diverse, including different data types (e.g., numerical, textual, and image) and complex hierarchical structures (e.g., gene-transcript-expression relationship in gene expression profiles). By analyzing the data structure, the inherent characteristics and patterns of the data are understood. Based on the data structure, features useful for the classification task are extracted from the biological sample data. The extracted classification features are then used to classify the biological sample. Ultimately, the intelligent classification model outputs a classification result for the biological sample, indicating the category or type to which the sample belongs, and ensuring the accuracy and reliability of the classification result.
[0035] In one possible implementation, step S500 further includes step S510, in which the sample identification unit is combined to extract biological tag features. The sample identification unit is responsible for identifying features from biological samples (such as DNA sequences, RNA sequences, protein sequences, image data, etc.) that can be used for subsequent analysis. Specifically, the sample identification unit is used to extract representative and discriminative features from the biological sample. These features can be specific patterns in the sequence, specific textures or shapes in the image, etc. For example, in gene expression analysis, gene expression levels may be extracted as features; in facial recognition, facial contours, the shape and position of parts such as the eyes and nose may be extracted as features.
[0036] In step S520, the biomarker features are subjected to block rotation processing to determine valid biomarker features. Block rotation processing is a feature processing process that aims to improve the effectiveness and discriminability of features through some transformation or reorganization. It typically includes steps such as feature selection, feature dimensionality reduction, and feature encoding. For example, in feature selection, features that are irrelevant or redundant to the classification target may be removed; in feature dimensionality reduction, high-dimensional features may be mapped to a low-dimensional space through methods such as principal component analysis (PCA) and linear discriminant analysis (LDA); and in feature encoding, categorical features may be converted to numerical features for subsequent processing. After block rotation processing, the resulting feature set may contain some features that are more effective and important for the classification task. These features are called valid biomarker features and will serve as input to the subsequent classification decision unit for sample classification.
[0037] In step S530, the valid biomarker features are transferred to the classification decision unit for bottom-up classification and determination of the sample classification result. The classification decision unit typically employs a hierarchical approach, starting from the bottom and gradually classifying samples into different categories based on the input features and the classifier's judgment rules. This process can be bottom-up, starting with the most specific categories and then gradually moving up to more general ones. After processing by the classification decision unit, a classification result is ultimately obtained for each sample. This result can be a specific category label or a probability distribution (indicating the likelihood that the sample belongs to each category).
[0038] In one possible implementation, step S530 further includes configuring the ports of the feature dimensionality reduction block, the feature granularity grading block, and the feature contrast enhancement block with feature state checkpoints based on the valid correlation features, and enabling block idle rounds. In the feature dimensionality reduction block, the feature state checkpoint based on the valid correlation features is used to screen and filter features, ensuring that only features with significant relevance and importance at different granularities are retained. In the feature granularity grading block, a feature state checkpoint based on the valid correlation features is also configured to ensure that enhanced features remain closely related to the prediction target and better reflect the differences between different categories. If the feature dimensionality reduction block, the feature granularity grading block, and the feature contrast enhancement block do not receive sufficient or qualified input features, they may temporarily cease operations and wait for more valid input, i.e., enter a "block idle round" state, to ensure the quality and validity of the input data.
[0039] In one possible implementation, step S530 further includes step S531, wherein the biometric tag feature is a fuzzy extracted feature. A fuzzy extracted feature generally refers to a feature extracted from a biometric signal that is random and stable and can tolerate slight changes in the input biometric feature to a certain extent.
[0040] Step S532: Based on the feature status checkpoints in the feature dimensionality reduction block, the biological label features are segmented to determine pre-processing label features and idle label features. In the feature dimensionality reduction block, the feature status checkpoint is used to evaluate each feature's importance, relevance, and association with the prediction target. For fuzzy extraction features, this checkpoint determines which features should be retained for subsequent processing and which features may be eliminated due to redundancy, noise, or irrelevance to the prediction target. The biological label features are segmented to determine pre-processing label features and idle label features. Specifically, pre-processing label features that pass the feature status checkpoint are considered important for subsequent processing and will directly enter the dimensionality reduction process. Idiosyncratic label features may be temporarily shelved due to not meeting specific conditions (such as insufficient relevance, excessive noise, etc.) and do not directly participate in the current dimensionality reduction process.
[0041] Step S533 traverses the preprocessed label features, using direct dimensionality reduction and lower-level decomposition dimensionality reduction as processing methods to determine the reduced-dimensionality label features. For the preprocessed label features, both direct dimensionality reduction and lower-level decomposition dimensionality reduction are used. Specifically, dimensionality reduction techniques such as PCA (principal component analysis) are used to directly reduce the dimensionality of the features to reduce the number of features while retaining key information. Lower-level decomposition dimensionality reduction involves more detailed analysis and decomposition of the features, breaking down complex features into simpler sub-features, and then reducing the dimensionality of these sub-features. This allows for more precise control of the dimensionality reduction process and determines the final set of reduced-dimensionality label features.
[0042] Step S534 integrates the reduced dimension label features with the idle wheel label features and performs block flow processing. After the dimensionality reduction process is complete, the reduced dimension label features are integrated with the previously set aside idle wheel label features. The importance of the idle wheel label features is reassessed to ensure their compatibility with the reduced dimension label features and their joint use in subsequent data processing or model training. The integrated feature set then undergoes block flow processing, i.e., subsequent steps such as feature encoding, model training, and prediction are performed according to the predetermined data processing flow.
[0043] In one possible implementation, step S530 further includes step S535, traversing the effective associated features and determining classification risk features based on the probability of misclassification, wherein the probability of misclassification is mined based on historical classification records. The effective associated features are traversed to further understand their impact on model performance and provide a basis for subsequent risk assessment and feature processing. Then, the classification risk features are determined based on the probability of misclassification, wherein the probability of misclassification refers to the probability that the model will mistakenly classify a sample into other categories during the classification process. Specifically, based on the mining of historical classification records, the probability of misclassification of each feature during the classification process is calculated. By comparing the probabilities of misclassification of different features, it can be determined which features have a higher classification risk, that is, changes in these features are more likely to lead to errors in the classification results, and then the classification risk features are determined, which usually refer to those features with a higher probability of misclassification and a greater impact on the classification results.
[0044] Step S536, based on the probability ratio, the classification risk feature is subjected to feature passivation. The probability ratio may refer to the ratio of the misclassification probability of the classification risk feature to other features, or the ratio of the misclassification probability of the classification risk feature between different categories. This ratio reflects the relative importance of the classification risk feature and its influence on the classification result. Feature passivation is a method to reduce the sensitivity and importance of features. For classification risk features, feature passivation can reduce their direct influence on the classification result, thereby reducing the risk of misclassification. Specifically, the importance of the features is reduced by adjusting the numerical range of the features, such as using maximum and minimum normalization, Z-score standardization, etc.; a part of the features that have a greater impact on the classification result is selected from the original feature set, while ignoring or reducing the weight of the classification risk features; nonlinear transformation or discretization is performed on the features to change their distribution and properties, thereby reducing their sensitivity; regularization terms are added to the loss function of the model to penalize the features to reduce the complexity of the model and prevent overfitting.
[0045] Step S600: Confidence verification is performed on the sample classification results, and the results are visualized on a display terminal. Confidence verification is a statistical method used to assess the reliability and accuracy of classification results. In biological sample classification, confidence verification can be performed based on historical classification results or known biological knowledge. Specifically, a confidence interval is determined based on historical classification data or prior knowledge, representing the credible range of the classification result. For example, if a classification label appears with a 95% confidence level in historical data, then when the current sample is classified as having that label, there is a 95% confidence that the classification is correct. For the current classification result, the confidence level of each possible category is calculated (e.g., probability, confidence interval, etc.), and the classification result of the current sample is compared with the confidence interval. If the classification result falls within the confidence interval, the classification result is considered qualified; if it does not fall within the confidence interval, the classification process or data quality may need to be reviewed. Result visualization is the process of displaying classification results in the form of graphics, images, or tables on a display terminal to facilitate intuitive understanding and analysis of the data. Visualization of biological sample classification results helps to quickly identify key information and discover potential patterns and trends. For example, directly displaying the classification label of each sample on a display terminal helps ensure the accuracy and reliability of the classification results and provides a more intuitive display of biological data.
[0046] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A method for intelligent identification and classification of biological samples, characterized in that: The method comprises: Obtaining a biological category map, screening classification quantitative labels, and reconstructing and determining the classification label map, wherein the classification quantitative labels are screened based on distinguishability, stability, and uniqueness; Traverse the classification label map, determine the strong correlation features of the labels, perform feature screening and validity processing, determine the effective correlation features, and measure the feature validity based on the classification decision requirements; Associating the classification label map with the effective association features to construct an intelligent classification model, wherein the intelligent classification model includes a sample identification unit and a classification decision unit, wherein the sample identification unit includes a feature dimensionality reduction block, a feature granularity classification block, and a feature contrast enhancement block; Reading biological sample data, wherein the biological sample data includes multi-source heterogeneous data; Transmitting the biological sample data to the intelligent classification model, performing feature extraction processing and classification based on the data structure, and determining the sample classification result; Performing confidence check on the sample classification results and visualizing the results on a display terminal; Determining the sample classification result includes: extracting biomarker features in combination with the sample identification unit; Performing block rotation processing on the bio-tag features to determine valid bio-tag features; Transferring the effective biomarker feature flow to the classification decision unit, performing classification from bottom to top, and determining the sample classification result; The ports of the feature dimension reduction block, the feature granularity classification block and the feature contrast enhancement block are provided with feature state checkpoints based on the effective associated features, and can execute block idle rounds.
2. The method for intelligent identification and classification of biological samples according to claim 1, wherein: The determining of the strong correlation features of the tags and performing feature screening includes: If it is less than the first correlation threshold, the first tag correlation feature is left vacant; If it is greater than the first correlation threshold and less than the second correlation threshold, configuring a first classification weight for the second tag association feature, wherein the second tag association feature is in an initialized feature state; If it is greater than the second association threshold, feature conversion is performed on the third label association feature and a second classification weight is configured, wherein the second classification weight is greater than the first classification weight, and the second association threshold is greater than the first association threshold.
3. The method for intelligent identification and classification of biological samples according to claim 2, wherein: The determining of the effective association features includes: Traversing the third tag associated features to determine classification feature elements; Based on the classification feature element, converting the mapped third label association feature to determine a conversion feature; The effective association feature is determined based on the second tag association feature and the conversion feature.
4. The method for intelligent identification and classification of biological samples according to claim 1, wherein: The block rotation processing is performed on the biological tag feature, including: The biological tag feature is a fuzzy extraction feature; Segmenting the biological label features based on the feature state level of the feature dimension reduction block to determine pre-processing label features and idle wheel label features; Traversing the pre-processed label features, determining the reduced dimension label features by direct dimensionality reduction and lower-level decomposition dimensionality reduction; The dimension reduction label feature and the idle wheel label feature are integrated to perform block flow processing.
5. The method for intelligent identification and classification of biological samples according to claim 1, wherein: After determining the effective biomarker characteristics, the method includes: Traversing the valid associated features, and determining classification risk features based on misclassification probabilities, wherein the misclassification probabilities are mined based on historical classification records; Based on the probability ratio, feature passivation processing is performed on the classification risk feature.
6. The method for intelligent identification and classification of biological samples according to claim 1, wherein: The intelligent classification model is extensible, including update extensions and new extensions; Interact with sample classification data of preset periods and locate classification defects based on abnormal classification data traceability; Based on the classification defects, the intelligent classification model is updated and learned.
Citation Information
Patent Citations
Method and system for classifying fine-grained images
CN108875827A
Subject classification method fusing multiple human brain atlases based on graph convolutional neural network
CN111563533A