A salt stress data intelligent processing method for a pineapple protein phosphatase gene

CN122551889APending Publication Date: 2026-08-11MINNAN NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-24
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提供一种面向菠萝蛋白磷酸酶基因的盐胁迫数据智能处理方法用于解决现有植物基因分析流程采用固定参数与统一规则,聚类划分不准、评价机制单一,最终数据分析质量与稳定性不足的问题

Benefits of technology

1、本发明提供一种面向菠萝蛋白磷酸酶基因的盐胁迫数据智能处理方法,整合多源信息构建跨物种分析数据集,依托蛋白质序列比对工具与隐马尔可夫模型完成联合检索,再通过双重筛选剔除无效序列,结合聚类分析、多序列比对与系统发育树完成基因亚家族划分;随后对高维特征开展降维与分析,最终基于转录组数据完成表达定量与加权打分。该方法流程衔接紧密、数据互通,针对菠萝蛋白磷酸酶基因的应用场景定制整体处理逻辑,摆脱流程环节割裂、通用化适配性差的弊端,从数据源到最终结果实现标准化、一体化分析,有效提升基因相关数据处理的整体规范性与运行效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551889A_ABST
    Figure CN122551889A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of bioinformatics data processing and discloses an intelligent data processing method for salt stress targeting bromelain phosphatase genes, including the following steps: S1, constructing a cross-species analysis dataset; S2, based on the cross-species analysis dataset, using protein sequence alignment tools and hidden Markov models to search the bromelain genome and obtain a candidate gene dataset; S3, sequentially performing conserved domain verification and physicochemical screening to remove invalid sequence data from the candidate genes and identify AcoPPs, members of the bromelain phosphatase family; S4, performing cluster analysis on the AcoPPs protein sequences, combined with multiple sequence alignment and phylogenetic trees, dividing AcoPPs into five subfamilies: PP2A, PP2C, DSP, PTP, and LMWP; S5, extracting high-dimensional features of genes from each subfamily and performing genomic and evolutionary feature data analysis; S6, processing and quantifying the expression of salt stress transcriptome data, and weighting and scoring according to the feature analysis results to screen for salt stress response candidate genes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bioinformatics data processing, specifically to an intelligent method for processing salt stress data targeting the bromelain phosphatase gene. Background Technology

[0002] Pineapple, a major tropical economic crop, is widely cultivated in many parts of South China. With the popularization of protected cultivation and multiple cropping models, soil salinization has become increasingly prominent, and salt stress has become a significant abiotic stress factor restricting the sustainable development of the pineapple industry. Protein phosphatases participate in multiple physiological processes in plant cells, including signal transduction, metabolism, and stress response, and are core functional proteins for plants to resist salt stress. Integrating and analyzing various data, including sequence, evolution, and transcriptional expression data of pineapple protein phosphatase genes, using bioinformatics techniques to screen for salt stress-related genes is a crucial preliminary step in elucidating the molecular mechanisms of pineapple salt tolerance and identifying superior gene resources.

[0003] Currently, bioinformatics analysis of plant gene families has become a standard procedure, mainly including sequence retrieval, clustering and typing, feature extraction, expression analysis, and functional evaluation. However, the existing analysis system still has several significant shortcomings. First, in the sequence retrieval stage, traditional tools such as BLAST and Hidden Markov Models use fixed thresholds for calculation, failing to consider the distribution differences of homologous sequences in different datasets, which easily introduces invalid sequence data and increases the workload of subsequent analysis. Second, in the gene clustering process, conventional unsupervised clustering algorithms use standard Euclidean distance for similarity calculation, assigning equal values ​​to all feature dimensions. This fails to reflect the actual contribution of different sequence features, causing the clustering results to deviate from the true evolutionary relationship and resulting in low accuracy in subfamily classification.

[0004] In addition, existing gene evaluation models are relatively simple: on the one hand, all genes within the same gene family are weighted and scored using a uniform weight, ignoring the differences in evolutionary characteristics and expression patterns of each subfamily, and the scoring results cannot truly reflect gene function; on the other hand, most schemes rely solely on the single indicator of expression level for judgment, without combining multi-dimensional data such as evolutionary conservation, protein-protein interaction, and functional enrichment for comprehensive evaluation, and also lack external dataset verification, resulting in poor stability of analysis results.

[0005] In summary, existing general-purpose data analysis algorithms and fixed processing rules cannot meet the data analysis needs of bromelain phosphatase genes. They suffer from problems such as large data bias, shallow feature mining, and one-sided evaluation system, and cannot accurately identify and classify salt stress response genes, thus limiting the development of subsequent basic research and applied work. Summary of the Invention

[0006] The purpose of this invention is to provide an intelligent data processing method for salt stress targeting the bromelain phosphatase gene, which solves the problems of existing plant gene analysis processes that use fixed parameters and uniform rules, resulting in inaccurate clustering, a single evaluation mechanism, and insufficient data analysis quality and stability.

[0007] To achieve the above objectives, the present invention adopts the following technical solution: A smart method for processing salt stress data targeting the bromelain phosphatase gene includes the following steps: S1. Construct a cross-species analysis dataset containing pineapple genome, gene annotations, protein sequences, and multi-species protein phosphatase reference sequence data; S2. Based on the cross-species analysis dataset, configure the search parameters, and use a protein sequence alignment tool and a hidden Markov model to search the bromelain group to obtain a candidate gene dataset. S3. Conserved domain verification and physicochemical screening were performed sequentially to eliminate invalid sequence data in candidate genes and identify AcoPPs, a member of the bromelain phosphatase family. S4. Cluster analysis of AcoPPs protein sequences, combined with multiple sequence alignment and phylogenetic tree, divides AcoPPs into five subfamilies: PP2A, PP2C, DSP, PTP, and LMWP. S5. Extract the high-dimensional features of genes from each subfamily, and after dimensionality reduction, perform genomic and evolutionary feature data analysis. S6. The transcriptome data of salt stress were processed and expressed quantitatively. Based on the feature analysis results, the data were weighted and scored to screen candidate genes for salt stress response.

[0008] Preferably, step S2 specifically includes: S21. Statistically analyze the distribution characteristics of sequences homologous to known plant protein phosphatases in the cross-species analysis dataset, and calculate the homologous sequence distribution density: ,in, The distribution density of homologous sequences, The number of homologous sequences in the dataset. This represents the total number of sequences in the dataset. S22. An adaptive threshold adjustment algorithm linked to the homology sequence distribution density is used to dynamically correct the retrieval parameters. The retrieval threshold is optimized by combining the homology adaptation correction coefficient. The correction formula is as follows: , , in, This represents the initial expected value for sequence alignment using a protein sequence alignment tool. This is the expected value for the corrected sequence alignment. To ensure initial amino acid sequence consistency, To ensure the consistency of the corrected amino acid sequence, For homologous adaptation correction coefficients; S23. A tiered joint search strategy is adopted. First, the protein sequence alignment tool is run with the corrected search parameters for initial screening. The sequence homology is initially compared and a preliminary screening sequence set is obtained. Then, the preliminary screening sequence set is input into the Hidden Markov Model to carry out fine identification of conserved structural domains. The preliminary screening results are screened and verified a second time to finally obtain a set of candidate protein phosphatase genes.

[0009] Preferably, step S3 specifically includes: S31. Cross-validation of conserved catalytic domains of candidate gene sequences was carried out using the Pfam database and NCBI Batch CD-Search. S32. Based on the pre-defined amino acid length range and protein hydrophilicity index that are adapted to the inherent properties of the bromelain phosphatase gene, the sequence data are screened. S33. Eliminate invalid sequence data that fail the above two-stage screening to identify AcoPPs, members of the bromelain phosphatase family.

[0010] Preferably, step S4 specifically includes: S41. Perform feature vectorization on all AcoPPs protein sequences to convert sequence information into multidimensional feature vectors. S42. The K-means unsupervised clustering algorithm is used to pre-group the feature vectors, and the improved weighted Euclidean distance is used as the sequence similarity metric. , in, , These are the feature vectors corresponding to two different protein sequences. , The eigenvector of the eigenvector 3D eigenvalues The total dimension of the feature vectors. These are the dimension weight coefficients; S43. Perform multiple sequence alignment on each protein sequence after clustering and grouping, and construct a phylogenetic tree of the corresponding taxa based on the alignment results. S44. Based on the combined clustering results and the evolutionary kinship of the phylogenetic tree, all AcoPPs are uniformly divided into five subfamilies: PP2A, PP2C, DSP, PTP, and LMWP.

[0011] Preferably, step S5 specifically includes: S51. Extract the gene structure, conserved motifs, cis-acting elements, and evolutionary pressure-related high-dimensional features corresponding to each subfamily, and construct the original high-dimensional feature dataset corresponding to each subfamily. S52. Use principal component analysis algorithm to perform dimensionality reduction transformation on the original high-dimensional feature dataset: , in, The original high-dimensional feature matrix, The matrix consists of orthogonal eigenvectors. This is the low-dimensional feature matrix obtained after dimensionality reduction; S53. Calculate the variance contribution rate of each principal component, and select and retain principal components whose cumulative variance contribution rate is not lower than the preset threshold as effective features. S54. Based on the effective features after screening, genomic analysis and evolutionary feature analysis were carried out for each subfamily to obtain complete feature analysis results for each subfamily.

[0012] Preferably, specifically: S61. Perform quality control, filtering and sequence alignment on the raw transcriptome data, and calculate the expression level of each gene based on the alignment results; S62. Calculate the fold change in gene expression based on gene expression levels, and screen for differentially expressed genes based on preset criteria. The formula for calculating the fold change in gene expression is as follows: , in, This represents the fold change in gene expression. This represents the gene expression levels in the stress-treated group. Gene expression levels in the blank control group; S63. Perform normalization operations on each feature using the range normalization method: , in, For the original numerical values ​​of the features, This represents the maximum value of the same feature across all gene samples. This represents the minimum value of the same feature across all gene samples. These are the normalized feature values, and the range of feature values ​​after normalization is unified into a fixed range; S64. Calculate the gene comprehensive score data using a multi-feature coupled weighted model: , in, The overall genetic score, , , , These represent the normalized differential expression fold, stress element copy number, evolutionary pressure value, and conserved motif integrity, respectively. , , , These are the weight coefficients for the corresponding features, and the sum of all weight coefficients equals 1. The characteristic coupling coefficient; S65. Sort the genes according to their comprehensive scores from high to low, and combine the feature analysis results obtained in step S5 to screen out the target salt stress response candidate genes.

[0013] Preferably, the specific process for calculating the gene expression level in step S61 is as follows: The number of sequencing fragments corresponding to each gene in the sequence alignment results and the effective gene length are statistically analyzed to calculate the gene expression level. , in, Gene expression level, To determine the number of sequencing fragments aligned to the target gene, The effective length of the target gene. To compare up to the first Number of sequencing fragments per gene For the first The effective length of each gene This is the sum of the ratios of the number of sequenced fragments of all genes to the effective length of the gene. This is the standardized magnification factor.

[0014] Preferably, before weighted scoring, step S64 further includes sub-family differential weight correction and hierarchical feature fusion calculation, specifically: Based on the five subfamilies divided in step S4, independent feature subsets are constructed for each family, and exclusive weight correction coefficients are configured. The normalized features are divided into an expressive feature layer and an evolutionarily conserved feature layer. The feature concatenation method is used to complete the hierarchical fusion. The dimension of the fused feature vector is consistent with the dimension of the reduced feature. By incorporating a specific weighting correction coefficient into the original weighting formula, a corrected comprehensive score is obtained and used as the core criterion for gene screening. , in, The corrected overall score, This is the weight correction coefficient corresponding to the subfamily.

[0015] Preferably, it further includes: S71. Based on the effective feature data obtained by dimensionality reduction in step S5, perform protein interaction network prediction, GO functional enrichment analysis and KEGG metabolic pathway enrichment analysis, and output the corresponding analysis results. S72. Retrieve the publicly available pineapple salt stress transcriptome dataset as an external validation set. Compare the expression patterns of the candidate genes selected in step S6 with the external validation set, calculate the expression pattern deviation rate, and remove genes with a deviation rate greater than a preset threshold. S73. A ternary ranking model was constructed by combining enrichment analysis results, protein interaction network connectivity, and gene comprehensive score. Based on the model calculation results, the salt stress response priority of the remaining candidate genes was divided. S74. Label each gene according to the priority of the division, output the candidate gene list, and complete the entire gene analysis and screening process.

[0016] Preferably, step S73 specifically includes: S731. Obtain the three evaluation indicators corresponding to the candidate genes: corrected comprehensive score, enrichment analysis significance, and protein interaction network connectivity. Use the range normalization method to normalize the three evaluation indicators, and combine the indicator characteristics to distinguish and process them. Perform reverse normalization operation on the enrichment analysis significance. S732. Assign corresponding weight coefficients to the three normalized evaluation indicators, and calculate the final ranking score of the candidate genes: , in, The final ranking score for candidate genes. , , These are the normalized gene-corrected composite score, the enrichment analysis significance score, and the protein-protein interaction network connectivity score, respectively. , , These are the weight coefficients corresponding to the three types of evaluation indicators, and the sum of each weight coefficient equals 1. This is a network correlation correction coefficient; S733. Based on the preset score range and the final ranking score of candidate genes, determine the priority of salt stress response at different levels.

[0017] By adopting the above technical solution, the present invention has the following advantages compared with the prior art: 1. This invention provides an intelligent data processing method for salt stress analysis of the bromelain phosphatase gene. It integrates multi-source information to construct a cross-species analysis dataset, utilizes protein sequence alignment tools and Hidden Markov Models for joint retrieval, and then employs a dual-screening process to eliminate invalid sequences. Finally, it combines cluster analysis, multiple sequence alignment, and phylogenetic tree analysis to classify the gene into subfamilies. Dimensionality reduction and analysis are then performed on high-dimensional features, and finally, expression quantification and weighted scoring are achieved based on transcriptome data. This method features a tightly integrated workflow and data interoperability. It customizes the overall processing logic for the application scenario of the bromelain phosphatase gene, overcoming the drawbacks of fragmented workflows and poor generalizability. From data source to final result, it achieves standardized and integrated analysis, effectively improving the overall standardization and operational efficiency of gene-related data processing.

[0018] 2. This invention provides an intelligent data processing method for salt stress targeting the bromelain phosphatase gene. It adds a homology sequence distribution density calculation module to the sequence retrieval stage and employs an adaptive threshold adjustment algorithm linked to homology density to dynamically correct retrieval parameters, replacing the fixed-parameter retrieval mode. This method can adaptively adjust the alignment threshold and consistency criteria based on the homology distribution characteristics of the dataset itself, effectively reducing the problem of interference sequences easily introduced by fixed parameters, improving the data purity in the joint retrieval stage, and providing a high-quality raw data foundation for subsequent analysis.

[0019] 3. This invention provides an intelligent data processing method for salt stress targeting the bromelain phosphatase gene. It uses a combination of dual verification and physicochemical index screening to remove invalid sequences, relies on professional databases to complete cross-validation of conserved domains, and combines species-specific physicochemical characteristics for screening. Dual quality control further filters out interference data remaining in the retrieval process, accurately retains the target gene sequence, and improves the accuracy of family member identification.

[0020] 4. This invention provides an intelligent data processing method for salt stress targeting the bromelain phosphatase gene. In the clustering stage, an improved weighted Euclidean distance is used as a similarity metric. Unlike the equal-weighted distance calculation method, it can assign corresponding weights according to the importance of different sequence features, effectively improving the accuracy of clustering, multiple sequence alignment and phylogenetic tree analysis, and making the gene subfamily classification results more consistent with the real evolutionary relationship.

[0021] 5. This invention provides an intelligent data processing method for salt stress targeting the bromelain phosphatase gene. It uses principal component analysis algorithm to reduce the dimensionality of high-dimensional features, thereby reducing the data dimensionality and computational load while retaining effective feature information, thus balancing analysis accuracy and computational efficiency. Based on the effective features after dimensionality reduction, genomic and evolutionary feature analysis can be carried out, which can fully explore the inherent data patterns of different subfamilies and deepen the analysis of gene features.

[0022] 6. This invention provides an intelligent data processing method for salt stress in the bromelain phosphatase gene. It constructs a multi-feature coupled weighted evaluation model, comprehensively calculating the overall score by integrating four types of features: differential expression fold, stress elements, evolutionary pressure, and conserved motifs. A feature coupling coefficient is set to balance the correlation effects of multiple features, overcoming the limitations of existing single-index evaluations. This allows the scoring results to comprehensively reflect the gene's salt stress response potential, resulting in a more complete evaluation dimension and more objective results. Furthermore, the gene scoring process incorporates subfamily differential weight correction and hierarchical feature fusion processing. Specific correction coefficients are configured for different subfamilies, and feature fusion is performed by distinguishing between expression features and evolutionarily conserved features. This fully considers the inherent differences between different subfamilies within a gene family, further optimizing the scoring results and enhancing the scientific rigor of the evaluation system. Attached Figure Description

[0023] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0025] Example Please refer to Figure 1 As shown, this invention discloses an intelligent data processing method for salt stress targeting the bromelain phosphatase gene, comprising the following steps: S1. Construct a cross-species analysis dataset containing pineapple genome, gene annotations, protein sequences, and reference sequences of protein phosphatases from multiple species.

[0026] This embodiment uses the latest complete pineapple telomere-to-telomere (T2T) reference genome as the core data source, derived from the pineapple genome database published by the Chinese Academy of Tropical Agricultural Sciences. Simultaneously, protein phosphatase sequences from three model crops—Arabidopsis thaliana, rice, and maize—were selected as cross-species reference data, with corresponding genomes, gene annotation files, and protein sequence files downloaded from biological databases such as NCBI, TAIR10, RGAP v5.0, and MaizeGDB. All raw data underwent standardized format conversion, repetitive sequence removal, and fragment cleanup to complete data standardization and regularization, ultimately constructing a complete cross-species analysis dataset to provide a reliable data source for subsequent retrieval operations.

[0027] S2. Based on the cross-species analysis dataset, configure the search parameters, and use the protein sequence alignment tool (BLASTP) and hidden Markov model (HMM) to search the bromelain group and obtain a candidate gene dataset.

[0028] S21. Statistically analyze the distribution characteristics of sequences homologous to known plant protein phosphatases in the cross-species analysis dataset, and calculate the homologous sequence distribution density: ,in, The distribution density of homologous sequences, The number of homologous sequences in the dataset. The total number of sequences in the dataset is used to calculate the homology distribution density by iterating through the data and comparing the number of homologous sequences in the dataset with the total number of sequences. This quantifies the overall homology characteristics of the dataset and provides a quantitative basis for dynamically adjusting retrieval parameters.

[0029] S22. An adaptive threshold adjustment algorithm linked to the homology sequence distribution density is used to dynamically correct the retrieval parameters. The retrieval threshold is optimized by combining the homology adaptation correction coefficient. The correction formula is as follows: , , in, For the initial E-value of BLASTP, For the corrected E-value, To ensure initial amino acid sequence consistency, To ensure the consistency of the corrected amino acid sequence, The homology adaptation correction coefficient is adaptively selected based on the overall homology features of the dataset. E-value represents the expected value of sequence alignment; the smaller the value, the higher the homology reliability. Amino acid similarity is used to assess the degree of sequence similarity. The initial fixed search parameters in this implementation are set to... With an amino acid sequence consistency of ≥45%, combined with the above adaptive algorithm and the dynamic correction parameters of the homology fitting coefficient, it adapts to the homology features of different datasets and avoids missed detections or false detections caused by fixed parameters.

[0030] S23. A tiered joint retrieval strategy is adopted. First, BLASTP is run with the corrected retrieval parameters for initial screening. Preliminary sequence homology is compared to obtain an initial screening sequence set. Then, the initial screening sequence set is input into a hidden Markov model to carry out fine identification of conserved structural domains. The initial screening results are screened and verified a second time to remove repetitive sequences and finally obtain a set of candidate protein phosphatase genes.

[0031] This implementation downloaded six core hidden Markov models of protein phosphatases—PF00149, PF00481, PF01451, PF00782, PF00102, and PF00287—from the Pfam database. A two-stage combined approach was used: BLASTP for coarse screening and HMMER (a sequence retrieval software for hidden Markov models) for fine screening. The results from the two sets of searches were merged and deduplicated to obtain a preliminary candidate gene sequence set.

[0032] S3. Conserved domain verification and physicochemical screening were performed sequentially to eliminate invalid sequence data in candidate genes and identify AcoPPs (bromelain phosphatase family genes), members of the bromelain phosphatase family.

[0033] S31. Cross-validation of conserved catalytic domains of candidate gene sequences was carried out using the Pfam database (conserved protein domain database) and NCBI (bioinformatics public database) Batch CD-Search (conserved domain batch search tool). S32. Based on the pre-defined amino acid length range and protein hydrophilicity index that are adapted to the inherent properties of the bromelain phosphatase gene, the sequence data are screened. S33. Eliminate invalid sequence data that fail the above two-stage screening to identify AcoPPs, members of the bromelain phosphatase family; S4. Cluster analysis of AcoPPs protein sequences, combined with multiple sequence alignment and phylogenetic tree analysis, classified AcoPPs into five subfamilies: PP2A (protein phosphatase 2A), PP2C (protein phosphatase 2C), DSP (bispecific phosphatase), PTP (protein tyrosine phosphatase), and LMWP (low molecular weight protein phosphatase). This five-subfamily classification system conforms to the general classification standards for plant protein phosphatases and is consistent with the classification rules of model crops such as rice, maize, and Arabidopsis thaliana, reflecting the evolutionary conservatism of this gene family. S41. Perform feature vectorization processing on all AcoPPs protein sequences to transform sequence information into multidimensional feature vectors. This process converts character-type protein sequences into computer-recognizable multidimensional numerical feature vectors, completing the structuring transformation of unstructured data and meeting the input requirements of clustering algorithms.

[0034] S42. Use the K-means unsupervised clustering algorithm to pre-group the feature vectors, and use the improved weighted Euclidean distance as the sequence similarity metric: , in, , These are the feature vectors corresponding to two different protein sequences. , The eigenvector of the eigenvector 3D eigenvalues The total dimension of the feature vectors. These are dimensional weight coefficients, assigned based on the importance of protein sequence features.

[0035] Existing Euclidean distance assigns equal values ​​to all feature dimensions. This scheme sets weights differently based on the importance of functional sites and conserved regions, making the similarity calculation results more consistent with actual biological laws.

[0036] The K-means unsupervised clustering algorithm sets the number of clusters and the maximum number of iterations. After each iteration, the cluster center offset is calculated. When the offset is less than the convergence threshold preset according to experimental requirements, the algorithm is considered to have converged, the clustering operation is terminated, and the grouping results are output. The number of clusters is configured according to the preset number of subfamilies, and a reasonable maximum number of iterations and convergence threshold are set. The cluster center offset is used as the convergence criterion to ensure that the clustering results are stable and do not change significantly.

[0037] S43. Perform multiple sequence alignment on each protein sequence after clustering and grouping, and construct a phylogenetic tree of the corresponding taxa based on the alignment results. Specifically, multiple sequence alignment is performed using multiple sequence alignment software (MAFFT v7.505), alignment result pruning software (TrimAI) removes low-quality regions from the alignment results, and then a maximum likelihood phylogenetic tree is constructed using phylogenetic tree construction software (IQ-TREE v3.1.1) combined with ultrafast bootstrapping. Finally, the visualization is completed on the evolutionary tree visualization platform (iTOL) to intuitively display the evolutionary kinship of the sequences.

[0038] S44. Based on the combined clustering results and the evolutionary kinship of the phylogenetic tree, all AcoPPs were uniformly divided into five subfamilies: PP2A, PP2C, DSP, PTP, and LMWP. Among them, PP2C is the dominant subfamily with the most members, followed by PP2A and DSP, while PTP and LMWP have the fewest members. The evolutionary characteristics of each subfamily are clearly distinguishable.

[0039] S5. Extract high-dimensional features from genes of each subfamily, perform dimensionality reduction processing, and then conduct genomic and evolutionary feature data analysis. High-dimensional features include six categories: gene structure, conserved motifs, promoter cis-acting elements, chromosome location, gene replication, and evolutionary selection pressure, covering the core analytical dimensions at the genomic and evolutionary levels.

[0040] S51. Extract the gene structure, conserved motifs, cis-acting elements, and evolutionary pressure-related high-dimensional features corresponding to each subfamily, and construct the original high-dimensional feature dataset corresponding to each subfamily. Specifically, the study utilizes the Bioinformatics Comprehensive Analysis Toolkit (TBtools) to extract gene structure and chromosome location information, the Conserved Motif Analysis Toolkit (MEME Suite) to identify conserved motifs, and the Plant Promoter Cis-Action Elements Database (PlantCARE) to predict promoter cis-action elements. Simultaneously, the study calculates the Ka (non-synonymous substitution rate) and Ks (synonymous substitution rate) values ​​of repetitive gene pairs to characterize evolutionary pressure, and integrates all information to construct a high-dimensional feature dataset.

[0041] S52. The Principal Component Analysis (PCA) algorithm is used to perform dimensionality reduction transformation on the original high-dimensional feature dataset, remove redundant information between features, reduce data dimensionality while retaining core effective information, reduce subsequent computation, and improve overall analysis efficiency. , in, The original high-dimensional feature matrix, The matrix consists of orthogonal eigenvectors. This is the low-dimensional feature matrix obtained after dimensionality reduction; S53. Calculate the variance contribution rate corresponding to each principal component, and select and retain principal components whose cumulative variance contribution rate is not lower than the threshold preset according to experimental requirements as effective features. The variance contribution rate represents the information proportion of a single feature. Select features according to the cumulative contribution rate to ensure that key biological information is not lost while simplifying the data.

[0042] S54. Based on the effective features after screening, genomic analysis and evolutionary feature analysis were carried out for each subfamily to obtain complete feature analysis results for each subfamily.

[0043] S6. The transcriptome data of salt stress were processed and expressed quantitatively. Based on the feature analysis results, the data were weighted and scored to screen candidate genes for salt stress response.

[0044] S61. Perform quality control, filtering, and sequence alignment on the raw transcriptome data, and calculate the expression level of each gene based on the alignment results.

[0045] Specifically, the sequencing data format conversion tool (SRA Toolkit) is used to convert the data format, the sequencing data quality assessment software (FastQC) is used to complete the quality assessment, the sequencing data filtering software (Trimmomatic) is used to remove adapters and low-quality bases, and then the sequencing sequence alignment software (HISAT2) is used to align the clean reads (high-quality sequencing reads) to the pineapple reference genome. Finally, the gene expression quantification software (StringTie) is used to complete the gene expression quantification.

[0046] S62. Calculate the fold change in gene expression based on gene expression levels, and screen for differentially expressed genes based on pre-set criteria according to experimental requirements. The formula for calculating the fold change in gene expression is: , in, This represents the fold change in gene expression. This represents the gene expression levels in the stress-treated group. Gene expression levels in the blank control group; screening criteria set for this implementation: Significance This is used to define differentially expressed genes.

[0047] S63. Perform normalization operations on each feature using the range normalization method: , in, For the original numerical values ​​of the features, This represents the maximum value of the same feature across all gene samples. This represents the minimum value of the same feature across all gene samples. The normalized feature values ​​are defined as a fixed range. Different indicators have different dimensions and numerical ranges. Range normalization unifies the numerical range, eliminates the influence of dimensions, and provides a unified data standard for the joint weighted calculation of multiple indicators.

[0048] S64. Calculate the gene comprehensive score data using a multi-feature coupled weighted model: , in, The overall genetic score, , , , These represent the normalized differential expression fold, stress element copy number, evolutionary pressure value, and conserved motif integrity, respectively. , , , These are the weight coefficients for the corresponding features, and the sum of all weight coefficients equals 1. This is the feature coupling coefficient, used to balance the correlation effects between multiple types of features.

[0049] By combining four core biological characteristics for joint scoring and integrating coupling coefficients to coordinate the intrinsic relationships between characteristics, the evaluation dimensions are comprehensive and the results are more in line with the actual stress response capabilities of genes.

[0050] S65. Sort the genes according to their comprehensive scores from high to low, and combine the results of the genomic and evolutionary characteristic analysis obtained in step S5 to screen out candidate genes for salt stress response.

[0051] The specific process for calculating gene expression levels in step S61 is as follows: The number of sequencing fragments corresponding to each gene in the sequence alignment results and the effective gene length are statistically analyzed to calculate the gene expression level. , in, Gene expression level, To determine the number of sequencing fragments aligned to the target gene, The effective length of the target gene. To compare up to the first Number of sequencing fragments per gene For the first The effective length of each gene This is the sum of the ratios of the number of sequenced fragments of all genes to the effective length of the gene. This is the standardized magnification factor.

[0052] Step S64, before weighted scoring, also includes sub-family differential weight correction and hierarchical feature fusion calculation, specifically: Based on the five subfamilies divided in step S4, independent feature subsets are constructed and exclusive weight correction coefficients are configured. Since the five major subfamilies PP2A, PP2C, DSP, PTP, and LMWP have natural differences in evolutionary patterns and stress response laws, correction coefficients are configured separately for each subfamilies, abandoning the unified evaluation standard.

[0053] The normalized features are divided into an expressive feature layer and an evolutionarily conserved feature layer. The feature concatenation method is used to complete the hierarchical fusion. The dimension of the fused feature vector is consistent with the dimension of the reduced feature. By incorporating a specific weighting correction coefficient into the original weighting formula, a corrected comprehensive score is obtained and used as the core criterion for gene screening. , in, The corrected overall score, This is the weight correction coefficient corresponding to the subfamily.

[0054] Subfamily-specific coefficients were used to correct the scores, ensuring that the evaluation criteria matched the biological characteristics of each subfamily. The corrected results showed that upregulated genes were mainly distributed in the PP2C, PP2A, and DSP subfamilies, while downregulated genes were members of the PP2C subfamily.

[0055] S71. Based on the effective feature data obtained from the dimensionality reduction in step S5, perform protein interaction network prediction, GO gene ontology functional enrichment analysis, and KEGG (metabolic pathway database) metabolic pathway enrichment analysis, and output the corresponding analysis results. The AcoPP protein sequence was submitted to the protein interaction network database (STRING) to predict protein interaction relationships. Functional annotation was completed using the gene function annotation tool (EggNog-mapper), followed by GO and KEGG enrichment analysis.

[0056] S72. Retrieve the publicly available pineapple salt stress transcriptome dataset as an external validation set. Compare the expression patterns of the candidate genes selected in step S6 with the external validation set, calculate the expression pattern deviation rate, and remove genes with a deviation rate greater than the pre-set value according to experimental requirements. S73. A ternary ranking model was constructed by combining enrichment analysis results, protein interaction network connectivity, and gene comprehensive score. Based on the model calculation results, the salt stress response priority of the remaining candidate genes was divided. The ternary ranking model is defined as a comprehensive evaluation model that uses three core evaluation indicators—corrected comprehensive score, enrichment analysis significance, and protein interaction network connectivity—as the core evaluation indicators. After normalization and weight allocation, the final ranking score is calculated and the response priority is determined based on the score.

[0057] S731. Obtain the three evaluation indicators corresponding to the candidate genes: corrected comprehensive score, enrichment analysis significance, and protein interaction network connectivity. Use the range normalization method to normalize the three evaluation indicators, and combine the indicator characteristics to distinguish and process them. Perform reverse normalization operation on the enrichment analysis significance. The smaller the significance value of enrichment analysis, the higher the degree of enrichment and the stronger the functional correlation. Therefore, reverse normalization is performed on this indicator to unify the evaluation logic of all indicators.

[0058] S732. Assign corresponding weight coefficients to the three normalized evaluation indicators, and calculate the final ranking score of the candidate genes: , in, The final ranking score for candidate genes. , , These are the normalized gene-corrected composite score, the enrichment analysis significance score, and the protein-protein interaction network connectivity score, respectively. , , These are the weight coefficients corresponding to the three types of evaluation indicators, and the sum of each weight coefficient equals 1. This is the network association correction coefficient, used to characterize the influence of protein-protein interaction network association strength; S733. Based on the score range set in advance according to the experimental requirements, and combined with the final ranking score of the candidate genes, different levels of salt stress response priorities are determined.

[0059] S74. Mark each gene according to the priority of the division, output the candidate gene list, complete the entire process of gene analysis and screening, and facilitate subsequent molecular biology verification, stress resistance mechanism research and breeding application.

[0060] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for intelligent processing of salt stress data oriented to a pineapple protein phosphatase gene, characterized by, Includes the following steps: S1. Construct a cross-species analysis dataset containing pineapple genome, gene annotations, protein sequences, and multi-species protein phosphatase reference sequence data; S2. Based on the cross-species analysis dataset, configure the search parameters, and use a protein sequence alignment tool and a hidden Markov model to search the bromelain group to obtain a candidate gene dataset. S3. Conserved domain verification and physicochemical screening were performed sequentially to eliminate invalid sequence data in candidate genes and identify AcoPPs, a member of the bromelain phosphatase family. S4. Cluster analysis of AcoPPs protein sequences, combined with multiple sequence alignment and phylogenetic tree, divides AcoPPs into five subfamilies: PP2A, PP2C, DSP, PTP, and LMWP. S5. Extract the high-dimensional features of genes from each subfamily, and after dimensionality reduction, perform genomic and evolutionary feature data analysis. S6. The transcriptome data of salt stress were processed and expressed quantitatively. Based on the feature analysis results, the data were weighted and scored to screen candidate genes for salt stress response.

2. The salt stress data intelligent processing method for a pineapple protein phosphatase gene according to claim 1, characterized in that, Step S2 is as follows: S21, statistics the distribution characteristics of sequences homologous to known plant protein phosphatase in the cross-species analysis dataset, and calculates the homologous sequence distribution density: wherein, is the homologous sequence distribution density, is the number of homologous sequences in the dataset, is the total number of all sequences in the dataset; S22. An adaptive threshold adjustment algorithm linked to the homology sequence distribution density is used to dynamically correct the retrieval parameters. The retrieval threshold is optimized by combining the homology adaptation correction coefficient. The correction formula is as follows: , , wherein, is the initial sequence alignment expectation for the protein sequence alignment tool, is the revised sequence alignment expectation, is the initial amino acid sequence identity, is the revised amino acid sequence identity, is the homology fit revision coefficient. S23. A tiered joint search strategy is adopted. First, the protein sequence alignment tool is run with the corrected search parameters for initial screening. The sequence homology is initially compared and a preliminary screening sequence set is obtained. Then, the preliminary screening sequence set is input into the Hidden Markov Model to carry out fine identification of conserved structural domains. The preliminary screening results are screened and verified a second time to finally obtain a set of candidate protein phosphatase genes.

3. The method for intelligent processing of salt stress data oriented to a pineapple protein phosphatase gene according to claim 1, wherein, Step S3 is as follows: S31. Cross-validation of conserved catalytic domains of candidate gene sequences was carried out using the Pfam database and NCBI Batch CD-Search. S32. Based on the pre-defined amino acid length range and protein hydrophilicity index that are adapted to the inherent properties of the bromelain phosphatase gene, the sequence data are screened. S33. Eliminate invalid sequence data that fail the above two-stage screening to identify AcoPPs, members of the bromelain phosphatase family.

4. The salt stress data intelligent processing method for a pineapple protein phosphatase gene according to claim 1, wherein, Step S4 is as follows: S41. Perform feature vectorization on all AcoPPs protein sequences to convert sequence information into multidimensional feature vectors. S42. The K-means unsupervised clustering algorithm is used to pre-group the feature vectors, and the improved weighted Euclidean distance is used as the sequence similarity metric. , in, , These are the feature vectors corresponding to two different protein sequences. , The eigenvector of the eigenvector 3D eigenvalues The total dimension of the feature vectors. These are the dimension weight coefficients; S43. Perform multiple sequence alignment on each protein sequence after clustering and grouping, and construct a phylogenetic tree of the corresponding taxa based on the alignment results; S44. Based on the combined clustering results and the evolutionary kinship of the phylogenetic tree, all AcoPPs are uniformly divided into five subfamilies: PP2A, PP2C, DSP, PTP, and LMWP.

5. The salt stress data intelligent processing method for a pineapple protein phosphatase gene according to claim 1, wherein, Step S5 is as follows: S51. Extract the gene structure, conserved motifs, cis-acting elements, and evolutionary pressure-related high-dimensional features corresponding to each subfamily, and construct the original high-dimensional feature dataset corresponding to each subfamily. S52. Use principal component analysis algorithm to perform dimensionality reduction transformation on the original high-dimensional feature dataset: , wherein, is the original high-dimensional feature matrix, is the orthogonal eigenvector matrix, is the low-dimensional feature matrix obtained after dimension reduction; S53. Calculate the variance contribution rate of each principal component, and select and retain principal components whose cumulative variance contribution rate is not lower than the preset threshold as effective features. S54. Based on the effective features after screening, genomic analysis and evolutionary feature analysis were carried out for each subfamily to obtain complete feature analysis results for each subfamily.

6. The salt stress data intelligent processing method for a pineapple protein phosphatase gene according to claim 1, wherein, Step S6 is as follows: S61. Perform quality control, filtering and sequence alignment on the raw transcriptome data, and calculate the expression level of each gene based on the alignment results; S62. Calculate the fold change in gene expression based on gene expression levels, and screen for differentially expressed genes based on preset criteria. The formula for calculating the fold change in gene expression is as follows: , wherein, is the gene differential expression fold, is the gene expression amount of the stress treatment group, is the gene expression amount of the blank control group; S63. Perform normalization operations on each feature using the range normalization method: , wherein, is a characteristic original value, is a maximum value of the same characteristic in the entire gene sample, is a minimum value of the same characteristic in the entire gene sample, is a normalized characteristic value, and the normalized characteristic value is in a fixed interval. S64. Calculate the gene comprehensive score data using a multi-feature coupled weighted model: , wherein, is the gene comprehensive score, , , , are the normalized differential expression fold, stress element copy number, evolutionary pressure value, and conserved motif integrity, respectively, , , , are the weight coefficients of the corresponding features, and the sum of each weight coefficient is equal to 1, is the feature coupling coefficient; S65. Sort the genes according to their comprehensive scores from high to low, and combine the feature analysis results obtained in step S5 to screen out the target salt stress response candidate genes.

7. The salt stress data intelligent processing method for a pineapple protein phosphatase gene according to claim 6, wherein, The specific process for calculating gene expression levels in step S61 is as follows: The number of sequencing fragments corresponding to each gene in the sequence alignment results and the effective gene length are statistically analyzed to calculate the gene expression level. , wherein, is the gene expression amount, is the number of sequencing fragments aligned to the target gene, is the effective length of the target gene, is the number of sequencing fragments aligned to the th gene, is the effective length of the th gene, is the sum of the ratio of the number of sequencing fragments to the effective length of all genes, is the standardization amplification factor.

8. The method of claim 6, wherein the salt stress data of the pineapple protein phosphatase gene is processed intelligently, and Step S64, before weighted scoring, also includes sub-family differential weight correction and hierarchical feature fusion calculation, specifically: ​ Based on the five subfamilies divided in step S4, independent feature subsets are constructed for each family, and exclusive weight correction coefficients are configured. The normalized features are divided into an expressive feature layer and an evolutionarily conserved feature layer. The feature concatenation method is used to complete the hierarchical fusion. The dimension of the fused feature vector is consistent with the dimension of the reduced feature. By incorporating a specific weighting correction coefficient into the original weighting formula, a corrected comprehensive score is obtained and used as the core criterion for gene screening. , wherein, is the corrected overall score, is the weight correction coefficient corresponding to the subfamily.

9. The salt stress data intelligent processing method for a pineapple protein phosphatase gene according to claim 1, wherein, Also includes: S71. Based on the effective feature data obtained by dimensionality reduction in step S5, perform protein interaction network prediction, GO functional enrichment analysis and KEGG metabolic pathway enrichment analysis, and output the corresponding analysis results. S72. Retrieve the publicly available pineapple salt stress transcriptome dataset as an external validation set. Compare the expression patterns of the candidate genes selected in step S6 with the external validation set, calculate the expression pattern deviation rate, and remove genes with a deviation rate greater than a preset threshold. S73. A ternary ranking model was constructed by combining enrichment analysis results, protein interaction network connectivity, and gene comprehensive score. Based on the model calculation results, the salt stress response priority of the remaining candidate genes was divided. S74. Label each gene according to the priority of the division, output the candidate gene list, and complete the entire gene analysis and screening process.

10. The salt stress data intelligent processing method for a pineapple protein phosphatase gene according to claim 9, wherein, Step S73 is as follows: S731. Obtain the three evaluation indicators corresponding to the candidate genes: corrected comprehensive score, enrichment analysis significance, and protein interaction network connectivity. Use the range normalization method to normalize the three evaluation indicators, and combine the indicator characteristics to distinguish and process them. Perform reverse normalization operation on the enrichment analysis significance. S732. Assign corresponding weight coefficients to the three normalized evaluation indicators, and calculate the final ranking score of the candidate genes: , wherein, is the final ranking score of the candidate gene, , , are the normalized gene correction comprehensive score, the enrichment analysis significance score, and the protein interaction network connectivity score, respectively, , , are the weight coefficients corresponding to the three types of evaluation indexes, respectively, and the sum of each weight coefficient is equal to 1, is the network correlation correction coefficient. S733、According to the preset score interval, combined with the final ranking score of the candidate gene, the priority of salt stress response of different levels is determined.