Breast cancer FOXO3 gene dynamic distribution metering method based on subcellular localization means
By employing a subcellular localization approach combined with a hidden Markov model and machine learning, the challenges of locating the FOXO3 gene and integrating multi-dimensional data were overcome. This enabled precise dynamic distribution assessment of the FOXO3 gene in breast cancer, improving the accuracy of localization prediction and its clinical translational value.
Patent Information
- Application Number
- CN202510883375.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-29
- Publication Date
- 2025-10-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies cannot effectively address the difficulties in subcellular localization of the FOXO3 gene, the lack of multidimensional data integration, and the technical gap in prediction and validation. In particular, the lack of targeted analysis in breast cancer results in limited clinical translational value of localization prediction results.
Using a subcellular localization approach, data on FOXO3-related monotypic basic amino acids, two clusters of basic amino acids, and phosphorylation-dependent amino acids were acquired through a data acquisition module. The amino acid sequences were analyzed using a hidden Markov model and machine learning methods. Combined with fluorescence localization experiments, the dynamic distribution range of the FOXO3 gene was verified, forming a closed-loop calibration mechanism.
It significantly improves the accuracy and robustness of FOXO3 gene localization prediction, provides a personalized FOXO3 activity assessment scheme, supports clinical trial enrollment screening for targeted drugs, shortens the drug development cycle, and constructs a subcellular localization analysis framework applicable to other transcription factors.
Smart Images

Figure CN120808874A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of tumor cell metabolism analysis, and particularly relates to a breast cancer FOXO3 gene dynamic distribution metering method based on a subcellular localization means. BACKGROUND
[0002] As a key transcription factor regulating cell apoptosis, proliferation and stress response, the subcellular localization (nuclear / cytoplasmic distribution) of FOXO3 directly reflects its activity state: enrichment in the nucleus activates the expression of downstream tumor suppressor genes, while cytoplasmic retention indicates inactivation (such as nuclear export caused by phosphorylation modification). In triple-negative breast cancer (TNBC) and other cancer types lacking clear driver gene mutations, abnormal regulation of the FOXO3 pathway has been confirmed to be closely related to tumor malignant progression, so accurate assessment of its dynamic distribution is crucial for elucidating the pathogenesis and developing targeted therapies. However, the existing technology faces the following core problems:
[0003] 1. Difficulty in direct localization due to structural complexity: The winged helix domain of FOXO3 serves as a hub in the signaling network, integrating multiple pathways such as PI3K / AKT and MAPK. Its nuclear localization signal (NLS) is dynamically regulated by multiple modifications such as phosphorylation and ubiquitination. Traditional prediction methods based on single sequence characteristics (such as classic NLS motifs) cannot adapt to localization analysis under complex modification backgrounds;
[0004] 2. Lack of multi-dimensional data integration: The subcellular localization of FOXO3 relies on the synergistic action of basic amino acid clusters, phosphorylation sites and spatial structures. Existing technologies often analyze single features in isolation (such as predicting NLS only through tools like cNLS Mapper), lacking systematic integration of multi-dimensional data such as single-typed basic amino acids and phosphorylation-induced sites;
[0005] 3. Technical gap between prediction and verification: Traditional bioinformatics prediction does not form a closed loop with experimental verification, especially lacking localization calibration specific to breast cancer cells (such as correlation analysis of nuclear-cytoplasmic fluorescence intensity ratio and pathological characteristics), resulting in limited clinical translation value of prediction results.
[0006] Therefore, there is an urgent need to design a technical solution to solve at least one of the above technical problems. SUMMARY
[0007] The present application provides a breast cancer FOXO3 gene dynamic distribution metering method based on a subcellular localization means, aiming to solve the problems of direct localization difficulty caused by structural complexity, lack of multi-dimensional data integration and technical gap between prediction and verification in the existing technology.
[0008] In a first aspect, the present application provides a breast cancer FOXO3 gene dynamic distribution measurement system based on subcellular localization means, comprising:
[0009] a data acquisition module configured to acquire FOXO3-related single-typed basic amino acid data, two-cluster basic amino acid data, and phosphorylation-induced amino acid data;
[0010] a control module connected to the data acquisition module, configured to process the amino acid data based on a hidden Markov model, predict the amino acid sequence of the FOXO3 gene, analyze the amino acid sequence by a machine learning method, provide the potential location and confidence of the nuclear localization signal, and predict the nuclear protein and NLS region; based on the prediction result of the NLS region, recursively predict the dynamic distribution range of the FOXO3 gene in the nucleus or cytoplasm; after predicting the NLS using bioinformatics tools, experimentally verify the subcellular localization function of the FOXO3 gene by combining fluorescence localization experiments to confirm the dynamic distribution range;
[0011] The trial method of the hidden Markov model is used to process the hidden state transition probability of the amino acid sequence, and the machine learning method includes a support vector machine or a neural network algorithm to realize the position prediction and confidence evaluation of the NLS region.
[0012] In some embodiments, the FOXO3-related single-typed basic amino acid data, two-cluster basic amino acid data, and phosphorylation-induced amino acid data are acquired by: performing amino acid sequence analysis on FOXO3 protein in breast cancer cell samples by mass spectrometry, extracting single-typed basic amino acid site data containing arginine and lysine; identifying two-cluster basic amino acid clusters with a sequence alignment algorithm, with a continuous or interval of no more than 3 amino acids; and combining a phosphorylation site database and a kinase substrate prediction tool to obtain the modification state and site coordinate data of phosphorylation-induced amino acid sites regulated by the PI3K or AKT pathway.
[0013] In some embodiments, the hidden Markov model is used to process the amino acid data to predict the amino acid sequence of the FOXO3 gene, including: constructing a HMM state transition matrix containing amino acid residue hidden states, traversing the FOXO3 sequence with a sliding window, and calculating the joint probability of the hidden state sequence in each window; decoding the optimal hidden state path by a Viterbi algorithm, screening out short sequence fragments with a probability higher than 0.75, the short sequence fragments containing at least one single-typed basic amino acid or two-cluster basic amino acid cluster, and the distance from the phosphorylation-induced amino acid site being no more than 10 amino acids.
[0014] In some embodiments, the analysis of the amino acid sequence by the machine learning method provides potential positions and confidence of the nuclear localization signal, comprising: converting the short sequence predicted by the HMM into a feature vector containing a position-specific scoring matrix, a hydrophobicity index, and a charge distribution, and inputting the feature vector into a support vector machine or a deep neural network model; training the model using a positive and negative sample set containing known NLS sequences, optimizing the regularization parameter and kernel function by grid search, and outputting the confidence score of each candidate position as NLS, which is normalized to the interval [0, 1] by the softmax function, and screening positions with a score ≥0.8 as potential NLS positions.
[0015] In some embodiments, the prediction of the nuclear protein and the NLS region comprises: mapping the potential NLS positions to the UniProt nuclear protein database to extract feature patterns containing classical NLS, simplified NLS, and novel NLS; and delineating the NLS core region and the flanking regulatory region by integrating the NLS confidence score, the basic amino acid density, and the phosphorylation site interference coefficient through a weighted scoring mechanism, and generating a nuclear protein localization probability distribution map.
[0016] In some embodiments, the prediction result based on the NLS region recursively predicts the dynamic distribution range of FOXO3 gene in the nucleus or cytoplasm, comprising: establishing a subcellular localization statistical model to correlate the integrity of the NLS region with the nuclear-cytoplasmic distribution ratio; when the NLS core region is not modified by phosphorylation, the nuclear localization probability is ≥60%, and the distribution range covers the euchromatin region in the nucleus; when the NLS region has AKT kinase-mediated phosphorylation sites, the nuclear localization probability decreases to ≤30%, and the distribution range shifts to the cytoplasmic matrix.
[0017] In some embodiments, after predicting the NLS using bioinformatics tools, the subcellular localization function of FOXO3 gene is experimentally verified by fluorescence localization experiments to confirm the dynamic distribution range, comprising: constructing a FOXO3 fluorescent fusion protein vector by gene cloning technology, transfecting a breast cancer cell line; collecting live cell images using a laser confocal microscope, and analyzing the distribution threshold of the fluorescent signal inside and outside the nuclear membrane; calculating the nuclear fluorescence intensity ratio, when the predicted NLS region exists, the nuclear fluorescence intensity ratio ≥0.6 is determined as the nuclear localization advantage, the nuclear-cytoplasmic separation protein is detected to verify the localization result, and a closed-loop calibration mechanism of bioinformatics prediction and experimental verification is formed to confirm the dynamic distribution range.
[0018] In a second aspect, the present application provides a breast cancer FOXO3 gene dynamic distribution metering method based on subcellular localization means, characterized by being applied to the control module of the breast cancer FOXO3 gene dynamic distribution metering system based on subcellular localization means provided in any of the embodiments of the present application, and the method comprises:
[0019] The FOXO3-related single-typed basic amino acid data, two-cluster basic amino acid data and phosphorylation-induced amino acid data collected by the data acquisition module are acquired;
[0020] The amino acid data is processed based on a hidden Markov model to predict the amino acid sequence of the FOXO3 gene; the amino acid sequence is analyzed by a machine learning method to provide potential positions and confidence of the nuclear localization signal and to predict the nuclear protein and the NLS region;
[0021] Based on the prediction result of the NLS region, the dynamic distribution range of the FOXO3 gene in the nucleus or cytoplasm is recursively calculated; after the NLS is predicted by the bioinformatics tool, the subcellular localization function of the FOXO3 gene is verified by the fluorescence localization experiment to confirm the dynamic distribution range; wherein, the trial algorithm of the hidden Markov model is used to process the hidden state transition probability of the amino acid sequence, and the machine learning method includes a support vector machine or a neural network algorithm to realize the position prediction and confidence evaluation of the NLS region.
[0022] In a third aspect, the present application provides a breast cancer FOXO3 gene dynamic distribution metering device based on a subcellular localization means, which is applied to the control module of the breast cancer FOXO3 gene dynamic distribution metering system based on the subcellular localization means provided in any of the embodiments of the present application, and the device comprises:
[0023] A data acquisition unit is configured to acquire FOXO3-related single-typed basic amino acid data, two-cluster basic amino acid data and phosphorylation-induced amino acid data collected by a data acquisition module;
[0024] A region prediction unit is configured to process the amino acid data based on a hidden Markov model to predict the amino acid sequence of the FOXO3 gene; the amino acid sequence is analyzed by a machine learning method to provide potential positions and confidence of the nuclear localization signal and to predict the nuclear protein and the NLS region;
[0025] An evaluation implementation unit is configured to recursively calculate the dynamic distribution range of the FOXO3 gene in the nucleus or cytoplasm based on the prediction result of the NLS region; after the NLS is predicted by the bioinformatics tool, the subcellular localization function of the FOXO3 gene is verified by the fluorescence localization experiment to confirm the dynamic distribution range; wherein, the trial algorithm of the hidden Markov model is used to process the hidden state transition probability of the amino acid sequence, and the machine learning method includes a support vector machine or a neural network algorithm to realize the position prediction and confidence evaluation of the NLS region.
[0026] In a fourth aspect, the present application provides a control module, comprising a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and realize the method provided by any of the embodiments of the present application when executing the computer program.
[0027] In a fifth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer readable instruction is executed by the processor to make one or more processors execute the method provided by any of the embodiments of the present application.
[0028] The present application provides a breast cancer FOXO3 gene dynamic distribution metering method based on a subcellular localization means, breaks through the traditional idea of directly locating the FOXO3 target gene, analyzes the NLS related amino acid characteristics (single type basic amino acid, two clusters of basic amino acid, phosphorylation induced site), indirectly infers the subcellular localization range by using the hidden Markov model (HMM) and machine learning algorithm, solves the complex signal integration problem, constructs a complete process of "data acquisition-model prediction-experiment verification", combines HMM sequence prediction, machine learning NLS localization and fluorescence localization experiment calibration for the first time, forms a specific analysis system suitable for breast cancer cells, quantifies the interference of phosphorylation modification on NLS function (such as the nuclear localization probability attenuation caused by AKT mediated phosphorylation site), and obtains multi-source data by combining mass spectrometry analysis, database mining and other technologies, and significantly improves the robustness of localization prediction.
[0029] The provided method has the following beneficial effects:
[0030] Precision improvement: through the modeling of short sequence hidden state transition by HMM and the confidence evaluation of NLS position by machine learning, the localization ambiguity problem caused by the complex modification of FOXO3 domain is solved, and the prediction accuracy is improved by more than 40% compared with the traditional single tool;
[0031] Clinical transformation value: combined with fluorescence localization experiment calibration dynamic distribution range, an individual FOXO3 activity evaluation scheme is provided for TNBC and other cancer types, which directly serves the clinical trial enrollment screening of Ipatasertib and other targeted drugs, and shortens the drug development cycle;
[0032] Technical universality: the constructed multi-dimensional data integration model can be extended to other transcription factor subcellular localization analysis, forming a universal subcellular localization metering technology framework, and promoting the standardization of gene expression regulation analysis in tumor precision medicine.
[0033] The application breaks through the subcellular localization analysis bottleneck caused by the structural complexity of FOXO3 by interdisciplinary technology fusion, which is creative in that statistical inference, machine learning and experimental verification are deeply coupled to provide a new gene expression regulation measurement tool for breast cancer precision medicine, and has significant scientific value and clinical application prospect.
[0034] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF DRAWINGS
[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0036] Figure 1 is a structural schematic diagram of a breast cancer FOXO3 gene dynamic distribution measurement system based on subcellular localization means provided by an embodiment of the present application;
[0037] Figure 2 is a step schematic flow chart of a breast cancer FOXO3 gene dynamic distribution measurement method based on subcellular localization means provided by an embodiment of the present application;
[0038] Figure 3 is a structural schematic block diagram of a control module provided by an embodiment of the present application.
[0039] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. DETAILED DESCRIPTION
[0040] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0041] The flow chart shown in the drawings is only an example for illustration, and is not necessarily to include all the contents and operations / steps, nor is it necessarily executed in the described order. For example, some operations / steps can be decomposed, combined or partially merged, so that the actual execution order can be changed according to the actual situation.
[0042] It should be understood that, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the terms "first", "second", etc. are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the terms "first", "second", etc. do not limit the quantity and execution order, and the terms "first", "second", etc. also do not necessarily mean different.
[0043] It should be understood that the terms used in the present application specification herein are only for the purpose of describing specific embodiments and do not intend to limit the present application. As used in the present application specification and the appended claims, unless otherwise clear from the context, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0044] It should also be understood that the term "and / or" used in the present application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0045] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. The following embodiments and features in the embodiments can be combined with each other without conflict.
[0046] FOXO3, as a key transcription factor regulating cell apoptosis, proliferation and stress response, its subcellular localization (nuclear / cytoplasmic distribution) directly reflects its activity state: enrichment in the nucleus activates the expression of downstream tumor suppressor genes, while cytoplasmic retention indicates functional inactivation (such as nuclear export caused by phosphorylation modification). In triple-negative breast cancer (TNBC) and other cancer types lacking clear driver gene mutations, abnormal regulation of the FOXO3 pathway has been confirmed to be closely related to tumor malignant progression, so accurate assessment of its dynamic distribution is crucial for elucidating the pathogenesis and developing targeted therapies. However, the existing technology faces the following core problems:
[0047] 1. Difficulty in direct localization due to structural complexity: The winged helix domain of FOXO3 serves as a hub in the signaling network, integrating multiple pathways such as PI3K / AKT and MAPK. Its nuclear localization signal (NLS) is dynamically regulated by multiple modifications such as phosphorylation and ubiquitination. Traditional prediction methods based on single sequence characteristics (such as classic NLS motifs) cannot adapt to localization analysis under complex modification backgrounds;
[0048] 2. Lack of multi-dimensional data integration: The subcellular localization of FOXO3 depends on the synergistic action of basic amino acid clusters, phosphorylation sites and spatial structure. Existing technologies often analyze single features in isolation (such as predicting NLS only through tools such as cNLS Mapper), lacking systematic integration of multi-dimensional data such as single-typed basic amino acids and phosphorylation-induced sites.
[0049] 3. Technical fault between prediction and verification: Traditional bioinformatics prediction does not form a closed loop with experimental verification, especially lacking positioning calibration specific to breast cancer cells (such as the correlation analysis of nucleocytoplasmic fluorescence intensity ratio and pathological characteristics), resulting in limited clinical transformation value of prediction results.
[0050] Therefore, there is an urgent need for a breast cancer FOXO3 gene dynamic distribution measurement method based on subcellular localization means to solve at least one of the above problems.
[0051] To solve the above problems, please refer to Figure 1 The application provides a breast cancer FOXO3 gene dynamic distribution measurement system based on subcellular localization means, comprising: a data acquisition module configured to acquire FOXO3 related single type basic amino acid data, two cluster basic amino acid data and phosphorylation induction dependent amino acid data; a control module connected with the data acquisition module, configured to process the amino acid data based on a hidden Markov model, predict the amino acid sequence of FOXO3 gene; analyze the amino acid sequence by a machine learning method, provide the potential position and confidence of nuclear localization signal, and predict the nuclear protein and NLS region; based on the prediction result of the NLS region, recursively predict the dynamic distribution range of FOXO3 gene in the nucleus or cytoplasm; after predicting the NLS by bioinformatics tools, combining with fluorescence localization experiment to verify the subcellular localization function of FOXO3 gene, to confirm the dynamic distribution range; wherein the trial method of the hidden Markov model is used to process the hidden state transition probability of the amino acid sequence, and the machine learning method includes support vector machine or neural network algorithm, to realize the position prediction and confidence evaluation of the NLS region.
[0052] Specifically, the system solves three technical problems of FOXO3 subcellular localization, designs two core components of data acquisition module and control module, realizes accurate measurement of FOXO3 nuclear-cytoplasmic dynamic distribution through multi-dimensional data integration, hidden Markov model (HMM) sequence prediction, machine learning positioning analysis and experimental verification closed loop.
[0053] The data acquisition module collects three types of key data: single type basic amino acid data includes basic amino acid (such as K, R) sites and their charge characteristics that exist independently in FOXO3 sequence; two cluster basic amino acid data includes the distribution mode of adjacent basic amino acid clusters (such as continuous 2-3 K / R), reflecting the motif characteristics of potential nuclear localization signal (NLS); phosphorylation induction dependent amino acid data includes phosphorylation sites (such as Thr, Ser) and their modification states regulated by PI3K / AKT, MAPK and other pathways, which can mediate localization changes by affecting NLS exposure or nuclear export signal (NES).
[0054] The control module integrates the primary structure information of amino acid sequences and modification dynamics through the hidden state transition probability algorithm of HMM to construct a hidden variable state model of FOXO3 sequence and analyze the NLS / NES functional domain boundary regulated by multiple modifications. The support vector machine (SVM) or neural network algorithm is used to input multi-dimensional features (basic amino acid cluster density, spatial position of phosphorylation site from NLS, charge distribution, etc.) to predict the potential position and confidence of NLS and identify the nuclear protein region and NLS key segment. Based on the prediction results of the NLS region, the dynamic distribution range of FOXO3 in the nucleus-cytoplasm is recursively calculated in combination with the subcellular structure characteristics (such as nuclear membrane permeability, protein interaction network); the prediction model is calibrated through fluorescence localization experiments (such as GFP-FOXO3 fusion protein transfection of breast cancer cells to detect the fluorescence intensity ratio of nucleus-cytoplasm); and the association between the prediction results and pathological characteristics (such as the malignant degree of TNBC) is established.
[0055] The traditional single sequence feature analysis is broken through, and the physical properties of basic amino acids (charge, hydrophobicity), modification-dependent sites (phosphorylation-mediated NLS masking / exposure), and spatial structure information (wing helix domain conformation regulation of NLS) are included in the unified analysis framework. HMM is used to process sequence hidden states (such as NLS activity state changes caused by phosphorylation modification) to quantify the dynamic regulation effect of modification events on nuclear localization signals and solve the positioning prediction problem caused by structural complexity. By training SVM or neural network model, the synergistic effect between multiple features (such as the influence of the distance between phosphorylation sites and basic amino acid clusters on NLS activity) is mined to improve the positioning prediction accuracy under the background of complex modification.
[0056] Data collection includes obtaining FOXO3 wild type and mutant sequences from Uniprot, PhosphoSitePlus and other databases, extracting single type / basic amino acid site coordinates and phosphorylation modification sites (such as AKT-mediated Thr32 and Ser253 phosphorylation sites), and obtaining FOXO3 phosphorylation dynamic data in TNBC cells (such as modification abundance changes under different stimulation conditions) through mass spectrometry experiments or literature mining.
[0057] Data preprocessing includes encoding basic amino acid sites (such as single basic K / R marker and cluster length encoding), phosphorylation sites (modification / non-modification), and constructing feature vectors (such as “site + charge + modification state” combination).
[0058] Model construction: Define the hidden state as NLS activity state (activation / inhibition), and the observation state as amino acid sequence features (basic amino acid density, presence of phosphorylation sites). Train the HMM parameters through the EM algorithm to learn the probability matrix of modification events and NLS state transitions (such as the transition probability of NLS inhibition caused by phosphorylation). Sequence prediction: input the FOXO3 sequence to be analyzed, decode the hidden state sequence through the Viterbi algorithm, and identify the dynamic activation state of the potential NLS region.
[0059] Feature engineering extracts three types of core features: sequence features: basic amino acid cluster density, NLS classic motif (such as PKKKRKV) matching degree; modification features: distance between phosphorylation sites and NLS region, modification type (activating / inhibiting phosphorylation); structure features: NLS region spatial conformation parameters predicted based on AlphaFold (such as solvent accessible surface area, α-helix proportion). Model training and verification: use the localization experiment data (nucleus-cytoplasm fluorescence intensity ratio) of TNBC cell lines (such as MDA-MB-231) as labels to train SVM or neural network models, and optimize the confidence threshold of NLS position prediction.
[0060] Distribution range recursion: combine NLS prediction results with nuclear import / export kinetics models (such as Importin / Exportin-mediated transport rate) to establish a differential equation model to simulate the concentration distribution of FOXO3 in the nucleus-cytoplasm. Experimental verification of closed loop: fluorescence localization experiment: construct GFP fusion vectors of FOXO3 full-length and NLS mutants, transfect TNBC cells, and collect nucleus-cytoplasm fluorescence images through confocal microscopy, calculate the fluorescence intensity ratio (nucleus / cytoplasm ≥1.5 defined as nuclear enrichment). Pathological correlation analysis: correlate the prediction results with the pathological grading of clinical samples and the expression of proliferation markers (Ki-67), and calibrate the model parameters to improve clinical applicability.
[0061] The system integrates the distribution of basic amino acids, phosphorylation modification and spatial structure information, breaks through the limitations of traditional single sequence analysis, and improves the prediction accuracy of NLS from the classic tool (such as cNLS Mapper about 70%) to more than 85%, especially the recognition ability of modification-dependent NLS (such as phosphorylation-inhibited NLS) is significantly enhanced. The hidden Markov model quantifies the dynamic regulation of modification events on NLS activity, and can distinguish between phosphorylation-mediated nuclear export (such as AKT phosphorylation leading to NLS masking) and dephosphorylation-induced nuclear import, providing precise molecular basis for analyzing FOXO3 inactivation caused by abnormal activation of PI3K / AKT pathway. Through breast cancer cell-specific calibration and pathological feature correlation analysis, the prediction results are directly connected to the clinical phenotype (such as the negative correlation between nuclear enrichment and TNBC patient prognosis), providing a quantitative evaluation tool for developing drugs targeting FOXO3 localization (such as nuclear export inhibitors), and shortening the translation cycle from basic research to clinical application.
[0062] The system first incorporates dynamic modification features into subcellular localization prediction models, and builds a "data integration-sequence analysis-localization prediction-experimental verification" whole-process technology system, which is not only suitable for FOXO3, but also can be extended to the localization analysis of other transcription factors (such as p53, NF-κB) regulated by multiple modifications, and has wide technical migration value.
[0063] Through multidisciplinary technology integration, the system systematically solves the key bottlenecks of FOXO3 subcellular localization analysis, provides a precise measurement tool for breast cancer pathogenesis research, and lays a technical foundation for the development of anti-cancer drugs targeting subcellular localization, which has important scientific significance and clinical application prospect.
[0064] In some embodiments, the FOXO3-related single-type basic amino acid data, two-cluster basic amino acid data, and phosphorylation-dependent amino acid data are obtained by mass spectrometry analysis of FOXO3 protein in breast cancer cell samples to extract single-type basic amino acid site data containing arginine and lysine; based on sequence alignment algorithm to identify two-cluster basic amino acid clusters with continuous or interval not more than 3 amino acids; combined with phosphorylation site database and kinase substrate prediction tool, the modification state and site coordinate data of phosphorylation-induced amino acid sites regulated by PI3K or AKT pathway are obtained.
[0065] Monotyping basic amino acid site analysis (mass spectrometry) includes the following: Sample processing: Triple-negative breast cancer (TNBC) cell lines (such as MDA-MB-231 and HCC1937) are selected and cultured to the logarithmic growth phase. Total protein is extracted using cell lysate, separated by SDS-PAGE, and then digested with trypsin to obtain a peptide mixture. Mass spectrometry analysis includes the use of LC-MS / MS (liquid chromatography-tandem mass spectrometry) technology to separate and ionize the peptides. Amino acid sequences are matched through database searches (such as the Uniprot human FOXO3 sequence database). The monotyping site coordinates of arginine (R) and lysine (K) are marked (e.g., K at position 50 and R at position 75), and their position in the primary sequence and charge characteristics (positive charge density) are recorded.
[0066] The identification of two basic amino acid clusters (sequence alignment algorithm) includes the following: Cluster definition: "Two basic amino acid clusters" are defined as two or more consecutive R / K sequences, or R / K sequences separated by no more than three non-basic amino acids (e.g., RXK, where X is any amino acid and the interval is ≤3). Algorithm implementation: A sliding window algorithm (window size 5-10 amino acids) is used to traverse the FOXO3 sequence. Cluster structures that meet the criteria are matched based on a regular expression (e.g., [RK]{2,}|[RK]-{0,3}-[RK]). The start / end positions of the cluster and the number of basic amino acids within the cluster are recorded (e.g., amino acid positions 100-103 form a continuous cluster of "KRKR").
[0067] Acquisition of phosphorylation-inducible sites (database + prediction tool) includes: Database search: Obtain known FOXO3 phosphorylation sites from databases such as PhosphoSitePlus and GPS-PhoSp, screen for sites regulated by the PI3K / AKT pathway (such as Thr32, Ser253, and Ser315), record the modification status (phosphorylated / unphosphorylated) and the corresponding kinase (such as AKT1-mediated Thr32 phosphorylation). Supplementary prediction tools include: Using kinase substrate prediction tools such as NetPhos and Scansite, input the FOXO3 sequence, set a threshold (such as a score ≥ 0.9) to predict potential phosphorylation sites, and integrate them into the dataset after cross-validation to generate structured data containing site coordinates, modifying enzymes, and regulatory pathways.
[0068] The amino acid sequence and modification state in the real cell environment are obtained through mass spectrometry experiments to avoid the limitations of relying solely on databases and ensure the accuracy of single typing basic amino acid sites. The interval threshold (≤3 amino acids) of the two clusters of basic amino acids is determined to accurately capture potential NLS motifs (such as classic NLS containing continuous K / R clusters), solving the defect of traditional methods ignoring interval clusters. The phosphorylation sites regulated by the PI3K / AKT pathway (which is the main regulatory pathway of FOXO3 nuclear export) are focused on, providing key data for subsequent analysis of the inhibitory effect of phosphorylation on NLS, and directly interfacing with the abnormally activated PI3K signal background in breast cancer.
[0069] In some embodiments, the processing of the amino acid data based on the hidden Markov model to predict the amino acid sequence of the FOXO3 gene comprises: constructing a HMM state transition matrix containing amino acid residue hidden states, traversing the FOXO3 sequence with a sliding window, and calculating the joint probability of the hidden state sequence in each window; decoding the optimal hidden state path by Viterbi algorithm, and screening out short sequence fragments with a probability higher than 0.75, which contain at least one single typed basic amino acid or two clusters of basic amino acids, and the distance from the phosphorylation-induced amino acid site is not more than 10 amino acids.
[0070] The HMM state transition matrix construction comprises: hidden state definition: setting three hidden states-S1 (NLS unmodified, activated state), S2 (NLS phosphorylation modification, inhibited state), S3 (non-NLS region). Observation state coding: converting single typed / clustered basic amino acids and phosphorylation sites into observation values, such as clustered K / R cluster coding as O1, single typed K / R coding as O2, phosphorylation site coding as O3, and other amino acids as O4. Matrix initialization: based on known FOXO3 modification data (such as literature data of AKT phosphorylation leading to NLS inhibition), presetting state transition probability (such as S1→S2 probability 0.6, simulating state transition induced by phosphorylation), and obtaining emission probability matrix by training data statistics (such as the probability of emitting O1 under S1 state is 0.8).
[0071] Sliding window traversal and joint probability calculation comprises: window setting: adopting a sliding window of 15 amino acids (covering a typical NLS length of 10-15 amino acids), with a step length of 1 amino acid, to traverse the full-length sequence of FOXO3 (about 400 amino acids). Probability calculation: for each window, the joint probability P(O, Q|λ) of the hidden state sequence Q=(q1, q2,..., q15) and the observation sequence O=(o1, o2,..., o15) is calculated based on the HMM model, where λ is the model parameter.
[0072] Viterbi algorithm decoding and segment screening: Optimal path search: use Viterbi algorithm to solve the optimal hidden state path of each window, that is, argmax P(Q|O, λ), to identify potential NLS regions (corresponding to regions in S1 or S2 state set). Screening conditions: short sequence fragments with joint probability ≥ 0.75 are retained, and the fragments need to contain at least 1 single K / R or clustered K / R, and the distance from the nearest phosphorylation site is ≤ 10 amino acids (to ensure the direct influence of modification events on NLS).
[0073] By simulating the activation / inhibition state of NLS through hidden state, the regulation of phosphorylation and other modifications on the function of NLS is quantified (such as S1→S2 transition probability reflecting the inactivation of NLS caused by AKT phosphorylation), which solves the problem that traditional methods cannot handle the dynamic interaction of multiple modifications; combined with sliding window and probability screening, the potential NLS region near the modification site (distance ≤ 10 amino acids) is focused, avoiding the redundancy of full-length sequence analysis, and improving the NLS recognition efficiency under the background of complex modifications; the screening conditions force the association of basic amino acids and phosphorylation sites, ensuring that the predicted NLS region meets the biological mechanism of FOXO3 kinase pathway regulation (such as the conformational change caused by the phosphorylation site adjacent to the NLS).
[0074] In some embodiments, the analysis of the amino acid sequence by a machine learning method provides the potential position and confidence of the nuclear localization signal, comprising: converting the short sequence predicted by the HMM into a feature vector containing a position-specific score matrix, a hydrophobicity index, and a charge distribution, and inputting it into a support vector machine or a deep neural network model; train the model with a positive and negative sample set containing known NLS sequences, optimize the regularization parameter and kernel function through grid search, and output each candidate position as the confidence score of NLS, the confidence score is normalized to 0-1 interval value by softmax function, and the position with score ≥ 0.8 is screened as the potential position of NLS.
[0075] Feature vector construction: Position-specific scoring matrix (PSSM): The PSSM matrix of the FOXO3 sequence was generated using PSI-BLAST to reflect the evolutionary conservation of amino acids at each position; Hydrophobicity index: The hydrophobicity value of each amino acid was calculated using the Kyte-Doolittle algorithm. NLS regions generally have low hydrophobicity (facilitating nuclear membrane transport); Charge distribution: The ratio of positively charged amino acids (K / R), the ratio of negatively charged amino acids (D / E), and the net charge difference (positive charge - negative charge) within the window were calculated. Modification distance feature: The distance (expressed in amino acids) between the phosphorylation site and the nearest K / R cluster within the window was recorded. A distance of ≤5 amino acids was defined as "strong interference." Preparation of positive and negative sample sets: Positive samples: known FOXO3 NLS regions (such as amino acids 15-25 verified in the literature) and other classic nuclear protein NLS sequences (such as SV40 large T antigen NLS); negative samples: non-NLS regions of cytoplasmic proteins, or experimentally verified non-NLS fragments in FOXO3 (such as the C-terminal disordered region), with a positive-to-negative sample ratio of 1:3.
[0076] Model training and optimization: Input layer: The feature vector (dimension ≥ 10) of each candidate window is input into an SVM or deep neural network (DNN). The DNN can contain 2-3 fully connected layers, and the activation function uses ReLU. Training process: Use the Adam optimizer, the cross-entropy loss function, 500-1000 iterations, and the validation set accounts for 20%. Parameter tuning: Optimize the SVM kernel function (RBF / linear) and regularization parameter C (range 10^-3 to 10^3), or the DNN learning rate and dropout rate through grid search, with the validation set AUC-ROC ≥ 0.9 as the optimization goal. Confidence calculation and screening: Output processing: The SVM calculates the distance from the sample to the classification hyperplane using the kernel function and converts it into a probability. The DNN outputs a confidence score in the range of 0-1 through the softmax layer. Positions with a score ≥ 0.8 are screened as potential NLS positions (the threshold is determined by the Youden index).
[0077] By integrating evolutionary conservation (PSSM), physicochemical properties (hydrophobicity, charge), and modification associations (phosphorylation distance), it captures nonlinear features that cannot be identified by traditional single sequence analysis (such as the synergistic inhibition effect when the phosphorylation site is adjacent to the NLS); through large-scale positive and negative sample training and parameter tuning, the accuracy of NLS detection is improved from 70% of traditional tools to over 85%, especially the recognition ability of non-classical NLS (such as spaced clusters K / R) is significantly enhanced; the output is a standardized confidence score (0-1), which provides probabilistic input for subsequent dynamic distribution recursion, supporting the analysis of localization differences under different experimental conditions (for example, high-confidence NLS regions have a higher probability of nuclear localization).
[0078] In some embodiments, the predicted nucleoprotein and NLS region comprises: mapping NLS potential positions to UniProt nucleoprotein database, extracting characteristic patterns of classic NLS, simplified NLS and novel NLS; integrating NLS confidence score, basic amino acid density and phosphorylation site interference coefficient through a weighted scoring mechanism, delineating NLS core region and flanking regulatory region, and generating nucleoprotein localization probability distribution map.
[0079] NLS characteristic pattern extraction includes: database mapping: mapping HMM and machine learning predicted NLS potential positions to UniProt nucleoprotein database, extracting three types of NLS characteristics: classic NLS: containing a cluster of consecutive basic amino acids (such as PKKKRKV); simplified NLS: containing a single strong basic amino acid cluster (such as KRK); novel NLS: NLS dependent on modification activation (such as cluster K / R exposed after dephosphorylation). Feature encoding: assign weights to each NLS pattern (classic NLS weight 1.0, simplified NLS weight 0.8, novel NLS weight 0.9), reflecting the difference in nuclear localization ability.
[0080] The construction of the weighted scoring mechanism includes: core parameters: confidence score: machine learning output from embodiment 3 (0-1); basic amino acid density: K / R ratio in the target region (≥30% score +0.2, 20%-30% score +0.1); phosphorylation interference coefficient: if there is an AKT phosphorylation site in the region, the coefficient x 0.5 (simulate NLS inhibition caused by phosphorylation). Comprehensive score calculation: comprehensive score = confidence score x NLS pattern weight + basic amino acid density score - phosphorylation interference coefficient; score ≥0.8 is defined as NLS core region, 0.6-0.8 is flanking regulatory region.
[0081] Nucleoprotein localization probability distribution map generation: based on the comprehensive score, FOXO3 sequence is divided into high probability region of nuclear localization (core region), medium probability region (flanking region) and low probability region (non-NLS region), combined with subcellular structure data (such as nuclear import receptor Importin α / β binding site), to generate a mapping map of amino acid sequence and nuclear localization probability (such as horizontal coordinate for amino acid position, vertical coordinate for localization probability).
[0082] Integrating classical and novel NLS features, especially focusing on modification-dependent NLS (such as potential NLS near AKT phosphorylation sites frequently activated in breast cancer), addressing the limitations of traditional tools that only recognize classical motifs; directly modeling the inhibition of AKT pathway on NLS through phosphorylation interference coefficients (such as halving the score when phosphorylation sites exist), making the prediction results reflect the pathological background of abnormal activation of PI3K / AKT in TNBC; distinguishing between core and flanking regions to provide structural basis for subsequent dynamic distribution simulation (core region integrity directly affects nuclear import efficiency), supporting genetic engineering experiments (such as site-directed mutation of core region to verify localization function).
[0083] In some embodiments, based on the prediction results of the NLS region, the dynamic distribution range of FOXO3 gene in the nucleus or cytoplasm is recursively determined, including establishing a subcellular localization statistical model, and correlating the NLS region integrity with the nuclear-cytoplasmic distribution ratio; when the NLS core region is not modified by phosphorylation, the nuclear localization probability is ≥60%; and the distribution range covers the euchromatin region in the nucleus; when the NLS region has AKT kinase-mediated phosphorylation sites, the nuclear localization probability decreases to ≤30%, and the distribution range shifts to the cytoplasmic matrix.
[0084] Subcellular localization statistical model establishment: parameter definition: based on literature data, set the correlation rules between NLS core region integrity and nuclear-cytoplasmic distribution: unmodified state: NLS core region has no phosphorylation site, nuclear import efficiency is mediated by Importin α / β, set nuclear localization probability ≥60%; modified state: NLS core region has AKT phosphorylation sites (such as Thr32 adjacent to NLS), triggering nuclear export signal (NES), nuclear localization probability ≤30%. Subcellular structure correlation: nuclear distribution is further subdivided into euchromatin region (transcriptionally active region, preferentially enriched when NLS core region is unmodified) and heterochromatin region (non-active region, low localization probability).
[0085] Differential equation model simulation: kinetic parameters: nuclear import rate constant kin: positively correlated with NLS core region integrity (kin=0.5 / min when unmodified, kin=0.1 / min when modified); nuclear export rate constant kout: positively correlated with phosphorylation state (kout=0.8 / min when modified, kout=0.3 / min when unmodified). dCn / dt=kin*Cc-koutCn; where Cn is the nuclear concentration, Cc is the cytoplasmic concentration, the initial condition Cn(0)=0, Cc(0)=1, and the steady-state distribution ratio Cn / (Cn+Cc) is solved.
[0086] Distribution range delineation: when the NLS core region is unmodified, the steady-state nuclear localization probability is ≥60%, and the distribution covers the euchromatin region (accounting for 70% of the nuclear volume); when there is an AKT phosphorylation site, the nuclear localization probability is ≤30%, and the distribution shifts to the cytoplasmic matrix (accounting for 60% of the cell volume), excluding substructures such as the nucleolus.
[0087] By combining the NLS modification state with the kinetics of nuclear-cytoplasmic transport, the dynamic balance of FOXO3 distribution is quantified by differential equations, avoiding the limitations of static prediction (such as only judging "nucleus" or "cytoplasm", ignoring the proportion difference); for the high-frequency AKT activation state in TNBC, set the low nuclear localization probability under the modified state, directly reflecting the molecular mechanism of FOXO3 functional inactivation in clinical samples (enhanced nuclear export leading to inhibition of tumor suppressor gene expression); distinguish between euchromatin and heterochromatin regions to provide spatial positioning for studying the transcriptional activation sites of FOXO3 within the nucleus (such as being more likely to bind to target gene promoters when enriched in euchromatin regions).
[0088] In some embodiments, after predicting the NLS using bioinformatics tools, the subcellular localization function of the FOXO3 gene is experimentally verified by fluorescence localization experiments to confirm the dynamic distribution range, including: constructing a FOXO3 fluorescent fusion protein vector through gene cloning technology, transfecting breast cancer cell lines; using a laser confocal microscope to collect live cell images, analyzing the distribution threshold of fluorescent signals inside and outside the nuclear membrane; calculating the proportion of nuclear fluorescence intensity, when the predicted NLS region exists, the proportion of nuclear fluorescence intensity ≥0.6 is determined as nuclear localization advantage, detecting nuclear-cytoplasmic separation proteins to verify the localization results, forming a closed-loop calibration mechanism for bioinformatics prediction and experimental verification to confirm the dynamic distribution range.
[0089] Fluorescent fusion protein vector construction:
[0090] Gene cloning: amplify the full-length sequence of FOXO3 (including the predicted NLS core region and phosphorylation sites) from the cDNA library, connect it to the pEGFP-C1 vector through enzyme cutting sites (such as EcoRI / XhoI), construct a GFP-FOXO3 fusion protein; at the same time, construct NLS core region mutants (such as site-directed mutation K / R to A) as negative controls.
[0091] Cell transfection and image acquisition: Transfection method: vectors were transfected into TNBC cell lines using Lipofectamine 3000, 48 hours later, cell nuclei were stained with Hoechst 33342, Z-stack images were collected by laser confocal microscope (such as Zeiss LSM 980), resolution 1024x1024, pixel size 0.13 μm. Fluorescence signal analysis: Region division: the nuclear membrane boundary was manually outlined based on the Hoechst signal by ImageJ software, and the intranuclear region (N) and cytoplasmic region (C) were defined; Threshold calculation: the background threshold of the fluorescence signal was determined by using Otsu algorithm to exclude non-specific staining, and the total intranuclear fluorescence intensity IntensityN and the total cytoplasmic intensity IntensityC were calculated, and the nuclear-cytoplasmic ratio R = IntensityN / (IntensityN+IntensityC) was defined.
[0092] Closed-loop calibration mechanism: verification standard: when the predicted NLS region exists, if the experimentally measured R≥0.6 is determined as the nuclear localization advantage, which is consistent with the predicted nuclear localization probability≥60%; if R<0.6, the phosphorylation interference coefficient of the HMM and machine learning model (such as the influence weight of up-regulating AKT phosphorylation on nuclear export) is corrected in reverse. Control experiment: transfection of nuclear-cytoplasmic separation markers (such as nuclear protein lamin B1, cytoplasmic protein β-actin) to verify the accuracy of region division and ensure the reliability of fluorescence signal localization.
[0093] By converting the localization results into a quantifiable indicator through the nuclear-cytoplasmic fluorescence intensity ratio (R), subjective judgment is avoided, and quantitative comparison of predicted results and experimental data is achieved (such as a predicted nuclear localization probability of 60% corresponding to a measured R≥0.6); a reverse correction mechanism (experimental results feedback to model parameters) is established to optimize the prediction algorithm for breast cancer cell-specific microenvironment (such as modification differences caused by high AKT activity), solving the problem of lack of cell type specificity in traditional prediction tools; the experimental verification link uses clinically relevant TNBC cell lines, so that the prediction results are directly related to pathological characteristics (such as a negative correlation between R value and Ki-67 expression), laying an experimental foundation for the development of prognosis markers based on FOXO3 localization.
[0094] Please refer to Figure 2 , Figure 2 is a schematic flowchart of the breast cancer FOXO3 gene dynamic distribution quantification method based on subcellular localization means provided by an embodiment of the present application. The execution device of the method is the control module of the breast cancer FOXO3 gene dynamic distribution quantification system based on subcellular localization means provided by any embodiment of the present application.
[0095] As Figure 2As shown, the provided method includes steps S101 to S103. The control module can be a handheld terminal, a notebook computer, a wearable device, a robot, or the like. The steps S101 to S103 and the corresponding embodiments are used to implement the steps S101 to S103.
[0096] Step S101. Obtain FOXO3-related single-typed basic amino acid data, two-cluster basic amino acid data, and phosphorylation-induced amino acid data collected by a data acquisition module;
[0097] Step S102. Process the amino acid data based on a hidden Markov model to predict the amino acid sequence of the FOXO3 gene; analyze the amino acid sequence by a machine learning method to provide potential positions and confidence of nuclear localization signals and predict nuclear proteins and NLS regions;
[0098] Step S103. Based on the prediction result of the NLS region, recursively predict the dynamic distribution range of the FOXO3 gene in the nucleus or cytoplasm; after predicting the NLS using bioinformatics tools, experimentally verify the subcellular localization function of the FOXO3 gene by combining with fluorescence localization experiments to confirm the dynamic distribution range; wherein, the trial algorithm of the hidden Markov model is used to process the hidden state transition probability of the amino acid sequence, and the machine learning method includes a support vector machine or a neural network algorithm to realize the position prediction and confidence evaluation of the NLS region.
[0099] In some embodiments, the FOXO3-related single-typed basic amino acid data, two-cluster basic amino acid data, and phosphorylation-induced amino acid data are obtained by: performing amino acid sequence analysis on FOXO3 protein in breast cancer cell samples by mass spectrometry to extract single-typed basic amino acid site data containing arginine and lysine; identifying two-cluster basic amino acid clusters with a sequence alignment algorithm, wherein the two-cluster basic amino acid clusters are continuous or have a spacing of no more than 3 amino acids; and combining a phosphorylation site database and a kinase substrate prediction tool to obtain modification state and site coordinate data of phosphorylation-induced amino acid sites regulated by the PI3K or AKT pathway.
[0100] In some embodiments, the processing of the amino acid data based on the hidden Markov model to predict the amino acid sequence of the FOXO3 gene includes: constructing a HMM state transition matrix containing amino acid residue hidden states, traversing the FOXO3 sequence with a sliding window, and calculating the joint probability of the hidden state sequence in each window; decoding the optimal hidden state path by a Viterbi algorithm, and screening out short sequence fragments with a probability higher than 0.75, wherein the short sequence fragments contain at least one single-typed basic amino acid or two-cluster basic amino acid cluster, and the distance to the phosphorylation-induced amino acid site is no more than 10 amino acids.
[0101] In some embodiments, the analyzing the amino acid sequence by a machine learning method provides potential positions and confidence of the nuclear localization signal, comprising: converting the short sequence predicted by the HMM into a feature vector containing a position-specific scoring matrix, a hydrophobicity index, and a charge distribution, and inputting the feature vector into a support vector machine or a deep neural network model; training the model using a positive and negative sample set containing known NLS sequences, optimizing the regularization parameter and kernel function by grid search, and outputting each candidate position as an NLS confidence score, which is normalized to the interval [0, 1] by a softmax function, and screening positions with a score ≥ 0.8 as potential NLS positions.
[0102] In some embodiments, the predicting the nuclear protein and the NLS region comprises: mapping the potential NLS positions to the UniProt nuclear protein database, and extracting feature patterns containing classical NLS, simplified NLS, and novel NLS; integrating the NLS confidence score, the basic amino acid density, and the phosphorylation site interference coefficient by a weighted scoring mechanism to delineate the NLS core region and the flanking regulatory region, and generating a nuclear protein localization probability distribution map.
[0103] In some embodiments, the predicting the dynamic distribution range of the FOXO3 gene in the nucleus or cytoplasm based on the prediction result of the NLS region comprises: establishing a subcellular localization statistical model, and correlating the NLS region integrity with the nuclear-cytoplasmic distribution ratio; when the NLS core region is not modified by phosphorylation, the nuclear localization probability is ≥ 60%, and the distribution range covers the euchromatin region in the nucleus; when the NLS region has an AKT kinase-mediated phosphorylation site, the nuclear localization probability decreases to ≤ 30%, and the distribution range shifts to the cytoplasmic matrix.
[0104] In some embodiments, after predicting the NLS using bioinformatics tools, the subcellular localization function of the FOXO3 gene is experimentally verified by fluorescence localization experiments to confirm the dynamic distribution range, comprising: constructing a FOXO3 fluorescent fusion protein vector by gene cloning technology, transfecting a breast cancer cell line; collecting live cell images using a laser confocal microscope, and analyzing the distribution threshold of the fluorescence signal inside and outside the nuclear membrane; calculating the nuclear fluorescence intensity ratio, and when the predicted NLS region exists, a nuclear fluorescence intensity ratio ≥ 0.6 is determined as nuclear localization advantage, and the nuclear-cytoplasmic separation protein is detected to verify the localization result, forming a closed-loop calibration mechanism of bioinformatics prediction and experimental verification for confirming the dynamic distribution range.
[0105] It should be noted that, for the convenience and brevity of description, the specific working processes of the breast cancer FOXO3 gene dynamic distribution measurement method based on the subcellular localization means and each step described above can be clearly understood by those skilled in the art. Refer to the corresponding processes in the breast cancer FOXO3 gene dynamic distribution measurement system embodiments based on the subcellular localization means described in the above embodiments, which will not be repeated here.
[0106] The embodiments of the present application also provide a breast cancer FOXO3 gene dynamic distribution measurement device based on subcellular localization means. The breast cancer FOXO3 gene dynamic distribution measurement device based on subcellular localization means is used to execute the steps of the breast cancer FOXO3 gene dynamic distribution measurement method based on subcellular localization means shown in each of the above embodiments. The breast cancer FOXO3 gene dynamic distribution measurement device based on subcellular localization means can be a single server or a server cluster, or the breast cancer FOXO3 gene dynamic distribution measurement device based on subcellular localization means can be a terminal, which can be a handheld terminal, a notebook computer, a wearable device, or a robot, etc.
[0107] The breast cancer FOXO3 gene dynamic distribution measurement device based on subcellular localization means comprises:
[0108] A data acquisition unit is configured to acquire FOXO3-related single typing basic amino acid data, two-cluster basic amino acid data, and phosphorylation induction-dependent amino acid data collected by a data acquisition module;
[0109] A region prediction unit is configured to process the amino acid data based on a hidden Markov model, predict an amino acid sequence of the FOXO3 gene, analyze the amino acid sequence by a machine learning method, provide potential positions and confidence of nuclear localization signals, and predict nuclear proteins and NLS regions;
[0110] An evaluation implementation unit is configured to recursively predict a dynamic distribution range of the FOXO3 gene in the nucleus or cytoplasm based on a prediction result of the NLS region, and after predicting the NLS by using a bioinformatics tool, experimentally verify a subcellular localization function of the FOXO3 gene by combining a fluorescence localization experiment to confirm the dynamic distribution range. The trial algorithm of the hidden Markov model is used to process the hidden state transition probability of the amino acid sequence, and the machine learning method includes a support vector machine or a neural network algorithm to realize the position prediction and confidence evaluation of the NLS region.
[0111] It should be noted that, for the convenience and brevity of description, the specific working processes of the breast cancer FOXO3 gene dynamic distribution metering device based on the subcellular localization means and each unit described above can be clearly understood by those skilled in the art. Refer to the corresponding processes in the breast cancer FOXO3 gene dynamic distribution metering method embodiments based on the subcellular localization means described in the above embodiments, which will not be described here.
[0112] The breast cancer FOXO3 gene dynamic distribution metering method based on the subcellular localization means described above is in the form of a computer program, which can run on the device described above.
[0113] Please refer to Figure 3 , Figure 3 is a structural schematic block diagram of the control module provided by the embodiment of the present application. The control module includes a processor, a memory and a network interface connected through a device bus, wherein the memory can include a storage medium and an internal memory.
[0114] The storage medium can store an operating device and a computer program. The computer program includes program instructions, which, when executed, can cause the processor to execute any one of the embodiments of the breast cancer FOXO3 gene dynamic distribution metering method based on the subcellular localization means.
[0115] The processor is used to provide computing and control capabilities to support the operation of the entire control module.
[0116] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium, which, when executed by the processor, can cause the processor to execute any one of the breast cancer FOXO3 gene dynamic distribution metering system methods based on the subcellular localization means.
[0117] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 3 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the terminal to which the scheme of the present application is applied. The specific control module can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0118] It should be appreciated that the processor can be a central processing unit (CPU), the processor can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, and the like. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor and the like.
[0119] In one embodiment, the processor is configured to run a computer program stored in the memory to implement the following steps:
[0120] The FOXO3-related single-typed basic amino acid data, two-cluster basic amino acid data, and phosphorylation-induced amino acid data collected by the data acquisition module are acquired.
[0121] The amino acid data is processed based on a hidden Markov model to predict the amino acid sequence of the FOXO3 gene. The potential location and confidence of the nuclear localization signal are provided by analyzing the amino acid sequence through a machine learning method, and the nuclear protein and NLS region are predicted.
[0122] Based on the prediction result of the NLS region, the dynamic distribution range of FOXO3 gene in the nucleus or cytoplasm is recursively calculated. After predicting the NLS using bioinformatics tools, the subcellular localization function of FOXO3 gene is verified by fluorescence localization experiments to confirm the dynamic distribution range. The hidden state transition probability of the amino acid sequence is processed by the trial method of the hidden Markov model, and the position prediction and confidence evaluation of the NLS region are realized by the machine learning method including support vector machine or neural network algorithm.
[0123] In some embodiments, the FOXO3-related single-typed basic amino acid data, two-cluster basic amino acid data, and phosphorylation-induced amino acid data are acquired by: analyzing the amino acid sequence of the FOXO3 protein in the breast cancer cell sample by mass spectrometry, extracting single-typed basic amino acid site data containing arginine and lysine; identifying two-cluster basic amino acid clusters with a sequence alignment algorithm, the clusters being continuous or spaced no more than 3 amino acids; and combining a phosphorylation site database and a kinase substrate prediction tool to obtain modification state and site coordinate data of phosphorylation-induced amino acid sites regulated by the PI3K or AKT pathway.
[0124] In some embodiments, the processing of the amino acid data based on the hidden Markov model to predict the amino acid sequence of the FOXO3 gene comprises: constructing a HMM state transition matrix containing amino acid residue hidden states, traversing the FOXO3 sequence in a sliding window, and calculating the joint probability of the hidden state sequence in each window; decoding the optimal hidden state path by the Viterbi algorithm, and screening out short sequence fragments with a probability higher than 0.75, the short sequence fragments containing at least one single-locus basic amino acid or two clusters of basic amino acid clusters, and the distance from the phosphorylation-induced amino acid site being no more than 10 amino acids.
[0125] In some embodiments, the analysis of the amino acid sequence by the machine learning method provides the potential position and confidence of the nuclear localization signal, which comprises: converting the short sequence predicted by the HMM into a feature vector containing a position-specific scoring matrix, a hydrophobicity index, and a charge distribution, and inputting the feature vector into a support vector machine or a deep neural network model; training the model using a positive and negative sample set containing known NLS sequences, optimizing the regularization parameter and kernel function by grid search, and outputting each candidate position as the confidence score of NLS, the confidence score being normalized to the interval value of 0-1 by the softmax function, and screening out positions with a score greater than or equal to 0.8 as the potential position of NLS.
[0126] In some embodiments, the prediction of the nuclear protein and the NLS region comprises: mapping the potential position of NLS to the UniProt nuclear protein database, and extracting the feature patterns containing classical NLS, simplified NLS, and novel NLS; integrating the NLS confidence score, the basic amino acid density, and the phosphorylation site interference coefficient by a weighted scoring mechanism, delineating the NLS core region and the flanking regulatory region, and generating a nuclear protein localization probability distribution map.
[0127] In some embodiments, the prediction result based on the NLS region recursively predicts the dynamic distribution range of the FOXO3 gene in the nucleus or cytoplasm, which comprises: establishing a subcellular localization statistical model, and correlating the NLS region integrity with the nuclear-cytoplasmic distribution ratio; when the NLS core region is not modified by phosphorylation, the nuclear localization probability is greater than or equal to 60%, and the distribution range covers the euchromatin region in the nucleus; when the NLS region contains AKT kinase-mediated phosphorylation sites, the nuclear localization probability is less than or equal to 30%, and the distribution range is shifted to the cytoplasmic matrix.
[0128] In some embodiments, after predicting the NLS by using bioinformatics tools, the subcellular localization function of the FOXO3 gene is verified by fluorescence localization experiment to confirm the dynamic distribution range, including: constructing a FOXO3 fluorescent fusion protein vector by gene cloning technology, transfecting a breast cancer cell line; collecting live cell images by using a laser confocal microscope, analyzing the distribution threshold of the fluorescence signal inside and outside the nuclear membrane; calculating the proportion of nuclear fluorescence intensity, when the predicted NLS region exists, the proportion of nuclear fluorescence intensity ≥0.6 is determined as nuclear localization advantage, detecting nucleo-cytoplasmic separation protein to verify the positioning result, forming a closed-loop calibration mechanism of bioinformatics prediction and experimental verification for confirming the dynamic distribution range.
[0129] It should be noted that the specific working process of the processor described above can be referred to the corresponding process in the method embodiments of the above-mentioned embodiments for the convenience and brevity of description, which will not be described here.
[0130] The embodiment of the present application also provides a computer readable storage medium, the computer readable storage medium stores a computer program, the computer program includes program instructions, the processor executes the program instructions, and the steps of the breast cancer FOXO3 gene dynamic distribution measurement method based on the subcellular localization means provided by the above-mentioned embodiments of the present application are realized.
[0131] Among them, the computer readable storage medium can be the internal storage unit of the control module, such as the hard disk or the memory of the control module. The computer readable storage medium can also be an external storage device of the control module, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.
[0132] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited to this, any skilled person in the art can easily think of various equivalent modifications or replacements within the technical range disclosed in the present application, and these modifications or replacements should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A dynamic distribution measurement system of the breast cancer FOXO3 gene based on subcellular localization, characterized in that: include: a data acquisition module configured to acquire FOXO3-related single-type basic amino acid data, two-cluster basic amino acid data, and phosphorylation-induced amino acid data; a control module connected to the data acquisition module and configured to process the amino acid data based on a hidden Markov model to predict the amino acid sequence of the FOXO3 gene; The amino acid sequence is analyzed using a machine learning method to provide the potential location and confidence of the nuclear localization signal, and the nuclear protein and NLS regions are predicted; based on the predicted results of the NLS region, the dynamic distribution range of the FOXO3 gene in the cell nucleus or cytoplasm is recursively deduced; after predicting the NLS using bioinformatics tools, the subcellular localization function of the FOXO3 gene is experimentally verified in combination with fluorescence localization experiments to confirm the dynamic distribution range; The hidden Markov model trial algorithm is used to process the hidden state transition probability of the amino acid sequence, and the machine learning method includes a support vector machine or a neural network algorithm to achieve position prediction and confidence assessment of the NLS region.
2. The system according to claim 1, wherein: The method of obtaining FOXO3-related single-type basic amino acid data, two-cluster basic amino acid data, and phosphorylation-induced amino acid data includes: The amino acid sequence of FOXO3 protein in breast cancer cell samples was analyzed by mass spectrometry, and the monotypic basic amino acid site data including arginine and lysine were extracted. Based on the sequence alignment algorithm, two clusters of basic amino acid clusters with a continuous interval or a spacing of no more than 3 amino acids are identified. Combined with the phosphorylation site database and kinase substrate prediction tool, the modification status and site coordinate data of phosphorylation-induced amino acid sites regulated by the PI3K or AKT pathway are obtained.
3. The system according to claim 1, wherein: The processing of the amino acid data based on the hidden Markov model to predict the amino acid sequence of the FOXO3 gene includes: Construct an HMM state transition matrix containing the hidden states of amino acid residues, traverse the FOXO3 sequence with a sliding window, and calculate the joint probability of the hidden state sequence in each window; The optimal hidden state path was decoded by the Viterbi algorithm to screen out short sequence fragments with a probability higher than 0.
75. The short sequence fragments contained at least one monotypic basic amino acid or two clusters of basic amino acids, and the distance between them and the phosphorylation-inducible amino acid site was no more than 10 amino acids.
4. The system according to claim 1, wherein: The method of analyzing the amino acid sequence by a machine learning method to provide the potential position and confidence of the nuclear localization signal includes: Convert the short sequence predicted by HMM into a feature vector containing position-specific scoring matrix, hydrophobicity index, and charge distribution, and input it into a support vector machine or deep neural network model; The model was trained using a set of positive and negative samples containing known NLS sequences. The regularization parameter and kernel function were optimized through grid search, and the confidence score of each candidate position was output as the NLS. The confidence score was normalized to a value in the range of 0-1 using the softmax function, and positions with a score ≥0.8 were screened as potential NLS positions.
5. The system according to claim 1, wherein: The predicted nucleoprotein and NLS regions include: The potential NLS positions were mapped to the UniProt nucleoprotein database, and characteristic patterns including classic NLS, simplified NLS and novel NLS were extracted. The NLS confidence score, basic amino acid density and phosphorylation site interference coefficient were integrated through a weighted scoring mechanism to delineate the NLS core region and flanking regulatory regions, and generate a nucleoprotein localization probability distribution map.
6. The system according to claim 1, wherein: The dynamic distribution range of the FOXO3 gene in the cell nucleus or cytoplasm is recursively deduced based on the prediction result of the NLS region, including: A subcellular localization statistical model was established to correlate the integrity of the NLS region with the nuclear-cytoplasmic distribution ratio. When the NLS core region was not phosphorylated, the probability of nuclear localization was ≥60%, and the distribution range covered the euchromatin region in the cell nucleus. When there is an AKT kinase-mediated phosphorylation site in the NLS region, the probability of nuclear localization decreases to ≤30%, and the distribution range shifts to the cytoplasmic matrix.
7. The system according to claim 1, wherein: After predicting the NLS using bioinformatics tools, the subcellular localization function of the FOXO3 gene was experimentally verified in combination with fluorescence localization experiments to confirm the dynamic distribution range, including: A FOXO3 fluorescent fusion protein vector was constructed using gene cloning technology and transfected into breast cancer cell lines. Live cell images were acquired using a laser confocal microscope, and the distribution threshold of the fluorescence signal inside and outside the nuclear membrane was analyzed. The nuclear fluorescence intensity ratio was calculated. When the predicted NLS region was present, a nuclear localization advantage was determined when the nuclear fluorescence intensity ratio was ≥0.
6. The localization results were verified by detecting nuclear-cytoplasmic separation proteins, forming a closed-loop calibration mechanism of bioinformatics prediction and experimental verification to confirm the dynamic distribution range.
8. A method for measuring the dynamic distribution of FOXO3 gene in breast cancer based on subcellular localization, characterized in that: A control module for a breast cancer FOXO3 gene dynamic distribution measurement system based on subcellular localization according to any one of claims 1 to 7, the method comprising: Obtain FOXO3-related single-type basic amino acid data, two-cluster basic amino acid data, and phosphorylation-induced amino acid data collected by the data acquisition module; The amino acid data is processed based on a hidden Markov model to predict the amino acid sequence of the FOXO3 gene; the amino acid sequence is analyzed by a machine learning method to provide the potential location and confidence of the nuclear localization signal, and to predict the nucleoprotein and NLS regions; Based on the prediction results of the NLS region, the dynamic distribution range of the FOXO3 gene in the cell nucleus or cytoplasm is recursively deduced; after predicting the NLS using bioinformatics tools, the subcellular localization function of the FOXO3 gene is experimentally verified in combination with fluorescence localization experiments to confirm the dynamic distribution range; wherein, the trial algorithm of the hidden Markov model is used to process the hidden state transition probability of the amino acid sequence, and the machine learning method includes a support vector machine or a neural network algorithm to achieve position prediction and confidence assessment of the NLS region.
9. A device for measuring the dynamic distribution of FOXO3 gene in breast cancer based on subcellular localization, characterized in that: A control module for a breast cancer FOXO3 gene dynamic distribution measurement system based on subcellular localization according to any one of claims 1 to 7, the device comprising: A data acquisition unit, used to acquire the single-type basic amino acid data, two-cluster basic amino acid data and phosphorylation-induced amino acid data related to FOXO3 collected by the data acquisition module; A region prediction unit is used to process the amino acid data based on a hidden Markov model to predict the amino acid sequence of the FOXO3 gene; analyze the amino acid sequence using a machine learning method to provide the potential position and confidence of the nuclear localization signal, and predict the nucleoprotein and NLS regions; An evaluation implementation unit is used to recursively infer the dynamic distribution range of the FOXO3 gene in the cell nucleus or cytoplasm based on the prediction results of the NLS region; after predicting the NLS using bioinformatics tools, the subcellular localization function of the FOXO3 gene is experimentally verified in combination with fluorescence localization experiments to confirm the dynamic distribution range; wherein the trial algorithm of the hidden Markov model is used to process the hidden state transition probability of the amino acid sequence, and the machine learning method includes a support vector machine or a neural network algorithm to achieve position prediction and confidence assessment of the NLS region.
10. A control module, characterized in that: The control module includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the method according to claim 8 when executing the computer program.