Utility sequence pattern mining method, system and application for multi-objective decision
By employing a utility sequence pattern mining method oriented towards multi-objective decision-making, and combining quantitative and utility evaluations with a pruning strategy to generate optimal DNA sequence patterns, this approach addresses the problem of incomplete sequence pattern analysis in existing technologies, improves mining efficiency and result quality, and supports large-scale data processing.
Patent Information
- Application Number
- CN202510965364.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-28
AI Technical Summary
Existing DNA sequence pattern mining methods only consider the frequency of sequence pattern occurrence, resulting in the mined sequence data containing only some common sequence fragments. This ignores the biological importance of DNA sequence patterns, such as gene function, regulatory mechanisms, and disease associations, and also has high computational complexity.
A utility sequence pattern mining method oriented towards multi-objective decision-making is adopted. High-frequency DNA sequence patterns are screened by quantitative indicators, and their biological importance is evaluated by utility functions. A quantitative utility array and a projection database are constructed, and the optimal solution sequence pattern is generated by combining pruning strategy.
It improves mining efficiency, ensures the biological significance and data representativeness of the results, supports large-scale DNA sequence data processing, enables deeper exploration of potential patterns, and provides valuable information for biological research.
Smart Images

Figure CN120853689A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of bioinformatics and data, specifically to a method, system, and application for mining utility sequence patterns for multi-objective decision-making. Background Technology
[0002] In the field of bioinformatics, the analysis of DNA sequence data is a crucial approach to revealing biological genetic information, understanding life phenomena, and exploring disease mechanisms. During DNA evolution, conserved regions in most sequences form specific sequence patterns, whose structure and function play vital roles. DNA sequence patterns are typically functionally specific fragments within a DNA sequence. Therefore, identifying these patterns is an important aspect of DNA sequence data analysis, helping to predict DNA sequence functions and explain evolutionary relationships between sequences. DNA sequence pattern mining can effectively discover these sequence patterns and identify genes and their functions.
[0003] With the development of high-throughput sequencing technology, the amount of biological sequence data has exploded. How to effectively mine DNA sequence patterns from massive amounts of DNA sequence data and discover valuable sequence patterns hidden behind the data that can better solve biological problems has important research and practical value.
[0004] Existing DNA sequence pattern mining methods suffer from three main problems: First, they only consider the frequency of sequence pattern occurrence, resulting in mined sequence data containing only a portion of common sequence fragments, which cannot comprehensively analyze and predict DNA sequence functions. Second, they often overlook the biological importance of DNA sequence patterns, such as gene function, regulatory mechanisms, and disease associations. These problems pose challenges to DNA sequence data analysis. Third, they generate a large number of candidate patterns, increasing computational complexity.
[0005] Therefore, how to analyze DNA sequence data more deeply and help researchers better combine the biological characteristics of DNA sequences to analyze and predict DNA sequence functions has become an urgent problem to be solved in the field of bioinformatics. Thus, it is necessary to study methods, systems, and applications for utility sequence pattern mining oriented towards multi-objective decision-making to solve the above-mentioned technical problems. Summary of the Invention
[0006] The purpose of this invention is to provide a method, system, and application for utility sequence pattern mining for multi-objective decision-making. In this method and system, not only are quantitative indicators used to screen out DNA sequence patterns that occur frequently in the dataset, but the biological importance of DNA sequence patterns is also evaluated through a utility function. When mining DNA sequence data, this invention can discover sequence patterns that have optimal solutions in both quantity and utility, enabling a deeper exploration of the potential patterns in DNA sequence data and providing more valuable information for biological research.
[0007] To achieve the above objectives, the technical solution adopted by this invention is as follows:
[0008] A method, system, and application for mining utility sequence patterns for multi-objective decision-making, including the following steps:
[0009] SF1. Scan the DNA sequence database, obtain DNA sequences that appear in the DNA database but are only for a single item, and mark them as T1 sequences. Construct a quantity utility array for the T1 sequences and mark it as a QU array.
[0010] SF2. Apply the SQUD pruning strategy to each T1 sequence. The SQUD pruning strategy is as follows: calculate the sequence weighted utility value of each T1 sequence based on the QU array and mark it as SWU, construct the maximum utility quantity array and mark it as MUQA, and then delete the T1 sequences whose SWU is less than the corresponding MUQA value.
[0011] SF3. Record the T1 sequence of MUQA after non-dominated sorting as an initialized sequence Pareto front array and label it as SPFA; generate a projection database for the T1 sequence in SPFA based on the QU array and label it as QUPro; traverse each T1 sequence in MUQA and generate an extended pattern for the traversed sequence pattern.
[0012] SF4. After the above expansion pattern is generated, the DNA data in each T1 sequence has a utility upper bound. The utility upper bound is calculated using the Projection Expansion Utility Method (PEU). The utility upper bound of the DNA data in the T1 sequence is calculated using PEU and evaluated. The DNA sequence in the T1 sequence is compared with the DNA sequence in the SPFA. When the utility upper bound or quantity index of the sequence in T1 is not lower than the corresponding value of any sequence in the SPFA, the expansion operation is triggered, and a new candidate sequence is generated. The sequence is expanded from the T1 sequence to T2.
[0013] SF5. Repeat steps SF3 and SF4 until the candidate sequence has been completely traversed and no new candidate sequence is generated.
[0014] SF6. Output all non-dominated sequences in SPFA as a set of utility sequence patterns.
[0015] Preferably, in step SF1, the QU array contains five parallel subarrays, which are as follows:
[0016] Item array: stores q - The name of each item in the sequence;
[0017] Utility array: records the local utility value of each item;
[0018] Quantity array: stores the quantity of each item;
[0019] Residual utility array: Stores the sum of residual utility from the current item index to the end of the sequence;
[0020] Element index table: marks the position of each item in the original sequence.
[0021] Preferably, the sequence-weighted utility value (SWU) calculation and the maximum utility quantity array (MUQA) construction method in step SF2 are as follows:
[0022] The sequence-weighted utility (SWU) for each T1 sequence is calculated using the following formula:
[0023]
[0024] In equation (01), T represents the DNA sequence to be expanded, S represents the full-length DNA sequence, D represents the DNA database, and u represents the utility value.
[0025] The formula for constructing the maximum utility quantity array (MUQA) is as follows:
[0026] MUQA(i)=max{U(T,D)|Q(T,D)=i} (02);
[0027] In equation (02), i represents different quantity values, U represents the utility value of the sequence, and Q represents the quantity value of the sequence.
[0028] Preferably, the calculation method of the projection extension utility method (PEU) in step SF4 is as follows:
[0029]
[0030] After calculating the PEU value of the current matching position of the DNA sequence according to formula (03), the maximum value of the PEU value is then calculated according to formula (04).
[0031]
[0032] Then, the PEU value is obtained according to equation (05):
[0033]
[0034] Where T represents the DNA sequence to be expanded, p represents the position of the DNA data; S represents the full-length DNA sequence; and I represents the sum of all sequence fragments after p.
[0035] Preferably, in step SF3, the projection database QUPro is a data structure used to efficiently mine efficient sequence patterns, and it is a sequence projection structure. The projection database QUPro expands the required sequence information through dynamic maintenance patterns to guide the utility calculation and pruning process.
[0036] The architecture of the projection database QUPro includes: a quantity utility array (QU-array) and an extension list (Extension-list);
[0037] Among them, the quantity utility array QU-array is a utility array that directly references the original sequence. It is used to store the utility value of each item in the sequence, retaining the complete utility information of the original data and avoiding duplicate calculations.
[0038] The Extension-list records the extension position index and utility value of the current pattern in the sequence. The Extension-list includes a position index module and a cumulative utility module. The position index module shows the specific position of the new item in the original sequence when the pattern is extended. The cumulative utility module shows the cumulative utility from the current pattern to the extension point, which is used to quickly filter high-potential candidate patterns.
[0039] The QUPro projection database is dynamically generated only during pattern expansion, rather than pre-compiling all possible patterns. After storing the original sequence pointers and the expanded list associated with the current pattern, the original sequence pointers are marked as QU-array.
[0040] Preferably, the specific steps for generating the SF3 extended mode include the following:
[0041] SF301. Based on the projection database QUPro constructed from sequence T, obtain the list of all connectable extension items.
[0042] SF302. Remove irrelevant extension items from the queue according to the IQUD pruning strategy;
[0043] SF303. Traverse the queue of the above extension items, perform a concatenation operation on each extension item, and obtain the extension sequence T′;
[0044] SF304, the QUPro projection database of scan sequence T, is used to prune and reduce candidate sequences according to the EQUD pruning strategy, and the QUPro of each extended sequence T′ is constructed.
[0045] SF305. Calculate the specific utility value and number of the remaining candidate sequences, determine whether they dominate the sequences in the SPFA, and if they dominate, add the candidate sequences to the updated SPFA.
[0046] Preferably, the IQUD strategy in step SF302 specifically includes the following steps:
[0047] A01. Traverse the QUPro projection database of the current candidate sequence to obtain all possible extensions;
[0048] A02. Calculate the compact sequence utility for each expansion item and label it as RSU and quantity;
[0049] A03. For each extension, check whether its RSU and quantity are dominated by any sequence already present in the sequence Pareto pre-SPFA.
[0050] A04. If there exists a sequence X∈SPFA such that the number of X is greater than or equal to the number of current extensions and the utility of X is greater than or equal to the RSU of the current extension, then the extension is determined to be dominated.
[0051] Preferably, the formula for calculating the RSU of compact sequence utility in step A02 is:
[0052]
[0053] Equation (06) above represents the compact sequence utility value for each matching sequence;
[0054]
[0055] Equation (07) above is the sum of the compact sequence utility values for each matching sequence.
[0056] Preferably, the EQUD strategy in step SF304 specifically includes the following steps:
[0057] B01. During the pattern expansion process, for the current candidate sequence T′, calculate its compact utility upper bound TRSU and quantity Q(T′,D). The compact utility upper bound TRSU is the utility value of a more compact sequence than RSU. The sequence T is generated by connecting the parent sequence T.
[0058] B02. Compare the upper bound of the compactness effect of T′, TRSU, and the quantity Q(T′,D) with all sequences in the sequence Pareto front SPFA.
[0059] B03. If there exists any sequence X∈SPFA such that Q(X,D)≥Q(T′,D) and U(X,D)≥TRSU(T′,D), then it is determined that T′ is dominated by X.
[0060] B04. If T′ is dominated, then directly prune T′ and all its extended sequences, stop generating its projection database QUPro, and terminate the recursive search of the current branch.
[0061] B05. If T′ is not dominated, add T′ to SPFA and remove the old sequence dominated by T′ in SPFA. Continue to recursively generate candidate sequences, generate the projection database QUPro of T′, and mine its extended sequences.
[0062] Preferably, the formula for calculating the upper bound of the tightening effect TRSU is:
[0063]
[0064] Equation (08) is the upper bound of the compact utility of each matching sequence;
[0065]
[0066] Equation (09) is the sum of the upper bounds of the compact utility of each matched sequence; ru represents the total utility of the sequence segment after the matching position.
[0067] The beneficial effects of the present invention are:
[0068] (1) Improve mining efficiency: By introducing utility upper bound and pruning strategy, unnecessary calculations are reduced and the efficiency of the algorithm is improved.
[0069] (2) Ensure the quality of results: Consider both quantity and utility, the sequence patterns mined are important in terms of both biological significance and data representativeness.
[0070] (3) Support for large-scale data processing: Optimized data structures (such as QU-array and QUPro) enable the algorithm to process large-scale DNA sequence data. Attached Figure Description
[0071] Figure 1 This is a flowchart of a utility sequence pattern mining method for multi-objective decision-making. Detailed Implementation
[0072] The present invention will now be described in detail with reference to the accompanying drawings:
[0073] Example 1
[0074] Referring to the attached diagrams in the instruction manual Figure 1 A utility sequence pattern mining method for multi-objective decision-making includes the following steps:
[0075] SF1. Scan the DNA sequence database, obtain DNA sequences that appear in the DNA database but are only for a single item, and mark them as T1 sequences. Construct a quantity utility array for the T1 sequences and mark it as a QU array.
[0076] SF2. Apply the SQUD pruning strategy to each T1 sequence. The SQUD pruning strategy is as follows: calculate the sequence weighted utility value of each T1 sequence based on the QU array and mark it as SWU, construct the maximum utility quantity array and mark it as MUQA, and then delete the T1 sequences whose SWU is less than the corresponding MUQA value.
[0077] SF3. Record the T1 sequence of MUQA after non-dominated sorting as an initialized sequence Pareto front array and label it as SPFA; generate a projection database for the T1 sequence in SPFA based on the QU array and label it as QUPro; traverse each T1 sequence in MUQA and generate an extended pattern for the traversed sequence pattern.
[0078] SF4. After the above expansion pattern is generated, the DNA data in each T1 sequence has a utility upper bound. The utility upper bound is calculated using the Projection Expansion Utility Method (PEU). The utility upper bound of the DNA data in the T1 sequence is calculated using PEU and evaluated. The DNA sequence in the T1 sequence is compared with the DNA sequence in the SPFA. When the utility upper bound or quantity index of the sequence in T1 is not lower than the corresponding value of any sequence in the SPFA, the expansion operation is triggered, and a new candidate sequence is generated. The sequence is expanded from the T1 sequence to T2.
[0079] SF5. Repeat steps SF3 and SF4 until the candidate sequence has been completely traversed and no new candidate sequence is generated.
[0080] SF6. Output all non-dominated sequences in SPFA as a set of utility sequence patterns.
[0081] In step SF1, the QU array contains five parallel subarrays, which are as follows:
[0082] Item array: Stores the name of each item in the q-sequence;
[0083] Utility array: records the local utility value of each item;
[0084] Quantity array: stores the quantity of each item;
[0085] Residual utility array: Stores the sum of residual utility from the current item index to the end of the sequence;
[0086] Element index table: marks the position of each item in the original sequence.
[0087] Example 2
[0088] Based on the above embodiments, this embodiment further discloses the following: The method for calculating the sequence weighted utility value (SWU) and constructing the maximum utility quantity array (MUQA) in step SF2 is as follows: The sequence weighted utility (SWU) for each T1 sequence is calculated using the following formula:
[0089]
[0090] In equation (01), T represents the DNA sequence to be expanded, S represents the full-length DNA sequence, D represents the DNA database, and u represents the utility value.
[0091] The formula for constructing the maximum utility quantity array (MUQA) is as follows:
[0092] MUQA(i)=max{U(T,D)|Q(T,D)=i} (02);
[0093] In equation (02), i represents different quantity values, U represents the utility value of the sequence, and Q represents the quantity value of the sequence.
[0094] The calculation method for the projection extension utility (PEU) in step SF4 is as follows:
[0095]
[0096] After calculating the PEU value of the current matching position of the DNA sequence according to formula (03), the maximum value of the PEU value is then calculated according to formula (04).
[0097]
[0098] Then, the PEU value is obtained according to equation (05):
[0099]
[0100] Where T represents the DNA sequence to be expanded, p represents the position of the DNA data; S represents the full-length DNA sequence; and I represents the sum of all sequence fragments after p.
[0101] Example 3
[0102] Based on the above embodiments, this embodiment further discloses the following:
[0103] QUPro is a projection database that is a data structure for efficiently mining efficient sequence patterns. It is a sequence projection structure. QUPro expands the required sequence information by dynamically maintaining the pattern, guiding the utility calculation and pruning process.
[0104] The architecture of the projection database QUPro includes: a quantity utility array (QU-array) and an extension list (Extension-list);
[0105] Among them, the quantity utility array QU-array is a utility array that directly references the original sequence. It is used to store the utility value of each item in the sequence, retaining the complete utility information of the original data and avoiding duplicate calculations.
[0106] The Extension-list records the extension position index and utility value of the current pattern in the sequence. The Extension-list includes a position index module and a cumulative utility module. The position index module shows the specific position of the new item in the original sequence when the pattern is extended. The cumulative utility module shows the cumulative utility from the current pattern to the extension point, which is used to quickly filter high-potential candidate patterns.
[0107] The QUPro projection database is dynamically generated only during pattern expansion, rather than pre-compiling all possible patterns. After storing the original sequence pointers and the expanded list associated with the current pattern, the original sequence pointers are marked as QU-array.
[0108] The specific steps for generating the SF3 extended mode include the following:
[0109] SF301. Based on the projection database QUPro constructed from sequence T, obtain the list of all connectable extension items.
[0110] SF302. Remove irrelevant extension items from the queue according to the IQUD pruning strategy;
[0111] SF303. Traverse the queue of the above extension items, perform a concatenation operation on each extension item, and obtain the extension sequence T′;
[0112] SF304, the QUPro projection database of scan sequence T, is used to prune and reduce candidate sequences according to the EQUD pruning strategy, and the QUPro of each extended sequence T′ is constructed.
[0113] SF305. Calculate the specific utility value and number of the remaining candidate sequences, determine whether they dominate the sequences in the SPFA, and if they dominate, add the candidate sequences to the updated SPFA.
[0114] Example 4
[0115] Based on the above embodiments, this embodiment further discloses the following:
[0116] In a utility sequence pattern mining system for multi-objective decision-making, the IQUD strategy is adopted, which specifically includes the following steps:
[0117] A01. Traverse the QUPro projection database of the current candidate sequence to obtain all possible extensions;
[0118] A02. Calculate the compact sequence utility for each expansion item and label it as RSU and quantity;
[0119] A03. For each extension, check whether its RSU and quantity are dominated by any sequence already present in the sequence Pareto pre-SPFA.
[0120] A04. If there exists a sequence X∈SPFA such that the number of X is greater than or equal to the number of current extensions and the utility of X is greater than or equal to the RSU of the current extension, then the extension is determined to be dominated.
[0121] The formula for calculating the RSU of compact sequence utility in step A02 is:
[0122]
[0123] Equation (06) above represents the compact sequence utility value for each matching sequence;
[0124]
[0125] Equation (07) above is the sum of the compact sequence utility values for each matching sequence.
[0126] Example 5
[0127] Based on the above embodiments, this embodiment further discloses the following:
[0128] The EQUD strategy in step SF304 specifically includes the following steps:
[0129] B01. During the pattern expansion process, for the current candidate sequence T′, calculate its compact utility upper bound TRSU and quantity Q(T′,D). The compact utility upper bound TRSU is the utility value of a more compact sequence than RSU. The sequence T is generated by connecting the parent sequence T.
[0130] B02. Compare the upper bound of the compactness effect of T′, TRSU, and the quantity Q(T′,D) with all sequences in the sequence Pareto front SPFA.
[0131] B03. If there exists any sequence X∈SPFA such that Q(X,D)≥Q(T′,D) and U(X,D)≥TRSU(T′,D), then it is determined that T′ is dominated by X.
[0132] B04. If T′ is dominated, then directly prune T′ and all its extended sequences, stop generating its projection database QUPro, and terminate the recursive search of the current branch.
[0133] B05. If T′ is not dominated, add T′ to SPFA and remove the old sequence dominated by T′ in SPFA. Continue to recursively generate candidate sequences, generate the projection database QUPro of T′, and mine its extended sequences.
[0134] The formula for calculating the upper bound of the tightening effect TRSU is as follows:
[0135]
[0136] Equation (08) is the upper bound of the compact utility of each matching sequence;
[0137]
[0138] Equation (09) is the sum of the upper bounds of the compact utility of each matched sequence; ru represents the total utility of the sequence segment after the matching position.
[0139] Example 6
[0140] Application of DNA utility sequence pattern mining for multi-objective decision-making (SQUSP-DNA algorithm)
[0141] Based on the above-described invention, this embodiment further discloses a multi-objective utility sequence pattern mining method applied to DNA sequence databases. The specific implementation employs the aforementioned SQUSP-DNA algorithm, and its complete mining process is as follows:
[0142] Step 601: Construct the initial utility structure:
[0143] Scan each sequence in the DNA database D to extract all item sequences of length 1 (i.e., T1T1 sequences). Construct a five-dimensional quantity-utility array (QU-array) for each T1 sequence, including: item name, utility value, quantity value, remaining utility, and element index.
[0144] Step 602: Execute the SQUD pruning strategy:
[0145] Calculate the weighted utility (SWU) of each T1 sequence, construct the maximum utility quantity array (MUQA), compare the SWU of each T1 sequence with the corresponding MUQA value, and if the SWU is less than the corresponding MUQA value, then prune and remove the sequence.
[0146] Step 603: Initialize the sequence Pareto front set (SPFA):
[0147] Perform non-dominated sorting on the remaining T1 sequences, retaining all T1 patterns that are not dominated in terms of quantity and utility as the initial Skyline pattern set SPFA; for each sequence in the SPFA, construct the corresponding projection database QUPro based on its QU-array.
[0148] Step 604: Recursively execute pattern expansion and pruning. For each unprocessed candidate sequence, perform the following operations:
[0149] First, expand and generate candidate patterns: based on the expansion position of the current pattern in QUPro, generate all connectable expansion items; perform I-connection (item connection) and S-connection (sequence connection) respectively to generate new candidate patterns T′.
[0150] Secondly, the IQUD pruning strategy is applied: the compact utility RSU and quantity of each extension are calculated; if its RSU and quantity are dominated by any sequence in SPFA, the extension is pruned.
[0151] Furthermore, the EQUD pruning strategy is applied: for each expansion candidate T′, its compact utility upper bound TRSU is calculated; if both TRSU and quantity are dominated by any sequence in SPFA, then T′ and all its expansions are terminated; otherwise, the expansion continues and its projection database QUPro is updated.
[0152] Finally, update the Skyline set SPFA: if T′ is not dominated, add it to the SPFA and remove all old sequences dominated by T′. This process is repeated until all candidate sequences have been processed or no new candidate sequences are generated.
[0153] In this embodiment, the system content code includes the following:
[0154]
[0155]
[0156] Example 7
[0157] Based on the above embodiments, this embodiment further discloses the following:
[0158] This patent uses DNA sequence data from a genome database as the data source for mining. Below is an example of a DNA sequence database after data preprocessing; sequence IDs represent different gene sequences, and letters represent bases or base combinations:
[0159] Table 1 DNA Sequence Database
[0160] Serial ID DNA sequence pattern SF1 <A(TG)(CG)A(AT)> SF2 <(AC)G(TA)(CG)> SF3 <(GC)(AT)(AG)TC> SF4 <GA(CT)TGA> SF5 <(TG)>
[0161] In Table 1, each sequence consists of a series of itemsets, and the items / itemsets in the sequence are ordered. An item represents a single base or combination of bases, such as item A representing adenine, item G representing guanine, item C representing cytosine, and item T representing thymine. An itemset represents a set of bases appearing at a specific position. For example, the DNA sequence pattern <(AC)G(TA)(CG)> for sequence ID SF2 contains four itemsets: (AC), G, (TA), and (CG), and four items: A, C, G, and T. If an itemset contains only one item, the parentheses can be omitted, such as itemset G.
[0162] When implementing the application, the specific content includes the following:
[0163] First, prepare the DNA sequence database; Table 1 shows the DNA sequence data from a certain genome database.
[0164] Secondly, utility and quantity calculations are performed; in this step, the utility and quantity of each DNA sequence pattern (T1 sequence) of length 1 are evaluated to construct the basic structure QU-array for mining, including: utility calculation method.
[0165] The utility of each sequence pattern is assessed using three biological dimensions: Conservation Score: This assesses the degree to which the fragment has been preserved during evolution through cross-species sequence alignment. Functional Position Score: This score is based on the pattern's presence in functional locations such as promoters, coding regions, and regulatory regions. Disease Association Score: This assesses the significant association between the pattern and known diseases.
[0166] The three scores mentioned above are denoted as u1, u2, and u3, respectively. The overall utility value is calculated using a weighted function:
[0167] u(p,S)=α·u1+β·u2+γ·u3
[0168] Where α,β,γ∈[0,1] are preset weight parameters that satisfy α+β+γ=1.
[0169] In the quantity calculation method, the quantity value q(p,S) represents the number of times the pattern appears in sequence s; the total quantity Q(p,D) is the cumulative number of times it appears in the entire database D.
[0170] Example 8
[0171] An example is provided based on the calculation method and sequence in Example 7:
[0172] In sequence mode<A(TG)> For example, it appears once in sequence SF1. Its conservation score is 0.9, its functional area localization score is 1.0, and its disease association score is 0.2. With parameters set to α = 0.4, β = 0.4, and γ = 0.2, its utility is calculated as follows:
[0173] u(<A(TG)> ,SF1)=0.4×0.9+0.4×1.0+0.2×0.2=0.36+0.4+0.04=0.8; its quantity value is 1, then the initial QU-array items are processed as follows:
[0174] First, perform pattern expansion and pruning:
[0175] Extended sequence pattern<A(TG)> ,generate<A(TG)C> Identify candidate patterns; calculate the utility value and TRSU of the candidate patterns, and then calculate the upper bound of utility; based on the upper bound of utility and pruning conditions, decide whether to retain the candidate patterns for further expansion.
[0176] Then, the algorithm execution result is:
[0177] The final set of skyline patterns includes<A(TG)> ,<G(AT)> Sequence patterns that have advantages in both quantity and utility.
[0178] Then it is applied to gene function prediction:
[0179] By utilizing the discovered sequence patterns, functional regions of genes can be predicted, aiding in gene annotation. Through the mining of DNA sequence data, this invention effectively discovers biologically significant sequence patterns, providing crucial support for gene function research.
[0180] In this embodiment, its application is as follows:
[0181] (1) Gene functional region identification: accurately identify key functional region sequences to support genome annotation and functional prediction.
[0182] (2) Disease mechanism research: Discover specific sequence patterns related to diseases and reveal the genetic mechanisms of diseases.
[0183] (3) Drug target discovery: Locating potential drug target sequences to assist in new drug development.
[0184] The DNA sequence pattern mining method based on the SQUSP algorithm in this embodiment enables more in-depth analysis of DNA sequence data, providing strong technical support for bioinformatics research.
[0185] Of course, the above description is not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention should also fall within the protection scope of the present invention.
Claims
1. A method, system, and application for utility sequence pattern mining in multi-objective decision-making, characterized in that, Includes the following steps: SF1. Scan the DNA sequence database, obtain DNA sequences that appear in the DNA database but are only for a single item, and mark them as T1 sequences. Construct a quantity utility array for the T1 sequences and mark it as a QU array. SF2. Apply the SQUD pruning strategy to each T1 sequence. The SQUD pruning strategy is as follows: calculate the sequence weighted utility value of each T1 sequence based on the QU array and mark it as SWU, construct the maximum utility quantity array and mark it as MUQA, and then delete the T1 sequences whose SWU is less than the corresponding MUQA value. SF3. Record the T1 sequence of MUQA after non-dominated sorting as an initialized sequence Pareto front array and label it as SPFA; generate a projection database for the T1 sequence in SPFA based on the QU array and label it as QUPro; traverse each T1 sequence in MUQA and generate an extended pattern for the traversed sequence pattern. SF4. After the above expansion pattern is generated, the DNA data in each T1 sequence has a utility upper bound. The utility upper bound is calculated using the Projection Expansion Utility Method (PEU). The utility upper bound of the DNA data in the T1 sequence is calculated using PEU and evaluated. The DNA sequence in the T1 sequence is compared with the DNA sequence in the SPFA. When the utility upper bound or quantity index of the sequence in T1 is not lower than the corresponding value of any sequence in the SPFA, the expansion operation is triggered, and a new candidate sequence is generated. The sequence is expanded from the T1 sequence to T2. SF5. Repeat steps SF3 and SF4 until the candidate sequence has been completely traversed and no new candidate sequence is generated. SF6. Output all non-dominated sequences in SPFA as a set of utility sequence patterns.
2. The utility sequence pattern mining method, system, and application for multi-objective decision-making according to claim 1, characterized in that, In step SF1, the QU array contains five parallel subarrays, which are as follows: Item array: Stores the name of each item in the q-sequence; Utility array: records the local utility value of each item; Count array: stores the quantity of each item; Residual utility array: Stores the sum of residual utility from the current item index to the end of the sequence; Element index table: marks the position of each item in the original sequence.
3. The utility sequence pattern mining method, system, and application for multi-objective decision-making according to claim 1, characterized in that, The sequence-weighted utility value (SWU) calculation and the maximum utility quantity (MUQA) array construction method in step SF2 are as follows: The sequence-weighted utility (SWU) for each T1 sequence is calculated using the following formula: In equation (01), T represents the DNA sequence to be expanded, S represents the full-length DNA sequence, D represents the DNA database, and u represents the utility value. The formula for constructing the maximum utility quantity array (MUQA) is as follows: MUQA(i)=max{U(T,D)|Q(T,D)=i} (02); In equation (02), i represents different quantity values, U represents the utility value of the sequence, and Q represents the quantity value of the sequence.
4. The utility sequence pattern mining method, system, and application for multi-objective decision-making according to claim 1, characterized in that, The calculation method for the projection extension utility (PEU) in step SF4 is as follows: After calculating the PEU value of the current matching position of the DNA sequence according to formula (03), the maximum value of the PEU value is then calculated according to formula (04). Then, the PEU value is obtained according to equation (05): Where T represents the DNA sequence to be expanded, p represents the position of the DNA data; S represents the full-length DNA sequence; and I represents the sum of all sequence fragments after p.
5. The utility sequence pattern mining method, system, and application for multi-objective decision-making according to claim 1, characterized in that, In step SF3, the projection database QUPro is a data structure used to efficiently mine efficient sequence patterns, and it is a sequence projection structure. The projection database QUPro expands the required sequence information through dynamic maintenance patterns to guide the utility calculation and pruning process. The architecture of the projection database QUPro includes: a quantity utility array (QU-array) and an extension list (Extension-list); Among them, the quantity utility array QU-array is a utility array that directly references the original sequence. It is used to store the utility value of each item in the sequence, retaining the complete utility information of the original data and avoiding duplicate calculations. The Extension-list records the extension position index and utility value of the current pattern in the sequence. The Extension-list includes a position index module and a cumulative utility module. The position index module shows the specific position of the new item in the original sequence when the pattern is extended. The cumulative utility module shows the cumulative utility from the current pattern to the extension point, which is used to quickly filter high-potential candidate patterns. The QUPro projection database is dynamically generated only during pattern expansion, rather than pre-compiling all possible patterns. After storing the original sequence pointers and the expanded list associated with the current pattern, the original sequence pointers are marked as QU-array.
6. The utility sequence pattern mining method, system, and application for multi-objective decision-making according to claim 1, characterized in that, The specific steps for generating the SF3 extended mode include the following: SF301. Based on the projection database QUPro constructed from sequence T, obtain the list of all connectable extension items. SF302. Remove irrelevant extension items from the queue according to the IQUD pruning strategy; SF303. Traverse the queue of the above extension items, perform a concatenation operation on each extension item, and obtain the extension sequence T′; SF304, the QUPro projection database of scan sequence T, is used to prune and reduce candidate sequences according to the EQUD pruning strategy, and the QUPro of each extended sequence T′ is constructed. SF305. Calculate the specific utility value and number of the remaining candidate sequences, determine whether they dominate the sequences in the SPFA, and if they dominate, add the candidate sequences to the updated SPFA.
7. The utility sequence pattern mining method, system, and application for multi-objective decision-making according to claim 6, characterized in that, The IQUD strategy in step SF302 specifically includes the following steps: A01. Traverse the QUPro projection database of the current candidate sequence to obtain all possible extensions; A02. Calculate the compact sequence utility for each expansion item and label it as RSU and quantity; A03. For each extension, check whether its RSU and quantity are dominated by any sequence already present in the sequence Pareto pre-SPFA. A04. If there exists a sequence X∈SPFA such that the number of X is greater than or equal to the number of current extensions and the utility of X is greater than or equal to the RSU of the current extension, then the extension is determined to be dominated.
8. The utility sequence pattern mining method, system, and application for multi-objective decision-making according to claim 7, characterized in that, The formula for calculating the RSU of compact sequence utility in step A02 is as follows: Equation (06) above represents the compact sequence utility value for each matching sequence; Equation (07) above is the sum of the compact sequence utility values for each matching sequence.
9. The utility sequence pattern mining method, system, and application for multi-objective decision-making according to claim 6, characterized in that, The EQUD strategy in step SF304 specifically includes the following steps: B01. During the pattern expansion process, for the current candidate sequence T′, calculate its compact utility upper bound TRSU and quantity Q(T′,D). The compact utility upper bound TRSU is the utility value of a more compact sequence than RSU. The sequence T is generated by connecting the parent sequence T. B02. Compare the upper bound of the compactness effect of T′, TRSU, and the quantity Q(T′,D) with all sequences in the sequence Pareto front SPFA. B03. If there exists any sequence X∈SPFA such that Q(X,D)≥Q(T′,D) and U(X,D)≥TRSU(T′,D), then it is determined that T′ is dominated by X. B04. If T′ is dominated, then directly prune T′ and all its extended sequences, stop generating its projection database QUPro, and terminate the recursive search of the current branch. B05. If T′ is not dominated, add T′ to SPFA and remove the old sequence dominated by T′ in SPFA. Continue to recursively generate candidate sequences, generate the projection database QUPro of T′, and mine its extended sequences.
10. The utility sequence pattern mining method, system, and application for multi-objective decision-making according to claim 9, characterized in that, The formula for calculating the upper bound of the tightening effect TRSU is as follows: Equation (08) is the upper bound of the compact utility of each matching sequence; Equation (09) is the sum of the upper bounds of the compact utility of each matched sequence; ru represents the total utility of the sequence segment after the matching position.