Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

41 results about "Sequence clustering" patented technology

In bioinformatics, sequence clustering algorithms attempt to group biological sequences that are somehow related. The sequences can be either of genomic, "transcriptomic" (ESTs) or protein origin. For proteins, homologous sequences are typically grouped into families. For EST data, clustering is important to group sequences originating from the same gene before the ESTs are assembled to reconstruct the original mRNA.

A student intelligent psychological state evaluation system

The application discloses a student intelligent psychological state evaluation system, and relates to the technical field of data communication. First, the quality of student published data is checked and automatically corrected to ensure the reliability of subsequent processing input. Then, data sequence clusters are divided through clustering analysis, and resource scheduling strategies are dynamically adjusted based on the similarity confidence of each cluster to realize the optimization of analysis frequency and sampling rate of different clusters. Meanwhile, the resource scheduling strategy in the cluster is dynamically adjusted, and the optimal task allocation plan is intelligently calculated and issued according to the scheduling exposure index and resource proportion index of each sequence cluster to realize the fine matching of computing power and manpower. Finally, the real-time monitoring of key performance indicators improves the real-time performance and accuracy of student psychological state evaluation, and significantly optimizes the resource utilization efficiency and stability of the system in a high-concurrency scenario.
Owner:HUNAN ANZHI NETWORK TECH CO LTD

Heuristic biological sequence clustering method based on semi-global comparison

PendingCN120299521ABiostatisticsSequence analysisSequence alignment algorithmSequence clustering
The invention provides a heuristic biological sequence clustering method based on semi-global comparison. The method comprises the following steps: firstly, reading all biological sequences, removing repetitive sequences, and carrying out descending sorting on the biological sequences according to sequence lengths; a first sequence is used as a representative sequence of a first category, then a next sequence is read, the similarity between the current sequence and the representative sequence is calculated by adopting a semi-global sequence comparison algorithm, each sequence can be compared to the most similar area in the representative sequence through semi-global sequence comparison, the similarity between the sequences is found to the maximum extent, and the similarity between the sequences and the representative sequence is calculated. Obtaining a relatively high comparison similarity value between the sequences; if the similarity meets a clustering threshold value, adding the similarity into a category with the same representative sequence, otherwise, taking the similarity as a new representative sequence, and generating a new category; repeating the steps until all the sequences are processed; the number of the final representative sequences is the number of clustering categories, and the categories which are the same as the representative sequences are member sequences of each category.
Owner:BAOJI UNIV OF ARTS & SCI

Compound bait data set construction method based on structure and sequence collaborative redundancy elimination

A complex bait data set construction method based on structure and sequence collaborative redundancy elimination belongs to the field of bioinformatics, and comprises the following steps: screening an initial protein complex structure set, removing entries containing nucleic acids, small molecules or non-protein chains, and selecting binary complexes meeting integrity and resolution requirements; secondly, structure clustering and sequence clustering are carried out based on three-dimensional structure similarity and sequence homology, combined comparison is carried out on the two results, and redundant compound entries which are highly similar in structure and sequence are removed; then, taking each cluster representative compound as a target, generating a plurality of groups of bait structures by using a molecular docking or prediction modeling method, and calculating a quality index; and finally, performing stratified sampling and proportion balance based on the score interval of the quality index, and constructing a high-quality protein complex bait data set with structure and sequence collaborative redundancy elimination and balanced quality distribution. The data set generated by the method has the advantages of low redundancy, high diversity and quality distribution controllability.
Owner:ZHEJIANG UNIV OF TECH

Ploughing layer capacity expansion regulation and control method fused with organic material carbon-nitrogen conversion model

The invention relates to the field of agricultural data analysis, in particular to a plough layer capacity expansion regulation and control method fused with an organic material carbon nitrogen conversion model, and the method comprises the steps: obtaining carbon nitrogen mineralization rate time sequence original data; performing coupling evaluation on mineralization energy gravity center moment difference and morphological overlapping characteristics to obtain a phase hysteresis correction factor; performing normalized attenuation mapping on the sequence local oscillation energy functional and the total sample average oscillation energy to obtain a transient excitation smoothing factor; performing optimization factor multiplicative correction on the Euclidean distance to obtain a corrected distance and performing clustering iterative calculation; and performing carbon nitrogen conversion model parameter matching and regulation strategy reverse derivation on a clustering result to obtain a plough layer capacity expansion regulation scheme so as to solve the problems of regulation type region division distortion and model matching precision reduction caused by inoculation hysteresis translation and excitation effect high-frequency oscillation in the existing Euclidean distance-based mineralization rate time sequence clustering.
Owner:JILIN ACAD OF AGRI SCI

Anomaly detection through clustering of time-series data subsequences and determination of adaptive thresholding

Computerized methodologies are disclosed that are directed to detecting anomalies within a time-series data set. An aspect of the anomaly detection process includes determining one or more seasonality patterns that correspond to a specific time-series data set by evaluating a set of candidate seasonality patterns (e.g., hourly, daily, weekly, day-start off-sets, etc.). The evaluation of a candidate seasonality pattern may include dividing the time-series data set into a collection of subsequences based on the particular candidate seasonality pattern. Further, the collection of subsequences may be divided into clusters and a silhouette score may be computed to measure the clustering quality of the candidate seasonality pattern. In some instances, the candidate seasonality pattern having the highest silhouette score is selected and utilized in anomaly detection process. In other instances, a plurality of seasonality patterns may be combined forming a time policy, where the time policy is utilized in anomaly detection process.
Owner:CISCO TECHNOLOGY INC

Depth time sequence clustering enhancement-based disease deterioration risk identification method and system

PendingCN121768650AImprove discrimination abilityImprove migration abilityHealth-index calculationMedical automated diagnosisLaboratory Test ResultDisease
The invention relates to the technical field of clinical medical treatment, and discloses a disease deterioration risk identification method and system based on depth time sequence clustering enhancement, and the method comprises the steps: 1, obtaining multi-modal clinical sequence data of a patient, including physiological indexes of a time sequence, a laboratory detection result, historical diseases and medication data; 2, performing feature extraction and classification on the patient sequences by adopting a knowledge enhanced sequence clustering method, and grouping the patient sequences according to future outcome distribution of the patient sequences; 3, enabling the model to quickly adapt to a prediction task of a new patient subgroup through a meta-training process; and step 4, based on the trained meta-model, carrying out rapid adaptation on the new patient subtype, and predicting the possibility that the new patient subtype has a deterioration event in a certain time window in the future. The method and the system can effectively learn the disease change mode of the patient under the condition of limited clinical data, improve the prediction accuracy of the new patient subgroup, and are especially suitable for clinical prediction scenes under the condition of small samples.
Owner:ZHONGBEI UNIV

Intelligent dialogue method and system for accompanying old people based on user portraits

The invention discloses an elder accompanying intelligent dialogue method and system based on a user portrait, and relates to the technical field of intelligent interaction, and the method comprises the steps: calculating user features based on multi-modal data, splicing the user features into a user portrait vector, generating a frame sequence from voice, calculating an interaction feature vector, constructing an interaction sequence, generating a prefix tree, and extracting all paths. Calculating a support degree, and generating a candidate mode set; extracting an intention label corresponding to each mode in the candidate mode set, generating a frequent mode set, calculating similarity, generating an associated intention label and confidence, extracting frequency features in user portrait vectors, splicing the frequency features into interest feature sub-vectors, calculating joint utility, and selecting a maximum value as an optimal intention. According to the method, through combination of the prefix tree and frequent pattern mining and interaction sequence clustering, the capability of capturing long-term interaction behavior rules of old people is enhanced, and through combination of interest utility and matching utility, the accuracy of intention inference is improved.
Owner:KUAISHANGYUN (SHANGHAI) NETWORK TECHNOLOGY CO LTD

Data loading and conversion processing method under lake-warehouse fusion architecture

The invention belongs to the technical field of data management and intelligent analysis, and particularly relates to a data loading and conversion processing method under a lake-warehouse fusion architecture, which comprises the following steps of: after an externally uploaded file is received, performing compatibility verification and content verification on the externally uploaded file in sequence; identifying that the tables and the partitions which are influenced by data loading and need to be synchronously updated are added, executing atomic data loading operation, synchronously updating the tables and the partitions in the table and the partitions, and generating corresponding version numbers for each table and each partition at the same time; for a data block set under the lake-warehouse fusion architecture, performing hierarchical adjustment according to the access popularity of the data block set; income-cost evaluation is carried out on each data block, and the data blocks needing to be subjected to Z-sequence clustering rearrangement are recognized and added into a requeuing column; and performing Z-sequence clustering rearrangement on each data block in the requeuing column.
Owner:BEIJING INST OF TECH +1

A method for regulating the expansion of the topsoil by integrating a carbon and nitrogen conversion model of organic materials

This invention relates to the field of agricultural data analysis, and particularly to a method for expanding and regulating the topsoil volume by integrating an organic material carbon and nitrogen conversion model. The method includes: acquiring raw time-series data of carbon and nitrogen mineralization rates; obtaining a phase hysteresis correction factor by coupling and evaluating the temporal differences and morphological overlap characteristics of the mineralization energy centroid; obtaining a transient excitation smoothing factor by normalizing and attenuating the local oscillation energy functional of the sequence with the average oscillation energy of the entire sample; obtaining a corrected distance by performing multiplicative optimization of the Euclidean distance and performing iterative clustering calculations; and obtaining a topsoil volume expansion and regulation scheme by matching carbon and nitrogen conversion model parameters and inversely deriving the regulation strategy from the clustering results. This addresses the problems of distortion in the division of regulation type regions and decreased model matching accuracy caused by inoculation hysteresis translation and high-frequency oscillations of excitation effects in existing Euclidean distance-based mineralization rate time-series clustering.
Owner:JILIN ACAD OF AGRI SCI

Atmospheric pollution time series data feature fragment mining method and system based on form

PendingCN120670803AAir quality improvementEngineeringSequence clustering
The invention discloses a form-based atmospheric pollution time series data feature fragment mining method and system. According to the method, clustering analysis is carried out by extracting time sequence sub-fragments (u-fragments), so that typical modes in the air pollution process are recognized. Specifically, the method proposes a discrete feature representation method to capture representative local morphological features in atmospheric pollution time series data, and at the same time, uses the similarity between extended Euclidean distance degree quantum sequences, and identifies sub-fragments with pollution identification degree through sub-sequence clustering. Based on the method, the invention further constructs a form-based atmospheric pollution time series data feature fragment mining system. The system can mine pollution fragments with identification degree from atmospheric pollution time sequence data, overcomes the limitation that a traditional statistics and frequency domain analysis method is weak in local feature capturing capability, and can be applied to the fields of urban pollution mode recognition, pollution traceability auxiliary decision making, environmental governance optimization and the like.
Owner:LANZHOU JIAOTONG UNIV

Nucleic acid sequence clustering method, apparatus, computer readable storage medium, terminal

ActiveCN115497567BBiostatisticsSequence analysisNucleic acid sequencingSequence clustering
The application discloses a nucleic acid sequence clustering method and device, a computer readable storage medium and a terminal. The terminal constructs a tree structure with multiple branches to search a specified interval of a nucleic acid sequence, thereby avoiding a large amount of time consumed by traditional calculation of an editing distance. In addition, the application adopts a node drift algorithm to resist interference caused by errors in the nucleic acid sequence. Compared with existing nucleic acid clustering algorithms, the method provided by the application can cluster a large number of unidentified nucleic acid sequences, and has the functions of automatically correcting and comparing the clustered nucleic acid sequences. The method can directly output the corrected nucleic acid original sequence, thereby greatly reducing the processing time after sequencing reading.
Owner:TIANJIN UNIV

Methods and systems for phasing sequencing strands and long-range sequencing

ActiveUS12637713B2Microbiological testing/measurementNucleotideSequence clustering
Described herein are methods synchronizing sequencing primers within a sequencing cluster and methods of generating long-range sequencing reads. The methods can include hybridizing primers to polynucleotide copies within a sequencing cluster; extending the primers through a first region of the polynucleotide copies using labeled nucleotides according to a sequencing flow order; extending the primers through a second region of the polynucleotide copies using one or more re-phasing flow steps that each include at least two different types of nucleotide bases; and extending the primers through a third region of the polynucleotide copies using labeled nucleotides according to the sequencing cycle. The rephasing flow steps may be initiated after a predetermined number of sequencing flow steps, after a measured sequencing signal strength falls below a predetermined sequencing signal strength threshold, or a measured sequencing signal-to-noise ratio falls below a sequencing signal-to-noise ratio threshold.
Owner:ULTIMA GENOMICS INC

Multi-system data connection and service fusion method and system based on front-end instruction

The invention provides a multi-system data communication and service fusion method and system based on a front-end instruction. According to the method, operation events such as clicking, inputting, skipping and data reading of a user in a plurality of service systems are collected at the front end of a browser, a cross-system operation track is constructed, key action extraction, sequence clustering and structural feature recognition are carried out on the operation track, and a reusable service mode template is automatically generated. Based on natural language service requirements of a user, system execution intention recognition and task element extraction, task requirements are mapped to a service template, structured task description is generated, a front-end instruction sequence and a cross-system dependency relationship are constructed according to the structured task description, and an executable instruction diagram is formed. The instruction graph is dynamically executed in the browser. According to the method, the original system interface does not need to be transformed, automatic communication and data collaboration of multi-system business processes can be realized, the manual operation cost is remarkably reduced, and the business processing efficiency and the data consistency are improved.
Owner:国网甘肃省电力公司嘉峪关供电公司

Three-phase load prediction method based on time sequence clustering

The invention discloses a three-phase load prediction method based on time sequence clustering, and the method comprises the steps: firstly collecting the data of an intelligent electric meter of a user in a target transformer area, and carrying out the preprocessing of the collected current load data; clustering the preprocessed load data based on a time sequence clustering algorithm to obtain K clusters of data of different load types; and clustering is completed by iteratively updating the clustering centroid. Constructing a space-time diagram neural network according to the clustering result and the three-phase phase information of the user; the space-time diagram neural network is used for carrying out space-time feature extraction and fusion on the clustered load data and outputting a three-phase load prediction result; and finally, training the constructed space-time diagram neural network by using the training data, and predicting the future three-phase load by using the trained model. According to the method, data with the same load type are classified into one class through a time sequence clustering algorithm, and a space-time diagram neural network is built according to three-phase phase information of a user to carry out three-phase load prediction.
Owner:HANGZHOU DIANZI UNIV +1

Abnormal behavior detection method and system for communication information flow

The invention provides an abnormal behavior detection method and system for a communication information flow, and relates to the technical field of network communication security, firstly, an original communication information flow set containing massive communication session recording units is captured in a network communication link, and each recording unit carries information such as source and destination internet protocol address tags; performing communication behavior pattern analysis on the original communication information flow set, and constructing a cross-session associated interaction behavior sequence cluster; secondly, extracting features of an interactive behavior time sequence chain, and processing the features through a semantic embedding layer and a time sequence convolution encoder to generate a joint embedded feature vector; and inputting the vector into an abnormal behavior prototype network for prototype matching to obtain an initial abnormal behavior mark. And finally, tracing the abnormal behavior propagation path according to the initial mark to obtain an abnormal behavior propagation path sequence set. According to the invention, abnormal behaviors in network communication can be effectively detected, and the detection accuracy and reliability are improved.
Owner:SHANGHAI MINGQI NETWORK TECH CO LTD

A method for analyzing macroviral group data

ActiveCN116682492BSequence analysisHybridisationMedicineSequence clustering
The application discloses a macrovirus group data analysis method, and belongs to the technical field of macrovirology. The method comprises the following steps: sequence quality control, sequence assembly, sequence clustering, virus sequence identification, virus sequence checking, virus abundance calculation, species annotation, virus lifestyle judgment, virus host prediction and virus auxiliary metabolism gene analysis. The application uses Trimmomatic software, BWA-MEN, Megahit software, CD-HIT and other tools to execute the analysis process of macrovirus group data. Practice proves that the application can accurately identify and annotate virus species, comprehensively and systematically deeply analyze and mine macrovirus group data, the steps are simple and clear, the analysis time is short, and the effect of macrovirus identification research is greatly improved.
Owner:JIANGNAN UNIV

Deep active time sequence clustering method fused with multi-level comparative learning

Aiming at the defects in a DATC-CT method, the invention provides an active time sequence clustering model (DATC-MC) based on multi-level comparative learning, the model fuses a multi-level comparative learning mechanism comprising instance-level comparison and class-cluster-level comparison on the basis of the DATC-CT method, the sample feature representation quality is optimized, and the distinguishing capability of boundaries between class clusters is enhanced. In the consolidation stage of active sampling, the DATC-MC calculates the density of each sample in a representation space by using a KNN algorithm, preferentially selects the sample with the maximum sample density, and stores the sample with the maximum sample density in an annotation class cluster set ACS or a corresponding auxiliary annotation set AAS, thereby generating high-quality ML and CL constraints. And the DATC-MC model uses comparison loss and constraint loss to jointly optimize coding network parameters and a class cluster center. A contrast experiment result on a public data set shows that the clustering performance of the proposed model is superior to that of an existing active time sequence clustering method.
Owner:CIVIL AVIATION UNIV OF CHINA

Metagenome sequence clustering method, system and equipment based on multi-modal feature fusion and storage medium

The invention relates to the technical field of metagenome sequence clustering, in particular to a method, a system and equipment for clustering metagenome sequences based on multi-modal feature fusion and a storage medium. According to the clustering method, a Dirichlet process Gaussian mixture model and a sparse affinity graph model are combined, and in the clustering process, the clustering time is shortened, and the clustering efficiency is improved. The system can dynamically adjust a processing strategy (for example, whether to transfer into a sparse affinity graph model) according to the clustering number, ensures that a final clustering result meets the requirement (K = 1), adapts to metagenome data sets of different scales, realizes efficient and accurate clustering of metagenome data, can fully explore information contained in metagenome sequencing data, and improves the clustering efficiency. Effective clustering is carried out by utilizing biological characteristics contained in the sequencing data to the greatest extent, and the calculation complexity is reduced and the clustering process is accelerated through a hierarchical clustering method.
Owner:XI AN JIAOTONG UNIV

Self-adaptive two-stage clustering method for maximizing distance between load sequence clusters and storage medium

The invention relates to a self-adaptive two-stage clustering method for maximizing the distance between load sequence clusters and a storage medium, and solves the problems of poor clustering effect and the like in the prior art, and adopts the technical scheme that the method comprises the following steps of: extracting load data, and establishing a load data set; preprocessing the load data set; obtaining a plurality of first regions of interest (ROI); carrying out weighted calculation on the average cluster spacing through DB Index to maximize the average cluster spacing and complete first-stage adaptive clustering; selecting clusters with relatively large average cluster spacing obtained by the first-stage clustering to obtain a plurality of second regions of interest (ROI); and carrying out weighted calculation on the average cluster spacing through DB-Index to maximize the average cluster spacing and complete second-stage adaptive clustering. The method has the technical effects that the complexity in the power load data is effectively dealt with by maximizing the inter-cluster distance of the load data, a typical load data mode can be captured, and a low-density mode with a special value can be identified.
Owner:ELECTRIC POWER RES INST OF STATE GRID ZHEJIANG ELECTRIC POWER COMAPNY

Data processing method and device, equipment and medium

The embodiment of the invention discloses a data processing method and device, equipment and a medium, and is applied to the technical field of computers. The method comprises the following steps: acquiring a plurality of operation sequence groups to be clustered; performing sequence clustering on the plurality of operation sequence groups based on the operation sequences in the plurality of operation sequence groups to obtain clustering clusters related to the plurality of operation sequence groups; scene matching operation matched with the clustering scene corresponding to the clustering cluster is determined based on service operation represented by the operation sequence contained in the clustering cluster; and generating an element positioning anchor point for performing element positioning on a page element of the scene matching operation, so as to position the page element in a task page of a scene task through the element positioning anchor point when the scene task related to the clustering scene is obtained, and executing the scene matching operation of the scene task through the page element. By adopting the embodiment of the invention, the definition efficiency and flexibility of the business process can be improved.
Owner:XINGIN INFORMATION TECH (SHANGHAI) CO LTD

Method, system, equipment and medium for assembling third-generation sequencing data based on clustering and graph construction

The invention discloses a method, a system, equipment and a medium for assembling third-generation sequencing data based on clustering and graph construction, and belongs to the technical field of biological sequence processing. The method comprises the following steps: obtaining sequence similarity based on third-generation sequencing data, and clustering sequences by using a clustering algorithm to obtain different clusters; assembling the sequences in the same cluster based on a graph construction method to obtain an intra-group consensus sequence; and combining all intra-group consensus sequences from different clusters, and assembling based on the graph construction method to obtain an inter-group consensus sequence. According to the technical scheme, through the strategy of sequence clustering, intra-group assembly and inter-group assembly, the assembly complexity is effectively reduced, the accuracy and integrity of the consensus sequence are ensured through the optimal overlap graph algorithm and depth pruning, and the method has remarkable technical advantages.
Owner:欣基(杭州)生物科技有限公司

Paroxysmal atrial fibrillation attack position marking method and paroxysmal atrial fibrillation attack position marking system

The invention relates to the technical field of electrocardiogram monitoring, in particular to a paroxysmal atrial fibrillation attack position marking method and system.The paroxysmal atrial fibrillation attack position marking method comprises the steps that electrocardiosignals to be marked are obtained and input into a trained segmentation network model to obtain an output prediction sequence; traversing the prediction sequence and calculating a merging likelihood value and a segmentation total likelihood value of the current data point to perform sequence clustering; performing binarization processing on the prediction sequence to screen potential atrial fibrillation points; traversing all the potential atrial fibrillation points, judging the continuity with the previous potential atrial fibrillation point, and generating a plurality of potential atrial fibrillation event segments based on the continuity; processing the potential atrial fibrillation event segments based on the distances between the potential atrial fibrillation event segments and the sizes of the potential atrial fibrillation event segments to obtain actual atrial fibrillation event segments, wherein the processing actions comprise merging and eliminating; and determining an atrial fibrillation starting position and an atrial fibrillation ending position based on the wave information corresponding to the actual atrial fibrillation event segment. The method has the effect of improving the accuracy of atrial fibrillation event annotation recognition.
Owner:HANGZHOU PROTON TECH CO LTD

Method and system for monitoring working state of surgical robot

InactiveCN120632603ACluster algorithmMedicine
The invention relates to the technical field of electric data processing, in particular to a method and system for monitoring the working state of a surgical robot, and the method comprises the steps: collecting the operation parameters of the surgical robot in a surgical process; for any operation, acquiring a plurality of execution actions of the surgical robot in the operation process and a parameter sequence corresponding to each execution action period; and for any execution action, all the parameter sequences of the execution action are clustered by using a DBSCAN clustering algorithm, and the size of the clustering radius is inversely correlated with the size of the abnormal trend factors of all the parameter sequences in the execution action. The DBSCAN clustering algorithm is adopted to construct the standard parameter sequence, and the clustering radius is dynamically adjusted by introducing the abnormal trend factor, so that the difference of abnormal data distribution in different execution actions can be considered, the clustering precision is improved, and the missing report is reduced.
Owner:GUANGZHOU YINGHUIXING TECH CO LTD

A high-throughput sequencing data sequence clustering method and system based on cyclic self-blast

ActiveCN122067616BBarcodeSequence clustering
The application discloses a high-throughput sequencing data sequence clustering method and system based on cyclic self-blast: the high-throughput sequencing data is de-duplicated; unique sequences with a number of repetitions lower than a threshold value a are filtered, and the filtered unique sequences are sorted according to the number of repetitions; the sorted unique sequences are divided into subsets and a parent set; the subsets are subjected to blast multiple alignment to obtain a de-redundant subset; the parent set and the de-redundant subset are subjected to cyclic blast alignment to obtain a de-redundant parent set; all representative sequences in a preliminary clustering set are combined, sorted according to the number of repetitions and distributed with identification tags to generate a final OTU clustering result set; in the prior art, when OTU clustering and ASV methods are used to process environmental DNA macro-barcode sequencing data, different OTU or ASV sequences are annotated to the same species, so that a large amount of redundancy still exists in the classified units after clustering, and the reliability of species analysis is improved.
Owner:NANJING NORMAL UNIVERSITY +1

Analysis method and analysis system for automatically processing metagenome high-throughput sequencing data binning to obtain viral genome and application

PendingCN121096424AEnsemble learningBiostatisticsOriginal dataSequence clustering
The invention discloses an analysis method for automatically processing metagenome high-throughput sequencing data binning to obtain a viral genome, which comprises the following steps: acquiring second-generation sequencing original data, and performing data quality control to obtain Reads to be analyzed; a plurality of Reads are assembled into an overlapping group through overlapping of fragments, statistics is conducted on the overlapping group, and a length distribution diagram is drawn; reconstructing a viral genome based on contig binning, and generating a box after sequence clustering; constructing a three-dimensional matrix of all boxes according to the attributes of the sequence, and predicting and identifying the species category of each box; carrying out quality identification on the virus box bin, and detecting and rejecting virus host genes; counting the abundance of the virus box bin in each sample to generate a virus box bin abundance table; constructing a virus annotation database, completing species annotation, constructing a virus evolutionary tree, and finally generating a report. The invention further discloses an analysis system and application for implementing the analysis method.
Owner:SHANGHAI OE BIOTECH CO LTD

High-throughput sequencing data sequence clustering method and system based on circulating self-blast

ActiveCN122067616ABiostatisticsSequence analysisBarcodeSequence clustering
The invention discloses a high-throughput sequencing data sequence clustering method and system based on circulating self-blast. The method comprises the following steps: carrying out duplicate removal processing on high-throughput sequencing data; filtering the unique sequences of which the repetition number is lower than a threshold value a, and sorting the filtered unique sequences according to the repetition number; dividing the sorted unique sequence into a subset and a mother set; performing blast multiple comparison on the subsets to obtain redundancy-removed subsets; performing cyclic blast comparison on the mother set and the redundancy-removed subset to obtain a redundancy-removed mother set; combining all representative sequences in the preliminary clustering set, sorting according to the number of repetitions and distributing identification labels, and generating a final OTU clustering result set; in the prior art, when an OTU clustering method and an ASV method are used for processing environment DNA macro bar code sequencing data, different OTU or ASV sequences are annotated to the same species, so that a large amount of redundancy still exists in a classified unit after clustering, and the reliability of species analysis is improved.
Owner:NANJING NORMAL UNIVERSITY +1

A Deep Learning-Based Trajectory Sequence Clustering Method

The present invention relates to the field of data mining, and specifically relates to a trajectory sequence clustering method based on deep learning, including the following steps: Step 1, pre-training layer: Use a sequence-to-sequence autoencoder model to learn the low-dimensional feature representation of trajectory data; Step 2, initial clustering layer: Perform the K-Means clustering algorithm multiple times on the trajectory feature representation obtained by the pre-training layer, and select the cluster centers in the optimal clustering result as the initial cluster centers. Step 3, joint training optimization layer: Combine the trajectory clustering and deep feature extraction methods, propose an optimized loss function that combines the reconstruction error and clustering error of the sequence-to-sequence autoencoder model, and map the trajectory feature representation to a feature space more suitable for clustering.
Owner:ZHEJIANG LAB +1

Artificial intelligence (AI) validation dataset generation systems and methods for generating ai validation datasets for validating ai-based protein sequence models

PCT designated stageWO2026006797A1BiostatisticsInstrumentsSequence clusteringSequence model
Artificial intelligence (Al) validation dataset generation systems and methods are described for generating Al validation datasets for validating Al-based protein sequence models. Such systems and methods include filtering a reference proteome dataset to generate a high completeness proteome data subset comprising reference proteomes. The high completeness proteome data subset is filtered to generate a protein confidence-based proteome data subset with reference proteomes above a protein confidence threshold. The plurality of protein sequences is clustered to generate a set of protein sequence clusters and a validation cluster subset is randomly allocated therefrom for creation of a protein sequence validation dataset. Validation metrics may then be generated by comparing (a) a set of predetermined values of the protein sequence validation dataset; and (b) the one or more predicted protein sequences as output by an Al-based protein sequence model.
Owner:AMGEN INC

Multivariable time sequence clustering method and device based on binary dual clustering factors and medium

The invention relates to the technical field of time series clustering, and provides a multivariable time series clustering method and device based on binary dual clustering factors and a medium, and the method comprises the steps: carrying out the data preprocessing of to-be-clustered MTS data; performing data enhancement on the MTS data after data preprocessing to obtain a plurality of subsequences; binary dual clustering is carried out on the subsequences, and a binary dual clustering factor is calculated based on a clustering result; obtaining a distance matrix based on the binary dual clustering factor, and predicting an optimal clustering cluster number based on the distance matrix; and outputting a final MTS data clustering result based on the optimal clustering cluster number in combination with the distance matrix. According to the method, the precision of predicting the unknown potential class cluster number based on MTS data multivariable parameter features can be improved, and the clustering accuracy of complex MTS data is improved.
Owner:SOUTHWEST CHINA RES INST OF ELECTRONICS EQUIP