Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

115 results about "Cluster labeling" patented technology

In natural language processing and information retrieval, cluster labeling is the problem of picking descriptive, human-readable labels for the clusters produced by a document clustering algorithm; standard clustering algorithms do not typically produce any such labels. Cluster labeling algorithms examine the contents of the documents per cluster to find a labeling that summarize the topic of each cluster and distinguish the clusters from each other.

Frequency access data storage migration method and device for big data

The invention belongs to the field of data storage, and discloses a frequency access data storage migration method and device for big data, and the method comprises the steps: obtaining a data access log, carrying out the preprocessing of the data access log, and obtaining a time window statistical matrix; carrying out timeliness attenuation weighting calculation on a plurality of data access times in the time window statistical matrix to obtain a weighted access frequency vector; obtaining a multi-dimensional feature matrix based on the weighted access frequency vector; clustering the feature vector of each time window slice in the multi-dimensional feature matrix through a Gaussian mixture model to obtain a clustering tag vector; inputting the clustering label vector into an access frequency prediction model to obtain an access frequency prediction value of the data in a future preset time step; the migration decision matrix is determined based on the access frequency predicted value and the data storage cost parameter, data migration is carried out based on the migration decision matrix, the data migration complexity can be reduced, and the access speed and the storage cost are balanced.
Owner:XIAMEN MEIYA YIAN INFORMATION TECH CO LTD

Communication equipment production intelligent management system based on machine learning

The invention relates to the technical field of communication production management, and discloses a communication equipment production intelligent management system based on machine learning. The system comprises a production data acquisition module, a feature engineering construction module, a dynamic clustering analysis module, an anomaly detection engine module and a production decision optimization module. The production data acquisition module acquires multi-source sensor data in real time and converts the multi-source sensor data into a standardized sequence with a unified timestamp; the feature engineering module extracts a time domain statistical feature, a frequency domain energy feature and an equipment state association feature to generate a high-dimensional feature vector set; the dynamic clustering module adopts an incremental algorithm to divide clusters online; the anomaly detection module establishes a multi-level Gaussian mixture model based on a clustering label, and quantifies an anomaly probability through a mahalanobis distance; and the production decision module integrates the results to generate an equipment maintenance priority sequence and a production takt adjustment instruction. According to the system, intelligent monitoring and dynamic optimization of the whole production process of the communication equipment are realized, and the real-time change requirement of a complex production environment is met.
Owner:HANGZHOU WEISHI INFORMATION TECH CO LTD

Power multi-source data slice processing method and system based on reinforcement learning

The invention relates to an electric power multi-source data slicing processing method and system based on reinforcement learning, and the method comprises the following steps: S1, obtaining electric power multi-source data, and carrying out the preprocessing of the electric power multi-source data, and obtaining a structured original data stream; s2, obtaining a multi-dimensional feature vector through feature extraction; s3, adopting a clustering algorithm to identify a behavior pattern to obtain a clustering label, calculating the distance from a sample to each clustering center, and when the distance is obviously higher than the average value plus threshold value beta times standard deviation of the cluster, additionally marking as abnormal; s4, intelligent slicing decision making is carried out based on reinforcement learning, and a slicing strategy, data and labels are obtained; s5, according to the obtained slicing strategy, data and label, through knowledge graph service rule verification, it is ensured that domain constraints are met; and S6, sending the verified strategy to distributed slice storage and management, and carrying out data slice storage and index construction according to the strategy. According to the method, the structured governance level and the high-value utilization capability of the power data assets are effectively improved.
Owner:STATE GRID JIANGXI ELECTRIC POWER CO LTD ECONOMIC & TECH RES INST +2

Manufacturing system risk control knowledge matching method based on semantic embedding and clustering analysis

The invention relates to a manufacturing system risk control knowledge matching method based on semantic embedding and clustering analysis, and the method comprises the following steps: collecting and preprocessing risk control text data: collecting unstructured text data of a manufacturing system history record, and obtaining preprocessed risk control text data, constructing a professional corpus for a discrete manufacturing scene; text semantic embedding generation; semantic clustering modeling: performing unsupervised clustering modeling on all semantic vectors, mining semantic association and potential structures between texts, obtaining semantic representations of risk control knowledge through a clustering algorithm, and assisting in generating clustering tags; and a risk knowledge matching mechanism.
Owner:TIANJIN UNIV

Public opinion video tag aggregation method and system based on artificial intelligence

The invention provides a public opinion video tag aggregation method and system based on artificial intelligence, and relates to the technical field of artificial intelligence. Comprising the following steps: acquiring pictures and text information in a short video, and performing semantic alignment; different large language models are adopted to generate preliminary labels for the pictures and the text information after semantic alignment; clustering the pictures and the texts after semantic alignment to obtain clusters; calculating the labeling probability of each primary label type in the current cluster by each large language model, and selecting the primary label with the highest probability sum as a clustering label of the current cluster; calculating the reliability weight of each large language model in the current cluster based on the clustering label of the current cluster; and based on the reliability weight, calculating the weighted support degree of all the large language models to different preliminary label types of each piece of data in the current cluster, calculating the weighted label of the current data, and further determining a final label. According to the method, the condition of few labels or no labels can be effectively processed, and the manual workload is greatly reduced.
Owner:SHANDONG DAZHONG INFORMATION IND CO LTD

Method for identifying traditional Chinese medicinal materials by combining infrared spectroscopy with clustering analysis

The invention discloses a method for identifying traditional Chinese medicinal materials by combining infrared spectroscopy with clustering analysis, which comprises the following steps of: acquiring original infrared spectral data, averaging the original infrared spectral data to obtain single-sample original spectral data, constructing a sample graph and a wavelength graph based on standardized spectral characteristics, and fusing the sample graph and the wavelength graph to obtain the single-sample original spectral data. Obtaining fusion image data, inputting the fusion image data into a pre-trained image neural network model, extracting a low-dimensional feature vector of a target sample through forward propagation, and inputting the low-dimensional feature vector into a pre-trained dynamic cluster diffusion module to obtain a clustering label and distance data; and determining and outputting the quality grade of the rhizome traditional Chinese medicinal material sample based on the clustering label and the distance data. Therefore, interference can be effectively reduced, key features can be extracted, associated features of fused graph data are extracted in combination with a graph neural network, and the quality grade of traditional Chinese medicinal materials can be accurately determined in combination with clustering analysis of a dynamic cluster diffusion module.
Owner:ZHEJIANG FORESTRY UNIVERSITY

Multi-modal false news detection method based on unsupervised clustering and frequency domain information

The invention provides a multi-modal false news detection method based on unsupervised clustering and frequency domain information, and relates to the technical field of image text data. The method comprises the following steps: based on multi-modal sample data, respectively extracting text features, visual features and image text features, obtaining text feature clustering labels according to text feature global semantic correlation, inputting the text features with the labels into an unsupervised clustering learning network, and obtaining an unsupervised clustering learning result; semantic consistency is enhanced through a bidirectional gating loop unit and a multi-head attention mechanism, and enhanced text features are obtained; extracting frequency domain features based on the visual features, and fusing the frequency domain features with the spatial domain features to obtain visual joint features; and after image text features are also enhanced by the unsupervised clustering learning network, the image text features are fused with text and visual features through an attention fusion module by taking the image text features as a bridge to obtain multi-modal fusion features, and the multi-modal fusion features are input into a full connection layer to complete false news classification. According to the method, the accuracy of false news detection is effectively improved.
Owner:SOUTHWEST PETROLEUM UNIV

Non-technical line loss state evaluation method based on multi-scale spatial-temporal feature fusion in low-voltage distribution network environment

The invention provides a non-technical line loss state evaluation method based on multi-scale spatial-temporal feature fusion in a low-voltage distribution network environment, and relates to the technical field of electric power big data analysis and intelligent operation and maintenance of a distribution network. According to the invention, a fusion architecture based on a multi-scale time convolution network and a long short-term memory network is constructed; extracting multi-scale spatio-temporal characteristics of a user from instantaneous electricity utilization abrupt change to a periodic load rule through an MSTBlock unit; designing a cluster balance constraint mechanism to ensure that rare and key non-technical line loss abnormal early warning signals are not covered by mass normal power utilization data; according to the data scale, adaptively selecting a graph segmentation or spectral clustering integration strategy to output a clustering label, and mapping the clustering label into a user power consumption behavior evolution track; according to the method, the power utilization abnormal level can be identified from the original load signal with random fluctuation interference, and the troubleshooting priority is calculated in combination with the transformer area correlation analysis, so that the accuracy and interpretability of the non-technical line loss unsupervised evaluation decision of the power distribution network are remarkably improved.
Owner:SHENYANG UNIVERSITY OF TECHNOLOGY

Personalized teaching course recommendation method and system based on artificial intelligence

The invention discloses a personalized teaching course recommendation method and system based on artificial intelligence, and relates to the technical field of artificial intelligence and education recommendation, and the method comprises the steps: collecting text data for preprocessing, and extracting a standard target set, a teaching target candidate set and a knowledge point candidate set; calculating cosine similarity and weight based on the standard target set and the teaching target candidate set, calculating knowledge point mastery degree and weight based on the knowledge point candidate set, and splicing the knowledge point mastery degree and weight to generate a learning feature vector; based on the learning feature vector, clustering is carried out by using a k-means + + algorithm, a clustering label and a clustering center set are output, after the clustering center is updated in combination with a Thompson Sampling algorithm and Monte Carlo, the posterior probability is recalculated, and a final recommendation result is output; the robustness and recommendation accuracy of personalized teaching course recommendation are effectively improved.
Owner:SHIHEZI UNIVERSITY

K-plane clustering roof segmentation method based on multiple geometric consistency constraints

The invention provides a k-plane clustering roof segmentation method based on multiple geometric consistency constraints. The k-plane clustering roof segmentation method comprises the steps of performing initialization preprocessing on input original building roof point cloud data; based on each point in the original building roof point cloud data and a candidate fitting plane cluster, establishing a clustering label optimization objective function fused with triple geometric constraints; iteratively updating the clustering label and the candidate fitting plane parameter of each point by adopting an alternating minimization optimization strategy, and solving a mixed integer non-convex optimization problem until the clustering label optimization objective function is converged; and based on the converged clustering tag optimization objective function, outputting a building roof point cloud instance segmentation result, the building roof point cloud instance segmentation result comprising the clustering tag of each point and the corresponding candidate fitting plane geometric parameters. According to the method, the segmentation precision and the calculation efficiency of the building roof point cloud instance are improved.
Owner:EAST CHINA NORMAL UNIV

Eye movement trajectory analysis method and system based on hybrid clustering and time constraint

The invention provides an eye movement trajectory analysis method and system based on hybrid clustering and time constraint, and belongs to the technical field of computer vision. Calculating a time difference and a moving speed between continuous original eye movement data points; comparing the moving speed with a speed threshold value, and classifying the moving speed into candidate fixation points and glancing points; extracting spatial features and time features of the candidate fixation points, performing standardization processing, and performing weighted fusion to obtain spatial-temporal feature vectors; clustering the spatio-temporal feature vectors to obtain preliminary clustering labels of the candidate fixation points; performing time constraint processing on each cluster; calculating the duration of the clustering cluster after the time constraint processing, and generating a final clustering label of the candidate fixation point; and outputting a final classification label of each original eye movement data point in combination with the classification result of the glancing points.
Owner:NAVAL AVIATION UNIV

Enterprise credit evaluation and analysis method

The invention discloses an enterprise credit evaluation analysis method, and relates to the field of data processing, and the method comprises the steps: obtaining the credit data of a plurality of to-be-evaluated enterprises; performing feature extraction on the credit investigation data through an improved stack type self-encoding neural network model to obtain feature data; the improved stack type self-encoding neural network model comprises an encoder, a decoder, an attention layer and a prior rule layer; the attention layer adopts a Scaled Dot-Product Attention mechanism to learn a weight feature matrix Ha, and the priori rule layer sets feature constraints according to a priori rule; clustering the feature data to obtain a cluster to which the enterprise belongs; according to the feature value distribution of the enterprises in the clustering clusters, performing feature scoring by using a prior rule to obtain clustering labels of the clustering clusters; according to the clustering label and the weight feature matrix Ha, generating an enterprise credit investigation portrait, and according to the enterprise credit investigation portrait, carrying out credit investigation evaluation analysis; aiming at the low enterprise credit evaluation precision in the prior art, the reliability of the evaluation result is improved.
Owner:BEIJING YONGFENG AGRICULTURAL PORT SUPPLY CHAIN MANAGEMENT DEVELOPMENT CO LTD +1

LLM-based scientific literature topic discovery method and device

PendingCN120780837ASemantic analysisInference methodsPattern recognitionDocument representation
The invention discloses a scientific literature topic discovery method and device based on LLM. The method comprises the following steps: 1) acquiring text representation of each scientific literature sample and encoding the text representation by using a text encoder to obtain a document representation matrix corresponding to the scientific literature sample; 2) clustering the scientific literature samples to obtain clustering results of different themes; calculating an entropy value of each scientific literature sample, and selecting a high-uncertainty sample; 3) calculating semantic similarity between each high-uncertainty sample and other scientific literature samples, and constructing a plurality of triple tasks; fine-tuning the text encoder by using each triple task through a comparative learning method; 4) encoding the text representation of each scientific literature sample by using a text encoder to obtain a document representation matrix corresponding to the scientific literature sample; and 5) performing topic clustering on each scientific literature sample by using the document representation matrix of each scientific literature sample, and generating a clustering tag and a topic division result of each scientific literature sample.
Owner:COMP NETWORK INFORMATION CENT CHINESE ACADEMY OF SCI

Large-scale unstructured data joint processing method and system

The invention discloses a large-scale unstructured data joint processing method and system. The method comprises the steps that S1, large-scale unstructured data is collected, and a metadata feature library containing dynamic updating and index retrieval optimization is constructed; s2, constructing a federated index based on the dynamic metadata feature library, constructing a cross-domain joint index on the basis of the federated index, and performing layered federated architecture optimization to obtain a cross-domain joint optimization index; s3, constructing a global correlation matrix based on the cross-domain joint optimization index to perform cross-domain correlation modeling, and obtaining a global correlation function through cross-domain correlation modeling; and S4, carrying out joint optimization solution based on the global correlation function to realize unstructured data joint processing, uniformly optimizing tasks such as clustering, label prediction and correlation rule mining by fusing a processing target function, and enabling the tasks to cooperate with each other, so that the whole data processing process is more intelligent and efficient, and the accuracy of a data processing result is improved.
Owner:SUZHOU HEINQI INFORMATION TECH CO LTD

A method and device for short-term early warning of rock failure based on acoustic emission clustering analysis

This invention provides a method and device for short-term early warning of rock failure based on acoustic emission clustering analysis, relating to the fields of rock mechanics and geotechnical engineering. The method includes: acquiring acoustic emission characteristic parameters during rock deformation and failure; performing clustering analysis on the acoustic emission characteristic parameters using the K-means++ algorithm to obtain cluster labels corresponding to the acoustic emission characteristic parameters; calculating the importance scores of the acoustic emission characteristic parameters based on the cluster labels using the random forest algorithm; constructing an early warning index set based on the acoustic emission characteristic parameters; constructing an initial CNN-LSTM model; establishing a sample dataset of conventional and precursor signals of rock failure based on the early warning index set; training the initial CNN-LSTM model using the sample dataset to obtain a trained CNN-LSTM model; acquiring real-time acoustic emission signals; and inputting the real-time acoustic emission signals into the trained CNN-LSTM model to provide early warning by identifying the acoustic emission signal category. This invention can provide early warning for complex rock failure processes.
Owner:JIANGXI UNIV OF SCI & TECH

A method and apparatus for customizing a cryptographic protocol cluster

The embodiment of the application discloses a self-defined encryption protocol clustering method and device, the method comprises the following steps: obtaining a target encryption protocol to be identified, and extracting traffic data and metadata of the target encryption protocol; constructing a label matrix and a value matrix based on the traffic data; and fusing the label matrix and the value matrix by using a multi-mode matching algorithm to obtain a clustering label. In this way, the self-defined encryption protocol clustering method provided by the application combines double-matrix combination with a pattern matching algorithm to realize clustering of self-defined encryption protocols. The method solves the problem that mixed self-defined protocols cannot be quickly and accurately clustered, and provides support for subsequent protocol analysis by quickly and accurately realizing clustering of self-defined protocols.
Owner:VIEWINTECH

Multi-view clustering method and device, equipment, storage medium and program product

The embodiment of the invention provides a multi-view clustering method and device, equipment, a storage medium and a program product, and relates to the field of financial science and technology. Obtaining a view data set corresponding to each view; based on the similarity between the feature elements in each view data set, constructing a similar matrix of each view data set; multiplying the similar matrixes of the plurality of view data sets to obtain a consistency matrix, and reconstructing the similar matrix of each view data set based on the consistency matrix to obtain a reconstructed similar matrix; and performing multi-view clustering based on the plurality of reconstructed similar matrixes to obtain a clustering result, the clustering result being used for indicating a clustering label corresponding to each sample. Through the above mode, the complementary information of the consistency matrix is utilized to maintain the stability of the clustering boundary, the misjudgment and missed judgment caused by the missing of the sample information in the view are reduced, and the accuracy of the multi-view clustering result is improved.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

New energy unit icing shutdown prediction method and system based on multi-feature interaction threshold

The invention discloses a new energy unit icing shutdown prediction method and system based on a multi-feature interaction threshold. The method comprises the steps of collecting and preprocessing multi-dimensional data of a new energy unit; based on meteorological and geographic features, generating a corresponding clustering label for each new energy unit by using a clustering algorithm; constructing a full-connection deep neural network shutdown prediction model, taking the preprocessed data and the clustering labels as input features, and training the icing shutdown probability of a model output unit; for a single feature, a feature fixing strategy is adopted, and a single feature threshold interval is determined; key features are selected for double-feature interaction analysis, a three-dimensional decision boundary is constructed, an interaction effect is quantified, and a multi-feature interaction rule is extracted based on a decision tree algorithm; and constructing a comprehensive discrimination rule, setting risk preference parameters, carrying out adaptive threshold updating, and finally outputting a shutdown prediction result. The method can significantly improve the shutdown prediction precision, and is suitable for different types of new energy equipment such as wind power and photovoltaic equipment.
Owner:STATE GRID HENAN ELECTRIC POWER ELECTRIC POWER SCI RES INST +1

Calculus risk prediction method based on machine learning and electronic equipment

The invention belongs to the technical field of calculus risk assessment, and provides a calculus risk prediction method based on machine learning and electronic equipment, and the method comprises the steps: to-be-detected original data collection and calculus risk prediction; the construction process of the stone risk prediction model comprises the steps of unmarked data collection, image feature extraction, index feature extraction, text semantic extraction, feature fusion, feature clustering analysis, iterative training, fine adjustment small sample data collection, meta-learning fine adjustment and cluster marking. According to the method, through a semi-supervised learning strategy of unmarked data clustering and small sample adjustment, the demand of marked data is greatly reduced; according to the method, clustering and unsupervised iteration are fused through DBSCAN and DPC double algorithms, so that pre-training under unmarked data is realized, and the reliability of unmarked training is improved; through combination of multi-modal data fusion, clustering iteration and meta-learning fine tuning, the robustness of the model is improved.
Owner:SHENZHEN LUOHU PEOPLELS HOSPITAL

Microservice architecture root cause positioning method based on heterogeneous graph modeling and active learning

The invention relates to a micro-service architecture root cause positioning method based on heterogeneous graph modeling and active learning, and the method comprises the following steps: 1, carrying out the feature extraction and topological structure of indexes, logs, call chains and deployment information, and forming unified heterogeneous graph modeling; and step 2, clustering, label diffusion and boundary sample selection are carried out on the heterogeneous graph model constructed in the step 1, semi-supervised training driven by active learning is used to continuously optimize a root cause positioning model, and the model is deployed in a production environment online to realize real-time fault detection and positioning. According to the method, the structural characteristics of the micro-service system can be fully utilized, and the root cause positioning method with low labeling requirements is provided, so that the balance between the task performance and the labeling overhead is realized, and a better solution is provided for efficient root cause positioning of the micro-service system.
Owner:STATE GRID TIANJIN ELECTRIC POWER COMPANY +1

Customer consumption cost prediction method and device

The invention discloses a customer consumption cost prediction method and device. The method comprises the following steps: acquiring first consumption data of a plurality of users in a first time period; performing statistical analysis on the first consumption data to obtain a multi-dimensional consumption data feature and a consumption data time sequence corresponding to each user; clustering the plurality of users based on the consumption data features to obtain a clustering label corresponding to each user, the clustering labels being used for reflecting consumption habits of the users; and for each user, determining an expense prediction model matched with the clustering tag of the user, and analyzing the consumption data time sequence of the user by using the expense prediction model to obtain second consumption data of the user in the second time period. The technical problem that a traditional consumption prediction scheme is difficult to accurately analyze and predict multi-dimensional consumption data of different users is solved.
Owner:CHINA TELECOM CORP LTD

Fast multi-view clustering method, system, storage medium and computer device based on logarithmic sparsity constrained anchor graph decomposition

The present invention discloses a fast multi-view clustering method, system, storage medium, and computer device based on logarithmic sparsity-constrained anchor graph decomposition. The method, which belongs to the field of machine learning and data analysis, mainly comprises the following steps: reading a vectorized multi-view dataset and generating anchor points in all views; constructing an anchor graph between the original data and the anchor point set formed by the anchor points; performing anchor graph decomposition on the anchor graph to obtain a basis matrix and a coefficient matrix, and applying log-norm constraints to the basis matrix to remove redundant information; and performing K-means on the coefficient matrix to obtain cluster labels and generate clustering results. The present invention can significantly improve the accuracy and efficiency of clustering when processing large-scale multi-view clustering tasks and can be widely used in the field of data analysis.
Owner:XI AN JIAOTONG UNIV

Skeletal muscle spasm-to-contracture evolution rule quantitative analysis method and system and application

The invention belongs to the technical field of medical data mining and rehabilitation evaluation, and provides a skeletal muscle spasm-to-contracture evolution rule quantitative analysis method and system and application, and the method comprises the steps: preprocessing spasm contracture time sequence data, and obtaining one-dimensional feature vector data; k-means clustering is carried out, state labeling of spasm and contracture is carried out, and clustering labels are obtained; determining the change rate of the torque characteristic and the adjacent angular velocity characteristic, and generating an eight-bit pseudo-sequential sequence representing the evolution rule from the spasm state to the contracture state in combination with a state transition threshold value determined by an ROC curve; constructing a Markov chain model to calculate a transition probability matrix from a spasm state to a contracture state, and identifying torque as a key driving feature of state evolution through grouping risk ratio; and constructing a dynamic correlation model based on the key driving features to complete quantitative analysis of the evolution rule from skeletal muscle spasm to contracture. According to the method, objective division from skeletal muscle spasm to contracture state can be realized, and the transition probability and key driving factors between the skeletal muscle spasm and the contracture state can be excavated.
Owner:JILIN UNIV FIRST HOSPITAL

Incremental multi-view data clustering method and system based on cross-time consensus graph

PendingCN121456530AInformaticsConsensus
The embodiment of the invention provides an incremental multi-view data clustering method and system based on a cross-time consensus graph, and belongs to the technical field of artificial intelligence. The method comprises the following steps: integrating historical knowledge with a consensus affinity matrix of current view information and learning time based on kernel induction expression, and constructing a dynamic consensus graph; spectrum embedding and discrete label learning are carried out; and alternately optimizing the consensus affinity matrix, the orthogonal rotation matrix, the consensus spectrum embedding matrix and the discrete clustering label matrix of the moments in the dynamic consensus graph by adopting staged updating variables. The method can efficiently adapt to incremental environment application, and is better in clustering precision, higher in calculation efficiency, higher in time sequence stability and better in robustness.
Owner:ANHUI NORMAL UNIV

Multifunction radar signal sorting method and system based on hypergraph

The application discloses a multifunctional radar signal sorting method and system based on a hypergraph, and the method steps are as follows: S1, for RIPS obtained through radar reconnaissance receiver interception, interference pulses in the RIPS are removed; S2, the remaining pulse sequence is subjected to screening rules and FCM clustering, and clustering labels of partial pulse sequences are obtained; S3, feature parameters and data potential energy values of the remaining pulse sequence are used to construct a hypergraph; and S4, the clustering labels of the partial pulse sequences are used to learn the constructed hypergraph, and a final radar signal sorting result is obtained. The application can sort multifunctional radar signals without a marked sample, can effectively alleviate the 'batch increase' problem, and has good sorting performance.
Owner:HANGZHOU DIANZI UNIV +1

A semi-supervised clustering method and its open-ended question answering text encoding method

This invention relates to the field of data representation technology, and discloses a semi-supervised clustering method and its open-ended question answer text encoding method. The semi-supervised clustering method includes: acquiring a dataset to be clustered, its labeled dataset, and its unlabeled dataset; mapping the dataset to be clustered to a spatial density map and / or a topological density map; and using the labeled and unlabeled datasets to cluster the data in the dataset to be clustered into several clusters, wherein each data point in each cluster has a cluster label, which can be an existing label or a new label. This invention uses a semi-supervised clustering method to efficiently and accurately cluster open-ended question answer text data, and can discover new classes with limited prior knowledge. After clustering, keywords can be extracted and encoded from the open-ended question answer text of each class, facilitating a quick and accurate understanding of the situation of the interviewed group and improving the efficiency and quality of diagnosis and treatment.
Owner:RENMIN UNIVERSITY OF CHINA

Unknown protocol clustering method and system based on frequent item extraction and bi-layer autoencoder

The present application relates to the technical field of information security, in particular to a kind of unknown protocol clustering method and system based on frequent item extraction and double-layer self-encoder, through the preprocessing process, the original data is converted from bit form to byte form, and then the first 32 bytes are intercepted, then the frequency of frequent item and the number of frequent item of each byte are obtained by carrying out frequent item statistics to the preprocessed data byte by byte;Frequency self-encoder and quantity self-encoder are used to extract features from the byte with larger frequency of frequent item and larger number of frequent item respectively;The features extracted by frequency self-encoder are coarsely clustered to obtain coarse clustering label, and the coarse clustering label and the features extracted by quantity self-encoder are merged and clustered to obtain the final fine clustering label.The present application can not only retain protocol-level clustering, but also realize class-level clustering, with small amount of calculation, while ensuring the real-time of clustering, it can effectively solve the under-partition problem of unknown protocol clustering and improve the performance of clustering.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

Model training method, system, equipment and medium

The invention provides a model training method, system and device and a medium, and belongs to the technical field of feature extraction, and the method comprises the following steps: collecting training statistical information and a feature clustering label of a client, and dividing the client into a main client and an auxiliary client according to the training statistical information and the feature clustering label of the client, the training statistical information is independently trained and determined by a local model which is distributed to each client by the server, and the feature clustering tag is used for extracting and clustering local data features of the client in advance; receiving the local model of the main body client and performing average aggregation to obtain a main body model; selecting the auxiliary client according to the training statistical information of the auxiliary client, and performing dynamic weighted aggregation on the local model of the selected auxiliary client to obtain an auxiliary model; and updating parameters of the main body model by using the auxiliary model to obtain a trained global model. The convergence speed of model training can be increased.
Owner:CHINA UNIV OF MINING & TECH

Drug recommendation method and device, electronic equipment and storage medium

The invention relates to the technical field of clinical auxiliary diagnosis, and provides a drug recommendation method and device, electronic equipment and a storage medium, and the method comprises the steps: carrying out the matching from a clustering result, obtaining a target class cluster to which detection data belongs, and taking a clustering tag of the dynamic change of the target class cluster as reconstructed detection data; based on the semantic similarity between the diagnosis description text and the standard diagnosis description text, performing text correction on the diagnosis description text to obtain a corrected diagnosis description; and performing drug recommendation based on the reconstructed detection data and the corrected diagnosis description. According to the method provided by the invention, the target class cluster to which the detection data belongs is obtained through matching from the clustering result, the detection data is abstracted based on the dynamically-changed clustering label, the reconstructed detection data which is high in interpretability, visual and accurate is obtained, and the diagnosis description text which is redundant and difficult to understand is subjected to text correction, so that the diagnosis accuracy is improved. Uniform expression of the standard diagnosis text is realized, the expression specification of the health record data is improved, and the accuracy of drug recommendation is further improved.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI

Intelligent wearable human body action recognition method based on multi-stage feature optimization clustering

ActiveCN121301854ANeural learning methodsSingular value decompositionLaplacian spectrum
The invention relates to an intelligent wearable human body action recognition method based on multi-stage feature optimization clustering, and belongs to the technical field of human body action recognition of intelligent wearable equipment. The method comprises the following steps: segmenting a multivariate time sequence of the intelligent wearable equipment by adopting a sliding window; discrete cosine transform is executed, low-frequency components of a plurality of previous proportions are reserved, singular value decomposition is introduced to remove redundant information across variables, low-rank representation is obtained and input to a depth embedding modeling module, and unified embedding representation is obtained; inputting the unified embedded representation into an improved dynamic movement-splitting-merging distance measurement model to obtain a dynamic movement-splitting-merging distance; a clustering label is obtained through construction of a similarity matrix, graph Laplacian spectrum decomposition and clustering operation in sequence; and human body action recognition is realized based on the clustering labels. The objective of the invention is to solve the technical problems of feature redundancy, insufficient time sequence dependence modeling and inaccurate similarity measurement of an existing action recognition method.
Owner:KUNMING UNIV OF SCI & TECH