Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

70 results about "Text cluster" patented technology

Text clustering method, device, electronic device and computer-readable storage medium

The present application provides a text clustering method, apparatus, electronic device and computer-readable storage medium. Based on anchor words corresponding to the full-text semantics of each training text, anchor word model features corresponding to the anchor words are obtained to avoid the introduction of additional noisy features. Then, a first clustering result and a second clustering result of each training text are obtained based on the anchor word model features, and a self-training objective function of multiple training texts is determined based on each first clustering result and each second clustering result, as well as a self-training target value of the self-training objective function. Finally, the text clustering model is updated based on the self-training target value until the text clustering model converges, and the converged text clustering model is applied to text clustering. After continuous training until the convergence of the text clustering model, the accuracy and stability of the text clustering model are continuously improved, and the accuracy of text clustering is avoided from being affected by additional noisy features.
Owner:JILIN UNIVERSITY

Text processing method, apparatus, device, storage medium, and program product

Embodiments of the present application provide a text processing method, device, equipment, storage medium and program product. The method comprises: obtaining a target text; processing the target text based on a vector dictionary to obtain a target vector; and processing the target vector based on a target task model to obtain a text processing result. In the embodiments of the present application, the target text is processed by the pre-constructed vector dictionary to obtain the target vector, and then the target vector is processed by the target task model. Since the target vector includes a semantic feature vector and a local difference vector, the accuracy of text processing can be improved. Moreover, the target task model is any one of a text classification model, a text clustering model and a sentiment analysis model in the call center customer service field. Compared with the use of a general large model, the efficiency of text processing can be improved.
Owner:BEIJING HOLLYCRM TECH

Methods, apparatus, devices, and storage media for aggregating cross-page test questions

This application relates to the field of computer technology and discloses a method, apparatus, device, and storage medium for cross-page test question aggregation. The method includes acquiring multiple test question images, each image including test question elements; extracting text blocks from the test question images and grouping text blocks belonging to the same test question and the same test question element into a text cluster, obtaining a text cluster set; searching for a baseline text cluster and multiple candidate text clusters in the text cluster set, where candidate text clusters are those related to the baseline text cluster; generating label codes for each text cluster in the baseline and candidate text clusters based on the test question element to which the text cluster belongs, the text blocks within the text cluster, and the position of the text blocks in the test question images; and searching for text clusters belonging to the same test question as the baseline text cluster in the candidate text clusters based on the label codes, obtaining a test question text cluster set. The method of this application can achieve automatic clustering of test question elements with high clustering accuracy.
Owner:HANGZHOU ZHIJUAN PLANET TECHNOLOGY CO LTD

A keyword reverse propagation algorithm based on citation network structure

The application provides a keyword reverse propagation algorithm based on a citation network structure, and comprises the following steps: a spring charge model is established, force-directed layout processing is performed, and a force-directed layout graph is established; a keyword propagation model is established by using a reverse propagation algorithm, and a keyword weight change contrast curve is obtained; a citation network model is constructed by using the force-directed layout graph; iterative calculation is performed on the force-directed layout graph until the energy state in the force-directed layout graph reaches a minimum value; while the citation network is iteratively calculated, the keyword weight in the citation network model is adjusted, and a converged citation network layout graph is calculated. The application has the beneficial effects that the entire network is taken as a main body, keywords owned by cited documents are selected at a certain probability and are propagated backward along the network to the citing documents, the weight of the keywords is increased while the text clustering idea is retained, and clear visual data effects are obtained through force-directed layout.
Owner:UNICLOUD TECH CO LTD

A text clustering method based on improved whale optimization algorithm

The application relates to a text clustering method based on an improved whale optimization algorithm, and belongs to the fields of big data mining and machine learning. The method comprises the following steps: S1: obtaining scientific research text data from a scientific research text database by using Spark, and performing data cleaning; S2: performing word segmentation, stop word removal, feature selection, vectorization and dimension reduction operations on the text data after data cleaning, so that unstructured text data is converted into structured numerical data; and S3: calculating initial clustering centers by using a K-means algorithm, optimizing the clustering centers by using an improved whale optimization algorithm, and outputting clustering results. The global optimization capability of the improved whale optimization algorithm is utilized, the problem that current text clustering algorithms are prone to falling into local optimization is solved, the accuracy of clustering of scientific research text data is improved, the data dimension is reduced, and the clustering effect is enhanced.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Text classification method, electronic equipment, storage medium and program product

The embodiment of the invention provides a text classification method, electronic equipment, a storage medium and a program product. The method comprises the following steps: processing a to-be-classified text according to a lexical element division rule to obtain a lexical element sequence; according to the lexical element sequence, adopting a pre-trained unsupervised topic model to obtain a topic distribution data set corresponding to the to-be-classified text; adopting a pre-trained unsupervised clustering model to obtain a clustering distribution data set corresponding to the to-be-classified text; splicing the topic distribution data set and the clustering distribution data set to obtain a text hidden topic; obtaining a plurality of text clusters and current cluster hidden topics corresponding to the text clusters, and calculating the similarity between the text hidden topics and the current cluster hidden topics; and determining a target text cluster from the plurality of text clusters, and adding the to-be-classified text to the target text cluster. Through a cluster classification mechanism of subject distribution and cluster distribution conjoint analysis, the accuracy and result stability of text classification are improved.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD +1

A cross-modal retrieval method based on prototype regularization learning

The application provides a cross-modal retrieval method based on prototype regularization learning, and belongs to the field of artificial intelligence and multi-modal information processing, and comprises the following steps: extracting visual features and text features through a visual encoder and a text encoder and calculating a basic contrast loss; dividing visual clusters and text clusters through a clustering algorithm and calculating visual prototypes and text prototypes, selecting text features most similar to the visual prototypes as anchor points, and constructing cross-modal prototypes; calculating a prototype-level discriminant loss according to the visual prototypes and the text prototypes, calculating an instance-level discriminant loss according to the visual clusters and the text clusters, and calculating a prototype projection loss according to the cross-modal semantic prototypes; constructing a joint optimization objective function by combining all the losses, training the visual encoder and the text encoder, and performing cross-modal retrieval after the training is completed. The application solves the problems of excessive intra-class variance and excessively high inter-class similarity caused by appearance bias in the existing cross-modal retrieval method, and causes the problems of insufficient retrieval accuracy and robustness.
Owner:SHANGHAI EYE DISEASE PREVENTION & TREATMENT CENTER

Text clustering method, text clustering device, electronic device, and storage medium

The application provides a text clustering method, a text clustering device, an electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence. An initial clustering center is selected from a plurality of original sentence embedding vectors, an initial problem text class is constructed based on the initial clustering center, clustering processing is performed on all original sentence embedding vectors based on the initial clustering center, a plurality of intermediate problem text classes are obtained, a clustering center is identified for each intermediate problem text class, a text clustering center is obtained, inter-class dispersion between the intermediate problem text classes and intra-class dispersion of each intermediate problem text class are calculated according to the text clustering center, the intermediate problem text classes are adjusted in text according to the inter-class dispersion and the intra-class dispersion, a target problem text class is obtained, semantic analysis is performed on all original sentence embedding vectors of the target problem text class, a standard problem text is obtained, and standard reply content is determined according to the standard problem text, so that the accuracy of problem answering can be improved.
Owner:CHINA PING AN LIFE INSURANCE CO LTD

Text clustering with heuristic and multi-metric control

Implementations generally relate to text clustering with heuristic and multi-metric control. In some implementations, a method includes receiving an electronic source document containing text. The method further includes dividing the text into text units, encoding the text units, and transforming the text units into numerical values. The method further includes generating a graph of the text units based on the numeric values, where the graph includes nodes corresponding to the text units and edges corresponding to pairs of the text units. The method further includes ordering the text units into text clusters based on the graph of the text units. The method further includes generating an electronic target document that presents the text clusters based on one or more preference heuristics.
Owner:JPMORGAN CHASE BANK NA

Text clustering processing method, text classification method and related device

The invention provides a text clustering processing method, a text classification method and a related device. The method comprises the following steps: clustering first texts in a first text set into a plurality of text clusters; performing text coding on the first text in each text cluster to obtain a coding result of the text cluster; determining intra-cluster compactness of the text clusters based on the coding result of each text cluster; based on the coding result of the first text cluster and the coding result of the second text cluster, determining an inter-cluster separation degree between the first text cluster and the second text cluster; the first text cluster and the second text cluster are any two text clusters in the plurality of text clusters; and determining a first clustering index value of the first text set based on the intra-cluster compactness and the inter-cluster separation degree. Through the method and the device, the accuracy of text clustering processing can be improved.
Owner:MASHANG CONSUMER FINANCE CO LTD

Cross-border ethnic text clustering method and device integrating domain knowledge graph

The present invention relates to a cross-border ethnic text clustering method and device that integrates domain knowledge graphs, and belongs to the field of natural language processing technology. Cross-border ethnic domain texts have their own unique cultural domain vocabulary. In order to solve the problem that the cultural background of cross-border ethnic cultural texts is missing and the cultural differences between different ethnic groups are large, resulting in poor text clustering effect, the present invention proposes a cross-border ethnic cultural text clustering method that integrates domain knowledge graphs, which mainly includes a cross-border ethnic cultural text data preprocessing part, a cross-border ethnic text data feature extraction part, and a cross-border ethnic cultural text clustering structure. According to the functional modularization of these three parts, a cross-border ethnic cultural text clustering device that integrates domain knowledge graphs is made. Compared with the general text clustering method, the present invention effectively alleviates the problem of poor cross-border ethnic cultural text clustering effect.
Owner:KUNMING UNIV OF SCI & TECH

A clustering-based text supervision checking method, device and computer equipment

The application relates to a clustering-based text supervision checking method and device and a computer device. Feature word sets corresponding to a plurality of question description texts in a question list are obtained by respectively extracting feature words from the question description texts. The word frequency of each feature word in the question description text, the first inverse document frequency in all question description texts and the second inverse document frequency in all checked units are respectively calculated to obtain a feature word weight set of the question description text of each checked unit. A plurality of question description text clusters and a plurality of initial high-correlation-degree feature word sets corresponding to the question description text clusters are obtained through clustering. Finally, the initial high-correlation-degree feature word set is optimized according to a chi-square statistic to obtain an accurate high-correlation-degree feature word set. The method saves the calculation resources while ensuring the accuracy of the results by optimizing the feature word weight calculation, adopting a clustering algorithm for preliminary clustering and optimizing the clustering results through chi-square statistics.
Owner:NAT UNIV OF DEFENSE TECH

A cross-domain technology fusion opportunity finding method based on science and technology fusion interaction

The application belongs to the technical field of data mining, and specifically discloses a cross-field technology fusion opportunity finding method based on science and technology fusion interaction, which comprises the following steps: obtaining scientific paper and technology patent data, identifying cross-field scientific fusion and technology fusion based on semantic anchor points and a similarity threshold; using a trained deep text clustering model to respectively perform deep clustering on scientific fusion and technology fusion, and obtaining corresponding clustering sets; obtaining text representation corresponding to each cluster, filtering through a threshold, and establishing association between two types of clusters through bipartite graph maximum weight matching to identify shared topics and exclusive topics; using a dynamic time warping method to align the time series of scientific and technological trends, judging topics with scientific convergence preceding technological convergence, and identifying technology fusion opportunities with growth potential based on the development trend of scientific papers. The application can deeply mine cross-field science and technology semantics, improve clustering accuracy, quantify science and technology association, and dynamically find forward-looking technology fusion opportunities.
Owner:SOUTHWEST JIAOTONG UNIV

Efficient and comprehensible alarm text clustering

Methods for management of alarm text data for received alarms, which can involve processing the alarm text data for clustering to generate processed alarm text data; clustering the processed alarm text data into first clusters through executing a first clustering algorithm, the first clustering algorithm configured to be facilitated via accelerated processing; for the first clusters not meeting a metric, executing a second clustering algorithm on the first clusters to form second clusters that are representative of subgroups of the first clusters with improved homogeneity, the second clustering algorithm being different from the first clustering algorithm; executing a tagging process on the second clusters and ones of the first clusters that meet the metric to tag the second clusters and the ones of the first clusters that meet the metric into groups to form tagged cluster groups; and providing the tagged cluster groups to be displayed on the user interface.
Owner:HITACHI ENERGY GERMANY AG

Text clustering method and device

The embodiment of the invention provides a text clustering method and device, and the method comprises the steps: obtaining a plurality of to-be-clustered texts, and calculating the text similarity between the to-be-clustered texts; determining a noise text in the plurality of to-be-clustered texts according to the text similarity, and deleting the noise text in the plurality of to-be-clustered texts to obtain a to-be-clustered text set; based on to-be-clustered texts in the to-be-clustered text set and the text similarity between the to-be-clustered texts, constructing a target text connected graph corresponding to the to-be-clustered text set; and determining a central text corresponding to the target text connected graph, deleting an edge text in the target text connected graph based on the central text, and determining a target text set according to a deletion result.
Owner:UC MOBILE CHINA CO LTD

Event text hot word extraction method based on similarity adjacent matrix clustering

The invention relates to the field of natural language processing, in particular to an event text hot word extraction method based on similarity adjacent matrix clustering, which comprises the following steps: acquiring a keyword set of each event text; calculating an improved Jaccard similarity value of any two event texts, and constructing a text similarity adjacent matrix; converting the text similarity adjacent matrix into a distance matrix, and dividing all texts into K text clusters based on the distance matrix; removing the K text clusters with abnormal text quantity by using a quartile method to obtain normal text clusters; word weight calculation is conducted on each text in the normal text cluster again to obtain a normal keyword set, the final comprehensive score of each keyword is calculated based on the normal keyword set, candidate keywords in the cluster are reordered according to the final comprehensive score, and the candidate keywords ranked in the front serve as final hot words. The hot word extraction accuracy and efficiency can be improved.
Owner:DIGITAL CHONGQING BIG DATA APPL DEV CO LTD

Text clustering method, apparatus, and electronic device

The application discloses a text clustering method and device and electronic equipment. It is related to the field of financial technology or other fields, and the method comprises the following steps: obtaining a plurality of digital vectors of a to-be-processed text, wherein each digital vector corresponds to part of the to-be-processed text; determining a first distance threshold and a second distance threshold based on the plurality of digital vectors, wherein the first distance threshold is the maximum limit value of the clustering range, and the second distance threshold is the minimum limit value of the clustering range; performing first clustering processing on the plurality of digital vectors based on the first distance threshold and the second distance threshold to obtain a clustering result; obtaining the number of clusters in the clustering result; performing second clustering processing on the plurality of digital vectors based on the number of clusters to obtain a target centroid vector corresponding to each cluster, wherein the target centroid vector represents the characteristics of the cluster corresponding to the target centroid vector. The application solves the technical problem of poor text clustering effect caused by the fact that the number of clusters cannot be accurately determined in the prior art.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

A method and system for product title clustering

The application discloses a kind of method and system for commodity title clustering, specifically related to text clustering technical field, for solving the technical scheme of general data is generally aimed at existing disclosure technology, there is also scheme is aimed at such data, but due to different application scenarios, such scheme is not completely applicable problem, its method includes data crawling, commodity title normalization, semantic vector conversion, similarity analysis, clustering analysis and search recommendation of similar goods, system is composed of hardware and software, software includes crawler module, processing module, semantic vector module, similarity calculation module, clustering module and recommendation module, hardware includes CPU, memory bank, storage and GPU;It is through the competitive product analysis of online commodity transaction information data, obtains the similarity of commodity title and carries out clustering analysis and search recommendation to the similarity of commodity title, to improve the accuracy of clustering and search recommendation.
Owner:ZHONGKE (XIAMEN) DATA INTELLIGENCE RES INST

Text clustering method, text clustering device and text clustering system

This application provides a text clustering method, a text clustering device, and a text clustering system. In this solution, the Word2vec model is combined with the TF-IDF algorithm to represent word vectors, enhancing the distinction between different texts. This approach not only takes advantage of word vectors but also incorporates the influence of words on text. The word vectors represented by the combination of the two serve as the input of the WMD algorithm, which is used as a similarity measurement algorithm in text clustering to improve the accuracy of text clustering.
Owner:中国邮政储蓄银行股份有限公司

Short text clustering method based on adaptive optimal transmission and three-level robust representation

The invention provides a short text clustering method based on adaptive optimal transmission and three-level robust representation, which comprises the following steps of: 1, coding an input original short text based on a pre-training language model in a server carrying a GPU (Graphics Processing Unit) to obtain a semantic representation of the original short text, and performing multi-type data enhancement processing on the semantic representation to obtain a semantic representation of the original short text; obtaining an enhanced sample representation; 2, performing pseudo label distribution on enhanced sample representation based on an optimal transmission algorithm; 3, learning and calculating a corresponding loss function based on the pseudo label and three-level robust representation, and updating model parameters of the loss function in a CPU (Central Processing Unit) of a server; 4, iteratively executing the steps 1-3, and outputting a target clustering result when a clustering result meets a convergence condition; the problems of data sparseness and class imbalance in short text clustering can be effectively relieved, and the stability and precision of a clustering result are improved.
Owner:YANCHENG INST OF TECH +1

Machining drawing processing information extraction method based on YOLO + OCR

The invention relates to the technical field of computer vision and intelligent information processing, in particular to a machining drawing information extraction method based on YOLO + OCR, which comprises the following steps: sequentially carrying out graying processing, Gaussian filtering noise reduction and Canny edge detection on a machining drawing, and extracting and screening drawing frames and drawing areas; detecting a text group in the drawing by adopting a YOLO-OBB algorithm; and performing discrete classification prediction on the rotation angle of the detected text group image. According to the method, text pixel loss caused by direct zooming is avoided through drawing preprocessing, the 360-degree text direction problem is solved by utilizing a YOLO-OBB algorithm and a four-classification neural network model, PMI information is aligned and recognized in combination with a special symbol model, global parameters are extracted through natural language processing, and finally errors are monitored and corrected. The problems of text pixel loss, insufficient direction processing, poor symbol and text association and the like in the prior art are solved.
Owner:BEIHANG UNIV

Intelligent safety risk control template automatic generation method

The invention provides a method for automatically generating an intelligent safety risk control template, which belongs to the technical field of template automatic generation and comprises the following steps of: 1, executing a template matching and auditing queue admission process automatically generated by the intelligent safety risk control template; step 2, carrying out text clustering based on multi-dimensional features; step 3, carrying out cluster text sampling; 4, template extraction driven by a large model is carried out; 5, performing template to-be-audited and alarm notification; step 6, after template to-be-audited and alarm notification is carried out, manual auditing and template effectiveness are executed; and step 7, after manual auditing and the template take effect, performing auditing queue message processing and template multiplexing. By fusing the text clustering technology and the text understanding ability of a large language model, an intelligent security risk control strategy for automatically generating a template is provided, and the goals of ensuring violation interception, efficiently generating the template and having no interference on customer experience are achieved.
Owner:SHANGHAI CHUANGLAN CULTURE COMM CO LTD

Event importance judgment method and device, storage medium and program product

PendingCN121327579ASocial mediaTarget weight
The invention provides an event importance judgment method and device, a storage medium and a program product, and the method comprises the steps: carrying out the event element extraction of a target event text cluster corresponding to a target event, and obtaining an event element corresponding to the target event; determining the propagation condition of the target event based on the media coverage and social media participation degree data corresponding to the target event; determining a target weight ratio and a target scoring rule corresponding to the event elements and the propagation conditions; scoring the event elements and the propagation condition based on a target scoring rule to obtain an event element score and a propagation condition score; and according to a target weight ratio, performing weighted summation on the event element score and the propagation condition score to obtain an importance degree score corresponding to the target event. The problem that the accuracy of the event importance degree judgment result is low possibly due to the fact that the importance degrees of events are judged and classified by means of keyword matching can be solved. The accuracy and the judgment efficiency of event importance judgment can be improved.
Owner:CHINA ELECTRONICS CYBERSPACE RESEARCH INSTITUTE CO LTD

Text key feature extraction system and method based on deep learning

The invention relates to the field of text extraction, in particular to a text key feature extraction system and method based on deep learning, and the system comprises a text screening module, a graph construction module, a deep learning module, a cluster feature module and an input fusion module, the text screening module is used for merging text clusters, the graph construction module is used for generating a stereogram model, and the deep learning module is used for inputting the stereogram model. The deep learning module is used for expanding the model through the cluster text, the cluster feature module is used for trimming the stereogram model, and the input fusion module is used for capturing new text features and generating a fused text.According to the invention, a visual interface can be provided for training complex texts, the coherence analysis ability of a machine to long texts is improved, the generalization of machine learning is enhanced, and the efficiency of the machine learning is improved. The method is suitable for processing a large amount of complex and high-dimensional text data, reduces external data requirements, adapts to data distribution changes, reduces model calculation burdens, and realizes language framework stabilization and accurate fusion of text features.
Owner:BEIJING CETEN EDUCATION TECH GRP CO LTD

Tutoring robot processing system based on natural language processing

The application discloses a tutoring robot processing system based on natural language processing and belongs to the technical field of robot processing. The system comprises a knowledge point text acquisition module, a knowledge point text preprocessing module, a knowledge point text vectorization processing module, a knowledge point text clustering module and a tutoring robot processing module. The application specifically introduces a chaotic adaptive wandering strategy to improve the dolphin swarm algorithm, obtains the part-of-speech influence, obtains the distribution concentration based on distribution entropy, generates the word importance feature vector, and splices the knowledge point text vector with the semantic feature vector. The initial clustering center and the dimension weight are constructed into a membership matrix according to the double-feature weighted distance, the clustering center is updated according to the credibility index, the dimension weight is updated by combining the clustering contribution, the part-of-speech influence, the distribution concentration and the variance, convergence is judged, and the clustering result is obtained, so that the recommendation result of the tutoring robot is stably focused on the core knowledge point and the accuracy of the tutoring robot recommendation is improved.
Owner:HEBEI HUAFA EDUCATION TECH CORP LTD

Text clustering method and device, computer readable medium and computer equipment

PendingCN120654011AFeature vectorEngineering
The embodiment of the invention provides a text clustering method and device, a computer readable medium and computer equipment. The text clustering method comprises the steps of obtaining a to-be-clustered text and clustering keywords respectively corresponding to at least two clustering categories; matching the text to be clustered with the clustering keywords respectively corresponding to the at least two clustering categories to obtain a matching degree between the text to be clustered and each clustering category; according to the matching degree between the to-be-clustered text and each clustering category, dividing the to-be-clustered text into the at least two clustering categories; and extracting a feature vector corresponding to the to-be-clustered text in each clustering category, and performing clustering processing on the to-be-clustered text in each clustering category according to the feature vector corresponding to the to-be-clustered text to obtain a clustering result of the to-be-clustered text. According to the technical scheme provided by the embodiment of the invention, the text clustering accuracy and the text clustering efficiency are effectively improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Knowledge base question and answer pair generation method and system combining large model analysis and text clustering

The invention provides a knowledge base question and answer pair generation method and system combining large model analysis and text clustering. The method comprises the following steps that each document is preprocessed and segmented sentence by sentence, and an obtained sentence set serves as a reference answer; according to the reference answers, respectively generating corresponding special question sentences through a plurality of large model platforms, performing text clustering on the special question sentences generated by the same reference answer, and selecting one question sentence as a representative question sentence in each cluster; each representative question-reference answer is audited and perfected, and candidate question-answer pairs are generated; comprehensively comparing the contents of all document candidate question and answer pairs, clustering question and answer pairs which may have information conflicts, analyzing, studying and judging clustering results, and generating credible knowledge base question and answer pairs; and comparing the credible knowledge base question and answer pairs with the existing question and answer pairs in the knowledge base one by one, and if no information conflict exists, storing and updating. According to the method, high-quality question and answer pair data is generated based on the document, and the application effect of the RAG technology is improved in an auxiliary mode.
Owner:FUJIAN NORMAL UNIV

Text clustering method and device, electronic equipment, storage medium and program product

The invention provides a text clustering method and device, electronic equipment, a storage medium and a program product, relates to the technical field of informatization software systems, and is used for improving the accuracy of text clustering. Inputting the abstract text into a first clustering model, and calculating a cosine similarity between the abstract text and a cluster center vector of each cluster in N clusters in the first clustering model through the first clustering model to obtain N similarity values; inputting the abstract text into a second clustering model, and calculating the matching degree between the abstract text and each cluster map in N cluster maps in the second clustering model through the second clustering model to obtain N first matching degree values; and determining a first cluster corresponding to the to-be-classified text based on the N similarity values and the N first matching degree values. The method is applied to a data classification scene.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD

Text database updating method and device, equipment and storage medium

The invention relates to a text database updating method and device, equipment and a storage medium. The method comprises the following steps: when first text data is detected by sliding a time window, calling second text data of a text temporary database, and respectively preprocessing the first text data and the second text data to obtain a first text preprocessing result and a second text preprocessing result; and determining a semantic relationship between the first text data and the second text data according to the two preprocessing results. And merging the first text data and the second text data according to the semantic relationship to obtain temporary updated text data. And after storing the temporary update text data in the text temporary database, calculating the data similarity between the temporary update text data and the cluster description data of each text cluster in the text historical database. And updating the temporary update text data corresponding to the maximum value of the data similarity to the text historical database. By adopting the method, the text database updating efficiency can be improved.
Owner:中央军委政法委员会侦查技术中心 +2

Short text clustering and fuzzy recognition algorithm based on large-scale network online subgraph sampling

This invention provides a short text clustering and fuzzy recognition algorithm for large-scale online subgraph sampling, comprising the following steps: Step S1, extraction and preprocessing of training samples; Step S2, construction of the neural network; Step S3, overall clustering prediction; Step S4, fuzzy sample recognition; Step S5, retraining of the neural network. This invention combines short text clustering with a large language model, which not only improves clustering accuracy but also enables the handling of clustering tasks with different themes and classification requirements, significantly reducing the manual cost of data annotation. Furthermore, this invention can annotate fuzzy samples for classification, using K-nearest neighbors combined with minimum spanning trees to assist subgraph sampling in the selection scheme. This utilizes sparse structure to reduce computational costs and exposes the fluctuations of boundary samples through spectral clustering, providing a more comprehensive perspective for fuzzy sample selection and improving the accuracy and interpretability of the clustering results.
Owner:RENMIN UNIVERSITY OF CHINA +1