Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

56 results about "Text cluster" patented technology

Text processing method, apparatus, device, storage medium, and program product

Embodiments of the present application provide a text processing method, device, equipment, storage medium and program product. The method comprises: obtaining a target text; processing the target text based on a vector dictionary to obtain a target vector; and processing the target vector based on a target task model to obtain a text processing result. In the embodiments of the present application, the target text is processed by the pre-constructed vector dictionary to obtain the target vector, and then the target vector is processed by the target task model. Since the target vector includes a semantic feature vector and a local difference vector, the accuracy of text processing can be improved. Moreover, the target task model is any one of a text classification model, a text clustering model and a sentiment analysis model in the call center customer service field. Compared with the use of a general large model, the efficiency of text processing can be improved.
Owner:BEIJING HOLLYCRM TECH

Methods, apparatus, devices, and storage media for aggregating cross-page test questions

This application relates to the field of computer technology and discloses a method, apparatus, device, and storage medium for cross-page test question aggregation. The method includes acquiring multiple test question images, each image including test question elements; extracting text blocks from the test question images and grouping text blocks belonging to the same test question and the same test question element into a text cluster, obtaining a text cluster set; searching for a baseline text cluster and multiple candidate text clusters in the text cluster set, where candidate text clusters are those related to the baseline text cluster; generating label codes for each text cluster in the baseline and candidate text clusters based on the test question element to which the text cluster belongs, the text blocks within the text cluster, and the position of the text blocks in the test question images; and searching for text clusters belonging to the same test question as the baseline text cluster in the candidate text clusters based on the label codes, obtaining a test question text cluster set. The method of this application can achieve automatic clustering of test question elements with high clustering accuracy.
Owner:HANGZHOU ZHIJUAN PLANET TECHNOLOGY CO LTD

A keyword reverse propagation algorithm based on citation network structure

The application provides a keyword reverse propagation algorithm based on a citation network structure, and comprises the following steps: a spring charge model is established, force-directed layout processing is performed, and a force-directed layout graph is established; a keyword propagation model is established by using a reverse propagation algorithm, and a keyword weight change contrast curve is obtained; a citation network model is constructed by using the force-directed layout graph; iterative calculation is performed on the force-directed layout graph until the energy state in the force-directed layout graph reaches a minimum value; while the citation network is iteratively calculated, the keyword weight in the citation network model is adjusted, and a converged citation network layout graph is calculated. The application has the beneficial effects that the entire network is taken as a main body, keywords owned by cited documents are selected at a certain probability and are propagated backward along the network to the citing documents, the weight of the keywords is increased while the text clustering idea is retained, and clear visual data effects are obtained through force-directed layout.
Owner:UNICLOUD TECH CO LTD

A text clustering method based on improved whale optimization algorithm

The application relates to a text clustering method based on an improved whale optimization algorithm, and belongs to the fields of big data mining and machine learning. The method comprises the following steps: S1: obtaining scientific research text data from a scientific research text database by using Spark, and performing data cleaning; S2: performing word segmentation, stop word removal, feature selection, vectorization and dimension reduction operations on the text data after data cleaning, so that unstructured text data is converted into structured numerical data; and S3: calculating initial clustering centers by using a K-means algorithm, optimizing the clustering centers by using an improved whale optimization algorithm, and outputting clustering results. The global optimization capability of the improved whale optimization algorithm is utilized, the problem that current text clustering algorithms are prone to falling into local optimization is solved, the accuracy of clustering of scientific research text data is improved, the data dimension is reduced, and the clustering effect is enhanced.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Text classification method, electronic equipment, storage medium and program product

The embodiment of the invention provides a text classification method, electronic equipment, a storage medium and a program product. The method comprises the following steps: processing a to-be-classified text according to a lexical element division rule to obtain a lexical element sequence; according to the lexical element sequence, adopting a pre-trained unsupervised topic model to obtain a topic distribution data set corresponding to the to-be-classified text; adopting a pre-trained unsupervised clustering model to obtain a clustering distribution data set corresponding to the to-be-classified text; splicing the topic distribution data set and the clustering distribution data set to obtain a text hidden topic; obtaining a plurality of text clusters and current cluster hidden topics corresponding to the text clusters, and calculating the similarity between the text hidden topics and the current cluster hidden topics; and determining a target text cluster from the plurality of text clusters, and adding the to-be-classified text to the target text cluster. Through a cluster classification mechanism of subject distribution and cluster distribution conjoint analysis, the accuracy and result stability of text classification are improved.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD +1

A cross-modal retrieval method based on prototype regularization learning

The application provides a cross-modal retrieval method based on prototype regularization learning, and belongs to the field of artificial intelligence and multi-modal information processing, and comprises the following steps: extracting visual features and text features through a visual encoder and a text encoder and calculating a basic contrast loss; dividing visual clusters and text clusters through a clustering algorithm and calculating visual prototypes and text prototypes, selecting text features most similar to the visual prototypes as anchor points, and constructing cross-modal prototypes; calculating a prototype-level discriminant loss according to the visual prototypes and the text prototypes, calculating an instance-level discriminant loss according to the visual clusters and the text clusters, and calculating a prototype projection loss according to the cross-modal semantic prototypes; constructing a joint optimization objective function by combining all the losses, training the visual encoder and the text encoder, and performing cross-modal retrieval after the training is completed. The application solves the problems of excessive intra-class variance and excessively high inter-class similarity caused by appearance bias in the existing cross-modal retrieval method, and causes the problems of insufficient retrieval accuracy and robustness.
Owner:SHANGHAI EYE DISEASE PREVENTION & TREATMENT CENTER

Text clustering method, text clustering device, electronic device, and storage medium

The application provides a text clustering method, a text clustering device, an electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence. An initial clustering center is selected from a plurality of original sentence embedding vectors, an initial problem text class is constructed based on the initial clustering center, clustering processing is performed on all original sentence embedding vectors based on the initial clustering center, a plurality of intermediate problem text classes are obtained, a clustering center is identified for each intermediate problem text class, a text clustering center is obtained, inter-class dispersion between the intermediate problem text classes and intra-class dispersion of each intermediate problem text class are calculated according to the text clustering center, the intermediate problem text classes are adjusted in text according to the inter-class dispersion and the intra-class dispersion, a target problem text class is obtained, semantic analysis is performed on all original sentence embedding vectors of the target problem text class, a standard problem text is obtained, and standard reply content is determined according to the standard problem text, so that the accuracy of problem answering can be improved.
Owner:CHINA PING AN LIFE INSURANCE CO LTD

Text clustering with heuristic and multi-metric control

Implementations generally relate to text clustering with heuristic and multi-metric control. In some implementations, a method includes receiving an electronic source document containing text. The method further includes dividing the text into text units, encoding the text units, and transforming the text units into numerical values. The method further includes generating a graph of the text units based on the numeric values, where the graph includes nodes corresponding to the text units and edges corresponding to pairs of the text units. The method further includes ordering the text units into text clusters based on the graph of the text units. The method further includes generating an electronic target document that presents the text clusters based on one or more preference heuristics.
Owner:JPMORGAN CHASE BANK NA

A clustering-based text supervision checking method, device and computer equipment

The application relates to a clustering-based text supervision checking method and device and a computer device. Feature word sets corresponding to a plurality of question description texts in a question list are obtained by respectively extracting feature words from the question description texts. The word frequency of each feature word in the question description text, the first inverse document frequency in all question description texts and the second inverse document frequency in all checked units are respectively calculated to obtain a feature word weight set of the question description text of each checked unit. A plurality of question description text clusters and a plurality of initial high-correlation-degree feature word sets corresponding to the question description text clusters are obtained through clustering. Finally, the initial high-correlation-degree feature word set is optimized according to a chi-square statistic to obtain an accurate high-correlation-degree feature word set. The method saves the calculation resources while ensuring the accuracy of the results by optimizing the feature word weight calculation, adopting a clustering algorithm for preliminary clustering and optimizing the clustering results through chi-square statistics.
Owner:NAT UNIV OF DEFENSE TECH

A cross-domain technology fusion opportunity finding method based on science and technology fusion interaction

The application belongs to the technical field of data mining, and specifically discloses a cross-field technology fusion opportunity finding method based on science and technology fusion interaction, which comprises the following steps: obtaining scientific paper and technology patent data, identifying cross-field scientific fusion and technology fusion based on semantic anchor points and a similarity threshold; using a trained deep text clustering model to respectively perform deep clustering on scientific fusion and technology fusion, and obtaining corresponding clustering sets; obtaining text representation corresponding to each cluster, filtering through a threshold, and establishing association between two types of clusters through bipartite graph maximum weight matching to identify shared topics and exclusive topics; using a dynamic time warping method to align the time series of scientific and technological trends, judging topics with scientific convergence preceding technological convergence, and identifying technology fusion opportunities with growth potential based on the development trend of scientific papers. The application can deeply mine cross-field science and technology semantics, improve clustering accuracy, quantify science and technology association, and dynamically find forward-looking technology fusion opportunities.
Owner:SOUTHWEST JIAOTONG UNIV

Text clustering method and device

The embodiment of the invention provides a text clustering method and device, and the method comprises the steps: obtaining a plurality of to-be-clustered texts, and calculating the text similarity between the to-be-clustered texts; determining a noise text in the plurality of to-be-clustered texts according to the text similarity, and deleting the noise text in the plurality of to-be-clustered texts to obtain a to-be-clustered text set; based on to-be-clustered texts in the to-be-clustered text set and the text similarity between the to-be-clustered texts, constructing a target text connected graph corresponding to the to-be-clustered text set; and determining a central text corresponding to the target text connected graph, deleting an edge text in the target text connected graph based on the central text, and determining a target text set according to a deletion result.
Owner:UC MOBILE CHINA CO LTD

Event text hot word extraction method based on similarity adjacent matrix clustering

The invention relates to the field of natural language processing, in particular to an event text hot word extraction method based on similarity adjacent matrix clustering, which comprises the following steps: acquiring a keyword set of each event text; calculating an improved Jaccard similarity value of any two event texts, and constructing a text similarity adjacent matrix; converting the text similarity adjacent matrix into a distance matrix, and dividing all texts into K text clusters based on the distance matrix; removing the K text clusters with abnormal text quantity by using a quartile method to obtain normal text clusters; word weight calculation is conducted on each text in the normal text cluster again to obtain a normal keyword set, the final comprehensive score of each keyword is calculated based on the normal keyword set, candidate keywords in the cluster are reordered according to the final comprehensive score, and the candidate keywords ranked in the front serve as final hot words. The hot word extraction accuracy and efficiency can be improved.
Owner:DIGITAL CHONGQING BIG DATA APPL DEV CO LTD

Text clustering method, apparatus, and electronic device

The application discloses a text clustering method and device and electronic equipment. It is related to the field of financial technology or other fields, and the method comprises the following steps: obtaining a plurality of digital vectors of a to-be-processed text, wherein each digital vector corresponds to part of the to-be-processed text; determining a first distance threshold and a second distance threshold based on the plurality of digital vectors, wherein the first distance threshold is the maximum limit value of the clustering range, and the second distance threshold is the minimum limit value of the clustering range; performing first clustering processing on the plurality of digital vectors based on the first distance threshold and the second distance threshold to obtain a clustering result; obtaining the number of clusters in the clustering result; performing second clustering processing on the plurality of digital vectors based on the number of clusters to obtain a target centroid vector corresponding to each cluster, wherein the target centroid vector represents the characteristics of the cluster corresponding to the target centroid vector. The application solves the technical problem of poor text clustering effect caused by the fact that the number of clusters cannot be accurately determined in the prior art.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

A method and system for product title clustering

The application discloses a kind of method and system for commodity title clustering, specifically related to text clustering technical field, for solving the technical scheme of general data is generally aimed at existing disclosure technology, there is also scheme is aimed at such data, but due to different application scenarios, such scheme is not completely applicable problem, its method includes data crawling, commodity title normalization, semantic vector conversion, similarity analysis, clustering analysis and search recommendation of similar goods, system is composed of hardware and software, software includes crawler module, processing module, semantic vector module, similarity calculation module, clustering module and recommendation module, hardware includes CPU, memory bank, storage and GPU;It is through the competitive product analysis of online commodity transaction information data, obtains the similarity of commodity title and carries out clustering analysis and search recommendation to the similarity of commodity title, to improve the accuracy of clustering and search recommendation.
Owner:ZHONGKE (XIAMEN) DATA INTELLIGENCE RES INST

Short text clustering method based on adaptive optimal transmission and three-level robust representation

The invention provides a short text clustering method based on adaptive optimal transmission and three-level robust representation, which comprises the following steps of: 1, coding an input original short text based on a pre-training language model in a server carrying a GPU (Graphics Processing Unit) to obtain a semantic representation of the original short text, and performing multi-type data enhancement processing on the semantic representation to obtain a semantic representation of the original short text; obtaining an enhanced sample representation; 2, performing pseudo label distribution on enhanced sample representation based on an optimal transmission algorithm; 3, learning and calculating a corresponding loss function based on the pseudo label and three-level robust representation, and updating model parameters of the loss function in a CPU (Central Processing Unit) of a server; 4, iteratively executing the steps 1-3, and outputting a target clustering result when a clustering result meets a convergence condition; the problems of data sparseness and class imbalance in short text clustering can be effectively relieved, and the stability and precision of a clustering result are improved.
Owner:YANCHENG INST OF TECH +1

Machining drawing processing information extraction method based on YOLO + OCR

The invention relates to the technical field of computer vision and intelligent information processing, in particular to a machining drawing information extraction method based on YOLO + OCR, which comprises the following steps: sequentially carrying out graying processing, Gaussian filtering noise reduction and Canny edge detection on a machining drawing, and extracting and screening drawing frames and drawing areas; detecting a text group in the drawing by adopting a YOLO-OBB algorithm; and performing discrete classification prediction on the rotation angle of the detected text group image. According to the method, text pixel loss caused by direct zooming is avoided through drawing preprocessing, the 360-degree text direction problem is solved by utilizing a YOLO-OBB algorithm and a four-classification neural network model, PMI information is aligned and recognized in combination with a special symbol model, global parameters are extracted through natural language processing, and finally errors are monitored and corrected. The problems of text pixel loss, insufficient direction processing, poor symbol and text association and the like in the prior art are solved.
Owner:BEIHANG UNIV

Intelligent safety risk control template automatic generation method

The invention provides a method for automatically generating an intelligent safety risk control template, which belongs to the technical field of template automatic generation and comprises the following steps of: 1, executing a template matching and auditing queue admission process automatically generated by the intelligent safety risk control template; step 2, carrying out text clustering based on multi-dimensional features; step 3, carrying out cluster text sampling; 4, template extraction driven by a large model is carried out; 5, performing template to-be-audited and alarm notification; step 6, after template to-be-audited and alarm notification is carried out, manual auditing and template effectiveness are executed; and step 7, after manual auditing and the template take effect, performing auditing queue message processing and template multiplexing. By fusing the text clustering technology and the text understanding ability of a large language model, an intelligent security risk control strategy for automatically generating a template is provided, and the goals of ensuring violation interception, efficiently generating the template and having no interference on customer experience are achieved.
Owner:SHANGHAI CHUANGLAN CULTURE COMM CO LTD

Event importance judgment method and device, storage medium and program product

PendingCN121327579ASocial mediaTarget weight
The invention provides an event importance judgment method and device, a storage medium and a program product, and the method comprises the steps: carrying out the event element extraction of a target event text cluster corresponding to a target event, and obtaining an event element corresponding to the target event; determining the propagation condition of the target event based on the media coverage and social media participation degree data corresponding to the target event; determining a target weight ratio and a target scoring rule corresponding to the event elements and the propagation conditions; scoring the event elements and the propagation condition based on a target scoring rule to obtain an event element score and a propagation condition score; and according to a target weight ratio, performing weighted summation on the event element score and the propagation condition score to obtain an importance degree score corresponding to the target event. The problem that the accuracy of the event importance degree judgment result is low possibly due to the fact that the importance degrees of events are judged and classified by means of keyword matching can be solved. The accuracy and the judgment efficiency of event importance judgment can be improved.
Owner:CHINA ELECTRONICS CYBERSPACE RESEARCH INSTITUTE CO LTD

Text key feature extraction system and method based on deep learning

The invention relates to the field of text extraction, in particular to a text key feature extraction system and method based on deep learning, and the system comprises a text screening module, a graph construction module, a deep learning module, a cluster feature module and an input fusion module, the text screening module is used for merging text clusters, the graph construction module is used for generating a stereogram model, and the deep learning module is used for inputting the stereogram model. The deep learning module is used for expanding the model through the cluster text, the cluster feature module is used for trimming the stereogram model, and the input fusion module is used for capturing new text features and generating a fused text.According to the invention, a visual interface can be provided for training complex texts, the coherence analysis ability of a machine to long texts is improved, the generalization of machine learning is enhanced, and the efficiency of the machine learning is improved. The method is suitable for processing a large amount of complex and high-dimensional text data, reduces external data requirements, adapts to data distribution changes, reduces model calculation burdens, and realizes language framework stabilization and accurate fusion of text features.
Owner:BEIJING CETEN EDUCATION TECH GRP CO LTD

Tutoring robot processing system based on natural language processing

The application discloses a tutoring robot processing system based on natural language processing and belongs to the technical field of robot processing. The system comprises a knowledge point text acquisition module, a knowledge point text preprocessing module, a knowledge point text vectorization processing module, a knowledge point text clustering module and a tutoring robot processing module. The application specifically introduces a chaotic adaptive wandering strategy to improve the dolphin swarm algorithm, obtains the part-of-speech influence, obtains the distribution concentration based on distribution entropy, generates the word importance feature vector, and splices the knowledge point text vector with the semantic feature vector. The initial clustering center and the dimension weight are constructed into a membership matrix according to the double-feature weighted distance, the clustering center is updated according to the credibility index, the dimension weight is updated by combining the clustering contribution, the part-of-speech influence, the distribution concentration and the variance, convergence is judged, and the clustering result is obtained, so that the recommendation result of the tutoring robot is stably focused on the core knowledge point and the accuracy of the tutoring robot recommendation is improved.
Owner:HEBEI HUAFA EDUCATION TECH CORP LTD

Text clustering method and device, electronic equipment, storage medium and program product

The invention provides a text clustering method and device, electronic equipment, a storage medium and a program product, relates to the technical field of informatization software systems, and is used for improving the accuracy of text clustering. Inputting the abstract text into a first clustering model, and calculating a cosine similarity between the abstract text and a cluster center vector of each cluster in N clusters in the first clustering model through the first clustering model to obtain N similarity values; inputting the abstract text into a second clustering model, and calculating the matching degree between the abstract text and each cluster map in N cluster maps in the second clustering model through the second clustering model to obtain N first matching degree values; and determining a first cluster corresponding to the to-be-classified text based on the N similarity values and the N first matching degree values. The method is applied to a data classification scene.
Owner:CHINA UNITED NETWORK COMM GRP CO LTD

Short text clustering and fuzzy recognition algorithm based on large-scale network online subgraph sampling

This invention provides a short text clustering and fuzzy recognition algorithm for large-scale online subgraph sampling, comprising the following steps: Step S1, extraction and preprocessing of training samples; Step S2, construction of the neural network; Step S3, overall clustering prediction; Step S4, fuzzy sample recognition; Step S5, retraining of the neural network. This invention combines short text clustering with a large language model, which not only improves clustering accuracy but also enables the handling of clustering tasks with different themes and classification requirements, significantly reducing the manual cost of data annotation. Furthermore, this invention can annotate fuzzy samples for classification, using K-nearest neighbors combined with minimum spanning trees to assist subgraph sampling in the selection scheme. This utilizes sparse structure to reduce computational costs and exposes the fluctuations of boundary samples through spectral clustering, providing a more comprehensive perspective for fuzzy sample selection and improving the accuracy and interpretability of the clustering results.
Owner:RENMIN UNIVERSITY OF CHINA +1

A method for automatically generating a first draft of a situation briefing

The application discloses a kind of situation briefing first draft automatic generation method, comprising the following steps: S1, abstract extraction: one or more sentences are extracted as abstract from each news material using the BertSum method based on Bert model;S2, text clustering: the abstract of each news material is carried out text clustering using Kmeans algorithm;S3, abstract generation: a summary is generated for the abstract of each class news material extracted using the abstract generation method based on T5 model as the subheading in situation briefing;S4, first draft generation: the results of abstract extraction, text clustering and abstract generation are sorted and combined according to the clustering category, form situation briefing first draft.The application automatically generates situation briefing, reduces the work burden of the author;Extraction type abstract as the main content of the briefing, ensures the authenticity of the briefing content.
Owner:10TH RES INST OF CETC

Text clustering method and device, nonvolatile storage medium and electronic device

The application discloses a text clustering method and device, a nonvolatile storage medium and an electronic device. The method comprises the following steps: clustering a text to be clustered according to a first algorithm to obtain a plurality of clustering clusters, and determining a keyword of the clustering cluster according to a second algorithm; determining a connection matrix according to the clustering cluster and the keyword, wherein the connection matrix is a symmetric matrix; performing normalization processing on the connection matrix to obtain a first target connection matrix, decomposing the first target connection matrix according to an average value of the first target connection matrix to obtain a second target connection matrix; and merging the second target connection matrix to generate a target clustering result. The application solves the technical problem that, in the existing clustering algorithm, a clustering model divides different expression methods of the same category into different categories due to a large number of clustering categories, thereby leading to inaccurate clustering.
Owner:CHINA TELECOM CORP LTD

Deep Learning-Based Text Key Feature Extraction System and Method

This invention relates to the field of text extraction, specifically to a deep learning-based system and method for extracting key text features. The system includes a text filtering module, a graph construction module, a deep learning module, a cluster feature module, and an input fusion module. The text filtering module merges text clusters, the graph construction module generates a stereo model, the deep learning module expands the model using clustered text, the cluster feature module prunes the stereo model, and the input fusion module captures new text features to generate fused text. This invention provides a visual interface for training complex text, enhances the machine's ability to analyze the coherence of long texts, improves the generalization ability of machine learning, is suitable for processing large amounts of complex, high-dimensional text data, reduces external data requirements, adapts to changes in data distribution, reduces the computational burden on the model, and achieves stable language frameworks and precise fusion of text features.
Owner:BEIJING CETEN EDUCATION TECH GRP CO LTD

A text clustering method, device, terminal equipment and storage medium

The application is suitable for the technical field of data processing, and provides a text clustering method and device, terminal equipment and computer readable storage medium, the method comprises: obtaining a text to be clustered; determining a one-hot encoding corresponding to each word in the text to be clustered; determining a paragraph vector corresponding to the text to be clustered according to the text to be clustered and the one-hot encoding corresponding to each word; and performing clustering processing on the text to be clustered according to the paragraph vector to obtain a clustering result of the text to be clustered. Compared with the prior art of clustering each text according to a keyword in each text, the method makes the paragraph vectors corresponding to different texts with the same word closer to each other according to the one-hot encoding of each word and the paragraph vectors corresponding to different texts with the same word determined according to the one-hot encoding of each word and different texts, thereby improving the accuracy and clustering effect of text clustering.
Owner:GREAT WALL MOTOR CO LTD

Public opinion dynamic monitoring and early warning system based on text clustering analysis

ActiveCN122173656BEarly warning systemEvent evolution
The present application relates to the technical field of public opinion monitoring, and particularly relates to a public opinion dynamic monitoring and early warning system based on text clustering analysis. In the clustering process, the semantic center of historical event clusters is introduced as a constraint condition, so that the semantic continuity of existing events can be fully considered when the public opinion text is classified, and event misjudgment or repeated generation caused by short-term expression changes can be effectively avoided. By constructing a drift suppression mechanism based on the semantic center change amplitude, the unstable event clusters are adjusted, the stability and accuracy of event evolution modeling are improved, the cross-time window event clusters are associated and determined, the stable event clusters are formed, the tracking management of public opinion events is realized, the text scale, the emotional mean, the emotional heterogeneity, the evolution characteristics of the scale and the emotion are introduced, and the low sample compensation mechanism is combined to model the public opinion event risk, so that the high-risk events are not ignored, the mature event risk is not excessively enlarged, and the timeliness and accuracy of public opinion early warning are improved.
Owner:KUNMING UNIV OF SCI & TECH

Document tagging method and system based on entity enhancement and multi-granularity fusion

The application provides a document labeling method and system based on entity enhancement and multi-granularity fusion, and relates to the technical field of artificial intelligence. The method comprises the following steps: cutting a to-be-processed document into multiple text blocks, and performing entity extraction on each text block to obtain a first entity list of each text block; based on the text blocks and the first entity list, text block-level labeling, text cluster-level labeling and full-text abstract-level labeling are respectively performed to generate text block labels, text cluster labels and full-text abstract labels; and the text block labels, the text cluster labels and the full-text abstract labels are fused to obtain labels of the to-be-processed document. The application improves the comprehensiveness and accuracy of document label generation.
Owner:DIGITAL ZHEJIANG TECH OPERATION CO LTD

Text clustering denoising method and system based on triple comparative learning and confidence false label

The invention discloses a text clustering denoising method and system based on triple contrast learning and a confidence false label, and belongs to the technical field of deep clustering. Constructing a dual-online network cooperative training to optimize a target network module and a confidence false label denoising module, and adopting a mode of selecting a false label by confidence to secondarily optimize the network to perform denoising and reduce false negative samples; and finally, inputting original text data, and clustering by using a cluster predictor. According to the method, the target network module is optimized by utilizing the cooperative training of the double online networks, so that the excessive dependence on large-scale batch sampling to obtain sufficient negative sample comparison signals is relieved, the risk of introducing noise samples due to random sampling is reduced, and the efficiency and robustness of model training are improved. Through a confidence coefficient pseudo label denoising module, reliable pseudo label data is selected for secondary optimization, the situation that semantics are similar but are mistakenly marked as negative samples, namely noise negative samples, is reduced, the representation quality of features learned by the model is improved, and the accuracy and stability of clustering results are enhanced.
Owner:HEFEI UNIV

Document marking method and system based on entity enhancement and multi-granularity fusion

The invention provides a document marking method and system based on entity enhancement and multi-granularity fusion, and relates to the technical field of artificial intelligence, the method comprises the following steps: segmenting a to-be-processed document into a plurality of text blocks, and performing entity extraction on each text block to obtain a first entity list of each text block; based on the text block and the first entity list, respectively performing text block level marking, text cluster level marking and full-text abstract level marking to generate a text block label, a text cluster label and a full-text abstract label; and fusing the text block tag, the text cluster tag and the full-text abstract tag to obtain a tag of the to-be-processed document. According to the invention, the comprehensiveness and accuracy of document label generation are improved.
Owner:DIGITAL ZHEJIANG TECH OPERATION CO LTD