Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

10 results about "Sentence clustering" patented technology

Usually sentence clustering is used to cluster sentences derived from different documents and can be considered as a transverse segmentation of the documents content. Thus, the number of clusters can exceed the number of documents.

Text abstract generation method, text processing method, device, equipment, medium and program product

The embodiment of the invention provides a method for generating a text abstract, a text processing method and device, equipment, a medium and a program product. The text abstract generation method comprises the steps of obtaining a to-be-processed text; performing clustering analysis on all sentences in the to-be-processed text to obtain at least one sentence cluster; a first parameter set and a second parameter set of each sentence cluster in the at least one sentence cluster are determined, and the first parameter set of the target sentence cluster is a set composed of position weights of sentences in the target sentence cluster; the second parameter set of the target sentence cluster is a set composed of semantic distances between sentences in the target sentence cluster and the clustering centroid of the target sentence cluster; determining representative sentences of each sentence cluster in the at least one sentence cluster according to the respective first parameter set and the respective second parameter set of each sentence cluster in the at least one sentence cluster; and combining the representative sentences corresponding to the at least one sentence cluster to obtain an abstract of the to-be-processed text.
Owner:SHENZHEN ZHICHENG SOFTWARE TECH SERVICE CO LTD +1

A text key phrase extraction method, storage medium and device integrating inter-sentence correlation relationship

The application belongs to the field of natural language processing, and particularly relates to a text key phrase extraction method, a storage medium and a device that integrate inter-sentence correlation, comprising: extracting a nominal phrase in a part-of-speech combination mode; filtering the nominal phrase by combining two Trie trees to obtain a candidate phrase set; calculating a global semantic similarity score of the candidate phrase; clustering each sentence to obtain a sentence cluster containing different semantic information; and sorting the candidate phrases in the sentence cluster according to the global semantic similarity score of the candidate phrase to obtain a key phrase set; the application integrates the inter-sentence relationship in the unsupervised key phrase extraction model based on embedding, improves the accuracy of the model, and constructs different Trie trees to calculate mutual information and left and right information entropy as a candidate phrase extraction method in the key phrase extraction model, thereby reducing the probability of incomplete semantic information appearing in the extracted candidate phrase set.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Method for generating text summary, text processing method, device, equipment, medium and program product

Embodiments of the present application provide a method for generating a text summary, a text processing method, an apparatus, a device, a medium and a program product. The method for generating a text summary comprises: obtaining a to-be-processed text; performing cluster analysis on all sentences in the to-be-processed text to obtain at least one sentence cluster; determining a first parameter set and a second parameter set of each sentence cluster in the at least one sentence cluster, the first parameter set of a target sentence cluster being a set composed of position weights of each sentence in the target sentence cluster, and the second parameter set of the target sentence cluster being a set composed of semantic distances between each sentence in the target sentence cluster and a cluster centroid of the target sentence cluster; determining a representative sentence of each sentence cluster in the at least one sentence cluster according to the first parameter set and the second parameter set of each sentence cluster; and combining the representative sentences corresponding to the at least one sentence cluster to obtain a summary of the to-be-processed text.
Owner:SHENZHEN ZHICHENG SOFTWARE TECH SERVICE CO LTD +1

Contract review method and device, electronic equipment and storage medium

The application relates to the technical field of natural language processing, and provides a contract examination method and device, electronic equipment and a storage medium, the method comprising: performing sentence clustering on a contract text based on semantic vectors of each sentence in the contract text, and lengths and / or structured types of the sentences, to obtain a plurality of semantic clusters; determining a plurality of semantic segments of the contract text based on the plurality of semantic clusters; searching for a target semantic segment associated with an examination item from the plurality of semantic segments; and examining the contract text based on the examination item and the target semantic segment. The method, device, electronic equipment and storage medium provided by the application can effectively guarantee the accuracy and reliability of contract examination while retaining long-distance semantic dependency relationships in the contract text, and can guarantee the traceability and checkability of examination conclusions based on the target semantic segment associated with the examination item.
Owner:IFLYTEK CO LTD

A method and apparatus for processing conversation information

This invention discloses a method and apparatus for processing conversation information, relating to the field of customer service technology. The method includes: extracting individual sentences from conversational statements generated by customer service; classifying the individual sentences using preset positive factor models and negative factor models; wherein the positive factor model is obtained by training a classification model through multiple positive clusters generated from positive key sentence clustering, and the negative factor model is obtained by training a classification model through multiple negative clusters generated from negative key sentence clustering; determining the target evaluation type of the individual sentence based on the evaluation types set by the classification results corresponding to the positive and negative factor models and the classification results of the individual sentences; and determining a service evaluation for customer service based on a preset evaluation system containing evaluation types and the target evaluation type of the individual sentences. This implementation can identify objective problems existing in conversational services based on the conversation and effectively improve the accuracy of conversation analysis.
Owner:JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD

Text classification apparatus, text classification system, text classification method, and text classification program

To automatically generate a label corresponding to each cluster obtained by clustering each sentence of a model document, and to automatically create sentence classification information which is teacher data.SOLUTION: The cluster configuration unit 130 clusters a plurality of sentence units 511 obtained as a result of decomposing the model document 52 into a plurality of clusters. The label setting unit 140 acquires a label setting result in which a label is set to each cluster based on the clustering result. The model learning unit 150 acquires the text classification information 53 in which the text units included in each of the plurality of clusters are associated with the labels based on the label setting result. Here, the text classification information 53 is training data used for learning of the classification model 60 that classifies text units included in the classification target document.SELECTED DRAWING: Figure 1
Owner:MITSUBISHI ELECTRIC CORP

An open-domain dialogue method based on dialogue structure graph constraint

The application discloses an open domain dialogue method based on dialogue structure graph constraint, which comprises the following steps: after obtaining the dialogue sentence vector representation of an encoder, a new comparative learning loss function is designed by using the features of dialogue sequence and correlation to further train, so that dialogue sentence vectors containing sufficient semantics are obtained; the newly obtained dialogue sentence vectors are clustered to obtain topic-level sentence clustering; finally, the transfer of topics in the dialogue data set is imitated by using imitation learning to construct a topic-level dialogue structure graph, i.e., the transfer between clusters, and the text generation of the autoregressive decoder is constrained by the dialogue structure graph. By using comparative learning to fully extract sentence meaning information and using imitation learning to obtain a dialogue structure graph and predict the next round of dialogue topics, the correlation between generated dialogue and topics is well constrained, and the overall dialogue fluency is improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Topic adding method and device, equipment, medium and program product

The invention discloses a topic adding method and device, equipment, a medium and a program product, and belongs to the field of data mining. The method comprises the steps that a first comment text is acquired; the first comment text is subjected to sentence segmentation, the first comment text is converted into one or more clause texts, and the clause texts are all texts or part of texts of the first comment text; dividing the plurality of clause texts into at least one sentence cluster; aiming at each sentence cluster in the at least one sentence cluster, respectively extracting a topic corresponding to each sentence cluster by adopting a pre-trained large language model; and storing the topic corresponding to each sentence cluster into a topic library. According to the method, a first comment text is converted into one or more clause texts, the clause texts are divided into sentence clusters, and topics corresponding to the sentence clusters are extracted and then stored in a topic library, so that topics are added into the topic library, and the number of topics in the topic library is increased.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method, device and system for obtaining training data, and storage medium

Some embodiments of the present application provide a method, device, system and storage medium for obtaining training data, the method comprising: obtaining semantic representation vectors of each sentence in a plurality of sentences according to a target semantic representation model; obtaining similarity values between any sentence and the remaining sentences in the plurality of sentences according to the semantic representation vectors and a similarity algorithm, to obtain a plurality of similarity values; if it is confirmed that the any sentence is similar to any reference sentence and the any sentence and the any reference sentence do not belong to the same sentence cluster according to the size relationship of the plurality of similarity values, then it is confirmed that the any sentence and the reference sentence form a set of negative sample data. Some embodiments of the present application can construct negative sample data with semantic matching level, and thus the text matching model trained by using the negative sample data has strong semantic matching ability.
Owner:阳光保险集团股份有限公司

Text script clustering method and device, computer device and storage medium

The application relates to the technical field of artificial intelligence, and discloses a text script clustering method and device, computer equipment and a storage medium. The method comprises the following steps: obtaining text data to be clustered, and randomly and evenly dividing the text data to be clustered into K groups; adopting an online greedy algorithm to perform intra-group clustering on the text data in each group respectively, so as to obtain an intra-group clustering result set of the text data to be clustered; adopting the online greedy algorithm to perform cross-group clustering on a center sentence in the intra-group clustering result set, so as to obtain a center sentence clustering result set; obtaining a corresponding clustering cluster of each center sentence in the center sentence clustering result set in the intra-group clustering result set, and merging clustering members of the clustering cluster, so as to obtain a clustering result of the text data to be clustered. The application can effectively improve the clustering accuracy of text scripts, reduce the clustering time of text scripts, and obtain a high-quality clustering result within a reasonable time range.
Owner:PING AN TECH (SHENZHEN) CO LTD