Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

30 results about "Paraphrase" patented technology

A paraphrase /ˈpærəfreɪz/ is a restatement of the meaning of a text or passage using other words. The term itself is derived via Latin paraphrasis from Greek παράφρασις, meaning "additional manner of expression". The act of paraphrasing is also called "paraphrasis".

A paraphrase sentence recognition method and system based on semantic primitive knowledge and abstract semantic representation

The application belongs to the field of natural language processing, and particularly relates to a method and system for paraphrase recognition based on semantic primitive knowledge and abstract semantic representation, which comprises the following steps: performing word segmentation on a sentence, and performing word-level vector representation and semantic primitive knowledge representation; performing mean value processing on the semantic primitive knowledge representation result, and extracting interactive attention feature information of the mean value processing result by using global semantic information to obtain global semantic primitive representation; performing abstract semantic analysis on a to-be-recognized paraphrase sentence from a sentence structure to obtain a single-root directed acyclic graph, and performing global semantic primitive representation and word-level vector representation; extracting global and local feature information in the order of the directed acyclic graph, and performing distance feature measurement on the information; inputting the distance feature measurement result into a neural network to obtain a recognition result; the application introduces external semantic primitive knowledge to perform semantic representation, the accuracy of the semantic primitive knowledge representation is assisted by global semantic information, and the abstract semantics of a Chinese paraphrase sentence is analyzed to obtain semantic relations, so that the accuracy of paraphrase recognition is improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Data visualization narrative generation method, device and equipment based on paraphrases and medium

The invention discloses a data visualization narrative generation method, device and equipment based on paramagnetism, and a medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: generating sentence-level parastyle features according to parastyle contents input by a user; generating corresponding plot information for the parastyle features; according to data input by a user, determining data feature information corresponding to the data feature name in the external features; generating visual structure information of features according to the data feature information; generating animation structure information of the parastyle features according to the parastyle features and the plot information, the data feature information and the visual structure information corresponding to the parastyle features; and obtaining to-be-rendered data based on the white features and the plot information, the data feature information, the visualization structure information and the animation structure information corresponding to the white features, and rendering the to-be-rendered data to generate a data visualization narrative picture. The visual narrative content with high personalization is generated according to the generation of the visual narrative driven by the user-defined parastaring.
Owner:ZHEJIANG HERYMED TECH CO LTD

Ground-truth-less performance prediction of generative question-answering systems

PendingUS20260017346A1Biological modelsMachine learningParaphraseData mining
Systems and techniques that facilitate ground-truth-less performance prediction of generative question-answering systems are provided. In various embodiments, a system can access a large language model (LLM) and a natural language question for which a ground-truth answer is unavailable. In various aspects, the system can generate, via a machine learning classifier that receives as input a set of properties associated with the natural language question, a classification label indicating whether or not the large language model will correctly answer the natural language question. In various instances, the set of properties can include a semantic category of the natural language question, a subject popularity of the natural language question, a semantic consistency exhibited by the LLM in response to repeated executions on the natural language question, or a semantic consistency exhibited by the LLM in response to execution on paraphrases of the natural language question.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Vocabulary memory training method and system based on handwriting interaction

PendingCN121685214AData processing applicationsHandwritingParaphrase
The invention relates to the technical field of vocabulary memory training, and discloses a vocabulary memory training method and system based on handwriting interaction, and the system comprises a dictation module which is configured to enter a handwriting training environment after determining a to-be-memorized vocabulary, play the voice of the to-be-memorized vocabulary and display the paraphrase of the to-be-memorized vocabulary, and then carry out the handwriting input of a user; the acquisition module is configured to determine a next-round initial training interval based on the analysis result; the judging module is configured to judge whether to adjust the initial training interval of the next round according to the writing characteristic parameters; the processing module is configured to determine an adjustment coefficient of a next-round initial training interval based on the historical training parameters and obtain a next-round final training interval; and the execution module is configured to train the vocabularies to be memorized at the following final training interval. According to the method, the next training interval can be scientifically and reasonably determined according to the handwriting input result of the user, the writing characteristic parameters and the historical training parameters of the vocabularies to be memorized, and personalized vocabulary memorizing training is achieved.
Owner:李昇杰

Metric for assessing a quality of one or more paraphrases

PendingUS20260093920A1Natural language translationSemantic analysisParaphraseGrammatical error
A computing system includes a memory; and processing circuitry in communication with the memory. The processing circuitry is configured to: receive a paraphrase comprising a paraphrase text sample corresponding to an original text sample; and calculate a paraphrase metric value corresponding to the paraphrase, wherein the paraphrase metric value is calculated based on an adequacy score, a novelty score, and a fluency score of the paraphrase, the adequacy score indicating an extent to which the paraphrase text sample preserves a meaning of the original text sample, the novelty score indicating a level of difference between words and characters of the paraphrase text sample and words and characters of the original text sample, and the fluency score indicating an extent to which the paraphrase text sample is devoid of repetition, spelling, and grammatical mistakes.
Owner:WELLS FARGO BANK NA

Book reading method and system based on artificial intelligence, medium and product

The invention discloses a book reading method and system based on artificial intelligence, a medium and a product, and relates to the field of data processing. In the method, a book text is segmented into a plurality of text sentence segments based on preset punctuation marks; performing text type judgment on each text sentence segment, and dividing the text sentence segment into a paraphrase text and a role dialogue text; performing sentiment analysis on the parastyle text and the role dialogue text to generate a corresponding parastyle sentiment identifier and a corresponding role sentiment identifier; calling a first speech synthesis strategy to perform speech synthesis on the paraphrase text according to the paraphrase emotion identifier to obtain a paraphrase speech segment, and calling a second speech synthesis strategy to perform speech synthesis on the role dialogue text according to the role emotion identifier to obtain a role speech segment; and according to the original sequence of the text sentence segments in the book text, splicing the white speech segments and the role speech segments to obtain a target leading audio. By implementing the technical scheme provided by the invention, the artificial intelligence reading experience feeling of the user is improved.
Owner:QUANLIAN BOOK PUBLISHING & DISTRIBUTION CO LTD

Text embedding model training method and related apparatuses

The text embedding model training method and the related device provided in the application, the data processing device obtains the paraphrase vector of the target vocabulary and the initial vector of the sample text, inputs the paraphrase vector of the target vocabulary and the initial vector of the sample text into the first neural network model to be trained for training, and obtains a text embedding model having an embedding vector conversion function. Since the sample text includes the target vocabulary, and the paraphrase vector of the target vocabulary is an embedding vector representing the paraphrase information of the target vocabulary, the paraphrase information is used as prior knowledge, so that the trained text embedding model has a better embedding vector conversion effect.
Owner:SHANGHAI ZHENGDA XIMALAYA NETWORK TECH CO LTD

Context semantic recognition method and system based on user behaviors

The invention relates to the technical field of natural language processing, and discloses a context semantic recognition method and system based on user behaviors, and the method comprises the steps: collecting the multi-source behavior data of a target user, and extracting the semantic features of the user behaviors; constructing a semantic association graph of the user behaviors, detecting semantic understanding defects of the user behaviors, and setting a defect correction mechanism of the semantic understanding defects; semantic modes of user behaviors in different scenes are identified, and real intentions of a target user in different scenes are analyzed; constructing a causal relationship graph of the user behaviors, analyzing driving factors of the user behaviors, and generating behavior intention paraphrases of the target user; and in combination with the semantic association map, the defect correction mechanism and the behavior intention paraphrasing, executing context semantic recognition processing of the user behavior to obtain a semantic recognition result. According to the method, the implicit intention of the user behavior can be understood, and the accuracy of context semantic recognition is improved.
Owner:CHINALIN SECURITIES CO LTD

Knowledge enrichment type question generation method and device for question and answer system robustness

The application provides a knowledge-rich question generation method and device for robustness of a question and answer system, obtains distilled fact descriptions, paraphrases and synonyms of a to-be-queried entity as injected knowledge, and generates knowledge-rich questions by rewriting existing questions by using an editing mechanism, so that different types of knowledge can be used to expand original questions without changing the meanings of the original questions, and more diversified and more meaningful knowledge-rich questions can be generated. Furthermore, the application also inspiringly provides "diagnosis" information for a question and answer model, and provides a dynamic weight for each injected knowledge, so that the question and answer model pays more attention to a question part containing clue information to predict a correct answer, and pays less attention to a question part containing irrelevant information, so that the performance of the question and answer model on knowledge-rich questions and original questions can be effectively improved by dynamically adjusting the weight.
Owner:FUDAN UNIVERSITY +1

A training method and device of an open information extraction model

The application provides a training method and device of an open information extraction model, comprising: obtaining a target data set with natural language sentences as samples; generating paraphrases of each natural language sentence in the target data set; performing structured knowledge recovery on the paraphrases of each natural language sentence in the target data set to obtain structured knowledge corresponding to each natural language sentence in the target data set; constructing a first data set with the paraphrases and the structured knowledge corresponding to all natural language sentences in the target data set; and training the open information extraction model by using the first data set and the target data set in a noise reduction training mode. The application constructs a syntax-robust training framework based on paraphrase generation and structured knowledge recovery, so that the open information extraction model can be trained on a data set with sufficient and accurate syntax distribution to adapt to real-world scenarios.
Owner:TSINGHUA UNIVERSITY

A paraphrase sentence generation method based on a controllable latent space diffusion model

ActiveCN118070899BParaphraseData set
This invention proposes a paraphrase generation method based on a controllable latent space diffusion model, comprising the following steps: Step 1, constructing a paraphrase model based on a latent space diffusion model, and training the paraphrase model based on an existing paraphrase text dataset; the existing paraphrase text dataset contains the original sentence and its paraphrase; Step 2, dividing the paraphrase source sentence into key information and non-key information, and constructing enhanced semantic fragment information; Step 3, constructing a controller; Step 4, training the controller; Step 5, combining the paraphrase model based on a latent space diffusion model trained in Step 1 and the controller trained in Step 4 for collaborative reasoning to generate paraphrased sentences, thus completing the paraphrase generation task based on a controllable latent space diffusion model.
Owner:NANJING UNIV

Paraphrase and aggregate with large language models for improved decisions

ActiveUS12664981B2Speech recognitionParaphraseData mining
A method of interpreting a verbal input, may include: assigning a meaning classification to the verbal input, and a confidence score to the meaning classification; and based on the confidence score corresponding to the meaning classification of the verbal input being less than or equal to a threshold, generating at least one paraphrase of the verbal input using at least one large language model (LLM); assigning the meaning classification to the at least one paraphrase, and the confidence score to the meaning classification; and concatenating the verbal input, the at least one paraphrase, the meaning classification, and the confidence score to generate a concatenated input; inputting the concatenated input into the at least one LLM.
Owner:SAMSUNG ELECTRONICS CO LTD

Metric for assessing a quality of one or more paraphrases

ActiveUS12524612B1Natural language translationSemantic analysisParaphraseGrammatical error
A computing system includes a memory; and processing circuitry in communication with the memory. The processing circuitry is configured to: receive a paraphrase comprising a paraphrase text sample corresponding to an original text sample; and calculate a paraphrase metric value corresponding to the paraphrase, wherein the paraphrase metric value is calculated based on an adequacy score, a novelty score, and a fluency score of the paraphrase, the adequacy score indicating an extent to which the paraphrase text sample preserves a meaning of the original text sample, the novelty score indicating a level of difference between words and characters of the paraphrase text sample and words and characters of the original text sample, and the fluency score indicating an extent to which the paraphrase text sample is devoid of repetition, spelling, and grammatical mistakes.
Owner:WELLS FARGO BANK NA

Dialogue understanding method, apparatus, readable medium and electronic device

ActiveUS12488193B2Semantic analysisProgramming languageParaphrase
A dialog understanding method and apparatus, a readable medium, and an electronic device. The method comprises: acquires dialog content and a preset dialog parsing template, the preset dialog parsing template comprising preset description and at least one c or at least one slot, where in the description information is used for describing a paraphrase of each candidate intention when the preset dialog parsing template comprises at least one candidate intention, and describes a paraphrase of each slot when the preset dialog parsing template at least one comprises slot; and uses the dialog content and the dialog parsing template as inputs of a pre-trained target dialog understanding model to obtain a dialog state corresponding to the dialog content.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Rewriting model training method and device, text rewriting method and device and electronic equipment

The invention provides a rewriting model training method and device, a text rewriting method and device and electronic equipment, relates to the technical field of computers, in particular to the technical fields of artificial intelligence, large models, natural language processing and the like, and can be applied to application scenes such as sentiment analysis, recommendation systems, generative search, man-machine conversation systems, intelligent assistants and virtual humans. According to the specific implementation scheme, a first text is obtained; using the initial rewriting model to obtain a first rewriting result for the first text; using the target evaluation model to obtain a first model evaluation result for the first rewriting result; and training the initial rewriting model based on the first model evaluation result to obtain a target rewriting model. By adopting the method and the device, the text rewriting capability of the target rewriting model can be improved, so that when the target rewriting model is utilized to obtain the rewriting result for the target text, the accuracy of the rewriting result can be improved.
Owner:BAIDU (CHINA) CO LTD

Translation method of software development kit

The invention relates to the technical field of data processing, in particular to a translation method of a software development kit, which comprises the following steps of: 1, determining a translation target and a translation range; 2, collecting a terminology library; 3, scanning data, obtaining translation paraphrases and establishing a translation space; step 4, generating a software function component and assembling and verifying complete software; 5, checking and modifying a translation result; and 6, generating a translation report. According to the method, a target software development kit is obtained, a software document, an example code, a configuration tool and a software function requirement document are extracted, a translation database is established, terms, professional nouns and emerging concepts in some specific fields are collected, and then the components are input into a function verification space; and assembling and function verification are performed according to the demand document, so that terms, professional nouns or emerging concepts in specific fields can be accurately translated, and the accuracy and reliability of important information translation results are ensured.
Owner:JINGZHOU WEIYUAN NETWORK TECH

Low-resource language machine translation method and device based on multilingual semantic understanding driving

PendingCN122655804AData setParaphrase
The application discloses a kind of low-resource language machine translation method and device based on multilingual semantic understanding driving, method includes the following steps: constructing data set and pre-processing data set;Build semantic representation model, use the paraphrase sentence in paraphrase corpus to train semantic representation model, the vectorization representation of target language monolingual corpus is carried out, and semantic retrieval library is established;Mask language modeling target and inter-sentence similarity constraint are introduced, and enhanced semantic retrieval model is obtained;With enhanced semantic retrieval model, retrieve the candidate sentence similar to target language sentence semantics in semantic retrieval library, and carry out sorting screening, retain Top-K candidate sentence;After screening, candidate sentence is respectively paired with source language sentence, and enhanced parallel corpus is constructed and merged with original parallel corpus, to form extended training set;Neural machine translation model is trained using extended training set, and output translation text.The method of the application can improve semantic consistency, translation quality and cross-language generalization ability.
Owner:XINJIANG UNIVERSITY

LLM term understanding method, system and device and storage medium

The invention discloses an LLM term understanding method, system and device and a storage medium, and the method comprises the steps: obtaining and recognizing professional terms in user input, carrying out the LLM input enhancement through constructing scene description for each professional term, sending the enhanced user input into a pre-trained LLM to complete semantic analysis, and carrying out the semantic analysis. And performing multi-dimensional reordering on a semantic analysis result in combination with LLM initial confidence, context alignment and user preference, taking the paraphrase with the highest score as the optimal semantics input this time, performing ambiguity processing on the optimal semantics, completing a semantic analysis process, and realizing iterative optimization based on user feedback. According to the method, the accuracy of preliminary understanding is ensured, the interpretability and the preintervention of the whole term understanding process are improved, the method can adapt to habits and requirements of different users, and continuous iterative optimization can be performed to keep pace with the times.
Owner:NANJING NARI NETWORK SECURITY TECH CO LTD

Method of interpreting verbal input and electronic device thereof

A method of interpreting a spoken input may include assigning a meaning classification to the spoken input and assigning a confidence score to the meaning classification; and generating at least one paraphrase of the verbal input using at least one large language model (LLM) based on the confidence score corresponding to the meaning classification of the verbal input being less than or equal to a threshold; assigning the meaning classification to the at least one paraphrase and assigning the confidence score to the meaning classification; and concatenating the verbal input, the at least one paraphrase, the meaning classification, and the confidence score to generate a concatenated input; and inputting the cascaded input into the at least one LLM.
Owner:SAMSUNG ELECTRONICS CO LTD

Entity relation labeling and correction method based on semantic understanding

PendingCN122389868AParaphraseConditional sentence
The present application relates to entity relation annotation and correction method based on semantic understanding, and the technology comprises: obtaining a to-be-processed text, performing sentence division and semantic unit segmentation on the to-be-processed text, identifying entity mention, candidate relation predicate and semantic trigger; generating a relation proposition unit for a candidate entity pair, the relation proposition unit comprising a subject entity, an object entity and a candidate relation type; performing relation establishment determination on the relation proposition unit based on the semantic trigger, and determining a relation state label; when the relation state label represents relation establishment, outputting corresponding structured relation data as an entity relation annotation result; the present application generates a relation proposition unit for a candidate entity pair based on semantic understanding, first judges whether the relation actually exists in the current text context, thereby reducing the probability of false annotation in negative, speculation, paraphrase, conditional sentence and other scenarios, and improving the accuracy and semantic consistency of the entity relation annotation result.

Unified low sample relation extraction method and device based on multiple selection matching network

The application discloses a unified low sample relation extraction method and device based on a multiple-choice matching network. The method comprises the following steps: a common coding and matching mechanism based on a pre-training language model and multiple-choice marking relation description and relation instance; a triple obtained through open information extraction of large-scale pure text and a paraphrase text generated through a generative pre-training language model, and a triple-paraphrase pre-training mode based on the same; and an online meta-learning training mode based on a small sample under a new task. The mechanism based on the multiple-choice matching network can model various scenes in the low sample relation extraction task, and provides an efficient and fast network architecture, so that the model is more in line with the multiple requirements of model performance and speed in actual application.
Owner:INST OF SOFTWARE - CHINESE ACAD OF SCI

Text summarization extraction method and system

The present disclosure relates to a text summary extraction method and system. The method comprises: selecting M sentences from a given document containing L sentences to construct N candidate summaries; concatenating each candidate summary with the given document and inputting into a PLM to obtain N output vectors; inputting the N output vectors into a text paraphrase ranking model to obtain N paraphrase probabilities; selecting the candidate summary corresponding to the highest probability from the N paraphrase probabilities as the extracted text summary of the given document. The present invention converts the summary extraction task into a text paraphrase problem between the candidate summary and the source text, narrows the training gap between the summary extraction task and the PLM, and can better tap the knowledge of the PLM to improve the model performance. Further, relevant knowledge is learned from the existing text paraphrase rich training dataset by using knowledge transfer, which assists the model to identify the candidate summary that can better paraphrase the core semantics of the document, and makes up for the problem of missing training supervision signal caused by small-scale dataset.
Owner:ALIBABA (CHINA) CO LTD

Quality controlled paraphrase generation

A computer-implemented method including: receiving, as input, a dataset comprising training pairs (s, t), wherein each training pair comprises (i) a source sentence s and (ii) a target paraphrase t of the source sentences; at a training stage, training a machine learning model on the dataset, to obtain a trained quality-controlled paraphrase generator model, wherein during the training stage, each of the training pairs is associated with a predicted control vector representing a predicted paraphrase quality of the source sentence in the training pair; and at an inference stage, inferencing the trained quality-controlled paraphrase generator model on an input sentence, wherein the input sentence is associated with an input quality control vector, to obtain an output paraphrase of the input sentence which conforms to the quality control vector.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

A participant-intention-based dialogue summary generation method

The present application belongs to the technical field of natural language processing, and particularly relates to a dialogue summary generation method based on participant intention. In order to make the generated summary more comprehensive and complete, the method comprises the following steps: constructing a training data set of a dialogue summary model; analyzing and semantically representing the intention of a dialogue participant to construct an intention-based enhanced pseudo paraphrase data set; constructing an intention attention-based dialogue summary model; training the dialogue summary model, and finally inputting a new dialogue text into the summary model to generate a summary. The method can not only narrow the gap between the format and language style of dialogue text and structured text, but also effectively condense the intention of dialogue participants, improve the deep understanding of the model on the transformation of roles and language information in the dialogue, and generate high-quality dialogue summaries.
Owner:SHANXI UNIV

Method and system for generating legal reasoning thinking chain data and electronic equipment

The invention provides a legal reasoning thinking chain data generation method and system and electronic equipment, the generation method is applied to a large language model, and the generation method comprises the following steps: obtaining case fact content; determining a thinking map corresponding to the case fact content according to a preset table; wherein the thinking map comprises a target request right and a target request right element associated with the case fact content, and a target defense right and a target defense right element associated with the target request right; and generating legal reasoning thinking chain data corresponding to the case fact content according to the thinking map and the paraphrase of the target defense right in the preset table, wherein the thinking chain data comprises a thinking chain and a recommended defense right. According to the method and the device, the preset table is constructed, and the thinking map corresponding to the case fact content is obtained according to the preset table, so that the high-quality thinking chain data is obtained according to the thinking map and the preset table.
Owner:HUA DATA TECH (SHANGHAI) CO LTD

Entity-conditioned sentence generation

An example operation may include one or more of tuning a language model based on dependencies between an original data set and a paraphrase data set of the original data set, parsing and annotating the paraphrase dataset with entity identifiers of predefined entities to generate an annotated paraphrase dataset, additionally tuning the language model based on entity dependencies between the original data set and the paraphrase data set based on the annotated paraphrase dataset, and storing the additionally tuned language model in a storage device.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Method, system, and medium for detecting generated text based on lightweight paraphrase transformations

The application discloses a text generation detection method and system based on lightweight rewriting conversion and a medium. The method comprises the following steps: performing AI removal conversion on the to-be-detected text through a forward rewriting model; inputting the to-be-detected text and the generated human-like equivalent text into a classifier after fusion, and outputting a classification result; the training process of the forward rewriting model is as follows: a paired sample set is constructed by using artificially written text and machine-generated text; the forward rewriting model is fine-tuned by minimizing cross-entropy loss; the machine-generated text is obtained by at least one large language model and filtering redundant samples; the generation text detection method, system and medium completely do not need to access any external, closed-source or proprietary large language model API, which not only reduces the cost and delay, but also greatly simplifies the system architecture and deployment process, and solves the problem that the watermark technology and the similarity regeneration method are limited by the closed-source model and are difficult to be actually deployed.
Owner:ANHUI PROVINCIAL HOSPITAL

A paraphrase generation method, device, equipment and storage medium

The application discloses a kind of paraphrase generation method, device, equipment and storage medium, method includes obtaining first paraphrase generation corpus and word segmentation processing, the input word sequence X_1 and label word sequence Y_1 obtained are used as pre-training data set to train neural network model M;Second paraphrase generation corpus is obtained and is constructed knowledge base by neural network model M, so that the paraphrase generation knowledge contained in the second paraphrase generation corpus comprising first paraphrase generation corpus and incremental paraphrase generation corpus with timeliness exists in the form of key-value pair in knowledge base, the input word sequence X_3 obtained by third paraphrase generation corpus word segmentation processing is input into neural network model M to predict, obtain neural network prediction result and query vector;Query vector is used to retrieve knowledge base, and obtain retrieval result;Fusion neural network prediction result and retrieval result, generate final paraphrase text.Knowledge base makes paraphrase system effective iteration update, and generates paraphrase text with decision basis.
Owner:NANJING UNIV

Data processing method and device, electronic equipment, computer readable storage medium and computer program product

The application provides a data processing method and device, electronic equipment, computer readable storage medium and computer program product; the method comprises: performing word embedding processing on the first dialogue text to obtain a first vector corresponding to the first dialogue text; determining an intent label of the first dialogue text, and determining a second vector and a third vector based on the intent label, the second vector corresponding to the intent label, and the third vector corresponding to the paraphrase information of the intent label; determining a first fusion vector of the first dialogue text based on the first vector, the second vector and the third vector; and performing meta information prediction on the first dialogue text based on the first fusion vector to obtain a prediction result. Through the application, the accuracy of meta information prediction on the first dialogue text can be improved through multi-dimensional information complementation, thereby improving the accuracy of the prediction result.
Owner:MASHANG CONSUMER FINANCE CO LTD

A paraphrase sentence generation method based on template sentence enhancement

ActiveCN116628133BPart of speechParaphrase
The application discloses a kind of based on template sentence enhancement's paraphrase sentence generation method, steps are as follows: step 1: obtain original sentence, and obtain the target paraphrase sentence of original sentence in parallel corpus, use the method of rule to retrieve 1 sentence from parallel predata with the highest similarity with target paraphrase sentence as optimal sentence as reference;Step 2: different part-of-speech special symbols are used to replace noun, verb, adjective and adverb in optimal sentence to construct template sentence;Step 3: original sentence and template sentence are spliced through special symbol and joint, and the spliced sentence is used as the input of paraphrase generation model, and the paraphrase generation model outputs several paraphrase sentences;Step 4: the similarity of several paraphrase sentences and original sentence in grammatical structure and semantics is calculated, and the highest similarity paraphrase sentence is used as the optimal paraphrase sentence of original sentence.The application avoids the information dependence of model training phase in template sentence by masking the words of related part-of-speech in template sentence.
Owner:JIANGSU UNIV OF SCI & TECH +1