Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

16 results about "Sentence pair" patented technology

Pair in a sentence A pair of. Pair of aces. was one of a pair. A pair of hawks. Very rare pairing. It was pairing time. paired up and synced. Paired off and cocky. Two pairs of. They fly in pairs.

A text fake information detection method and system fusing contradictory features

ActiveCN120744125BPattern recognitionSentence pair
The application discloses a text false information detection method and system fusing contradictory features, and the method comprises the following steps: performing data preprocessing on a given input text, extracting text features, and extracting sentence pairs with a similarity higher than a threshold to form a similar sentence pair dataset; extracting contradictory word vector features, contradictory scene features and contradictory semantic features in the given input text based on the similar sentence pair dataset; fusing the text features, the contradictory word vector features, the contradictory scene features and the contradictory semantic features, weighting through a self-attention mechanism, and obtaining a weighted and distributed feature fusion vector; and performing false information detection based on the feature fusion vector, and obtaining a false information detection result. The application can fuse contradictory features and style statistical features of the text to perform false information detection, and can effectively improve the accuracy of text-based false information detection.
Owner:NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP

Low-resource language translation method and system based on inference and retrieval fusion large model

PendingCN122452589AAlgorithmSentence pair
The application provides a large model low-resource language translation method and system based on reasoning and retrieval fusion, which comprises the following steps: obtaining enhancement information, which contains retrieved similar parallel sentence pairs and extracted auxiliary knowledge; performing first-stage implicit correction, constructing a comparative demonstration context containing retrieved example source sentences, auxiliary knowledge, initial translation and standard reference translation, inputting the large model with the to-be-translated source sentence and initial translation, and generating first-stage corrected translation; performing second-stage explicit correction, performing error detection and labeling on the first-stage corrected translation based on the large model according to multi-dimensional quality measurement standards, generating structured feedback, and inputting the large model again to generate second-stage corrected translation. The application effectively improves the accuracy and robustness of low-resource language translation.
Owner:XINJIANG UNIVERSITY

Semantic reasoning method and device, electronic equipment, and storage medium

This application provides a semantic reasoning method, apparatus, electronic device, and storage medium, belonging to the field of artificial intelligence technology. The method includes: acquiring target text information; dividing the target text information into sentences; encoding each sentence to obtain multiple lexical features; performing semantic extraction based on the multiple lexical features of each sentence to obtain a first semantic feature of each sentence; reasoning based on the first semantic features of each sentence using a large language model to obtain a second semantic feature corresponding to each sentence; decoding the second semantic features of each sentence to obtain a semantic reasoning result for each sentence; and using the semantic reasoning results of each sentence as the semantic reasoning result for the target text information. The semantic reasoning method, apparatus, electronic device, and storage medium provided in this application can improve the computational efficiency of large language model reasoning.
Owner:BEIJING UNIV OF POSTS & TELECOMM

A subtitle processing method, system, electronic device and storage medium

The application provides a subtitle processing method and system, electronic equipment and a storage medium, which are applied to the technical field of data processing, and multiple sentences and original subtitle sentence time lengths thereof are generated according to a subtitle file; wherein a sentence is composed of at least one piece of subtitle, and the original subtitle sentence time length is the sum of the original subtitle time lengths of all subtitles constituting the sentence; the sentence is translated to obtain a translation of the sentence; based on the sentence and the original subtitle sentence time length thereof, the translation of the sentence is split with minimum deviation to obtain a splitting scheme of the sentence; wherein the splitting scheme comprises a translation fragment of each piece of subtitle corresponding to the sentence; the translation fragment of each piece of subtitle is dubbed, and the dubbing speed of the translation fragment dubbing of each piece of subtitle is adjusted based on the original subtitle time length of the subtitle, so that the purpose of high-precision time alignment is achieved under the condition that the sentence is not fragmented, the semantics is coherent, the semantics is complete, and the watching process is ensured.
Owner:HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD

A subtitle line level alignment method, electronic device and storage medium

PendingCN122334192ASemantic vectorAlgorithm
This application provides a subtitle line-level alignment method, electronic device, and storage medium, including comparing the number of sentences in two text segments; intelligently segmenting the target language text according to its linguistic habits; using a pre-trained large-scale bilingual semantic embedding model to convert each sentence of the source and target language texts into a high-dimensional semantic vector; calculating the semantic similarity matrix between all sentence pairs and finding an alignment path with optimal semantic similarity globally; introducing a length ratio penalty factor to balance semantic similarity and sentence length reasonableness; if blank lines exist, locating the text line that produces the blank line and its adjacent areas, segmenting the preceding sentence of the target language text according to common punctuation marks of that language to organize sentence fragments and eliminate blank lines. This method adopts a "coarse-to-fine, progressively layered" processing strategy, ultimately outputting an alignment result without blank lines and with reasonable semantic matching.
Owner:FUZHOU CHANGXIN INFORMATION TECH CO LTD

A semantic fault-tolerant large-scale exploration and development BERT classification method

The application provides a large-scale exploration and development BERT classification method based on semantic fault tolerance, which comprises the following steps: applying a certain proportion of random noise to literature and expanding the corpus; adopting a BERT algorithm to realize context-related first-order classification according to the expanded corpus, and obtaining classified sentences; and adopting an open-source Jieba word segmentation module to perform word segmentation on the classified sentences. Through input text, the number of sentences is expanded while the text chapter structure is kept unchanged, and the BERT algorithm is adopted to take a sentence pair as input, so that the memory of the 1-type grammar to the context before and after the sentence is realized, and the understanding of the chapter structure knowledge is indirectly realized.
Owner:CHINA PETROLEUM & CHEMICAL CORP +1

Method for constructing myanmar-chinese parallel corpus based on multi-step thinking large model

This invention relates to a method for constructing a large-scale Burmese-Chinese parallel corpus based on multi-step thinking. The invention includes: translating existing Chinese-English parallel corpora using currently available translation models to obtain original English-Burmese and Chinese-Burmese parallel sentence pairs; calculating double-confidence intervals for semantic similarity and perplexity between aligned Burmese-Chinese sentence pairs based on publicly available high-quality Burmese-Chinese parallel corpora; using the selected double-confidence intervals to perform preliminary screening of Burmese-Chinese parallel sentence pairs, forming pre-processed pseudo-parallel sentence pairs; designing a multi-step thinking chain to guide the large-scale model to progressively optimize the pre-processed pseudo-parallel sentence pairs, thereby generating high-quality Burmese-Chinese parallel corpora for training the translation model, thus effectively improving the performance of Burmese-Chinese machine translation. This invention significantly enhances the ability of large language models to construct corpora in Burmese, a low-resource language, and provides an interpretable and transferable technical paradigm for corpus construction in other low-resource languages.
Owner:KUNMING UNIV OF SCI & TECH +4

Sentence semantic clustering compression method based on heterogeneous kv cache, electronic device, program product

ActiveCN121980031BSemantic representationContextual reasoning
The application discloses a sentence semantic clustering compression method based on a heterogeneous KV cache, electronic equipment and a program product. The method comprises the following steps: reserving the first T tokens of a given query on a GPU, and splitting the remaining tokens to obtain S sentences; for each sentence, taking the mean of the Key vector corresponding to the token as the sentence center, calculating the similarity between the Key vector corresponding to each token and the sentence center, calculating the GSA weight of each token according to the similarity, and calculating the semantic representation according to the GSA weight; and clustering the semantic representations of the S sentences to obtain C cluster representations. The application improves the accuracy and efficiency of long context reasoning, and reduces the long sequence reasoning delay and memory pressure under the limited KV budget.
Owner:UESTC (SHENZHEN) ADVANCED RES INST

A method, apparatus, device, medium and product for attributing session satisfaction

This invention discloses an attribution method, apparatus, device, medium, and product for session satisfaction. The method includes: determining the session satisfaction of the session data to be tested based on the session data to be tested and a pre-trained session satisfaction model; determining the influence weight of each sentence in the session data to be tested on the session satisfaction based on the session satisfaction model; filtering the session data to be tested based on the influence weight of each sentence to obtain at least two filtered sentences; determining a target satisfaction factor for each filtered sentence based on the at least two filtered sentences and a pre-trained satisfaction factor model; and determining the attribution data corresponding to the session data to be tested based on the influence weight of the at least two filtered sentences and the target satisfaction factor. This method solves the problem of unclear guidance direction of attribution data leading to scattered optimization resources and improves the optimization efficiency of session service quality.
Owner:JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD +1

Machine translation polysemous word translation evaluation method based on semantic item trigger word replacement

PendingCN122088522AReveal errors effectivelyImprove targetingNatural language translationSemantic analysisSentence pairWord sense
Aiming at the defects of an existing test method in the field of word sense disambiguation test of machine translation in the aspects of triggering effective word sense conversion, maintaining original sense and maintaining semantic consistency, the invention provides a machine translation polysemous word translation evaluation method based on semantic item trigger word replacement. The method comprises the following steps: firstly, constructing a polysemy semantic item library, and screening source language sentences containing polysemy from a parallel corpus according to the semantic item library; performing natural language processing on the sentence, and identifying a trigger word of a polysemy special definition item in the sentence; replacing the trigger word with a replacement word to generate a variant sentence; performing machine translation and alignment on the original sentences and the variant sentences to obtain expressions of polysemy words in translations; and calculating the similarity of the polysemy translation expression, and judging whether a polysemy disambiguation error exists or not. Through a trigger word replacement strategy, a semantic controlled contrast sentence pair can be constructed, potential errors of a system in polysemous word translation are effectively revealed, and the pertinence and effectiveness of evaluation are improved.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

A grammar error correction large model training method and device based on reinforcement learning, equipment and medium

PendingCN122114048ASemantic analysisBiological modelsGrammatical errorAlgorithm
The application provides a grammar error correction large model training method and device based on reinforcement learning, equipment and medium, relating to the technical field of data processing. The method comprises the following steps: obtaining an initial grammar correction corpus containing error sentences and corresponding correct sentence pairs, processing the initial grammar correction corpus to generate a reasoning correction training set, adjusting a first preset language model according to the reasoning correction training set to obtain an initial strategy model, performing reinforcement learning training on the initial strategy model based on a preset composite function by using a reinforcement learning algorithm, and finally obtaining a target grammar correction model. The target grammar correction model can fully utilize the reasoning ability, effectively improve the performance of the model in terms of precision and recall, better meet the actual needs of grammar correction, and generate correction results containing reasoning processes, which provides more transparent and interpretable basis for the correction process.
Owner:BEIJING FOUNDER ELECTRONICS CO LTD

Generation of imposition type lists for input text

Systems and methods are provided that include a processor executing a program to receive input text, divide the input text into sentences, generate and output sentence embeddings using a sentence embeddings encoder based on the sentences, identify matches in the input text with geographical jurisdictions listed in a table comprising imposition types and their corresponding geographical jurisdictions, generate and output an imposition types list based on the matches identified in the input text, generate and output candidate embeddings, perform a cosine similarity search between the candidate embeddings and the sentence embeddings to generate and output a scored sentence list of sentences comprising cosine similarity scores corresponding to respective top scoring imposition types for each of the sentences in the input text, aggregate the cosine similarity scores in the scored sentence list to determine and output the top scoring imposition types in the input text.
Owner:VERTEX INC

Method for deriving new rule for data augmentation in natural language inference task and system therefor

PendingUS20260148104A1Biological modelsKnowledge representationData setNatural language inference
A method and system are provided for deriving a new rule for data augmentation in a natural language inference task. The method for deriving a new rule according to some embodiments may include acquiring a base dataset for a natural language inference task, acquiring a rule detection model that detects an existing rule conforming to a premise-hypothesis sentence pair input from an existing rule set, and selecting a plurality of premise-hypothesis sentence pairs that does not conform to the existing rule set from the base dataset by performing out-of-distribution (OOD) detection based on the rule detection model on the base dataset. In this case, the selected premise-hypothesis sentence pairs may be used to derive the new rule set, and a high-quality augmented dataset for the natural language inference task can be easily generated through this new rule set.
Owner:KNU IND COOPERATION FOUND

A Summary Reordering Method and System Combining Extractive and Generative Approaches

This invention relates to a summary reordering method and system combining extractive and generative methods, comprising the following steps: obtaining a first article containing multiple first sentences and multiple original summary sentences; extracting each first sentence to obtain multiple first extracted sentences, the first extracted sentences being iconic sentences in the first article; determining a first summary sentence corresponding to each first extracted sentence based on each first extracted sentence and each original summary sentence, the first summary sentence being a summary predicted based on the first extracted sentence; determining a token corresponding to each first summary sentence based on each first summary sentence, the token being a character or a word; and determining a first target summary based on each token. This application achieves higher accuracy in the generated summaries.
Owner:GUILIN UNIV OF ELECTRONIC TECH