Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

6 results about "Single sentence" patented technology

A single sentence is a meaningful collection of words starting with a word with a capital first letter, having neccessary punctuation and ending with a full stop, an exclamation sign or a question mark.

A corpus processing method and system for large model search engines

PendingCN122332635ASingle sentenceSemantics
This application discloses a corpus processing method and system for a large-scale model search engine, relating to the field of data processing technology. The corpus processing method for a large-scale model search engine includes: obtaining a target word segmentation result based on a first sentence; obtaining the semantic contribution degree corresponding to each word element based on the target word segmentation result; filtering out word elements in the target word segmentation result whose semantic contribution degree is lower than a first preset value to obtain an optimized word segmentation result; obtaining the priority of the first sentence based on the optimized word segmentation result; and determining whether to input the first sentence as valid corpus into the large-scale model based on the priority. This application overcomes the limitations of traditional word segmentation and filtering, accurately identifying high-quality, low-frequency corpus with reliable sources and core semantics, preventing it from being overwhelmed by high-frequency, low-quality information, and simultaneously eliminating false, low-quality corpus based on keyword stuffing from the source, thereby improving the quality of the input corpus for the large-scale model.
Owner:ALIBABA TECHNOLOGY (GUANGZHOU) CO LTD

False keyword detection method, device, storage medium and program product

PendingCN122452578AFeature extractionAlgorithm
The application provides a false keyword detection method, device, storage medium and program product. The method comprises: based on a semantic boundary, performing splitting processing on a to-be-detected text to obtain a plurality of sentence units; performing feature extraction processing on each sentence unit to obtain a single sentence feature vector corresponding to each sentence unit; based on the single sentence feature vector corresponding to each sentence unit, determining a sentence sequence level hidden feature corresponding to each sentence unit and a full-text global feature of the to-be-detected text; determining a contradiction degree feature vector corresponding to each sentence unit according to the sentence sequence level hidden feature of each sentence unit; and determining a false keyword corresponding to the to-be-detected text according to the full-text global feature of the to-be-detected text and the contradiction degree feature vector corresponding to each sentence unit. The method improves the reliability of long text false keyword detection.
Owner:DAWNING CLOUD COMPUTING TECH CO LTD +1

Data processing method and device, equipment, storage medium and program product

PendingCN122334290ALinguistic modelSingle sentence
This application discloses a data processing method, apparatus, device, storage medium, and program product, relating to the field of artificial intelligence technology. The data processing method includes acquiring service dialogue data between customer service representatives and users, wherein the service dialogue data includes at least two consecutive messages from the same role in both the customer service representative and the user role; inputting the service dialogue data and a first prompt word into a first large language model to obtain a standard dialogue corpus output by the first large language model; using the first prompt word to constrain the first large language model to semantically integrate the at least two consecutive messages from the same role, forming a single-sentence dialogue text of alternating question-and-answer between the two roles; and constructing a training dialogue corpus based on the results of business topic verification and semantic structure verification in the standard dialogue corpus. This enables the rapid construction of a training dialogue corpus, solving the problem of low training efficiency.
Owner:CHINA UNIONPAY

A fake news oriented detection method and related device

The application provides a fake news detection method and related equipment, and belongs to the technical field of natural language processing and information authenticity detection. The method comprises the following steps: preprocessing and segmenting an original input text to obtain a sentence set and score a candidate claim set obtained; converting the sentence into a single sentence declarative expression to obtain a normalized claim, and structurally constructing the normalized claim; simultaneously performing information retrieval according to the structural construction mode to obtain a candidate evidence, performing multidimensional consistency verification on the structured claim and the candidate evidence, and scoring to obtain a total verification score; and correcting the total verification score by using the matching degree of a news site and an IP address, and outputting a verification result. The application prepositions noise processing and fact positioning, adopts a structured claim for fine-grained verification, and corrects through cross-information source matching of an IP address and a news site, so that the accuracy, interpretability and identification ability for numerical / time tampering of fake news detection are improved.
Owner:SOUTH CHINA UNIV OF TECH

A knowledge graph construction method based on fine-grained retrieval and reverse restoration self-correction

PendingCN122364468ALinguistic modelSorting algorithm
The application discloses a kind of knowledge graph construction methods based on fine-grained retrieval and reverse restoration self-error correction, specifically: first, standard reference case library is constructed, and standard reference vector is calculated.Then, the long text S to be measured is disassembled into sentences to be processed;Each sentence is vectorized, and the similarity score of its standard reference vector is calculated, and the first standard reference case of each sentence is screened;Through large language model, the non-standard logical relationship contained in each sentence is extracted and vectorized, the cosine similarity is calculated, and the candidate mapping is obtained by descending arrangement, after summarizing, double-feature reordering is carried out using multi-round greedy reordering algorithm, and the first standard reference case is screened out, and the single sentence is spliced into large language model, and triple extraction is carried out;Finally, the relationship triple set of all sentences constitutes knowledge graph.The application improves the precision and recall rate of triple extraction under the premise of avoiding redundancy.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Method and device for identifying malicious query intent of large model

PendingCN122451889ASingle sentenceTheoretical computer science
Embodiments of the present specification provide a method and device for malicious query intention recognition of a large model, the method comprising: obtaining a query sequence of a target user interacting with a large model, wherein the query sequence comprises a plurality of query sentences; for each query sentence, determining a single-sentence jump degree indicator according to a first similarity with an adjacent query sentence and a maximum similarity with the query sequence; determining a first indicator value reflecting logical coherence of the query sequence according to the jump degree indicator of each query sentence; dividing the query sequence into a plurality of topic clusters, any topic cluster comprising a plurality of continuous query sentences; determining a second indicator value reflecting an average follow-up depth of the query sequence according to at least the cluster size of each topic cluster; and determining whether the query sequence has a malicious query intention for the large model according to the first indicator value and the second indicator value. The effectiveness of malicious query intention recognition for a large model can be improved.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD