Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

21 results about "Sentence extraction" patented technology

Sentence extraction is a technique used for automatic summarization of a text. In this shallow approach, statistical heuristics are used to identify the most salient sentences of a text. Sentence extraction is a low-cost approach compared to more knowledge-intensive deeper approaches which require additional knowledge bases such as ontologies or linguistic knowledge. In short "sentence extraction" works as a filter which allows only important sentences to pass.

Keyword guidance and large language model near-end strategy optimization combined key sentence extraction method

The invention provides a key sentence extraction method combining keyword guidance and large language model near-end strategy optimization. The key sentence extraction method comprises the following steps: constructing a keyword-key sentence pair; the relevancy is evaluated by using a joint matching model, and a reward value is generated; introducing a KL divergence to measure the difference between the training model and the reference model, and estimating the value score of the current state in combination with a state value network; and optimizing the guidance model through a near-end strategy to realize extraction of key sentences. According to the method provided by the embodiment of the invention, the keyword-key sentence pair is constructed, the relevancy of the keyword-key sentence pair is evaluated by using the joint matching model, the reward value is generated, the KL divergence is introduced to measure the difference between the training model and the reference model, the value score of the current state is estimated by combining the state value network, and the guidance model is optimized through the near-end strategy. According to the method, the key sentences are extracted, the problems of high dependence on annotation data and high training cost of a large language model are solved to a great extent, and the key sentence extraction effect is remarkably improved.
Owner:BEIJING INFORMATION SCI & TECH UNIV

System and method of test case processing framework

A system and method are disclosed of automatically generating test cases for automation testing. Embodiments include receiving input data written in plain text as input from one or more users, the input data comprising one or more sentences, extracting one or more sentences and storing each sentence into an array of a database, extracting parameters and instructions from the sentences, determining if there are regular expressions of the input data and actions to be performed on a browser stored in the database and matching the regular expressions with the actions.
Owner:BLUE YONDER GROUP INC

Audio text abstract generation method and device and computer equipment

PendingCN121388615ASemantic analysisBiological modelsSentence extractionFeature data
The invention relates to the technical field of text abstract generation, in particular to an audio text abstract generation method, which comprises the following steps of: performing multi-scale segmentation on acquired audio, performing text generation by using acoustic characteristic data of audio clips subjected to multi-scale segmentation, and performing text abstract generation by adopting a text-time alignment optimization mechanism based on attention weight. According to the method, the generated text is aligned and adjusted, the aligned and adjusted text is obtained, keyword extraction is performed on the aligned and adjusted text, target sentence extraction is performed according to the extracted keywords, abstract text generation is performed by utilizing the extracted keywords and the target sentences, and the abstract text generation accuracy is improved.
Owner:SOUTH CHINA NORMAL UNIV

Text summarization extraction method based on joint training and corresponding device

The application provides a text summary extraction method based on joint training and a corresponding device, which are used to improve the problem of insufficient semantic correctness of the extracted summary text. The method comprises the following steps: obtaining a to-be-processed text, and performing sentence segmentation on the to-be-processed text to obtain a plurality of to-be-processed sentences; using a vector extraction layer in a summary extraction model to perform vectorization representation on the plurality of to-be-processed sentences to obtain word vectors and sentence vectors corresponding to the plurality of to-be-processed sentences; using a feature extraction layer in the summary extraction model to perform feature extraction on the word vectors and the sentence vectors corresponding to the plurality of to-be-processed sentences to obtain core feature vectors and similar feature vectors; and using a sentence extraction layer in the summary extraction model to extract the plurality of to-be-processed sentences according to the core feature vectors and the similar feature vectors to obtain a summary text corresponding to the to-be-processed text.
Owner:ZHONGKE DINGFU BEIJING TECH DEV

A structured outline driven scientific report automatic generation method and system

The application provides a structured outline driven scientific report automatic generation method and system, and relates to the technical field of natural language processing, the method comprises the following steps: performing semantic analysis and matching on a user input research topic to generate a first-level structure node, extracting and dividing research problems from multiple standard scientific literatures to construct a second-level structure node, performing opinion sentence extraction and refinement to generate a third-level structure node, generating a structured outline according to the first-level, second-level and third-level structure nodes and corresponding explicit binding relationships, performing task analysis, generating a task set according to the hierarchical order of the structured outline, generating a target scientific report and outputting the target scientific report by performing text generation, assembling and integrating. The technical problems of unstable structure, easy divergence of theme, difficulty in corresponding paragraphs and evidence, and easy repetition of chapter content in the existing scientific report automatic generation are solved. The technical effect of stable structure, focused theme and traceable content of the scientific report automatic generation is achieved.
Owner:DOCUMENT & INFORMATION CENT OF CHINESE ACAD OF SCI

A method and system for analyzing appeals of passenger normalization problems of urban rail transit, and a storage medium

PendingCN122334658AMulti source dataSentence extraction
This invention relates to a method for analyzing routine passenger complaints in urban rail transit, comprising: S1, multi-source data collection; S2, passenger complaint analysis; and S3, early warning mechanism construction and automatic departmental push notifications. This invention also relates to a system for analyzing routine passenger complaints in urban rail transit, comprising a data collection module, a passenger complaint analysis module, and an early warning module. This invention addresses the problem of insufficient data dimensions in existing systems, significantly improving the comprehensiveness of analysis results; based on large language models and deep semantic analysis technology, it can accurately identify the true content and sentiment of passenger complaints, achieving high-precision classification and key sentence extraction; it improves the efficiency of operating units in handling complaints; and it presents overall complaint trends through visual charts, providing real-time decision-making basis for operation management.
Owner:TIANJIN ACAD OF TRANSPORTATION SCI

Space-time knowledge graph construction method based on large language model and Beidou grid coding

The invention discloses a space-time mapping knowledge domain construction method based on a large language model and Beidou grid coding. The method comprises the steps that semantic partitioning is conducted on an input text through the large language model LLM, fact declarative sentence extraction is conducted on each block, and triple extraction, time entity extraction and space geographic entity extraction are conducted on the fact declarative sentences respectively; performing time attribute coding on the time entity; constructing a gridded Beidou satellite geographic map and obtaining a Beidou grid coding data set, and performing associated grid matching coding on spatial geographic entities to form corresponding spatial attributes; and collecting the time entity containing the time attribute, the space geographic entity containing the space attribute and the triple to construct the space-time knowledge graph. According to the method, attribute coding is carried out on the time entities and the space geographic entities, the coded time entities, the coded space geographic entities and the triple are collected and constructed to obtain the space-time knowledge graph, space-time information in the entities can be deeply mined, and the space-time reasoning capability is improved.
Owner:ZHEJIANG SHIZIZHIZI BIG DATA CO LTD

Image subtitle algorithm based on structured semantic extraction and geometric feature fusion

The invention provides an image subtitle algorithm based on structured semantic extraction and geometric feature fusion. The image subtitle algorithm comprises a visual semantic feature extraction module and a subtitle generation module, the visual semantic feature extraction module comprises a regional feature, a network feature, a structured semantic feature and the subtitle generation module, firstly, the most similar text sentence is retrieved from an image through a CLIP model, concept semantics and attribute features are extracted, and after Top-K text features are retrieved through cosine similarity, multi-level clustering is carried out, and finally, the subtitle generation module is used for generating subtitles; and the isolation problem of semantic information extraction is fundamentally solved. And secondly, gradually fusing the grid features, the regional features and the structured semantic features by designing a geometric perception semantic enhancement encoder, and introducing geometric coordinate information in the fusion process, thereby enhancing the expression ability of the structured semantic features. According to the process, the modeling capability of the model for the image space relationship is remarkably enhanced, and the understanding of global semantics is improved.
Owner:WUHAN ZHENGYUAN ELECTRIC

Information retrieval and management method for construction of land management policy knowledge graph

The invention discloses an information retrieval and management method for construction of a land management policy knowledge graph. The method comprises the following steps: performing clause and fact declarative sentence extraction on policy clauses in a structured land policy database by utilizing a large language model, and taking a time entity and a spatial geographic entity as attributes of the policy clauses; constructing a time entity containing a time attribute, a space geographic entity containing a space attribute and a space-time knowledge graph database of a triple; performing semantic partitioning on the query data by using a large language model to obtain a query combination; and creating a plurality of parallel retrieval paths matched with the query combination by utilizing a space-time retrieval planner, summarizing search results, and then calling and summarizing the retrieval results containing a plurality of policy terms from the structured land policy database. According to the method, the precision and generalization of complex space-time semantic understanding are improved, multi-dimensional collaborative retrieval optimization is realized, and the method has remarkable effects on policy query, execution feedback and management efficiency.
Owner:NINGBO NATURAL RESOURCES & PLANNING RESEARCH CENTER +1

Premium collection call script extraction and analysis method, apparatus, device, and medium

ActiveCN114281995BFinanceBiological modelsEngineeringSentence extraction
The application provides a premium collection script extraction and analysis method, device, equipment and medium, text data of premium collection telephone recording is obtained, and the text data is classified to screen out conversation text of no renewal willingness category; each sentence in the conversation text is subject classified to analyze theme missing rate and sequence difference between conversation texts of each seat with different success rates; meaningless sentences are deleted, keywords of each sentence are extracted to analyze word difference of keywords in the same theme in conversation texts of each seat with different success rates, and / or conversation texts of each seat with different success rates under each theme are scored to summarize high-score conversation texts. The application can more easily locate excellent scripts and words in the renewal retention business, thereby improving the success rate of seat telephone renewal.
Owner:AIA LIFE INSURANCE CO LTD

Extraction device, extraction method, and extraction program

It makes it easy to use content including text and images as input for chatbot systems that utilize RAG. [Solution] In the extraction device (10), a separation unit (15b) separates content to be processed into sentence objects and image objects and acquires position information for each of the sentence objects and image objects. A sentence extraction unit (15c) extracts text data from the sentence objects. An image extraction unit (15d) extracts text data from the image objects. A generation unit (15e) generates structured sentences including text data arranged in order according to the position information for each of the sentence objects and image objects.
Owner:NTT DOCOMO BUSINESS INC

Long-term care certification support system, long-term care certification support method, and long-term care certification support program

To provide a long-term care certification support system, a long-term care certification support method, and a long-term care certification support program that can reduce workload in a long-term care certification survey and increase accuracy of determination.SOLUTION: A long-term care certification support system comprises a terminal for inputting a certification survey form consisting of a "basic survey" and "special notes" and for outputting a notification of a condition of a subject, an artificial intelligence unit for analyzing the condition of the subject, and a server for estimating the condition of the subject. The artificial intelligence unit includes: a morphological analysis unit for performing morphological analysis of an analysis target in the "special notes"; an inferred basis sentence extraction unit for extracting sentences that serve as the basis for estimation from sentences in the "special notes"; a normalization unit for normalizing expressions of the sentences extracted by the inferred basis sentence extraction unit; a classification estimation unit for estimating a classification number based on the sentences normalized by the normalization unit; and an answer generation unit for generating an answer that corresponds to criteria of the certification survey form based on the classification number estimated by the classification estimation unit.SELECTED DRAWING: Figure 1
Owner:株式会社NTTデータ东北

Role recognition method and device based on speaker segmentation

The present invention provides a method and device for role recognition based on speaker segmentation. The method comprises: converting the conversational speech to be recognized into text data, segmenting the text data into multiple sentences based on a sentence segmentation model, and extracting the text features of each sentence; segmenting the conversational speech to be recognized, obtaining an audio segment corresponding to each sentence, and extracting the acoustic features of each audio segment; aligning the text features and acoustic features corresponding to each sentence based on an attention mechanism to generate an alignment vector corresponding to each sentence; and obtaining the speaker category corresponding to each sentence based on a classification model based on the alignment vector, text features, and acoustic features corresponding to each sentence. The present invention utilizes the interactive features between text and audio for role recognition, improving the accuracy of role recognition.
Owner:CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1

Generation of media segments from larger media content for media content navigation

PendingUS20250234074A1Sentence extractionSemantic similarity
A method is described and includes obtaining a list of utterances comprising captions from an item of content; computing sentence transformer embeddings for each of the utterances; dividing the utterances into sentences and extracting a sentence embedding for each sentence; computing a semantic similarity between adjacent sentences; and merging the adjacent sentences into a block comprising a segment if the semantic similarity between the adjacent sentences is greater than a predetermined threshold.
Owner:ROKU INC

File similarity detection method based on Simhash fusion keyword and key sentence extraction

The invention belongs to the technical field of computer information processing, and discloses a file similarity detection method based on Simhash fusion keyword and key sentence extraction, which comprises the following steps of: firstly, carrying out text preprocessing and dual-granularity feature extraction to respectively obtain keywords and key sentences with weights; secondly, providing a weight fusion strategy, and contributing keyword weights to key sentences according to the occurrence frequency to form fusion weights; secondly, optimized SimHash feature coding is carried out, 64-bit fingerprints are independently generated for keywords and key sentences by adopting grouping hash functions based on different seeds, and the 64-bit fingerprints and the key sentences are spliced into a 128-bit final feature fingerprint; and finally, the similarity is judged by calculating the Hamming distance between fingerprints of different files. According to the method, through word and sentence dual-granularity feature fusion and grouped Hash coding, the detection precision and robustness of an algorithm on synonymous replacement, word order adjustment and paragraph recombination are remarkably improved.
Owner:SHENZHEN SHIXI TECH CO LTD

Entity spatio-temporal information and topological relation retrieval method based on spatio-temporal knowledge graph

The invention discloses an entity spatio-temporal information and topological relation retrieval method based on a spatio-temporal knowledge graph, and the method comprises the steps: constructing a knowledge source database, and carrying out spatio-temporal knowledge graph processing to obtain a spatio-temporal knowledge graph database, the spatio-temporal knowledge graph comprising a time entity, a spatial geographic entity and a triple; the method comprises the following steps of: carrying out semantic partitioning on input query data by utilizing a large language model (LLM), carrying out fact declarative sentence extraction on each block, and extracting a query combination comprising a triple, a time entity and a spatial geographic entity from the fact declarative sentences; and a space-time retrieval planner is utilized to create a multi-path parallel retrieval path matched with the query combination, the multi-path parallel retrieval path comprises triple retrieval, time retrieval and spatial geographic retrieval, and the multi-path parallel retrieval path is utilized to perform multi-path parallel search in the space-time knowledge graph database and summarize and output retrieval results in a layered manner. According to the method, the precision and generalization of complex space-time semantic understanding are improved, and multi-dimensional collaborative retrieval optimization is realized.
Owner:ZHEJIANG SHIZIZHIZI BIG DATA CO LTD

Method, system and equipment for generating report lecture based on large model and medium

The invention provides a report lecture generation method, system and device based on a large model and a medium, and belongs to the technical field of natural language process.The method comprises the steps that a data document and a discourse document input by a user are received; analyzing the data document and calling a large model to generate a first segment in combination with a first segment cue word template; integrating the discourse document and calling a large model to extract a central sentence in combination with a central sentence extraction cue word template; performing RAG retrieval on the discourse document based on the central sentence, determining a related fragment, combining with a paragraph to generate a cue word template, and calling a large model to generate a paragraph; the first segment and the paragraph are spliced, an ending prompt word template is combined, the large model is called to generate an ending, and the first segment, the paragraph and the ending are spliced to obtain a lecture; and performing fine tuning on the large model based on the historical conference text and the superior department problem data set, inputting the lecture into the fine-tuned large model, outputting a simulation problem and modifying the lecture. The method improves the writing efficiency of the report lecture, reduces the manual workload, and improves the accuracy and speciality of the lecture.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

A Summary Reordering Method and System Combining Extractive and Generative Approaches

This invention relates to a summary reordering method and system combining extractive and generative methods, comprising the following steps: obtaining a first article containing multiple first sentences and multiple original summary sentences; extracting each first sentence to obtain multiple first extracted sentences, the first extracted sentences being iconic sentences in the first article; determining a first summary sentence corresponding to each first extracted sentence based on each first extracted sentence and each original summary sentence, the first summary sentence being a summary predicted based on the first extracted sentence; determining a token corresponding to each first summary sentence based on each first summary sentence, the token being a character or a word; and determining a first target summary based on each token. This application achieves higher accuracy in the generated summaries.
Owner:GUILIN UNIV OF ELECTRONIC TECH

Electronic apparatus for comparing similarity between texts on basis of sub-sentence extraction and method for comparing similarity between texts by using same

An electronic apparatus for comparing a similarity between texts on the basis of sub-sentence extraction, according to the present disclosure, may comprise: a memory storing at least one instruction; and at least one processor that executes the at least one instruction. The at least one processor may: perform analysis of dependence between a plurality of morphemes included in at least two pieces of text data; generate a dependent relationship graph for the plurality of morphemes, on the basis of the dependence analysis; extract at least one sub-sentence by using the dependent relationship graph; and calculate a similarity between the at least two pieces of text data by cross-comparing the at least one sub-sentence.
Owner:MUHAYU INC

Document-level Relation Extraction Method and System Based on Heuristic Evidence Sentence Extraction and Entity Representation Enhancement

The present invention belongs to the technical field of natural language processing, and particularly relates to a document-level relation extraction method and system based on heuristic evidence sentence extraction and entity representation enhancement. First, according to predefined heuristic rules, mentions in the original document that interact with the head and tail entities of the target entity pair are selected, the sentences where the entity mentions are located are used as the evidence sentences for the target entity pair, and the evidence sentences are constructed into a pseudo-document in the order of the original document; a pre-trained language model is used to learn the context information related to the target entity pair in the pseudo-document; the mentions of the head and tail entities interacting with the target entity pair and the relevant context are used to learn different entity representations of the same entity in different entity pairs; for different entity representations, an activation function is used to predict the relation type of the target entity pair. The present invention uses simple predefined rules to extract the evidence sentences for entity relation prediction and constructs them into a pseudo-document in the order of the original document, reducing the complexity of entity relation prediction and improving the performance of document-level relation extraction.
Owner:Chinese People's Liberation Army Cyberspace Force Information Engineering University

Sentence generation device, sentence generation learning device, sentence generation method, sentence generation learning method, and program

To enable information in need of consideration when a sentence is generated to be added as a text.SOLUTION: A sentence generation device includes a content selection unit and a generation unit. The content selection unit extracts a set of words based on an input sentence, a degree of importance of each word included in the input sentence, and an output length. The generation unit generates an output sentence based on the input sentence and the set of words by, with the input sentence and the set of words as input, inputting information corresponding to the input sentence and the set of words into a machine learning model in which supervised learning has been performed so as to generate the output sentence based on the input sentence and the set of words.SELECTED DRAWING: Figure 2
Owner:NIPPON TELEGRAPH & TELEPHONE CORP