Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

30 results about "Paragraph" patented technology

A paragraph (from the Ancient Greek παράγραφος paragraphos, "to write beside" or "written beside") is a self-contained unit of a discourse in writing dealing with a particular point or idea. A paragraph consists of one or more sentences. Though not required by the syntax of any language, paragraphs are usually an expected part of formal writing, used to organize longer prose.

Global fluctuation analysis methods, devices, and storage media for text

This application discloses a method, device, and storage medium for global fluctuation analysis of text. The method, relating to the field of text parsing technology, includes: processing the input text using a first model to obtain a first parsing result, and processing the input text using a second model to obtain a second parsing result; determining the sentence-level raw evaluation value of each sentence in the input text; determining the paragraph-level raw evaluation value of each paragraph in the input text based on the first parsing result, the second parsing result, and the sentence-level raw evaluation value; connecting the paragraph-level raw evaluation values ​​of each paragraph into an expression intensity curve using the paragraph sequence of the input text as a time axis; and identifying and labeling four geomorphic features—peak, valley, data slope, and data plain—based on the expression intensity curve to generate a visualized scoring geomorphic map. This application achieves accurate prediction and hierarchical selection of training corpus quality before model training.
Owner:LINGE TECHNOLOGY CO LTD

A PDF document content recognition method, device, equipment and storage medium

The application provides a PDF document content recognition method and device, equipment and a storage medium, the method comprises the following steps: obtaining an access link of an unanalyzed document; downloading a target PDF document from an object storage service according to the access link; when the content area corresponding to the page type in the target PDF document is a text area, the content area is divided into a text area, a table area and an image area according to the page type corresponding to each content area; when the content area is a text area, the original text character sequence is extracted from the text area, and the semantic coherent paragraph text is generated by similarity calculation on the original text character sequence. Based on the cosine similarity, the application automatically divides the sentences with similar semantics in the text area into the same text block, so that the document analysis obtains semantic coherent and complete text content, solves the problem of broken text paragraphs after document recognition, and effectively improves the semantic coherence of the document content.
Owner:SHENZHEN ISSMART SCI & TECH CO LTD

Synonym generation method and device, equipment and computer readable storage medium

ActiveCN116227473BPart of speechEngineering
The present disclosure provides a synonym generation method, device and equipment and a computer readable storage medium. The method comprises: segmenting historical text according to multiple themes to obtain multiple theme paragraph texts; each theme paragraph text corresponds to a theme; according to a combination of parts of speech in a preset business label word, filtering theme core words matching the combination of parts of speech from the multiple theme paragraph texts, each theme core word being used to represent a text feature of a theme paragraph text; and generating a synonym library corresponding to the preset business label word according to the theme core words corresponding to each theme paragraph text. According to the embodiment of the present disclosure, the accuracy of synonym mining can be improved.
Owner:MASHANG CONSUMER FINANCE CO LTD

Constructing a document hierarchy tree using machine learning models

In various examples, a technique for generating a hierarchical representation of a document includes generating, via a first machine learning model, a hierarchical structure associated with a document, wherein the hierarchical structure includes one or more headings and one or more paragraphs. The technique also includes identifying, via the first machine learning model, heading text included in the document and associated with each of the one or more headings and paragraph text included in the document and associated with each of the one or more paragraphs. The technique further includes generating, via a second machine learning model and based at least on the identified heading text included in the document, a formatted listing including the heading text associated with each of the one or more headings and generating a hierarchical document based at least on the formatted listing and the paragraph text associated with the one or more paragraphs.
Owner:NVIDIA CORP

Intelligent display voice-to-text method and system, and medium

ActiveCN116320614BSearch wordsDocument transformation
The application discloses an intelligent display voice-to-text method and system and a medium, and the method comprises the following steps: when an audio and video file is captured, the audio and video file is converted into text; each sentence of text is segmented according to a pre-agreed punctuation mark to obtain segmented sentence text information; a pre-established dictionary table is searched; if historical data corresponding to the segmented sentence text information is found in the dictionary table, a corresponding previously combined text paragraph is obtained from the dictionary table; the text paragraph is rendered; single characters in the rendered text paragraph content are processed according to sensitive words, forbidden words or search words; and the processed single characters are synchronized with the text paragraph on a page for display. Through the application, a user can quickly find out where a violation is located, and can also quickly jump to a corresponding progress for manual auditing to confirm whether a problem exists, thereby greatly reducing the workload of the user for compliance processing.
Owner:SHENZHEN CRAFTSMAN NETWORK TECH CO LTD

A long document outline extraction method, system and computer device

This invention discloses a method, system, and computer device for extracting outlines from long documents, belonging to the field of document processing technology. The method includes: S100, reading the document to be processed; S200, processing the document to obtain text blocks and recording the positions of the text blocks within the document; S300, inputting the text blocks into a large language model to generate corresponding result blocks, where each result block is the outline of the current text block; S400, repeating step S300 until all result blocks corresponding to all text blocks are obtained, and then concatenating all result blocks to obtain the document's outline. This invention employs a multi-level processing strategy, segmenting at the paragraph level, and, when encountering extremely long paragraphs, performing precise segmentation at the sentence level. This avoids truncation and information loss due to excessively long text, and also prevents resource waste caused by excessively short text blocks.
Owner:SHIP INFORMATION RES CENT (NO 714 RES INST OF CHINA STATE SHIPBUILDING CORP)

Method for detecting generated text based on a maximum mean discrepancy of depth

This invention discloses a paragraph-level generated text detection method based on deep maximum mean difference (MDD) metric. The method includes: modeling generated text detection as a nonparametric two-sample test problem; designing a deep composite kernel to fuse shallow syntactic features and deep semantic features; training a deep kernel network with the goal of maximizing unbiased estimation of test power; in the testing phase, introducing a Wild Bootstrap mechanism to generate random weight sequences with autoregressive structures; performing a weighted quadratic transformation on the MMD kernel difference matrix; preserving paragraph sentence sequence correlation; accurately estimating the null distribution of the test statistic; and completing the detection by comparing the observed statistic with the test threshold. This method improves sensitivity to high-dimensional semantic differences, alleviates the problem of decreased test power caused by sequence correlation, and can efficiently distinguish between human and AI-generated text. It is suitable for generated text detection in scenarios such as news and academic writing.
Owner:SOUTH CHINA UNIV OF TECH +1

A structured outline driven scientific report automatic generation method and system

The application provides a structured outline driven scientific report automatic generation method and system, and relates to the technical field of natural language processing, the method comprises the following steps: performing semantic analysis and matching on a user input research topic to generate a first-level structure node, extracting and dividing research problems from multiple standard scientific literatures to construct a second-level structure node, performing opinion sentence extraction and refinement to generate a third-level structure node, generating a structured outline according to the first-level, second-level and third-level structure nodes and corresponding explicit binding relationships, performing task analysis, generating a task set according to the hierarchical order of the structured outline, generating a target scientific report and outputting the target scientific report by performing text generation, assembling and integrating. The technical problems of unstable structure, easy divergence of theme, difficulty in corresponding paragraphs and evidence, and easy repetition of chapter content in the existing scientific report automatic generation are solved. The technical effect of stable structure, focused theme and traceable content of the scientific report automatic generation is achieved.
Owner:DOCUMENT & INFORMATION CENT OF CHINESE ACAD OF SCI

A multi-language patent intelligent translation method and system with IPC classification number embedding

PendingCN122452588ASemantic vectorTransformation of text
The application discloses a kind of multi-language patent intelligent translation method and system of fusion IPC classification number embedding, belong to natural language processing field.The method includes the following steps: obtaining and parsing the XML structure of the patent file to be translated, identify paragraph type identifier to identify paragraph type, and extract the IPC classification number mark corresponding to each paragraph;According to the paragraph type and corresponding IPC classification number mark, match the corresponding word segmentation strategy for each paragraph and execute;Using the multi-language BERT model of fusion IPC classification number embedding vector, convert the text after word segmentation into semantic vector;Using decoding model to decode the semantic vector, generate translation output text;After decoding generation, the translation output text is compared with the preset patent standard terminology library, when non-standard terminology is detected, trigger the fallback mechanism to regenerate translation output text.The application improves the accuracy and standardization of patent translation.
Owner:SANDIANSHUI NEW ENERGY TECH (ANHUI) CO LTD

Enhanced search generation method and apparatus, electronic device, and storage medium

The application discloses an enhanced retrieval generation method and device, electronic equipment and storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: receiving a query sentence input by a user, and parsing the query sentence to obtain a subgraph matching template; performing subgraph isomorphism matching on a hypergraph structure based on the subgraph matching template, so as to obtain one or more target hyperedges conforming to the subgraph matching template; wherein the hypergraph structure comprises a plurality of hyperedges, each hyperedge connects a plurality of entity nodes to represent a composite event; determining a target paragraph list associated with the target hyperedge; constructing a target input sequence based on a target paragraph pointed by the target paragraph list and the query sentence, inputting the target input sequence into a natural language model, and generating a target answer text matched with the query sentence. The application can improve the retrieval accuracy and reliability, and also takes into account the retrieval efficiency.
Owner:IFLYTEK CO LTD

Method and device for executing reasoning task of LLM, equipment cluster and product

The embodiment of the invention discloses an execution method and device of an LLM reasoning task, an equipment cluster and a product, and belongs to the technical field of AI. In the embodiment of the invention, a reasoning task and text data are obtained, and the text data comprise a plurality of paragraphs; according to the relevancy between the paragraph summary of each paragraph and the reasoning task, the relevancy between each paragraph and the reasoning task is determined, and the content of the paragraph summary is part of content extracted from each paragraph; determining a target number of paragraphs from the plurality of paragraphs based on the relevancy between each paragraph and the reasoning task; and inputting the target number of paragraphs and the reasoning task into LLM, so that the LLM outputs a reasoning result of the reasoning task based on the target number of paragraphs. On one hand, computing resources can be saved, video memory occupation is reduced, and throughput is improved; on the other hand, the total consumed time of the end-to-end process can be shortened, and the situation that the user needs to wait for a long time to obtain the answer after initiating the question and answer task is avoided.
Owner:HUAWEI TECH CO LTD

A method, system, terminal and medium for detecting machine-generated Chinese text

The application discloses a method for detecting machine-generated Chinese text, comprising: splitting the Chinese text to be detected according to a set step length to obtain a list of N text paragraphs; traversing the list of N text paragraphs, sampling each text paragraph at a set sampling rate, masking M times to obtain M masked text paragraphs, and inputting the M masked text paragraphs into a T5 model in sequence for decoding to obtain a list of unmasked text paragraphs; calculating the confidence score of each text paragraph according to the negative log-likelihood function score of each text paragraph and the negative log-likelihood function score of each element; comparing the confidence score with a set threshold value to determine whether the text paragraph is artificially written or machine-generated; and comparing the average confidence score with a set threshold value to determine whether the Chinese text to be detected is machine-generated or artificially written. The method is simple to implement and can quickly and accurately detect whether the Chinese text is machine-generated.
Owner:CHONGQING JUEXIAO TECH CO LTD

A text duplicate detection method and system

PendingCN122452536AFeature vectorTyping Classification
The application provides a text duplication checking method and system, which comprises the following steps: extracting semantic features of each paragraph in a paragraph set of a to-be-checked text and a source text, obtaining semantic feature vector sets of the to-be-checked text and the source text, screening, from each source text, a paragraph pair with a similarity greater than a preset similarity threshold to each paragraph of the to-be-checked text based on vector similarity retrieval, and constructing an initial candidate paragraph pair set; extracting multi-dimensional discriminant features of each paragraph pair in the candidate paragraph pair set, inputting the multi-dimensional discriminant features into a pre-trained plagiarism type classification model, outputting a plagiarism type label and a corresponding classification confidence of the current paragraph pair, and weighting a current paragraph similarity based on a preset weight corresponding to the plagiarism type label and the classification confidence, to obtain a type-weighted similarity; the method realizes accurate differentiation of plagiarism types and scientific quantification of comprehensive similarity through multi-dimensional feature fusion and a dynamic weighting algorithm, and improves the recognition accuracy.
Owner:BEIJING FUNSHION ONLINE TECH LTD +1

A hierarchical search method, storage medium and computer device

This invention discloses a hierarchical retrieval method, storage medium, and computer device, belonging to the field of document processing technology. It includes: S100, obtaining the hierarchical paragraph parsing results of the input document; S200, retrieving the title paragraphs in the parsing results according to the user's query, generating a set of title paragraphs based on the retrieval results, determining whether the title paragraph set is empty, and if not, outputting a semantically complete paragraph; if so, proceeding to step S300; S300, retrieving the body text paragraphs in the parsing results according to the user's query, and outputting a semantically complete paragraph or summary based on the retrieval results. This invention, by combining the hierarchical parsing results of the document, can improve the completeness of text retrieval results in large-scale model retrieval and enhance document question answering, thereby improving the correctness and completeness of the generated answers in document question answering.
Owner:SHIP INFORMATION RES CENT (NO 714 RES INST OF CHINA STATE SHIPBUILDING CORP) +1

A subjective question intelligent marking method and system supporting multi-text mixing

The application discloses a subjective question intelligent marking method and system supporting multi-text mixing, and relates to the technical field of online education. Through multi-granularity alignment and nonlinear fusion strategy, cross-paragraph information fusion and logic are effectively captured, and the scoring accuracy of complex answers is greatly improved. Combined with differentiated quality evaluation of Chinese and English, accurate scoring of Chinese coherence and structural integrity is realized, and English sentence-by-sentence fine revision can accurately mark grammatical errors and optimize vocabulary collocation. At the same time, the output of explainable scoring evidence can generate personalized comments and optimize model texts, helping students to improve accurately. The calibration mechanism and artificial review dynamically update the parameters to ensure the reliability of scoring, which not only improves the marking efficiency and reduces the human bias, but also promotes the upgrading of intelligent marking from simple scoring to accurate evaluation and personalized guidance. The problems of insufficient processing of multi-text mixed answers, lack of personalized comments and optimized model texts, and absence of English sentence-by-sentence fine revision in existing systems are effectively solved.
Owner:SHANDONG SHIJIJINBANG SCI & EDUCATION & CULTURE

Method, device and readable storage medium for writing beautification

Embodiments of the present application provide a writing beautification method, device and equipment and readable storage medium. In the writing process, the electronic device obtains strokes of each writing operation to obtain a stroke set and displays an original writing track on the touch screen. When a beautification instruction is recognized, the electronic device determines a detection frame of each category in multiple categories, such as a character detection frame, a text line detection frame and a paragraph detection frame, and determines strokes in each detection frame. Then, the electronic device generates a beautified board according to the strokes in each detection frame of each category in the multiple categories. By using this scheme, the writing track is beautified from the granularity of characters, text lines and paragraphs, and the text beautification strategy from fine to coarse greatly improves the writing aesthetics.
Owner:GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1

A text information summarization system based on LLM

PendingCN122451140AWeb pageParagraph
The application discloses a text information summarization system based on an LLM, which comprises the following steps: configuring a crawler cluster to perform a web text information crawling task, comparing the crawled text information with old text information based on content fingerprints, and storing the crawled text information in association with the old text information; generating index labels of natural paragraphs and sentences of the crawled text information, and generating index-based structured metadata text information; constructing a prompt word engineering library, extracting natural paragraphs and sentences from the index-based structured metadata text information by using a large language model (LLM), and generating a purified structured abstract through induction and rewriting; performing logical consistency verification on the purified structured abstract and the natural paragraphs and sentences extracted by the LLM, and generating a provenance mapping table of the text information; after a user subscribes to the text information, the system pushes a card of the purified information item to the user, and marks the confidence of each purified information item.
Owner:GUANGZHOU COLLEGE OF COMMERCE

Textbook question answering method and system based on multi-level attention

A textbook text question and answer method and system based on multi-level attention, the method comprising: inputting the question text and the corresponding chapter context paragraph into the first attention model after tokenization and coding, performing self-attention calculation and pooling within the sentence to obtain the sentence representation vector; calculating the cosine similarity of the question sentence representation vector and the representation vector of all context paragraph sentences, retaining only the representation vector corresponding to the sentence with the maximum similarity of each paragraph as the context representation vector of each paragraph corresponding to the question; inputting the question and the context representation vector of each paragraph corresponding to the question into the second attention model, performing self-attention calculation and pooling between the paragraphs and the question to obtain the answer representation vector corresponding to the question as the input of the classifier; outputting the answer options by the classifier, and obtaining the context paragraph of the answer text corresponding to the chapter. The application can select more accurate answer options for textbook text questions.
Owner:XI AN JIAOTONG UNIV

A document generation method and device based on a multi-agent large model system

The application provides a document generation method and device based on a multi-agent large model system, and relates to the technical field of artificial intelligence. The method comprises the following steps: an agent plans an outline according to an obtained writing task, and sends the outline to a writing agent; the outline indicates a plurality of writing paragraphs, a writing theme of each writing paragraph, and a writing order between the writing paragraphs; the following process is repeatedly executed until all the writing paragraphs are completed: the writing agent determines a current writing paragraph from the writing paragraphs according to the writing order, generates first content according to the writing theme of the current writing paragraph, and sends the first content to an auditing agent; the auditing agent audits the quality of the first content, and determines whether the current writing paragraph is completed according to an auditing result. In this embodiment, a plurality of agents work together, can write and audit paragraph by paragraph according to the outline, effectively reduce the problems of content deviating from the theme and logical confusion, and improve the overall quality.
Owner:BEIJING ACAD OF ARTIFICIAL INTELLLIGENCE

A legal term matching and translating method based on a semantic graph

PendingCN122334296AAttribute grammarSyntax
The application discloses a kind of legal term matching translation methods based on semantic atlas, belong to natural language processing technical field, specifically include: input source language legal document full text and be segmented into paragraph unit, calculate the paragraph emotional keynote vector and semantic field vector of each paragraph;Position legal term in paragraph and intercept its sentence level context window;Analysis window syntax structure, identify legal validity attribute syntax mark word and extract modifying phrase;The above features are input into cross-language emotional potential field model, search potential target paragraph in target corpus with continuous emotional and semantic transition;Lock the candidate term equivalent to source term concept, calculate potential matching degree and calibrate in conjunction with the grammatical compatibility of modifying phrase;Output final target language legal term, simultaneously generate potential matching degree report and context transitional explanation.The application can ensure that term concept is accurate while maintaining the coherence of emotional keynote and semantic field of legal text.
Owner:LIAONING UNIVERSITY

Long text generation method and system of large language model

The embodiment of the specification provides a long text generation method and system of a large language model. In the method, a long text generation system of a large language model acquires an initial text paragraph generated in a long text output process of the large language model in a piecewise generation manner, and performs repetitive text detection on the initial text paragraph. When the repetitive text detection fails, the initial text paragraph is modified until a target text paragraph is obtained. The repetitive text detection of the target text paragraph passes. Further, the long text generation system of the large language model returns the target text paragraph to the first large language model, so that the first large language model outputs the target text paragraph to the user.
Owner:ZHEJIANG ANT SECRET TECH CO LTD

Method for constructing document hierarchical index based on layout visual features and path constraints

The application discloses a document hierarchical index construction method based on layout visual features and path constraints. The method comprises the following steps: obtaining document data, performing paragraph segmentation, and obtaining a paragraph sequence; extracting layout visual features according to the paragraph sequence and encoding the layout visual features into a structured layout visual feature vector; extracting paragraph text according to the paragraph sequence and encoding the paragraph text into a semantic vector; fusing the layout visual feature vector and the semantic vector, inputting the fused vector into a pre-trained sequence model to predict an initial hierarchical label of each paragraph; generating a path vector according to the initial hierarchical label, an ancestor node set vector, a brother node set vector and a current paragraph fused vector, and establishing an index according to a hierarchical consistency constraint. The document hierarchical index construction method based on layout visual features and path constraints provided by the application combines layout visual features and semantic information to realize accurate hierarchical indexing of a document, and can make subsequent retrieval results more accurate and complete.
Owner:BEIJING JIAODA SIYUAN SCI & TECH

A document semantic comparison method, device, equipment, medium and product

PendingCN122287592ALinguistic modelEngineering
This application provides a document semantic comparison method, apparatus, device, medium, and product, comprising: performing structural parsing on the documents to be compared to obtain a structured paragraph set; vectorizing and encoding each paragraph to obtain an encoding result, and performing an approximate nearest neighbor search from a historical document database based on the encoding result to obtain candidate paragraphs; using each paragraph in the structured paragraph set as a source paragraph, forming paragraph pairs with the candidate paragraphs, and performing semantic comparison on the paragraph pairs to obtain a semantic similarity score; calculating a dynamic threshold based on meta-information tags, comparing the dynamic threshold with the semantic similarity score to determine the reuse risk level, and inputting high-risk paragraph pairs into a large language model to obtain an evidence chain risk report. This improves the accuracy, processing efficiency, and interpretability of document reuse risk identification.
Owner:CHINA MOBILE ZIJIN INNOVATION INST CO LTD +2

Paragraph logical relationship perception-based super-long document retrieval method, system and product

This invention provides a method, system, and product for retrieving ultra-long documents based on paragraph logical relationship awareness, belonging to the field of natural language processing and information retrieval technology. It includes: S1, offline construction of an internal paragraph logical dependency graph of the ultra-long document; S2, when a user initiates a query, retrieving multiple sets of ultra-long documents with pre-constructed internal paragraph logical dependency graphs, returning matching paragraphs and completing their logical dependencies. This invention solves the problem of missing key information caused by neglecting deep logic in traditional retrieval systems by constructing logical dependencies between paragraphs within the document and dynamically completing the context during the retrieval stage.
Owner:KYLIN CORP

Sentence ranking method and device based on information entropy and electronic equipment

The disclosure provides a sentence ranking method and device based on information entropy and electronic equipment, and belongs to the computer field. The sentence ranking method comprises: obtaining a plurality of sentences to be ranked; determining the probability of each sentence in the plurality of sentences at each position in the paragraph; determining the information entropy of each ranking path based on the probability of each sentence in the plurality of sentences at each position in the paragraph, wherein the information entropy of the ranking path is negatively correlated with the semantic coherence corresponding to the ranking path; and determining the ranking result of the plurality of sentences based on the information entropy of each ranking path. By using the disclosure, the sentence ranking can be realized to obtain a semantically coherent paragraph.
Owner:BEIJING CENTURY TAL EDUCATION TECH CO LTD

Dynamic report generation method and system based on large language model and configurable paragraphs

The invention discloses a dynamic report generation method and system based on a large language model and configurable paragraphs, and realizes fine management of report structures, content sources and generation logic by decomposing a report template into a plurality of independent configurable paragraph modules. Each paragraph module can be independently configured with a text template, a data retrieval logic and a prompt for interaction with a large language model; the system integrates a uniform data retrieval abstraction layer, can be seamlessly connected with a multi-source heterogeneous business data source, dynamically executes data retrieval for each paragraph module according to a template definition when a report is generated, combines the obtained data with a preset prompt template, calls a large language model to generate a natural language text, and performs data retrieval for each paragraph module according to the template definition. And finally, all paragraph contents are synthesized into a service report which is complete in structure, accurate in data and smooth in language, so that the efficiency of generating the service report is improved.
Owner:HUNAN ELECTRIC POWER DISPATCH HIGH TECH DEV

A long text intelligent processing method and device based on domain adaptation

PendingCN122432339ASemantic vectorText mining
The application belongs to the cross field of artificial intelligence and text mining, and specifically relates to a long text intelligent processing method and device based on field self-adaption, which comprises the following steps: inputting a long text to be processed into a pre-trained field classification model, identifying and outputting a field category to which the long text belongs; based on the field category, dynamically loading and generating a target Prompt adapted to the field category from a pre-constructed hierarchical Prompt template library; inputting the long text to be processed and the target Prompt into a dynamic segmentation engine, adaptively segmenting the long text based on semantic coherence by the dynamic segmentation engine, and outputting a plurality of semantically coherent paragraphs and corresponding paragraph semantic vectors; based on the paragraph semantic vectors, combining multi-dimensional importance evaluation indexes, extracting key information from the plurality of paragraphs, and generating an abstract for the long text. The generated abstract contains key information and maintains semantic coherence.
Owner:INSPUR GENERSOFT CO LTD

Method and system for providing intent-based information insights

PendingUS20260203507A1User inputDegree of similarity
A method and system for providing intent-based information insights is disclosed. NLP techniques and a first LLM are applied to identify an intent vector associated with an input question received from a user. A plurality of paragraphs are retrieved from a document dataset based on relevancy to the intent vector. A graph representation of the plurality of paragraphs is generated. Importance score for each paragraph is determined based on centrality values derived from the graph representation, and similarity measures between the intent vector and the encoded paragraph vectors corresponding to each paragraph. A subset of sentences from the plurality of paragraphs is selected based on the determined importance scores through an iterative optimization process. A natural language summary is generated based on the subset of sentences using a second LLM and is presented to the user with explanatory indicators.
Owner:LTIMINDTREE LTD

A multi-granularity controversy focus automatic induction method for legal texts

The application discloses a kind of legal text-oriented multi-granularity dispute focus automatic induction method, belong to natural language processing and legal science and technology field.The method is first structured to legal documents, and divided into structured block with clear stand label;Second, extract dispute expression unit from dispute block sentence, and convert it into semantic vector using legal field semantic understanding model;Then, based on semantic vector and stand label clustering is carried out, and paragraph-level dispute focus theme is generated by cross-stand theme alignment;Then, the logical correlation strength between themes is calculated, the logical subordinative relationship is analyzed, and the logical tree of chapter-level dispute focus is constructed;Finally, the complete dispute focus system that fuses sentence-level, paragraph-level and chapter-level information is output.The application overcomes the problem of single granularity and lack of logical hierarchy in the prior art induction results, and realizes the systematization, precision and structured automatic induction of dispute focus from microevidence to macrologic.
Owner:EAST CHINA UNIVERSITY OF POLITICAL SCIENCE AND LAW