Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

339 results about "Paragraph" patented technology

A paragraph (from the Ancient Greek παράγραφος paragraphos, "to write beside" or "written beside") is a self-contained unit of a discourse in writing dealing with a particular point or idea. A paragraph consists of one or more sentences. Though not required by the syntax of any language, paragraphs are usually an expected part of formal writing, used to organize longer prose.

Semantic enhancement adaptive partitioning method and system for natural resource large model questions and answers

The invention provides a semantic enhancement adaptive partitioning method and system for natural resource large model questions and answers, and aims to solve the problems of difficulty in term boundary recognition, damage to semantic integrity and the like. According to the method, three core technologies including theme perception coarse-grained paragraph division, self-adaptive sliding window theme hierarchy division and embedded perception context self-adaptive text segmentation are fused. The system firstly analyzes a natural resource long text structure, identifies titles and theme levels and aligns associated contents; paragraphs are extracted according to a theme perception strategy and are subdivided into sentence sets according to grammar rules; an improved sliding window mechanism is adopted to divide sentences into window sentence block groups. The method is characterized in that a dynamic aggregation threshold mechanism is introduced, the semantic association degree between adjacent sentence blocks is calculated through an embedded perception context semantic segmentation technology, whether the sentence blocks are combined or not is judged by combining a similarity distribution change trend and a dynamic adjustment threshold, self-adaptive delimitation of semantic boundaries is achieved, and text blocks which are clear in structure and coherent in semantics are generated.
Owner:HUBEI PROVINCIAL DEPT OF NATURAL RESOURCES INFORMATION CENT +1

Method and system for retrieving DOCX document content based on keywords

The invention belongs to the technical field of text processing, and particularly relates to a method and system for retrieving DOCX document content based on keywords, which comprises the following steps: analyzing an Office Open XML structure of a DOCX document, combining with multi-dimensional features such as style names, and utilizing a title classification score model to accurately distinguish a title and a text, so that a semantic hierarchical structure of the document is effectively reserved; and secondly, a multi-level semantic extension mechanism is introduced, and a Sension-BERT, a HowNet knowledge base and a Word2Vec model are fused, so that intelligent extension of synonyms and synonyms of keywords is realized, and the recall rate and semantic understanding ability of retrieval are remarkably improved. And in addition, a BM25 model is combined with paragraph length normalization and structure position weight to calculate a correlation score, so that retrieval results are sorted more accurately and reasonably. The construction of the reverse index is combined with the position coding and compression optimization strategy, and the retrieval efficiency and the storage performance are both considered.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

AI reading control method and system based on artificial intelligence

The invention relates to the technical field of natural language processing, in particular to an AI reading management and control method and system based on artificial intelligence, and the method comprises the following steps: carrying out text word segmentation processing on an input original text, segmenting the text into independent sentences, recognizing basic word units in each sentence, analyzing semantic adjacency relationships and syntactic structure features among vocabularies, and carrying out word segmentation processing on the word units; screening and extracting potential phrases representing paragraph meanings, and establishing a candidate semantic unit set; according to the method, through word segmentation, syntactic structure recognition and semantic adjacency analysis of the original text, potential phrases capable of representing paragraph significance are extracted, the candidate semantic unit set is constructed, and modeling of the semantic structure in the text is achieved. And associating the set with the reading fixation duration and the playback action of the user sentence by sentence to obtain a reading behavior response of each semantic fragment, and executing semantic weighting and reading behavior cross analysis according to the reading behavior response. Through the linkage mode, the semantic focus actually focused by the user at present can be recognized.
Owner:SHENZHEN JOYAR SMART MFG TECH LTD

Cross-language text fusion intelligent alignment method and system

The invention relates to the technical field of cross-language information processing, and provides a cross-language text fusion intelligent alignment method and system.The cross-language text fusion intelligent alignment method comprises the steps that a hierarchical alignment model is constructed through preprocessing and label recognition of multi-coding-type texts, deep semantic feature extraction and labeling of a multi-language pre-training model and text semantic and format information analysis; transform is taken as a core, cross-language semantic association is enhanced through a multi-head attention mechanism, character-level and paragraph-level format collaboration is realized through label weight allocation and condition constraint, the format alignment accuracy is obviously improved in a multi-language mixed typesetting scene, and the multi-language mixed typesetting efficiency is improved. The analysis capability of the model on the structured information can be extended to layout relation processing of texts, images and tables, so that the comprehensive alignment efficiency in a multi-modal fusion scene is remarkably improved. The accuracy of text and label fusion is guaranteed, the problem of confusion of messy codes and labels is avoided, and high-precision cross-language text alignment from semantics to formats, from single mode to multiple modes and from semantics to formats is achieved.
Owner:SHANGHAI MEGALIN SOFTWARE TECH CO LTD

Systems, apparatuses, methods, and non-transitory computer-readable storage media for adaptive information retrieval for question-answering

Methods and systems for retrieving relevant information in response to an input question. The method includes obtaining text content related to the input question and partitioning the content into one or more paragraphs based on predefined rules. The method further involves extracting one or more evidence spans that are relevant to the input question by inputting the text content and the question into a trained language model. A semantic search is then performed on both the paragraphs and the extracted evidence spans, ranking the candidate passages based on their relevance to the input question. Each candidate passage may comprise either a paragraph or an evidence span that addresses the question. The disclosed methods and systems improve the quality and relevance of retrieved information by combining heuristic-based content partitioning with machine learning-based evidence extraction.
Owner:HUAWEI TECH CO LTD

Industrial instruction fine tuning data set automatic generation method and system

The invention discloses an industry instruction fine tuning data set automatic generation method and system, and belongs to the technical field of artificial intelligence, S10: selecting industry standard document data, and analyzing the data into a plurality of paragraph texts as data sources according to chapters; s20, generating different types of questions for the paragraph text by using a large language model; s30, scoring the generated questions according to a scoring rule, and filtering the questions with scores lower than a preset threshold value; s40, the solvability and difficulty of the question are adjusted through a large language model, and the quality of the question is optimized, and S50, a first version answer of the question is generated by using the large language model, and the answer is optimized by using a hierarchical region optimization search algorithm; s60, constructing a sample pool through global and local selection methods, calculating a sample compression ratio and screening a data set; the method has the beneficial effects that through a highly automatic process and an optimization algorithm, the data set generation efficiency and quality are improved, the labeling cost is reduced, and support is provided for rapid deployment of a large language model in industry application.
Owner:GUANGZHOU ZHONGKE YIDE TECH CO LTD

Official document content generation and compliance verification system based on large language model

The invention relates to the technical field of large language models, in particular to an official document content generation and compliance verification system based on a large language model. Comprising an official document extraction module used for analyzing a model essay by using a large language model, extracting a writing style and a fixed sentence pattern, constructing a knowledge graph, storing format specifications, common expressions and logic structures, mining and constructing a core intention library, and storing typical intention prototype vectors; the content generation module is used for generating paragraphs through context sensing, expanding and writing keywords, analyzing behavior data, predicting writing intention vectors and performing condition control; the compliance verification module is used for warning improper words in real time, retrieving policy verification logic and verification formats, and performing pre-verification based on intention vectors; and the writing auxiliary module is used for providing real-time modification suggestions and forward-looking prompts based on writing intention vectors. According to the technical scheme, document writing efficiency and quality can be improved.
Owner:CHONGQING BORA INTELLIGENT COMPUTING TECHNOLOGY CO LTD

Text processing method and system based on semantic density

The invention provides a text processing method and system based on semantic density, which are applied to the technical field of text processing, and the method comprises the following steps: obtaining a target text, and carrying out feature extraction on the target text to obtain a plurality of text features of each paragraph of the target text; calculating a semantic density score of each paragraph according to the text feature of each paragraph; classifying each paragraph according to the semantic density score of each paragraph to obtain a paragraph category of each paragraph; if the paragraph category of the paragraph is high in density, performing block processing on the paragraph according to the text type of the target text and the semantic density score of the paragraph to obtain a plurality of text blocks corresponding to the paragraph; and if the paragraph category of the paragraph is low density, processing the paragraph according to the text type of the target text, the paragraph and the semantic density score thereof, and each other paragraph and the semantic density score thereof, so that information splitting and context loss can be avoided, redundant calculation and resource waste are reduced, and the field adaptability is improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Intelligent construction special scheme paragraph intelligent compilation method based on information network

The invention belongs to the technical field of construction technical scheme compilation, and particularly relates to an intelligent construction special scheme paragraph intelligent compilation method based on an information network, which comprises the following steps of: 1, dividing content processing into two categories according to different input contents: project general content and project characteristic content; step 2, vector embedding; step 3, index construction; step 4, index storage; step 5, carrying out similarity retrieval; and 6, outputting a retrieval result. According to the intelligent construction special scheme paragraph intelligent compilation method based on the information network, through automatic word segmentation, vector embedding and index construction, the construction special scheme compilation efficiency is greatly improved, and the time needed by manual compilation is shortened; historical schemes and standard requirements are utilized to ensure that the generated scheme content conforms to industrial standards and specifications, and the standardization degree of the schemes is improved; by distinguishing the project general content and the project characteristic content, a customized scheme can be generated according to the characteristics and requirements of a specific project.
Owner:CCCC SECOND HIGHWAY ENG CO LTD

Local knowledge base RAG method and device based on Bayesian reasoning

ActiveCN120218247AMathematical modelsSemantic analysisQuestion generationBayesian formulation
The invention discloses a local knowledge base RAG method based on Bayesian reasoning, and belongs to the technical field of information retrieval. The method comprises the steps of obtaining all knowledge data of a local knowledge base, performing paragraph classification and coding on the knowledge data to obtain a paragraph set, and calculating the semantic probability of each paragraph in the paragraph set; obtaining a target professional vocabulary set of the target question, and calculating the semantic probability of each vocabulary in the target professional vocabulary set in each paragraph and the semantic probability of all vocabulary in the target professional vocabulary set in each paragraph in the paragraph set Calculating the conditional probability of each vocabulary in the target professional vocabulary set in each paragraph based on the occurrence frequency; and calculating the conditional probability of the target professional vocabulary set in each paragraph by adopting a Bayesian formula, sorting to obtain a conditional probability sorting result, selecting the paragraphs of which the total word number does not exceed a preset threshold value as a target paragraph set, and generating target retrieval information based on the target paragraph set and the target question. According to the method, the accuracy and reliability of question retrieval are improved.
Owner:HUBEI TAIYUE SATELLITE TECH DEV CO LTD

Intelligent RAG enhanced large model auxiliary writing system

The invention discloses an intelligent RAG enhanced large model auxiliary writing system, and relates to the technical field of artificial intelligence auxiliary writing. In order to overcome the defects of an existing writing auxiliary tool, the adopted scheme comprises an interface and user layer for receiving a writing demand and redisplaying a generation result; the service layer is used for deploying three types of intelligent agents including writing, interactive writing enhancement and checking; the platform layer is used for carrying out industry adaptation analysis, mixed retrieval and RAG enhancement on writing demands and then sending the writing demands to the large model layer; three intelligent agents of writing, interactive writing enhancement and checking are called in sequence based on the output of the large model layer, and then a text is redisplayed to the interface and user layer after fact accuracy verification is completed through a tense knowledge graph; the data layer is used for providing a multi-type data storage system; the large model layer outputs a paragraph-level text and a reference index based on a writing large model, the paragraph-level text is returned to the platform layer, and reference index information is fed back to the interpretability and security layer for tracing; and the infrastructure layer is used for providing bottom layer resource support for each layer.
Owner:INSPUR QILU SOFTWARE IND

Method for coherent, unsupervised, transcript-based, extractive summarisation of long videos of spoken content

Disclosed is a video summarisation method that includes converting an input video file to a video transcript using speech recognition, segmenting the video transcript into a plurality of paragraphs using a transformer based model, performing relevance ranking of the plurality of paragraphs by assigning a paragraph importance score to each paragraph, based on relevance of each paragraph to an overall content of the video transcript, creating a plurality of candidate summaries from the plurality of paragraphs, each candidate summary comprising two or more paragraphs, performing coherence reranking of the plurality of candidate summaries based on a combination of coherence and relevance of each candidate summary, and selecting a summary based on the coherence reranking.
Owner:UNIV COLLEGE DUBLIN NAT UNIV OF IRELAND DUBLIN

Multi-modal digital publishing intelligent checking system and method based on large model

The invention relates to the technical field of document intelligent review, in particular to a multi-mode digital publishing intelligent review system and method based on a large model, and the system comprises a sample construction module, a multi-mode analysis module, a structure recognition module, a term verification module and an annotation output module. According to the method, a model configuration data set is generated by analyzing a subject-object combination relationship in a text and jointly screening part-of-speech, sentence pattern and semantic structure, the accuracy of semantic recognition and structure judgment is improved, and local features of semantic offset and image-text expression separation are recognized by combining a comparison relationship between an image action path and a text behavior object; a paragraph logic structure and a theme connection mode are analyzed based on semantic offset information, the problems of content dislocation and theme jump between paragraphs are effectively revealed, matching change tracks and part-of-speech continuation fluctuation in term cross-paragraph contexts are recognized, label overlapping and coverage redundancy conditions are analyzed from sentence group hierarchy, and the term cross-paragraph contexts are obtained. And forming a label combination suggestion and constructing an annotation structure record.
Owner:PEOPLES HEALTH ELECTRONIC AUDIO VISUAL PUBLISHING HOUSE CO LTD

Collaborative question and answer method for complex hierarchical table and large language model

The invention discloses a complex hierarchy table and large language model collaborative question and answer method. The method comprises the steps of S1, obtaining a user question, a complex hierarchy table and an unstructured text paragraph associated with the complex hierarchy table; s2, preprocessing the complex hierarchical table to obtain a multi-modal input sequence; s3, performing joint coding on the sequence by adopting a LayoutLM model, and obtaining a two-dimensional table with a reduced structure through a graph convolutional network; s4, based on the user question, performing fine-grained evidence retrieval on the two-dimensional table and the unstructured text paragraph to obtain a mixed fine-grained evidence set; and S5, taking the user question, the mixed fine-grained evidence set and the specific task prompt template as input of a large language model, and outputting to obtain a final answer of the user question. According to the method, the problems that a complex hierarchical table structure is difficult to understand and the numerical reasoning performance of the mixed context of the table and the text is insufficient are solved, and the high-precision and generalizable table question and answer ability is achieved.
Owner:CHONGQING JIAOTONG UNIV

Method for editing handwritten notes and electronic equipment

The invention provides a method for editing handwritten notes and electronic equipment, relates to the technical field of terminals, and is used for solving the problem of how to realize tidy and attractive handwritten note layout in a complicated handwritten note scene. According to the scheme, the method comprises the steps that a note interface is displayed on a display screen, handwritten notes are displayed on the note interface, and the handwritten notes comprise one or more paragraphs; receiving an editing instruction which acts on the note interface and aims at the handwritten note; in response to the editing instruction, the character position in at least one paragraph is adjusted to obtain an edited handwritten note, the line angle in any one paragraph in the at least one paragraph in the edited handwritten note is the same as the angle of the paragraph, and / or the line angle in any one paragraph in the edited handwritten note is the same as the line angle in any one paragraph in the at least one paragraph in the edited handwritten note. The position of each row of characters in any paragraph in the paragraph meets a first requirement, and the at least one paragraph is a paragraph acted by the editing instruction; and displaying the edited handwritten note.
Owner:HUAWEI TECH CO LTD

Paragraph division method and system based on text semantic information fusion

The invention discloses a paragraph division method, system and device based on text semantic information fusion, a medium and a program. The method comprises the following steps: identifying a character image to be identified to obtain textboxes, traversing each textbox, and merging the textboxes into lines according to relative positions to obtain a position information merged text; according to the distance between the lines in the position information merging text, identifying the spatial position of the text, and according to an identification result, performing paragraph merging on the text lines to obtain a paragraph information merging text; performing text semantic information fusion processing on the paragraph information merging text based on a semantic analysis model to obtain paragraph text information; and traversing each row in the paragraph character information, and carrying out paragraph calculation layout to obtain divided paragraphs. According to the method, text fusion and paragraph division can be rapidly carried out, missing paragraph division is supplemented by combining semantic information of the text needing paragraph division, and the accuracy is improved.
Owner:XIAN TPRI THERMAL CONTROL TECH +1

Language translation data processing method and system based on large model

The invention provides a language translation data processing method and system based on a large model, and the method comprises the steps: carrying out the coding segmentation of language translation data, and obtaining a plurality of text segments; performing dependency syntactic analysis on each text fragment, determining a dependency tree structure of each text fragment, and determining an explicit semantic dependency relationship between every two text fragments according to the semantic similarity between every two text fragments and the dependency tree structure of the corresponding text fragment; performing dependency analysis on an implicit semantic relationship between every two text segments according to the implicit semantic association degree between the text segments and the segmentation loss of the text segments to obtain an implicit semantic dependency relationship between every two text segments; and constructing a paragraph tag of the language translation data through a dependency relationship between explicit semantics and implicit semantics between every two text fragments, and translating the to-be-translated data based on the paragraph tags. By the adoption of the scheme, cross-segment semantic guidance translation of the long complex text can be achieved.
Owner:HUNAN COMM POLYTECHNIC

Document writing method and system based on government affair industry large model and medium

The invention discloses an official document writing method and system based on a government affair industry large model and a medium, mainly relates to the technical field of official document writing, and is used for solving the problems that in the prior art, in a long official document generation process, a model may lack logic consistency among different paragraphs, contexts are not coherent, and expression does not conform to official document specifications. Comprising the steps of obtaining a trained pre-training model; inputting the obtained cue word into a trained pre-training model, outputting an official document outline, and extracting a keyword corresponding to the official document outline; determining whether the keyword and the cue word have semantic consistency, and when the keyword and the cue word have semantic consistency, determining that the obtained document outline is qualified; expanding the qualified official document outline into an initial official document text in a manner of combining paragraph abstracts and context constraints; and performing preset normative auditing on the initial official document text, and modifying contents which do not accord with the preset normative to contents which accord with the preset normative to obtain a final official document.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Contract retrieval enhancement optimization method, equipment and medium

The invention discloses a contract retrieval enhancement optimization method and device and a medium, and belongs to the technical field of contract retrieval. The method comprises the steps that output of a large language model is constructed and optimized based on a hypothetical question and answer data set, so that the large language model generates hypothetical answers conforming to a specific format; obtaining a paragraph semantic data set of the contract text; training a preset neural network model based on the paragraph semantic data set to obtain a paragraph semantic model; inputting the contract text into a paragraph semantic model to generate a summary vector and a keyword vector, and storing the summary vector and the keyword vector to a preset vector database; processing a query problem of the user through the optimized large language model to obtain a query vector; calculating the similarity between the query vector and vectors in a vector database and selecting a first number of paragraphs; and constructing a prompt, and inputting the prompt into the optimized large language model so as to output a retrieval answer. According to the method, the technical effect of improving the accuracy and flexibility of contract retrieval is achieved.
Owner:INSPUR GENERSOFT CO LTD

Contract risk auditing method and device

The invention relates to the field of artificial intelligence, and particularly provides a contract risk auditing method and device, and the method comprises the following steps: S1, splitting a contract content paragraph, and decomposing a contract text into a plurality of independent paragraph units with a logic relation; s2, converting the contract text into a structured data unit, and performing text analysis and risk factor extraction; s3, table textualization: converting table data in the contract into a plain text format which can be processed by a large model; s4, constructing a list table to systematically organize key information in the contract; s5, performing text semantic analysis to deeply understand the contract text; s6, element checking: carefully reviewing and verifying key elements in the contract text; and S7, risk information summarization: summarizing and analyzing potential risks in the contract text. Compared with the prior art, the method has the advantages that potential risks in the text can be recognized and predicted through learning of a large amount of text data, and therefore the efficiency and accuracy of contract risk auditing are remarkably improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Multi-dimensional engineering paper abstract evaluation method

PendingCN121031578ASemantic analysisBiological modelsBasic languageEngineering
The invention discloses a multi-dimensional engineering paper abstract evaluation method, and belongs to the field of natural language processing. The implementation method comprises the following steps of: performing text correctness evaluation on a basic language specification in an abstract text, wherein a text correctness evaluation dimension comprises two sub-dimensions of grammar correctness and an expression specification degree; a pre-training language model GPT-2 is used as an evaluation base model, training fine tuning is carried out by using an abstract text in an RAAMove training set, and fluency dimension in abstract language paragraph fluency is evaluated; introducing a speech step theory, and constructing a speech step classification model based on comparative learning; performing double representation learning on sentence feature representation and speech step tag feature representation by adopting a supervised comparative learning method; performing coherence evaluation in paragraph fluency according to a speech step transfer similarity index; a ROUGE 1F1 value is calculated, and semantic similarity evaluation is carried out; and introducing a weighted summation strategy to carry out weighted summation on the scores of the dimensions to obtain a comprehensive evaluation score of the abstract text, namely realizing comprehensive evaluation on the abstract text.
Owner:BEIJING INST OF TECH

Text review method and device, storage medium and program product

The invention discloses a text review method and device, a storage medium and a program product, and relates to the technical field of artificial intelligence, the method comprises the steps that a domain knowledge base is constructed in advance, the domain knowledge base comprises a high-risk word set and element related knowledge, and the high-risk word set comprises violation words, violation types corresponding to the violation words and words needing element checking; the element-related knowledge includes provisions for essential elements of the compliance text, and compliance expression suggestions when the compliance text relates to a target description. When a target text is examined, paragraphs which contain high-risk words and contain the high-risk words meeting preset dangerous conditions are determined in the target text based on a high-risk word set to serve as candidate paragraphs, then element checking is conducted on each candidate paragraph based on element related knowledge, and an element checking result of each candidate paragraph is obtained; and generating an examination result of each candidate paragraph. Through two-stage screening, the examination efficiency is improved, and meanwhile, the examination accuracy is ensured.
Owner:ANHUI IFLYTEK INTELLIGENT SYST

A text analysis method, system, device and medium based on text structure

The present invention relates to a text analysis method, system, device and medium based on text structure, which comprises the following steps: parsing the acquired text to be analyzed to obtain its text structure; respectively performing machine reading on each text structure of the text to be analyzed to obtain the embedding vectors corresponding to the respective text structures; fusing the obtained embedding vectors to obtain a fused article embedding vector; and obtaining a text analysis result based on the fused article embedding vector. The present invention takes into account the important significance of the article structure for machine understanding, and parses according to the structure of abstract - paragraph {paragraph title - paragraph content}, enabling the model to have the ability of reading by sub - structure. Therefore, the present invention can be widely applied to the field of text analysis.
Owner:RENMIN UNIVERSITY OF CHINA

Method for generating Boolean logic samples and computing equipment

The embodiment of the invention relates to a Boolean logic sample generation method and computing equipment, and the method comprises the steps: firstly, inputting a plurality of text paragraphs related to a target topic into a large language model, indicating the large language model to generate atomic questions corresponding to the text paragraphs, and indicating the large language model to generate disjunction questions; wherein the disjunction question corresponds to the plurality of text paragraphs; then, generating a first Boolean problem according to at least one target problem in the plurality of atomic problems and disjunction problems and a first logic relationship, and taking the first Boolean problem as a Boolean problem in a first Boolean logic sample; and then, according to the text paragraph corresponding to the target question and the first logic relationship, determining the text paragraph corresponding to the first Boolean question as the text paragraph in a first Boolean logic sample.
Owner:SASI DIGITAL TECHNOLOGY (BEIJING) CO LTD +1

Document segmentation method and device based on large language model, equipment and storage medium

The invention provides a document segmentation method and device based on a large language model, equipment and a storage medium, and relates to the technical field of text processing. The method comprises the steps of inputting a to-be-segmented target document into a pre-trained large language model, and executing the following operations through the large language model: performing text layout analysis on the target document, and identifying titles and all paragraphs of each level in the target document; for each paragraph, inserting an associated title related to the paragraph in all titles into an initial position of the paragraph to obtain a corresponding target paragraph; and sorting all the target paragraphs based on the semantic similarity among all the target paragraphs, and determining a segmentation result of the target document based on all the sorted target paragraphs. By the adoption of the technical scheme, when document segmentation is carried out, semantic loss in the document segmentation process can be effectively reduced, and therefore the document segmentation effect is improved.
Owner:CHINA LIFE ASSET MANAGEMENT CO LTD

Intelligent agent construction method and system based on large language model, equipment and medium

The invention provides an agent construction method and system based on a large language model, equipment and a medium, and belongs to the technical field of artificial intelligence. The method comprises the steps that a user interaction module is called to receive a first target document input by a user; calling a document analysis module to analyze the first target document to obtain a plurality of first chapter titles; calling an automatic mapping module to classify the plurality of first chapter titles based on a document classification model, and constructing a mapping relationship between the plurality of first chapter titles and a plurality of second chapter titles of a second target document based on a classification result; calling a large language model to respectively extract and summarize paragraph contents of the plurality of first chapter titles to obtain paragraph contents of a second chapter title; and calling a document output module to output the second target document according to a preset format. According to the agent construction method and system based on the large language model, the equipment and the medium provided by the invention, the accuracy of generating the review report by the agent can be improved.
Owner:ZHONGJIAO ROAD & BRIDGE (HEBEI) CO LTD

Automatic long text fine tuning instruction set construction method based on large language model

The invention provides an automatic long text fine-tuning instruction set construction method based on a large language model, which comprises the following steps: acquiring an input text, and segmenting the input text by adopting a recursive character segmentation method to generate a paragraph set; aiming at the generated paragraph set, the large language model generates a question set and an answer set by adopting a self-guidance learning method through a preset question type set and a prompt template according to the task type; and generating an instruction set based on the generated question set and answer set, evaluating the quality of the instruction set in multiple dimensions, and optimizing the instruction set according to an evaluation result to obtain an optimized instruction set. According to the method, the high-quality long text fine tuning instruction set can be automatically generated, so that the performance of a large language model is improved, and meanwhile, the challenge of long context processing is solved. By automatically constructing the long text fine-tuning instruction set, the requirement of manual annotation is reduced, the cost is reduced, and meanwhile, the efficiency of the fine-tuning process and the long text processing capacity of the model are improved.
Owner:INST OF AUTOMATION CHINESE ACAD OF SCI +3

Entity enhancement and context-aware paragraph retrieval method for RAG system

The invention provides an entity enhancement and context-aware paragraph retrieval method for an RAG system for solving the problems of fuzzy query intention and insufficient paragraph context modeling in RAG retrieval. The method comprises the following steps: firstly, identifying and extracting a key entity by using a named entity, carrying out weighted fusion on the key entity and question representation obtained by a pre-training model to form an enhanced query vector, and accurately describing a core semantic intention of the question; secondly, performing semantic modeling on paragraphs in a document library, mining a potential semantic association relationship between the paragraphs, and constructing a context interaction model between the paragraphs based on a graph neural network and a gating loop unit mechanism, so as to obtain paragraph vector representation with complete semantics and clear hierarchy; and finally, calculating the similarity between the enhanced query and the paragraph vector, and completing high-precision paragraph-level retrieval. According to the method, the correlation and the recall rate are remarkably improved, insufficient entity utilization and weak context modeling are relieved, and good expansibility and cross-domain applicability are achieved.
Owner:SOUTHEAST UNIV +1

Question and answer query method and system

The invention provides a question and answer query method and system. The method comprises the following steps: receiving a to-be-queried question; the to-be-queried question is matched with a knowledge base, when at least one target paragraph vector matched with the to-be-queried question is obtained, a target paragraph corresponding to the target paragraph vector is determined, and a plurality of paragraphs and paragraph vectors corresponding to the paragraphs are stored in the knowledge base; and according to the to-be-queried question and the target paragraph, generating a query prompt, inputting the query prompt into a large language model, and obtaining an answer of the to-be-queried question output by the large language model. The question-answering quality of the question-answering system can be improved.
Owner:HITACHI LTD

Text localization method and apparatus, and device and storage medium

PCT designated stage expiredWO2025140051A1Character and pattern recognitionPattern recognitionMedicine
Disclosed in the present application are a text localization method and apparatus, and a device and a storage medium. The method comprises: detecting a fingertip location of a fingertip on a text page, and on the basis of the fingertip location, determining, from the text page, a regional image of text to be recognized; recognizing, from among paragraph text boxes corresponding to paragraphs in the regional image, a first text box having the maximum area; and performing screening to obtain one or more paragraphs in which the paragraph center points are located in a target text box, and determining, as the text to be recognized, text which is included in paragraphs obtained after screening.
Owner:ZHUHAI MOJIE TECH CO LTD