Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

215 results about "Paragraph" patented technology

A paragraph (from the Ancient Greek παράγραφος paragraphos, "to write beside" or "written beside") is a self-contained unit of a discourse in writing dealing with a particular point or idea. A paragraph consists of one or more sentences. Though not required by the syntax of any language, paragraphs are usually an expected part of formal writing, used to organize longer prose.

Semantic enhancement adaptive partitioning method and system for natural resource large model questions and answers

The invention provides a semantic enhancement adaptive partitioning method and system for natural resource large model questions and answers, and aims to solve the problems of difficulty in term boundary recognition, damage to semantic integrity and the like. According to the method, three core technologies including theme perception coarse-grained paragraph division, self-adaptive sliding window theme hierarchy division and embedded perception context self-adaptive text segmentation are fused. The system firstly analyzes a natural resource long text structure, identifies titles and theme levels and aligns associated contents; paragraphs are extracted according to a theme perception strategy and are subdivided into sentence sets according to grammar rules; an improved sliding window mechanism is adopted to divide sentences into window sentence block groups. The method is characterized in that a dynamic aggregation threshold mechanism is introduced, the semantic association degree between adjacent sentence blocks is calculated through an embedded perception context semantic segmentation technology, whether the sentence blocks are combined or not is judged by combining a similarity distribution change trend and a dynamic adjustment threshold, self-adaptive delimitation of semantic boundaries is achieved, and text blocks which are clear in structure and coherent in semantics are generated.
Owner:HUBEI PROVINCIAL DEPT OF NATURAL RESOURCES INFORMATION CENT +1

Method and system for retrieving DOCX document content based on keywords

The invention belongs to the technical field of text processing, and particularly relates to a method and system for retrieving DOCX document content based on keywords, which comprises the following steps: analyzing an Office Open XML structure of a DOCX document, combining with multi-dimensional features such as style names, and utilizing a title classification score model to accurately distinguish a title and a text, so that a semantic hierarchical structure of the document is effectively reserved; and secondly, a multi-level semantic extension mechanism is introduced, and a Sension-BERT, a HowNet knowledge base and a Word2Vec model are fused, so that intelligent extension of synonyms and synonyms of keywords is realized, and the recall rate and semantic understanding ability of retrieval are remarkably improved. And in addition, a BM25 model is combined with paragraph length normalization and structure position weight to calculate a correlation score, so that retrieval results are sorted more accurately and reasonably. The construction of the reverse index is combined with the position coding and compression optimization strategy, and the retrieval efficiency and the storage performance are both considered.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Systems, apparatuses, methods, and non-transitory computer-readable storage media for adaptive information retrieval for question-answering

Methods and systems for retrieving relevant information in response to an input question. The method includes obtaining text content related to the input question and partitioning the content into one or more paragraphs based on predefined rules. The method further involves extracting one or more evidence spans that are relevant to the input question by inputting the text content and the question into a trained language model. A semantic search is then performed on both the paragraphs and the extracted evidence spans, ranking the candidate passages based on their relevance to the input question. Each candidate passage may comprise either a paragraph or an evidence span that addresses the question. The disclosed methods and systems improve the quality and relevance of retrieved information by combining heuristic-based content partitioning with machine learning-based evidence extraction.
Owner:HUAWEI TECH CO LTD

Official document content generation and compliance verification system based on large language model

The invention relates to the technical field of large language models, in particular to an official document content generation and compliance verification system based on a large language model. Comprising an official document extraction module used for analyzing a model essay by using a large language model, extracting a writing style and a fixed sentence pattern, constructing a knowledge graph, storing format specifications, common expressions and logic structures, mining and constructing a core intention library, and storing typical intention prototype vectors; the content generation module is used for generating paragraphs through context sensing, expanding and writing keywords, analyzing behavior data, predicting writing intention vectors and performing condition control; the compliance verification module is used for warning improper words in real time, retrieving policy verification logic and verification formats, and performing pre-verification based on intention vectors; and the writing auxiliary module is used for providing real-time modification suggestions and forward-looking prompts based on writing intention vectors. According to the technical scheme, document writing efficiency and quality can be improved.
Owner:CHONGQING BORA INTELLIGENT COMPUTING TECHNOLOGY CO LTD

Intelligent RAG enhanced large model auxiliary writing system

The invention discloses an intelligent RAG enhanced large model auxiliary writing system, and relates to the technical field of artificial intelligence auxiliary writing. In order to overcome the defects of an existing writing auxiliary tool, the adopted scheme comprises an interface and user layer for receiving a writing demand and redisplaying a generation result; the service layer is used for deploying three types of intelligent agents including writing, interactive writing enhancement and checking; the platform layer is used for carrying out industry adaptation analysis, mixed retrieval and RAG enhancement on writing demands and then sending the writing demands to the large model layer; three intelligent agents of writing, interactive writing enhancement and checking are called in sequence based on the output of the large model layer, and then a text is redisplayed to the interface and user layer after fact accuracy verification is completed through a tense knowledge graph; the data layer is used for providing a multi-type data storage system; the large model layer outputs a paragraph-level text and a reference index based on a writing large model, the paragraph-level text is returned to the platform layer, and reference index information is fed back to the interpretability and security layer for tracing; and the infrastructure layer is used for providing bottom layer resource support for each layer.
Owner:INSPUR QILU SOFTWARE IND

Multi-modal digital publishing intelligent checking system and method based on large model

The invention relates to the technical field of document intelligent review, in particular to a multi-mode digital publishing intelligent review system and method based on a large model, and the system comprises a sample construction module, a multi-mode analysis module, a structure recognition module, a term verification module and an annotation output module. According to the method, a model configuration data set is generated by analyzing a subject-object combination relationship in a text and jointly screening part-of-speech, sentence pattern and semantic structure, the accuracy of semantic recognition and structure judgment is improved, and local features of semantic offset and image-text expression separation are recognized by combining a comparison relationship between an image action path and a text behavior object; a paragraph logic structure and a theme connection mode are analyzed based on semantic offset information, the problems of content dislocation and theme jump between paragraphs are effectively revealed, matching change tracks and part-of-speech continuation fluctuation in term cross-paragraph contexts are recognized, label overlapping and coverage redundancy conditions are analyzed from sentence group hierarchy, and the term cross-paragraph contexts are obtained. And forming a label combination suggestion and constructing an annotation structure record.
Owner:PEOPLES HEALTH ELECTRONIC AUDIO VISUAL PUBLISHING HOUSE CO LTD

Collaborative question and answer method for complex hierarchical table and large language model

The invention discloses a complex hierarchy table and large language model collaborative question and answer method. The method comprises the steps of S1, obtaining a user question, a complex hierarchy table and an unstructured text paragraph associated with the complex hierarchy table; s2, preprocessing the complex hierarchical table to obtain a multi-modal input sequence; s3, performing joint coding on the sequence by adopting a LayoutLM model, and obtaining a two-dimensional table with a reduced structure through a graph convolutional network; s4, based on the user question, performing fine-grained evidence retrieval on the two-dimensional table and the unstructured text paragraph to obtain a mixed fine-grained evidence set; and S5, taking the user question, the mixed fine-grained evidence set and the specific task prompt template as input of a large language model, and outputting to obtain a final answer of the user question. According to the method, the problems that a complex hierarchical table structure is difficult to understand and the numerical reasoning performance of the mixed context of the table and the text is insufficient are solved, and the high-precision and generalizable table question and answer ability is achieved.
Owner:CHONGQING JIAOTONG UNIV

Method for editing handwritten notes and electronic equipment

The invention provides a method for editing handwritten notes and electronic equipment, relates to the technical field of terminals, and is used for solving the problem of how to realize tidy and attractive handwritten note layout in a complicated handwritten note scene. According to the scheme, the method comprises the steps that a note interface is displayed on a display screen, handwritten notes are displayed on the note interface, and the handwritten notes comprise one or more paragraphs; receiving an editing instruction which acts on the note interface and aims at the handwritten note; in response to the editing instruction, the character position in at least one paragraph is adjusted to obtain an edited handwritten note, the line angle in any one paragraph in the at least one paragraph in the edited handwritten note is the same as the angle of the paragraph, and / or the line angle in any one paragraph in the edited handwritten note is the same as the line angle in any one paragraph in the at least one paragraph in the edited handwritten note. The position of each row of characters in any paragraph in the paragraph meets a first requirement, and the at least one paragraph is a paragraph acted by the editing instruction; and displaying the edited handwritten note.
Owner:HUAWEI TECH CO LTD

Multi-dimensional engineering paper abstract evaluation method

PendingCN121031578ASemantic analysisBiological modelsBasic languageEngineering
The invention discloses a multi-dimensional engineering paper abstract evaluation method, and belongs to the field of natural language processing. The implementation method comprises the following steps of: performing text correctness evaluation on a basic language specification in an abstract text, wherein a text correctness evaluation dimension comprises two sub-dimensions of grammar correctness and an expression specification degree; a pre-training language model GPT-2 is used as an evaluation base model, training fine tuning is carried out by using an abstract text in an RAAMove training set, and fluency dimension in abstract language paragraph fluency is evaluated; introducing a speech step theory, and constructing a speech step classification model based on comparative learning; performing double representation learning on sentence feature representation and speech step tag feature representation by adopting a supervised comparative learning method; performing coherence evaluation in paragraph fluency according to a speech step transfer similarity index; a ROUGE 1F1 value is calculated, and semantic similarity evaluation is carried out; and introducing a weighted summation strategy to carry out weighted summation on the scores of the dimensions to obtain a comprehensive evaluation score of the abstract text, namely realizing comprehensive evaluation on the abstract text.
Owner:BEIJING INST OF TECH

Text review method and device, storage medium and program product

The invention discloses a text review method and device, a storage medium and a program product, and relates to the technical field of artificial intelligence, the method comprises the steps that a domain knowledge base is constructed in advance, the domain knowledge base comprises a high-risk word set and element related knowledge, and the high-risk word set comprises violation words, violation types corresponding to the violation words and words needing element checking; the element-related knowledge includes provisions for essential elements of the compliance text, and compliance expression suggestions when the compliance text relates to a target description. When a target text is examined, paragraphs which contain high-risk words and contain the high-risk words meeting preset dangerous conditions are determined in the target text based on a high-risk word set to serve as candidate paragraphs, then element checking is conducted on each candidate paragraph based on element related knowledge, and an element checking result of each candidate paragraph is obtained; and generating an examination result of each candidate paragraph. Through two-stage screening, the examination efficiency is improved, and meanwhile, the examination accuracy is ensured.
Owner:ANHUI IFLYTEK INTELLIGENT SYST

Document segmentation method and device based on large language model, equipment and storage medium

The invention provides a document segmentation method and device based on a large language model, equipment and a storage medium, and relates to the technical field of text processing. The method comprises the steps of inputting a to-be-segmented target document into a pre-trained large language model, and executing the following operations through the large language model: performing text layout analysis on the target document, and identifying titles and all paragraphs of each level in the target document; for each paragraph, inserting an associated title related to the paragraph in all titles into an initial position of the paragraph to obtain a corresponding target paragraph; and sorting all the target paragraphs based on the semantic similarity among all the target paragraphs, and determining a segmentation result of the target document based on all the sorted target paragraphs. By the adoption of the technical scheme, when document segmentation is carried out, semantic loss in the document segmentation process can be effectively reduced, and therefore the document segmentation effect is improved.
Owner:CHINA LIFE ASSET MANAGEMENT CO LTD

Intelligent agent construction method and system based on large language model, equipment and medium

The invention provides an agent construction method and system based on a large language model, equipment and a medium, and belongs to the technical field of artificial intelligence. The method comprises the steps that a user interaction module is called to receive a first target document input by a user; calling a document analysis module to analyze the first target document to obtain a plurality of first chapter titles; calling an automatic mapping module to classify the plurality of first chapter titles based on a document classification model, and constructing a mapping relationship between the plurality of first chapter titles and a plurality of second chapter titles of a second target document based on a classification result; calling a large language model to respectively extract and summarize paragraph contents of the plurality of first chapter titles to obtain paragraph contents of a second chapter title; and calling a document output module to output the second target document according to a preset format. According to the agent construction method and system based on the large language model, the equipment and the medium provided by the invention, the accuracy of generating the review report by the agent can be improved.
Owner:ZHONGJIAO ROAD & BRIDGE (HEBEI) CO LTD

Entity enhancement and context-aware paragraph retrieval method for RAG system

The invention provides an entity enhancement and context-aware paragraph retrieval method for an RAG system for solving the problems of fuzzy query intention and insufficient paragraph context modeling in RAG retrieval. The method comprises the following steps: firstly, identifying and extracting a key entity by using a named entity, carrying out weighted fusion on the key entity and question representation obtained by a pre-training model to form an enhanced query vector, and accurately describing a core semantic intention of the question; secondly, performing semantic modeling on paragraphs in a document library, mining a potential semantic association relationship between the paragraphs, and constructing a context interaction model between the paragraphs based on a graph neural network and a gating loop unit mechanism, so as to obtain paragraph vector representation with complete semantics and clear hierarchy; and finally, calculating the similarity between the enhanced query and the paragraph vector, and completing high-precision paragraph-level retrieval. According to the method, the correlation and the recall rate are remarkably improved, insufficient entity utilization and weak context modeling are relieved, and good expansibility and cross-domain applicability are achieved.
Owner:SOUTHEAST UNIV +1

RAG question and answer traceability positioning method and device based on text coordinate index

The invention discloses an RAG question and answer traceability positioning method and device based on a text coordinate index, relates to the technical field of RAG and traceability positioning, and solves the problems that in the prior art, coordinate positioning precision is poor, and traceability in an original text is difficult after deconstruction. The method comprises the following steps of: analyzing and storing a text and corresponding coordinates into a mixed retrieval mode of Elasticsearch, meanwhile, carrying out vectorization storage on text fragments, deconstructing an original text paragraph or keyword in advance, storing the original text paragraph or keyword into mongoDB, carrying out vector similarity matching with Elasticsearch to find the original text when in use, and carrying out uniform format conversion on an input file, so that the extraction precision of the original text coordinates is higher, and the extraction efficiency is improved. The problem of tracing difficulty caused by RAG model text reconstruction is avoided to the greatest extent by utilizing an Elasticsearch technology, so that the most accurate tracing result can be obtained.
Owner:ZHEJIANG CONSTR INVESTMENT INNOVATION TECH CO LTD

Text slicing and recall method and device, electronic equipment and storage medium

The invention relates to a text slicing and recall method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a to-be-processed document which comprises at least one original paragraph; for each original paragraph, semantic segmentation is carried out on the original paragraph to obtain at least one text block, and the word number of the text block is smaller than or equal to the maximum word number limit; according to a title structure of the to-be-processed document, a title path of each original paragraph is extracted, feature extraction is carried out on the title paths, title feature vectors are obtained, and the title paths are used for indicating titles of each level to which the original paragraphs belong; and for each text block, taking the title feature vector of the original paragraph to which the text block belongs as an index field of the text block to obtain a text slicing result, and using the index field to retrieve the text block. In this way, the semantically related text blocks are associated during retrieval, and the situation that related information is dispersed is reduced, so that the retrieval accuracy is improved, and the reliability of generating answers by the RAG system is improved.
Owner:BEIJING JINGTOU ZHUOYUE TECH DEV CO LTD

Multi-modal mixed document OCR (Optical Character Recognition) and structured extraction method

The invention relates to the technical field of content extraction, in particular to a multi-modal hybrid document OCR (Optical Character Recognition) and structured extraction method, which comprises the following steps of: acquiring an image text region bounding box and classifying a style, extracting a font or stroke sequence to generate a character positioning structure, dividing paragraph and sentence groups to classify semantic fields, and calculating a field matching relationship to generate structural mapping. According to the method, logic mapping is constructed through character two-dimensional coordinate sorting and paragraph contours, complete reconstruction of a page structure is enhanced, semantic field categories are extracted by using syntactic density of sentence paragraph division and inter-paragraph features, the accuracy of field classification is improved, and the method has the advantages of being simple in structure, convenient to operate and high in practicability. A field mapping relation is established through Jaccard similarity and part-of-speech consistency analysis between a head word and a standard field keyword, a field path index and a structure node link are clarified, field semantic affiliation and structure position output are unified, and the document structure reduction degree and field extraction accuracy are improved.
Owner:HANGZHOU JINGSHENG HANGXING TECH CO LTD

English teaching story intelligent generation method and device, equipment and medium

The invention provides an English teaching story intelligent generation method and device, equipment and a medium, and relates to the technical field of the English teaching story intelligent generation method and device, and the method comprises the steps: obtaining the input content of a user; performing semantic analysis on the input content, and identifying a semantic intention vector and a target learning vocabulary of the user; based on the semantic intention vector, utilizing a retrieval enhancement generation technology to execute query retrieval in a preset corpus, and extracting text paragraphs from the query retrieval; constructing a retrieval result containing the target learning vocabulary based on the text paragraph; a pragmatic repeated distribution constraint mechanism is applied, and the pragmatic repeated distribution constraint mechanism is used for controlling the occurrence frequency and position distribution of the grammatical structure of the target learning vocabulary in the generated content; and taking the semantic intention vector and the retrieval result as input, and generating target English story content in combination with a pragmatic repetitive distribution constraint mechanism. According to the method, the distribution of the target vocabularies can be controlled, so that the accuracy of English teaching story generation is improved.
Owner:ZHEJIANG RONGQI MANUFACTURING TECHNOLOGY CO LTD

Professional process enhancement generation method, electronic equipment and computer program product

The invention provides a professional process enhancement generation method, electronic equipment and a computer program product. The professional process enhancement generation method comprises the following steps: acquiring associated context information of a paragraph to be generated from a memory according to reported state information of the paragraph to be generated; constructing a cue word of a paragraph to be generated, wherein the cue word comprises the associated context information; inputting the cue word into a large language model, and enabling the large language model to generate the content of the paragraph to be generated according to the associated context information; and extracting context information according to the content of the paragraph to be generated, and updating the context information into the memory. According to the method, the stability of the hierarchical structure of the paragraph of the generated content can be ensured, the style consistency and the logic coherence in long text generation are improved, and the quality of a vertical professional field report generated by a large language model is improved.
Owner:BEIJING JIAYUE DIGITAL INTELLIGENCE TECHNOLOGY CO LTD +1

Text information identification method and device, equipment and storage medium

The invention provides a text information recognition method and device, equipment and a storage medium. In some embodiments of the present disclosure, a first paragraph text and a second paragraph text of an audit report text are obtained; performing sentence segmentation processing on the first paragraph text and the second paragraph text to obtain a first sentence of the first paragraph text and a second sentence of the second paragraph text; combining the first clause and the second clause pairwise to obtain a sentence pair; encoding each sentence pair to obtain a sentence vector corresponding to each sentence pair; performing semantic emotion recognition on the sentence vector corresponding to each sentence pair to obtain a consistency result of each sentence pair; determining consistency information of the first paragraph text and the second paragraph text according to a consistency result of each sentence pair; based on semantic emotion recognition in the field of natural language processing, the audit report text is subjected to consistency auditing automatically, the labor cost is reduced, the auditing efficiency is improved, and the auditing accuracy is improved.
Owner:PICC INFORMATION TECH CO LTD

Research report generation method based on natural language processing

PendingCN122655689Aavoid enteringEliminate distraction issuesTimestampEngineering
The application discloses a research report generation method based on natural language processing and relates to the technical field of intelligent document analysis and generation, which comprises the following steps: version storage of source documents and generation of timestamp snapshots; paragraph-level and sentence-level segmentation of the timestamp snapshots and extraction of evidence segments carrying original text positioning indexes; core claim extraction of the evidence segments to generate conclusion units containing claim texts and confidence; establishment of the association relationship between the conclusion units and the evidence segments; generation of chapter texts according to the associated conclusion units and calculation of evidence segment coverage; and structural adjustment and regeneration of the texts when the evidence segment coverage is lower than a preset threshold. The method realizes evidence traceability, multi-type semantic association and automatic optimization based on coverage feedback.
Owner:BEIJING ZHENLI HENGYUAN TECHNOLOGY CO LTD

An ERNIE_CN-GRU step automatic recognition method, system, device and medium

The application provides an ERNIE_CN-GRU sentence step automatic identification method, system, device and medium, complete data of a paragraph is acquired, and a data set is constructed; an ERNIE pre-training model is built, a multi-head self-attention mechanism is fused to learn text semantics to obtain a word vector feature matrix of the multi-head attention mechanism; CN-GRU feature identification network training is carried out based on the ERNIE pre-training model and the word vector feature matrix, and an ERNIE_CN-GRU model is formed; the data set is input into the built ERNIE_CN-GRU model, a Softmax classifier is accessed to realize sentence step identification, and an identification label is output; the ERNIE pre-training model combining large-scale text content and a knowledge graph is used to learn deep text semantics, the disadvantages that a traditional machine learning does not sufficiently mine and utilize the internal relationship and features between words are improved, and good transferability and robustness are obtained.
Owner:XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY

Key point information generation method and device, equipment and storage medium

The invention provides a key point information generation method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence, in particular to the technical fields of deep learning, natural language processing, intelligent documents and the like. The method comprises the steps of generating a segmentation outline according to document contents of paragraphs in a to-be-processed document; according to preset theme information and the segmentation outline, paragraphs in the to-be-processed document are aggregated, and multiple aggregated theme contents are obtained; the theme content is segmented according to preset sub-themes, multiple segmented sub-theme content is obtained, and the sub-themes are sub-classes of themes obtained by dividing the theme information; and generating key point information for the to-be-processed document according to the information of the paragraphs in the multiple sub-theme contents. According to the method, the generation time of the document key point information is shortened, and the generation efficiency of the key point information is improved.
Owner:BEIJING DUYOU INFORMATION TECH CO LTD

System and method for identifying variations of phrases in text paragraphs

A system and method of identifying, by at least one processor, an occurrence of a semantic variant of a phrase in a paragraph may include calculating a phrase embedding vector representing a semantic meaning of the phrase; extracting at least one hierarchical set of nested sequences of words from the textual representation of the paragraph; calculating, for each sequence, a corresponding sequence embedding vector representing the semantic meaning of the sequence; calculating, for one or more sequence embedding vectors, a corresponding vector similarity value representing a similarity of the sequence embedding vector and the phrase embedding vector; identifying a sequence corresponding to a maximum vector similarity value of the one or more vector similarity values; and determining the identified sequence as a semantic variant of the phrase based on the maximum vector similarity value.
Owner:GENESIS CLOUD SERVICES CO LTD

Enhanced retrieval generation method and device, electronic equipment and storage medium

The invention discloses an enhanced retrieval generation method and device, electronic equipment and a storage medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: receiving a query statement input by a user, and analyzing the query statement to obtain a sub-graph matching template; performing sub-graph isomorphic matching on the hypergraph structure based on the sub-graph matching template to obtain one or more target hyperedges conforming to the sub-graph matching template; wherein the hypergraph structure comprises a plurality of hyperedges, and each hyperedge is connected with a plurality of entity nodes to represent a composite event; determining a target paragraph list associated with the target hyperedge; and based on a target paragraph pointed by the target paragraph list and the query statement, constructing a target input sequence, inputting the target input sequence into the natural language model, and generating a target answer text matched with the query statement. According to the method, the retrieval accuracy and reliability can be improved, and the retrieval efficiency is considered.
Owner:IFLYTEK CO LTD

Global fluctuation analysis methods, devices, and storage media for text

This application discloses a method, device, and storage medium for global fluctuation analysis of text. The method, relating to the field of text parsing technology, includes: processing the input text using a first model to obtain a first parsing result, and processing the input text using a second model to obtain a second parsing result; determining the sentence-level raw evaluation value of each sentence in the input text; determining the paragraph-level raw evaluation value of each paragraph in the input text based on the first parsing result, the second parsing result, and the sentence-level raw evaluation value; connecting the paragraph-level raw evaluation values ​​of each paragraph into an expression intensity curve using the paragraph sequence of the input text as a time axis; and identifying and labeling four geomorphic features—peak, valley, data slope, and data plain—based on the expression intensity curve to generate a visualized scoring geomorphic map. This application achieves accurate prediction and hierarchical selection of training corpus quality before model training.
Owner:LINGE TECHNOLOGY CO LTD

A text segment retrieval method, an electronic device, and a storage medium

This invention provides a method, electronic device, and storage medium for recalling text fragments, relating to the field of text fragment recall technology. The method includes: acquiring a target semantic vector T0 and a preset text fragment semantic vector library T; determining initial text fragment semantic vectors; clustering text fragment semantic vectors with the same article number into the same cluster and arranging them sequentially according to the paragraph order corresponding to the text fragment semantic vectors to obtain a cluster list A; and, based on f(j), retrieving text fragments from N... j A was identified in the corresponding text fragment. j The corresponding number of recalled text fragments; if N j For the pre-defined multi-event article numbering, then based on f(j) and N j The degree of correlation between any two text segments in the corresponding text fragment, from N j A was identified in the corresponding text fragment. j The invention provides several corresponding recall text fragments; it can significantly improve the actual relevance between the recalled text fragments and the target question, thereby increasing the accuracy of subsequent answer generation.
Owner:MOBILE TECH COMPANY CHINA TRAVELSKY HLDG

Prompt interference construction and optimization method and device based on fragment semantic cross combination

The present application relates to the technical field of artificial intelligence and natural language processing, and particularly relates to a prompt interference construction and optimization method and device based on fragment semantic cross combination. The method comprises the following steps: constructing candidate prompt combinations based on paragraph-level prompt fragments and sentence-level prompt fragments, screening initial effective prompt combinations based on semantic consistency with original prompt text, splicing the initial effective prompt combinations with attack target instructions to form a risk prompt, obtaining a risk response and harmfulness score corresponding to the risk prompt through a language model to calculate an initial Δ-TRDS score; performing a neutral content replacement operation on all fragments in the initial effective prompt combinations in sequence, and calculating a target Δ-TRDS score corresponding to a target effective prompt combination after replacement; and determining a prompt interference fragment in the initial effective prompt combination according to the difference between the Δ-TRDS scores before and after replacement. The present application solves the problem of lack of systematicness and scalability of the traditional manual prompt construction method.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

A machine reading comprehension method based on repeated span prediction

The present application relates to a kind of machine reading comprehension method based on repeated span prediction, belong to natural language processing machine reading comprehension field.The method includes: for the reading comprehension task of span prediction form, increase the task of predicting repeated span.The task first finds all repeated spans in text paragraph based on greedy algorithm, then the short span contained in long span is filtered, after obtaining repeated span set, for each group of repeated span, randomly select one as answer span, and other spans in group are replaced by mask.Processed text is input into pre-training model, and representation vector is obtained, and based on softmax, which span in paragraph should the mask position point to be predicted.In the task, the obtained model is further fine-tuned on target task.The method considers the problem of lack of span knowledge of pre-training model, and constructs data in unsupervised manner, so that the model can better learn span representation, and improve the performance of model in span prediction.
Owner:BEIJING INST OF TECH

Semantic segmentation method and device for restoring article structure and electronic equipment

PendingCN122634124AData setSemantics
The application is applied to the technical field of natural language processing, and discloses a semantic segmentation method and device for restoring article structure and electronic equipment. The method comprises the following steps: obtaining a text to be segmented; extracting a hierarchical structure of the text to be segmented to obtain a structured dictionary corresponding to the text to be segmented; the structured dictionary stores a chapter title set, a paragraph set, a list set and a table data set of the text to be segmented; performing sentence-level segmentation on the text stored in the structured dictionary to obtain a first candidate segmentation sequence; the first candidate segmentation sequence stores a plurality of first candidate text segments; and the first candidate text segments in the first candidate segmentation sequence are merged according to semantic coherence between adjacent first candidate text segments to obtain a text segment. In this way, the text segment can retain context-related information to the greatest extent, the semantics between segments are relatively independent, the semantic segmentation of the text to be segmented is realized, and the subsequent retrieval accuracy can be improved.
Owner:LESHAN POWER SUPPLY COMPANY STATE GRID SICHUAN ELECTRIC POWER

Full-text retrieval method and system, computer equipment and computer readable storage medium

The invention belongs to the field of information retrieval, particularly relates to a full-text retrieval method and system, computer equipment and a computer readable storage medium, and aims to solve the problem of improving the full-text retrieval accuracy. The method comprises the following steps: segmenting a document entering a corpus; calculating paragraph weights of words contained in each segmented document in the segmented document; calculating the document weight of the word in the document according to the paragraph weight of the word; calculating the query weight of the query word, wherein the calculation method of the query weight is the same as the calculation method of the paragraph weight; determining a corresponding target word in a corpus according to the query word; respectively calculating one or more query relevancy according to the query weight and the document weight of the target word; and taking the document corresponding to the document weight corresponding to the maximum n query relevancy as a query result. Semantic features are introduced into weight calculation, and document segmentation processing is combined, so that the full-text retrieval accuracy is effectively improved.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD +1