Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

442 results about "IntraText" patented technology

IntraText is a digital library that offers an interface while meeting formal requirements. Texts are displayed in a hypertextual way, based on a Tablet PC interface. By linking words in the text, it provides Concordances, word lists, statistics and links to cited works. Most content is available under a Creative Commons license It also offers publishing services that enable similar advantages.

Document interpretation and report generation method and device, equipment and medium

The invention relates to the technical field of natural language processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a document interpretation and report generation method, device, equipment and medium, which comprises the following steps: receiving an original document set to generate a structured document object, executing optical character recognition on an image content set to generate a recognition text set, the recognition text set and the text content set are combined into a unified text sequence, element item extraction is executed based on the interpretation template parameter set to generate an interpretation element set, a retrieval enhancement context is retrieved and generated from the domain knowledge base, and the unified text sequence, the interpretation template parameter set and the retrieval enhancement context are input into a language model to generate an interpretation result. And generating report content based on the historical report template set. According to the method, automatic closed loop of document interpretation and report generation is realized through multi-modal unified processing and semantic enhanced reasoning, the efficiency is improved, and the manual dependence and compliance risk are reduced.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Document format processing method and system based on large language model, terminal and medium

The invention relates to the field of document processing, and particularly provides a document format processing method, system, terminal and medium based on a large language model.The method comprises the steps that a generation instruction containing a target document type and a source material are received, and original text content is generated through the large language model integrating domain knowledge; obtaining a matched structured format template, and analyzing the style rule into a format instruction set; identifying logic elements and hierarchical relationships thereof in the original text based on a natural language processing technology; performing association mapping on the format instruction and the logic element, and performing automatic style rendering by calling a document object model interface to generate an intermediate document with a standard format; and finally, outputting after quality verification. According to the method, an automatic process of content generation and intelligent formatting is constructed, and the document writing efficiency and normalization are improved.
Owner:浪潮智慧科技有限公司 +2

Systems, apparatuses, methods, and non-transitory computer-readable storage media for adaptive information retrieval for question-answering

Methods and systems for retrieving relevant information in response to an input question. The method includes obtaining text content related to the input question and partitioning the content into one or more paragraphs based on predefined rules. The method further involves extracting one or more evidence spans that are relevant to the input question by inputting the text content and the question into a trained language model. A semantic search is then performed on both the paragraphs and the extracted evidence spans, ranking the candidate passages based on their relevance to the input question. Each candidate passage may comprise either a paragraph or an evidence span that addresses the question. The disclosed methods and systems improve the quality and relevance of retrieved information by combining heuristic-based content partitioning with machine learning-based evidence extraction.
Owner:HUAWEI TECH CO LTD

Knowledge document-oriented method for structurally and stably presenting large language model output

The invention discloses a knowledge document-oriented method for structurally and stably presenting large language model output. The method comprises the following steps of: preprocessing a knowledge document and converting the knowledge document into a text by a system; the back-end service module adopts a prompt word engineering technology based on a preset JSON mode, constructs a structured prompt containing clear instructions and output format constraints, calls a large language model to carry out analysis and information extraction on text contents, and forces the model to generate structured JSON data following the preset mode; the back-end service provides access to the structured JSON data for the front-end application; and the front end executes templated mapping rendering according to a preset user interface component template isomorphic to the JSON mode, and fills the content of each field of the JSON into a corresponding visual component. According to the method, through the cooperation of the back-end constraint prompt word engineering and the front-end modularized template rendering, the defects that the output content format of a large language model is unstable and is difficult to be directly applied to standardized interface presentation are relieved.
Owner:BOBAN ZHIJIE (BEIJING) TECHNOLOGY CO LTD

Generating probabilistic data structures for lookup tables in computer memory for multi-token searching

Methods, systems, and non-transitory computer readable storage media are disclosed for optimizing computer memory usage for lookup lists in computer memory via probabilistic data structures. For example, the disclosed system generates a probabilistic data structure (e.g., a Bloom filter) to represent data in a lookup list including multi-token items by hashing items of the lookup list to sets of bit values in a bit vector. The disclosed system classifies text content in a digital document by utilizing a maximum number of tokens from multi-token items in the lookup list to select and compare sets of sequential tokens in the digital document to the probabilistic data structure. The disclosed system also iteratively reduces the number of tokens in sets of sequential tokens for subsequent comparisons. Furthermore, in some aspects, the disclosed system causes a computing device to modify a digital document and / or database operations based on the classifications.
Owner:ONETRUST LLC

AIGC content security monitoring system and method based on dynamic reasoning and context awareness

The invention relates to the technical field of natural language processing, in particular to an AIGC content safety monitoring system and method based on dynamic reasoning and context awareness, and the system comprises a semantic graph construction module which is used for extracting entity nouns and predicate verbs in a text according to a received AIGC interaction text flow, generating a semantic concept node set, and sending the semantic concept node set to a database; and performing directed connection and hierarchical nesting on the concepts in the semantic concept node set according to a logic direction according to a subject-predicate-object dependency relationship rule. According to the method, the curvature value of the semantic track is calculated, and the similarity between the direction vector and the center of the sensitive semantic cluster is combined for double verification, so that sudden turning of an intention in a dialogue process or progressive induction to a sensitive field can be perceived, and abnormal mutation can be recognized through a curvature pulse form; therefore, hostile attack behaviors are accurately captured in real time in dynamic interaction, and the defense capability for context dependent attacks and implicit induction behaviors is improved.
Owner:XINGXUAN DIGITAL TECHNOLOGY (SHANGHAI) CO LTD

Heterogeneous document structured data extraction system and method based on multi-modal fusion

The invention discloses a heterogeneous document structured data extraction system and method based on multi-modal fusion, and the method comprises the following steps: S1, receiving a heterogeneous document, and carrying out the preprocessing; s2, extracting visual and semantic modal information based on the visual backbone network and an OCR module, and fusing spatial coordinates, text content and layout features; s3, receiving a dynamic target mode Schema, executing field-level semantic matching, and calculating semantic similarity; s4, performing consistency verification, verifying and filling the field values, and correcting the associated fields; s5, converting the format into a data file in a specified format, and reserving fields to be mapped with an original document semantic block; s6, automatically switching a cloud mode and a local mode according to a deployment environment; and S7, outputting the data file of the target field. According to the method, high-precision and low-delay structured data extraction of the heterogeneous format document can be realized, and the automatic processing efficiency and the data credibility are improved.
Owner:ANHUI HANGTIAN INFORMATION CO LTD

Electronic file content anti-counterfeiting and tampering detection method adopting locked signature

The invention discloses an electronic file content anti-counterfeiting and tampering detection method adopting a locked signature, and belongs to the field of information security. According to the method, file content is automatically segmented according to semantics through a context coding self-adaptive segmentation algorithm, and each segment of content fingerprint is generated by applying a content embedding disturbance weighted digest algorithm. And constructing each segmented abstract into a directed relation chain by adopting an embedded relation chain transformation signature mechanism, and carrying out digital signature overall locking. During detection, the segmented abstract and the relation chain structure are compared again, and precise positioning and tracing of tiny tampering, segmented change and structure adjustment of the text content are achieved. And through a signature duration chain entropy analysis algorithm, performing comparison and evolution path analysis on previous version states, and outputting a detailed traceability report. According to the method, authenticity verification and tampering detection of the electronic file content are greatly improved, and the method is suitable for the fields of data anti-counterfeiting, compliance supervision, electronic evidence management and the like.
Owner:CHINA ELECTRONICS STANDARDIZATION INST

Document adaptive conversion method and device based on multi-modal large model and medium

The invention provides a document adaptive conversion method and device based on a multi-modal large model and a medium, and belongs to the technical field of data processing. The method comprises the following steps: inputting a multi-modal document into a pre-constructed multi-modal large model, wherein the multi-modal large model comprises a multi-modal joint framework fusing a ViT visual model and an LLM model; using a ViT visual model and an LLM model to respectively extract multi-scale visual features corresponding to non-text content and semantic features corresponding to text content, and using a cross-modal attention mechanism to bidirectionally align the multi-scale visual features and the semantic features to obtain a multi-modal document; carrying out self-adaptive blocking on the multi-modal document by using a self-adaptive blocking strategy, and modeling a relative position relationship between blocks; establishing a dynamic mapping rule from a document element to an HTML tag, and mapping the multi-modal document into a front-end interaction component in combination with a natural language instruction of a user; and dynamically rendering the front-end interaction component in an on-demand mounting and resource isolation mode. The problem of insufficient document restoration capability in the prior art can be solved.
Owner:SHANDONG INSPUR DIGITAL BUSINESS TECHNOLOGY CO LTD

Document translation method and system, computer equipment and storage medium

The invention provides a document translation method and system, computer equipment and a storage medium, and the method comprises the steps: taking a document associated with steel as a target document, then carrying out the structured fragmentation of the target document, and obtaining a more precise semantic unit; for the content of each fragment, retrieving professional terms from a preset translation term contrast library, constructing prompt words, and performing translation term matching on the text content of each fragment according to the preset translation term contrast library, so that the translation terms of the same text content are the same; inputting the term cues and the original text into a multi-stage large language model, and respectively carrying out literal translation processing, quality evaluation and optimization suggestion, so as to generate interpretation content which is strong in specialty and naturally expresses; and finally, integrating and outputting all the fragmented contents into a structured complete translation document. The translation quality of technical data in the iron and steel industry can be remarkably improved, and the problems that an existing general translation system is not uniform in terms, split in structure, stiff in expression and the like are solved.
Owner:CISDI RES & DEV CO LTD

Generation of Interactive Data Visualizations and Textual Content

A system can be used to generate interactive data visualizations and textual content. The system receives a layout of a content item. The content item can be a data visualization dashboard or an article. The system can generate a content structure tree based on the layout. The content structure tree represents hierarchical and semantic relationships between sections of the content item. The system can receive input adding content elements, such as data visualizations or textual paragraphs, to sections within the content item. The system can identify text roles for different sections based on the content structure tree and the content elements. The system can generate text suggestions for the content item based on the text roles. The system can also generate text content for a text suggestion using a large language model, display the generated text content within the content item, and iteratively refine the content item.
Owner:SALESFORCE INC

Tokenization systems and methods for redaction

A tokenization system receives a request for redaction of sensitive textual content in a document, identifies a portion of the document as the sensitive textual content, and edits the document, including replacing the sensitive textual content thus identified with tokens, each token having a token value and a pattern that identifies a start and an end of the token value. The editing produces a transformed version of the document with the tokens and without the sensitive textual content. The tokenization system may then communicates the transformed version of the document with the tokens and without the sensitive textual content to the client computing system, an automated recognition service, or a redaction plug-in to a frontend application.
Owner:OPEN TEXT CORPORATION

Method and system for generating executive summaries and data visualizations for annual product quality reports

The present disclosure provides a system and method for generating text summary and data visualizations for executive summaries of annual product quality reports based on an input query from a user. The method for generating text summary for annual product quality reports (APQR) of an enterprise, comprising: providing one or more indexes corresponding to one or more categories of data stored in a database, receiving a query from one or more users, wherein the query is a text and based on the at least one of the indexes of the one or more categories of data, extracting one or more relevant textual content from the query using natural language processing, querying the database for retrieving the one or more indexes of the relevant textual content, providing one or more few shot prompts for generating an output in a predetermined format based on the query received from the one or more users, inputting, the one or more few shot prompts and the retrieved one or more indexes to a large language model (LLM), generating a text summary based on the input to the LLM and rendering the generated text summary at a user device.
Owner:HONEYWELL INTERNATIONAL INC

Machine learning techniques for improved content generation

Techniques for content generation using machine learning. A device may access textual content from a user device associated with a user profile, access content preferences associated with the user profile, and provide the textual content to a classification model to identify one or more topics of the textual content. The classification model may be trained to identify topics within text. The device may access, from a content repository, supplementary textual content associated with the one or more topics; form, based on the content preferences, an instruction prompt for a generative model; and provide the topics, the textual content, and the supplementary textual content to the generative model to obtain additional content. The device may identify, from the user profile, one or more additional user profiles having a relationship with the user profile; and provide the additional content to an external server.
Owner:VAN WIE DAVID +4

Intelligent text generation method and system based on natural language processing

The invention discloses an intelligent text generation method and system based on natural language processing, and relates to the technical field of natural language processing. The method comprises the following steps: firstly, combining a natural language processing technology and a deep learning algorithm, and carrying out multi-task identification analysis on an original text to obtain text original element information including styles, styles, emotions, entities and topics; performing fusion processing on the text generation control information input by the user and the style, the genre and the emotion in the element information to obtain a text generation control vector containing a target style, a target genre, a target emotion and a forbidden word set, and performing retrieval to obtain background knowledge and / or fact data of an original entity and a theme of the text; and finally, integrating the control vector and the retrieval result into a cue word, and importing the cue word into a large language model to output and obtain a new text, so that text contents with high quality, high correlation, high fact consistency and high controllability can be generated by deeply analyzing the original text and combining conditions specified by a user.
Owner:CHONGQING SILICON LIFE TECHNOLOGY CO LTD

A method for automatically analyzing an evaluation report

The application discloses a kind of methods for automatically analyzing and evaluating report, it is related to computer software and information processing technical field, including: based on graph neural network and visual-textual dual modal feature extractor, realize the adaptive analysis of complex non-standard structure report version, and the document is parsed into two-dimensional data table containing different classification dimensions;Based on BERT vectorization, etc. Adaptive clustering identifies the cluster structure of uneven semantic distribution, adopts the agglomerative hierarchical clustering to construct tree-shaped multi-granularity semantic hierarchy, realizes semantic aggregation;The application realizes the text content of non-standard form, the identification and processing of complex semantics;Overcome the defects of low efficiency, easy to make mistakes, subjective influence and only read specific format or specific location of text content, lack of flexibility and unable to handle complex semantics in prior art manual input.
Owner:中国华电集团有限公司北京数字科技分公司 +1

Providing suggested prompts for generating artificial intelligence (AI) content in a workspace

ActiveUS12499417B2Natural language translationOffice automationGenerative processWorkspace
A method for suggesting prompts on a page of a workspace includes receiving an input and displaying a prompt block configured to initiate a generative process to create in-block content in response to the input. The prompt block is embedded as an in-page object on the page. The method includes causing a large language model (LLM) system to create a set of suggested prompts. Each prompt includes instructions configured to create generative content of a respective type by the LLM system. The set of suggested prompts is created based on in-page text content or a relative location of the prompt block on the page. The method includes displaying the set of suggested prompts as a set of control items of the workspace. Each of the set of control items is selectable to input as a prompt for generating content based on existing content of the workspace.
Owner:NOTION LABS INC

Multi-agent cooperation and dynamic feedback long text generation system and method

PendingCN121413623ASemantic analysisBiological modelsEvaluation resultContinual improvement process
The invention provides a multi-agent cooperation and dynamic feedback long text generation system and method, and the system comprises a performance agent which is used for receiving a theme or an outline inputted by a user, planning and scheduling a long text generation process, coordinating the interaction of an author agent and an evaluation agent, and dynamically adjusting a generation strategy based on an evaluation result; the author agent is used for generating text content according to an instruction of the director agent and correcting a generation result; the evaluation agent is used for performing multi-dimensional quality evaluation on the text content generated by the author agent and returning optimization suggestions; and the memory bank is used for storing the historical generation information and the global setting information and providing context support in subsequent text generation. Through the synergistic effect of the director agent, the author agent, the evaluation agent and the memory bank, a dynamic feedback closed-loop mechanism is established, so that the text generation process has continuous improvement capability, and the overall quality and stability of long text generation are improved.
Owner:BEIJING JIBU QIANLI TECHNOLOGY CO LTD

Ultra-long text generation method and device, equipment and storage medium

The invention discloses a super-long text generation method and device, equipment and a storage medium, and the method comprises the steps: responding to a document generation instruction, and generating an initial text outline of a target text through a large model; wherein the initial document outline comprises a plurality of chapter titles; based on the keyword information of the initial document outline, the text complexity of the target text is obtained through calculation, and the text complexity is associated with the subject breadth of keywords and the expected word number of the target text; if the text complexity of the target text is greater than a preset threshold value, determining the initial text outline as a target text outline; and on the basis of the target text outline, the text content corresponding to each chapter title is generated in parallel and spliced to obtain the target text, so that the generation efficiency, the text coherence and the overall text quality of the generated super-long text can be improved, and the generation cost of the super-long text is reduced.
Owner:SHENZHEN YUEHUA EXPRESS CO LTD

Bidding document compiling method and system based on artificial intelligence, medium and product

The invention discloses a bid invitation file compiling method and system based on artificial intelligence, a medium and a product, and relates to the field of intelligent contract processing. The method comprises the steps that a file compiling system divides a natural language text into first text content containing computational logic and common second text content through semantic analysis; then, the first text content is converted into an executable formalized rule expression, and an association mapping relation between the executable formalized rule expression and the first text content is established; and calculating and generating a derivation record containing parameters, calculation steps and dependency relationships according to the expression. And finally, embedding a unique tracing identifier in the generated clause text, and querying and displaying a derivation record through the association mapping relationship by triggering the identifier by a user. According to the method, a complete traceable link from the natural language to the calculation result is realized, so that auditing personnel can clearly understand the logic source of the calculation result, the interpretability and the credibility of the generated calculation terms are improved, and the verification efficiency and the accuracy of the bid invitation file are improved.
Owner:GONGCHENG MANAGEMENT CONSULTING

Providing suggested prompts for generating artificial intelligence (AI) content in a workspace

PendingUS20260065228A1Natural language translationOffice automationGenerative processWorkspace
A method for suggesting prompts on a page of a workspace includes receiving an input and displaying a prompt block configured to initiate a generative process to create in-block content in response to the input. The prompt block is embedded as an in-page object on the page. The method includes causing a large language model (LLM) system to create a set of suggested prompts. Each prompt includes instructions configured to create generative content of a respective type by the LLM system. The set of suggested prompts is created based on in-page text content or a relative location of the prompt block on the page. The method includes displaying the set of suggested prompts as a set of control items of the workspace. Each of the set of control items is selectable to input as a prompt for generating content based on existing content of the workspace.
Owner:NOTION LABS INC

Sign language animation generation method and device based on semantic analysis, equipment and medium

The invention relates to the technical field of voice semantics, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a sign language animation generation method, device, equipment and medium based on semantic parse. The sign language animation generation method comprises the steps that voice input is received and recognized as text content, field semantic parse is conducted on the text content to generate a field semantic template, and the field semantic template is used for generating a sign language animation; and converting the domain semantic template into a sign language intermediate representation sequence, generating a three-dimensional sign language action sequence based on the sign language intermediate representation sequence, rendering the three-dimensional sign language action sequence into a virtual image sign language animation, and displaying the virtual image sign language animation. According to the invention, by fusing speech recognition, semantic analysis and three-dimensional action rendering, direct conversion from spoken language content to sign language animation is realized, and a complete visual expression link from speech to sign language is formed, so that a user can intuitively understand the speech content in a sign language form through a virtual image, and the user experience is improved. Therefore, the barrier-free performance of human-computer interaction and the accuracy of information transmission are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Inference based on different portions of a training set using a single inference model

Systems, methods and non-transitory computer readable media for inference based on different portions of a training set using a single inference model are provided. Textual inputs may be received, each of which may include a source-identifying-keyword. An inference model may be a result of training a machine learning model using a plurality of training examples. Each training example may include a respective textual content and a respective media content. The training examples may be grouped based on source-identifying-keywords included in the textual contents. Different parameters of the inference model may be based on different groups, and thereby be associated with different source-identifying-keywords. When generating new media content using the inference model and a textual input, parameters associated with the source-identifying-keyword included in the textual input may be used.
Owner:BRIA ARTIFICIAL INTELLIGENCE LTD

Adaptive RAG memory enhancement method

The invention belongs to the technical field of RAG memory enhancement, and particularly relates to a self-adaptive RAG memory enhancement method, which comprises the following steps of: 1, determining an application scene and task requirements; 2, a backbone LLM is initialized and configured; step 3, attribute mining and identification; step 4, attribute labeling granularity selection and labeling; step 5, evaluating and sorting attribute priorities; step 6, constructing memory instance embedding representation; step 7, implementing a retrieval strategy based on attributes and embedding; step 8, integrating and optimizing retrieval results; and step 9, performing experimental evaluation and continuous improvement, namely determining the specific application scene of the self-adaptive RAG memory enhancement technology, retrieving appropriate text content for recommendation according to the theme and requirements input by the user, and determining the specific requirements of the task according to different scenes. According to the method, memory information is enriched by deeply mining and marking key attributes, and meanwhile, two efficient retrieval means are adopted to ensure quick positioning of related information.
Owner:XIAODUO INTELLIGENT TECH (BEIJING) CO LTD

Text content review system based on artificial intelligence

The artificial intelligence-based text content review system comprises an analysis unit, a feature extraction unit, a rule optimization unit, a reasoning unit and a conflict detection unit, and the analysis unit is configured with a text-to-semantic association strategy and is used for performing collaborative analysis on multiple versions of a to-be-reviewed text or a context association text; the feature extraction unit executes multi-level extraction of text semantics through a configured feature tower dynamic weighting strategy; the rule optimization unit is used for improving the matching precision of an examination rule and a text feature through self-iterative optimization of a rule vector space; the reasoning unit is used for improving the reliability of a text review result through dual-path reasoning and consistency verification; the conflict detection unit improves the accuracy of text conflict detection through multi-scale positioning. The method remarkably improves the precision, efficiency and field adaptability of text content review, and is particularly suitable for text review scenes with high professional thresholds such as law and medical treatment.
Owner:BEIJING YUNQING INTELLIGENT TECHNOLOGY CO LTD

Automatic poster generation method and related equipment

The embodiment of the invention provides an automatic poster generation method and related equipment, and belongs to the cross technical field of computer vision and natural language processing. The method comprises the steps that firstly, through a detail insensitive demand analysis (DIPR) mechanism, user demands of different detailed degrees are analyzed into a structured design blueprint in a unified mode through a large language model subjected to instruction fine adjustment; secondly, based on the style description in the blueprint, a diffusion model subjected to style-based fine adjustment is called to generate a high-quality background image; and finally, fusing the text content and the background image by using a multi-modal large language model, and directly outputting an editable HTML format poster. According to the method, the problems of poor text rendering accuracy, weak demand understanding ability and non-editable output results in the prior art are fundamentally solved, the text accuracy is remarkably improved, full-process automatic generation is realized, and the method has excellent commercial practicability.
Owner:SOUTH CHINA UNIV OF TECH

Text matching method and device based on probability distribution, equipment and storage medium

The invention discloses a text matching method and device based on probability distribution, equipment and a storage medium, and relates to the technical field of natural language processing, and the method comprises the steps: obtaining semantic feature distribution of each professional knowledge text from a professional knowledge base, obtaining a knowledge probability distribution set, abstracting the intrinsic features of the text through probability distribution, and obtaining a knowledge probability distribution set; and standardized representation of the knowledge base is realized. The text input by the user is obtained, the corresponding user text probability distribution is calculated, and the fault tolerance of text matching is remarkably improved. The similarity distance between the user text probability distribution and each distribution in the knowledge probability distribution set is calculated, and the overall semantic similarity instead of local matching is measured through the distribution distance, so that the comparison process can tolerate the distribution offset. Finally, a text matching result is determined according to the minimum value of the similarity distance, the effect of stably retrieving semantic related professional knowledge under high noise is achieved, and fault tolerance and matching precision are remarkably improved.
Owner:HUBEI TAIYUE SATELLITE TECH DEV CO LTD

Text content translation with style preservation using attention heads

The present disclosure relates to systems, non-transitory computer-readable media, and methods for generating stylized translated text using attention heads from a transformer neural network. In particular, in some embodiments, the disclosed systems obtain an input text string in a first language, the input text string comprising a style formatting element. Additionally, in some embodiments, the disclosed systems generate, using a transformer neural network to process the input text string, a translated text string in a second language different from the first language. Moreover, in some embodiments, the disclosed systems determine attention head values generated by the transformer neural network for words of the input text string as part of generating the translated text string in the second language. Furthermore, in some embodiments, the disclosed systems generate a translated style formatting element for the translated text string based on the attention head values for the words of the input text string.
Owner:ADOBE INC

Method and device for improving RAG recall effect

The invention provides a method and device for improving an RAG recall effect, and belongs to the technical field of computers, and the method comprises the following steps: file input and typesetting structure analysis: identifying a typesetting unit for an input file, extracting text content in the typesetting unit, and generating associated data of typesetting and content; constructing a two-dimensional relation graph: performing semantic segmentation based on the typesetting units, and extracting a logic relation of the typesetting units; constructing a two-dimensional relation graph of the semantic relation and the typesetting relation; and multi-dimensional information fusion recall: receiving user query and performing semantic analysis, recalling similar semantic slices from a semantic community and a typesetting community, executing double-graph cross validation, dynamically adjusting weights, and calculating and obtaining a final recall result. According to the method, the knowledge base construction mode of the RAG is optimized from the perspective of typesetting, the multi-dimensional relation between text semantics and typesetting logic is fused, information association in the knowledge base is more comprehensive, and the retrieval recall can be based on the semantic similarity and the typesetting logic at the same time, so that the recall effect is remarkably improved.
Owner:KYLIN CORP

Document retrieval and conflict detection system based on retrieval enhancement generation

The invention discloses a document retrieval and conflict detection system based on retrieval enhancement generation, and the system comprises a document loading module which is used for analyzing a multi-format document and extracting text content; the text segmentation module is used for dividing the text content into semantic segments; the vector index module is used for converting the semantic fragments into vectors and constructing indexes; the retrieval module is used for receiving user query and retrieving related semantic fragments; the conflict detection module is used for inputting the semantic fragments queried and retrieved by the user into a large language model for comparative analysis so as to identify logic or fact conflicts; the streaming output module is used for returning an intermediate reasoning process and a final conclusion of the large language model in real time; and the batch processing module is used for automatically processing the documents of the specified directory and batch query and generating a structured result file. According to the method, the system transparency and the user credibility are improved, and the robustness and efficiency in practical application are improved.
Owner:BEIJING TECH & BUSINESS UNIV