Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

327 results about "Document analysis" patented technology

Document analysis is used to determine requirements by analyzing the existing documents. This process also identifies the types of information that are important to the requirements. There are numerous types of documents that are analyzed in project management to draw out the important requirements.

Document analysis and query method and device based on knowledge graph, equipment and medium

The invention discloses a document analysis and query method based on a knowledge graph, and the method comprises the steps: carrying out the part-of-speech tagging of a received to-be-analyzed document, and obtaining a part-of-speech tagging result; extracting knowledge element information from the part-of-speech tagging result based on a preset power grid domain ontology knowledge base, and constructing an initial knowledge graph based on the knowledge element information; combining nodes in the initial knowledge graph to obtain a fused knowledge graph; constructing a mapping table according to the fused knowledge graph, and generating a candidate query template set based on the mapping table; receiving a natural language query statement input by a user, selecting a target query template from the candidate query template set based on the natural language query statement, and generating a target query statement; querying from the fused knowledge graph by using the target query statement, and outputting a query result; according to the method, the accuracy and comprehensiveness of information analysis can be effectively improved, the query intention of the user can be accurately understood, and the accurate query and analysis requirements of professionals on project documents are met.
Owner:STATE GRID ECONOMIC TECH RES INST CO LTD

Education data report content interaction method and system based on retrieval enhancement generation

The invention relates to the technical field of artificial intelligence, and discloses an education data report content interaction method and system generated based on retrieval enhancement, and the method comprises the steps: judging whether a natural language problem is an education field problem or not through a large language model, and if yes, carrying out semantic analysis to generate a structured query instruction; when the problem relates to cross-document association analysis, retrieving the structured semantic index database to generate a retrieval result set; if policy association analysis is involved, matching a policy knowledge graph by combining semantic similarity calculation and an entity linking technology, and then performing cross-modal fusion processing to obtain a retrieval result set; and inputting the retrieval result set into the retrieval enhancement generation model, and calling an education field language model to generate an analysis report. According to the method, the industrial pain points of inaccurate intention recognition, low cross-document analysis efficiency, incapability of dynamically combining with latest policies and the like in a traditional interaction mode can be solved, the efficiency and quality of data report interaction in the education field are remarkably improved, and the user interaction experience is optimized.
Owner:MYCOS DATA CORP CO LTD

Enterprise data link treatment and value management method and system

The invention discloses an enterprise data link management and value management method and system, and the method comprises the steps: carrying out the real-time capturing of multi-source heterogeneous data of all business systems in an enterprise through a distributed data collection engine, and recognizing the types of original data in different formats through a preset data source adapter; if structured data is detected, a relational database connector is adopted for extraction, and if the structured data is recognized as unstructured data, a document analysis module is started for content extraction, and an initial data set containing metadata tags is obtained; performing format conversion and field mapping on the initial data set according to a pre-established data standardization rule base, eliminating duplicate records and abnormal values through a data cleaning algorithm, performing automatic evaluation on a data quality grade by adopting a naive Bayes classifier, if a data quality score is lower than a preset threshold value, triggering a data recovery process, and if the data quality score is lower than the preset threshold value, performing data recovery. And standardized data meeting a unified standard is obtained. The normativity and value utilization efficiency of data management are effectively improved.
Owner:FRIENDSHIP INT ENG CONSULTING CO LTD

Document analysis method based on dynamic knowledge graph and RAG model

The invention discloses a document analysis method based on an RAG model and a dynamic knowledge graph, and relates to the technical field of artificial intelligence. The method is combined with an RAG model and a dynamic knowledge graph technology, and is realized by the following steps of: performing entity relationship joint extraction on an input document, generating a structural triple, and constructing a dynamically updatable knowledge graph; based on the knowledge graph, mapping entities and relationships into low-dimensional vectors by adopting a graph embedding model, and constructing a local vector knowledge base with a topological structure; receiving user questions in real time, encoding the user questions into query vectors, executing approximate nearest neighbor search based on the vector knowledge base, and matching related map fragments; and combining the retrieved graph fragments with the large language model, and generating a structured answer through path constraint of the injection knowledge graph. The method is used for solving the problem that in the prior art, a model cannot capture document deep semantics and dynamic relations insufficiently.
Owner:ECONOMIC TECH RES INST OF STATE GRID ANHUI ELECTRIC POWER

Document analysis method and device, equipment and storage medium

The invention provides a document analysis method and device, equipment and a storage medium, and relates to the technical field of computers, in particular to the technical field of deep learning, data processing and document analysis. According to the specific implementation scheme, at least one layout area divided by an article to which the document image belongs is determined according to the layout of the document image; identifying element contents of a plurality of layout elements in the document image; sorting the reading sequence of the layout elements in the same layout area; and obtaining structured document information according to the sorting result and the corresponding element content. According to the technical scheme, different article areas on the same page can be accurately identified and separated, the respective reading sequence is correctly reconstructed on the basis, and the content in the document image is converted into structured information.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Method for improving long text processing efficiency and accuracy

The invention discloses a method for improving long text processing efficiency and accuracy, and relates to the technical field of natural language processing and large language models.According to the method, text word segmentation embedding, sliding block preprocessing, YaRN position code injection, dynamic sparse attention calculation, multi-level attention fusion, graded KV cache management and output generation are sequentially executed; position drift is inhibited through logarithmic scaling, and key contexts are adaptively screened according to the attention activeness, so that the attention calculation complexity is close to linearity; in million-level Token reasoning, the video memory occupation of the method is reduced, the remote dependency recall rate is improved, and the method is suitable for scenes such as document analysis, code auditing and multi-mode streaming understanding.
Owner:BEI JING JING YUE KE JI YOU XIAN GONG SI

Document information extraction using visual question answering and document type specific adapters

Document type specific adapters of a document analysis system are used to provide for additional document types or document specialization when generating answers to user submitted questions targeting information included in a document image provided with the user submitted question. The document analysis system receives a visual question answering (VQA) prompt comprising a document image and a question defining information to be extracted from the document image, generates tokens based on the document image and the question, and adjusts encoding of the tokens using document type specific adapters of a transformer model to extract the information from the document image. A classifier of the document analysis system may determine whether the document image matches a document type supported by the document type specific adapters.
Owner:AMAZON TECH INC

Multi-element document analysis method and system

The invention provides a multivariate document analysis method and system, relates to the field of computer information processing, and solves the technical problems of low information extraction efficiency and low accuracy caused by incapability of uniformly analyzing and integrating formats of various types of documents. The method comprises the following steps: identifying the type of a to-be-processed document; the document types comprise a table type, a text type and a demonstration type; the text class comprises a Word format and a PDF format; calling a corresponding analysis function according to the type of the to-be-processed document, and analyzing the to-be-processed document to obtain an analysis result, the analysis result comprising the extracted structure information and content data of the to-be-processed document; and converting an analysis result into a standard JSON format and outputting the standard JSON format. The method and the device are used in a document analysis process.
Owner:HEFEI HUIQI INTELLIGENT TECHNOLOGY CO LTD

Large language model dynamic adaptation method and system based on Java

The invention discloses a Java-based large language model dynamic adaptation method and system, belongs to the technical field of artificial intelligence, and aims to solve the technical problems of overcoming the interface difference of multiple model interfaces, simplifying the development process and providing standardized and high-expansibility large model management. Comprising the following steps: providing a model registration and dynamic loading service, a protocol conversion and parameter standardization service and a load balancing and failover service; a streaming transmission protocol service, a function call dynamic injection service and a global error processing mechanism are provided; when a user provides a document analysis and partitioning service, a vectorization index construction service and an RAG enhanced generation service to ask questions, relevant document blocks in the Pinecone are retrieved through the RAG enhanced generation service to serve as contexts to be injected into cue words of the large language model, and answers generated by the large language model are returned.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Computer-aided system for multidimensional generative value assessment and applicant selection

ActiveDE202025107568U1InstrumentsData packData stream
A computer-implemented system for multidimensional generative value assessment and applicant selection, consisting of: a data collection unit configured to electronically receive applicant data consisting of structured academic records, work experience records, digital documentation, and unstructured narrative responses generated from generative self-assessment instruments and contextual interviews; a feature extraction unit coupled to the data acquisition unit, configured to apply computer-assisted text processing, semantic analysis, and token-level attribute identification to transform narrative responses and structured data into multidimensional feature vectors that represent generative indicators of innovation, mentoring, collaborative performance, resilience, social contribution, ethical consistency, and predicted institutional impact; a weighting calculation unit configured to assign weight values ​​to the extracted feature vectors based on a digital generative profile definition matrix that includes dimensions, sub-criteria, indicators, documentation requirements and importance coefficients, with the weighting being distributed across the generative dimensions defined in the digital matrix and configurable according to the institutional context; a quantitative rating unit configured to calculate a generative rating score by aggregating weighted feature vectors derived from self-assessment inputs, interview-based ratings, document analyses, and authenticity predictions, with the aggregation including normalization, nonlinearity correction, conflict handling, and artifact frequency balancing to obtain a consolidated score; a proof verification unit configured to electronically validate referenced digital evidence by performing content extraction, metadata verification, pattern matching, and cross-document correlation to determine authenticity, credibility, and contextual relevance with respect to the calculated feature vectors; a classification determination unit configured to assign a classification level to an applicant by comparing the generative assessment score with a set of system-defined calculation thresholds, including at least a lower threshold, a middle threshold and an upper threshold, the classification levels representing different generative maturity states and determining subsequent eligibility for selection; a decision generation unit configured to produce a digital output data set that includes classification level, feature aggregation summaries, evidence validation results, and recommended organizational actions, wherein the decision generation unit encodes the data set in a digitally signed, tamper-proof format and stores it on a non-volatile storage medium; and A system control unit acts as an operational interface to all other units and is configured to orchestrate data flow, scheduling, process state transitions, and event logging to ensure verifiable traceability, consistency, and auditability of the evaluation and selection processes.
Owner:BERNARDO OHIGGINS UNIVERSITY +3

Multi-modal PDF document analysis method and device, equipment and medium

The invention provides a multi-mode PDF (Portable Document Format) document analysis method, device and equipment and a medium, and the method comprises the following steps: loading a PDF document and carrying out preprocessing, including page splitting and content cleaning, to generate standardized document data; dynamically extracting contents in the standardized document data, wherein the contents comprise paragraphs, pictures and table elements; processing the dynamically extracted pictures, including shielding meaningless pictures based on a preset rule, and analyzing picture contents by using a multi-modal model to generate readable picture information; processing the dynamically extracted table, including optimizing and merging the table structure into a single element format to generate structured table data; combining paragraphs, readable picture information and structured table data, converting the paragraphs, the readable picture information and the structured table data into a complete structured document format, and inserting in key positions to enhance coherence context description; and outputting the complete structured document format as a final analysis result. According to the method, the analysis precision and the knowledge extraction efficiency of the complex PDF document can be remarkably improved.
Owner:深圳市和讯华谷信息技术有限公司

Document analysis evaluation method and system based on multi-modal semantic consistency

The invention discloses a document analysis evaluation method and system based on multi-modal semantic consistency. The method comprises the following steps: acquiring multi-modal contract document data, and preprocessing the multi-modal contract document data to obtain preprocessed data; extracting contract element features from the preprocessed data to generate a fusion feature vector; based on the fusion feature vector, evaluating the semantic consistency of the document elements by using a double-constraint loss function and a three-level judgment mechanism to obtain a cross-modal semantic consistency quantification result; based on the cross-modal semantic consistency quantification result, calculating a basis and semantic enhancement index, and integrating the basis and semantic enhancement index into a comprehensive evaluation index; and outputting the comprehensive evaluation index and the corresponding related suggestions, and guiding contract intelligent analysis and algorithm optimization. By implementing the method provided by the invention, intelligent and accurate document analysis can be realized, and the problems of dependence on geometric positioning, missing of cross-modal association and insufficient logic verification in the existing document analysis technology are solved.
Owner:TIANGU INFORMATION SCI TECH HANGZHOU

WEB end intelligent bidding document structured processing system based on hybrid AI analysis engine

The invention belongs to the technical field of intelligent document processing, and provides a WEB-end intelligent bidding document structured processing system based on a hybrid AI analysis engine, comprising: a multi-modal document analysis module extracts key information of a bidding document by using a fuzzy starvation game algorithm, and dynamically sorts core terms according to semantic association; the bidding document blind box analysis module safely disassembles the encrypted bidding document through a block chain technology to generate a structured review matrix; the cloud collaborative review module integrates the multi-dimensional data board, simulates bid evaluation by using a Monte Carlo algorithm, and generates a risk thermodynamic diagram; the risk early warning module identifies dispute points, provides compliance suggestions in combination with a knowledge base, and synchronizes the compliance suggestions to all review terminals; according to the method, automatic and structured processing of the WEB end bidding document is realized by fusing the mixed AI analysis engine, so that the efficiency and accuracy of bidding document review are improved, manual intervention is reduced, the review process is accelerated, and the risk of misjudgment and omission is reduced.
Owner:BEIJING ZHIHAN TECHNOLOGY CO LTD

Credit report generation method and device, equipment and storage medium

The embodiment of the invention relates to a credit review report generation method and device, equipment and a storage medium. The method comprises the steps that document analysis is conducted on an investigation document to obtain a document analysis result; extracting an abstract of the survey document according to a document analysis result to obtain a document abstract, vectorizing the survey document to obtain a document vector, and / or constructing questions and answers related to the survey document to obtain document questions and answers, and storing the document questions and answers in a knowledge base; obtaining a target report type of the to-be-generated credit review report, and determining a retrieval condition and a cue word corresponding to the target report type; performing retrieval in the knowledge base by utilizing the retrieval condition to obtain a retrieval result; inputting the retrieval result and the cue word into a pre-trained first large language model, and obtaining a plurality of report segments of a to-be-generated credit review report output by the first large language model; and generating the to-be-generated credit report according to the plurality of report fragments. According to the embodiment of the invention, the accuracy and efficiency of credit review report generation are improved.
Owner:北京中科闻歌科技股份有限公司 +1

Parsing and editing system and device for high-frame-rate rendering of PDF (Portable Document Format) document

The invention discloses an analyzing and editing system and device for high-frame-rate rendering of a PDF document, and the system comprises a file list uploading module which is used for uploading a pdf file and displaying a version historical record list and information; the document analysis logic module is used for reading and analyzing document information, determining a block level sequence of obtained data, deleting useless elements, and carrying out content sorting and splicing according to a layout; the pdf original document display module is used for executing pdf document title display operation, document display page cutting operation, document directory generation operation and document analysis display result operation in original document uploading, and the document directory generation operation comprises the steps of capturing multi-level digital numbers of documents, classifying number types according to matching results of capture groups in regular expressions, and storing the classified numbers in a database; and then numbering processing is carried out through a specified hierarchy mapping rule, and finally a file directory number is obtained. The problems that a traditional tool is inaccurate in analysis, low in editing efficiency and poor in professional adaptation can be effectively solved.
Owner:CHINA AUTOMOTIVE SOFTWARE (SHENZHEN) CO LTD

Document generation method and system integrating document analysis and cognitive reasoning

The invention provides an official document generation method and system fusing document analysis and cognitive reasoning, and the method comprises the steps: carrying out the multi-modal feature extraction of an input original document through employing an analysis engine combining a convolutional neural network and a graph attention network, and obtaining the multi-modal data in the original document; based on the multi-modal data, generating a classification decision result of the multi-modal data by using a hierarchical classifier enhanced by a knowledge graph; based on the multi-modal data and the classification decision result, generating a task decomposition result by combining a cognitive inference engine; and based on the classification decision result and the task decomposition result of the multi-modal data, generating an official document by combining compliance constraint with an official document format template, so that multi-modal accurate analysis of a complex document can be realized, and dynamic semantic understanding and classification are realized by utilizing a hierarchical classifier enhanced by a knowledge graph. The classification accuracy in the fuzzy semantic scene can be improved, so that the recognition error rate can be reduced, and the accuracy can be improved.
Owner:STATE GRID INFORMATION & TELECOMM BRANCH

Chip document automatic generation system

The invention relates to the technical field of chips, in particular to an automatic chip document generation system which comprises M chip document analysis modules {A1, A2,..., Am,..., AM}, a chip document template library and a chip document generation module, and Am is the mth file analysis module; the Am is used for obtaining and analyzing the mth type of chip files to generate corresponding document basic data Bm; the chip document template library is used for storing document template parameters {C1, C2,..., Cn,..., CN} corresponding to chip document templates, and Cn is the document template parameter corresponding to the nth chip document template; and the document generation module is used for acquiring the Bm from the Am, acquiring the target Cn corresponding to the Bm from the chip document template library, and generating a corresponding chip document based on the Bm and the target Cn corresponding to the Bm. According to the invention, the generation cost of the chip document is reduced, and the generation efficiency of the chip document is improved.
Owner:METAX INTEGRATED CIRCUITS (SHANGHAI) CO LTD

Knowledge question-answering method and system based on large language model and semantic abstract

The invention discloses a knowledge question-answering method and system based on a large language model and a semantic abstract. The system comprises a tree structure abstract generation subsystem which comprises a knowledge base configuration module, a file management module, a document analysis module and a knowledge block management module and is used for automatically generating tree structure hierarchical knowledge blocks for large manual documents or multi-chapter manual documents; the dual-channel retrieval engine subsystem comprises a question rewriting module, a semantic abstract retrieval module, a detail knowledge block retrieval module, a father-son retrieval module and a prompt word construction module, and all the modules work cooperatively to ensure that semantic abstract knowledge blocks and detail knowledge blocks related to the question of the user are retrieved. According to the method, by introducing the semantic abstract based on the tree structure and the two-channel retrieval, perception of global knowledge and accurate extraction of local fine-grained knowledge can be achieved at the same time, and then the ability of a knowledge base question-answering system in processing the general and inductive problems is improved.
Owner:NANJING SCIYON AUTOMATION GRP

Table structure identification method and system based on self-adaptive anchor frame, terminal and medium

The invention relates to the technical field of computer vision and document analysis, and discloses a table structure identification method and system based on a self-adaptive anchor frame, a terminal and a medium. The method comprises the following steps: acquiring multiple types of typical table image samples, labeling structured data and merging types of table cells, and constructing a travel channel data set and a column channel data set through multi-modal processing; constructing a table structure recognition model, wherein the model comprises a trunk feature extraction network, a double-branch detection head and a merging cell classification module; training a table structure recognition model based on the row channel data set and the column channel data set; and executing a structure identification task on an input table image by utilizing the trained model, and generating structured data containing physical coordinates and logic position relationships of the cells according to an identification result. According to the method, the problem of accumulative errors caused by asymmetry of row and column detection logic in a traditional method can be effectively solved, and meanwhile the problem that the row width and the column height are inconsistent due to complete dynamic adjustment is avoided.
Owner:HEFEI DAZHIHUI CAIHUI DATA TECH CO LTD

Vehicle after-sales maintenance document analysis and maintenance knowledge acquisition system and method based on AD-RAG

The invention relates to an AD-RAG-based vehicle after-sales maintenance document analysis and maintenance knowledge acquisition system and method, and the system comprises a document analysis module which carries out the layout analysis and document recognition of after-sales maintenance document information, and obtains structural data; the knowledge organization module is used for establishing an association relationship between the text data and the structural relationship in the structured data to obtain a maintenance knowledge graph; the knowledge retrieval enhancement module is used for analyzing and expanding the vehicle fault query request based on an AD-RAG model to obtain an expanded query request, and querying in a maintenance knowledge graph based on the expanded query request to obtain document fragment information; and the task decomposition and reasoning optimization module analyzes and decomposes the extended query request to obtain at least one fault sub-problem sequence, and performs fault reasoning through a preset fault reasoning model based on the fault sub-problem sequence and the document fragment information to obtain a fault maintenance suggestion result. The retrieval requirement can be better met, and the obtaining efficiency of the maintenance knowledge can be improved.
Owner:AIDONG SUPER AI

Bidding document data processing method and system

The invention discloses a bidding file data processing method and system, and the system employs a MoE mixed expert model architecture, and comprises a multi-type bidding file recognition module, a routing network, a professional sub-model library, a bidding file analysis engine module, and an automatic bidding file generation module. Wherein the bid invitation file is input into the multi-type bid invitation file identification module so as to identify the type of the bid invitation file; according to the type of the bid invitation file, calling a corresponding sub-model in a professional sub-model library; then the bid invitation file analysis engine module analyzes the bid invitation file; and the automatic bidding document generation module generates a bidding bidding document.
Owner:BEIJING XJ ELECTRIC +1

Software variation test system and method based on structure-semantic fusion

The invention discloses a software variation testing system and method based on structure-semantic fusion, and the system comprises a document analysis module, a structure diagram construction module, an agent module, a defect injection module and an influence domain verification module.The method comprises the steps that a function document or natural language description serves as an entry, and after a code structure diagram (such as a function call diagram) is constructed through static analysis, the influence domain is verified; an agent autonomously explores semantic information of a code under the guidance of a structure diagram, precise positioning from function description to a target code position is completed, and defect injection and influence domain verification are carried out.
Owner:SHANGHAI JIAOTONG UNIV +1

Method and system for generating engineering design specification based on large language model

The invention discloses a method and system for generating an engineering design specification based on a large language model, and the method comprises the steps: converting a plurality of formats of engineering documents into structural data through a non-structural document analysis module, and carrying out the multi-modal storage through a vector database and a relational database; a document recall module is combined with an expert predefined generation strategy and user input information to intelligently retrieve related materials from a vector database; and finally, calling a large language model through a document generation module to rewrite, expand or independently generate recalled contents, and outputting a highly customizable design specification which conforms to engineering specifications. According to the method, the efficiency, accuracy and specialty of engineering design document generation are remarkably improved, multi-section concurrent generation and manual fine adjustment are supported, and the method is suitable for multiple engineering stages such as project proposal, feasibility research report and construction drawing design.
Owner:TIANJIN MUNICIPAL ENGINEERING DESIGN & RESEARCH INSTITUTE CO LTD

Picture type PDF document analysis method based on convolutional neural network, multi-modal model and regular expression

The invention discloses a picture type PDF document analysis method based on a convolutional neural network, a multi-modal model and a regular expression, and belongs to the technical field of artificial intelligence and text processing. The method comprises the following steps: detecting types of layout elements of a preprocessed PDF document to obtain bounding box coordinates of each layout element; performing content identification on each layout element according to the type of the layout element; using a regular rule engine and a large language model to perform structured information extraction on the identification content, and extracting to obtain a plurality of predefined first business fields corresponding to each layout element; the character recognition result and the table recognition result are combined, and the combined result serves as content needing to be extracted; and taking a proofreading result as an analysis result of the scanned PDF document. According to the method, high-precision structured extraction of paragraphs, tables, formulas and other contents in the PDF document is realized, and the method has good universality, expandability and automation capability.
Owner:MILITARY SCI INFORMATION RES CENT ACAD OF MILITARY SCI OF THE CHINESE PEOPLES LIBERATION ARMY

Artificial Intelligence-Empowered Artist Management Platform with Integrated Career Optimization System

A system and method for managing a musical artist's career through artificial intelligence and machine learning technologies. The system employs a distributed computing architecture that integrates multiple data sources through secure APIs, including social media platforms, streaming services, and venue databases. Machine learning algorithms analyze collected data to generate personalized career recommendations through a specialized chatbot interface. The system implements continuous feedback loops for recommendation refinement and includes integrated modules for health monitoring, emergency response, financial management, tour optimization, legal document analysis, and merchandise management. Real-time processing capabilities enable immediate insights and adaptive career strategies through synchronized data collection and analysis across digital platforms.
Owner:ROBINSON LASHION

Hybrid PDF (Portable Document Format) document analysis and knowledge fragment construction method and device

The invention discloses a hybrid PDF (Portable Document Format) document analysis and knowledge fragment construction method and device, and belongs to the technical field of digital document intelligent processing. The method comprises the following steps: firstly, carrying out layout element identification and semantic classification on a PDF page through a deep learning model, and calculating the confidence of the PDF page; when the confidence coefficient is lower than or equal to a threshold value, a correction module based on vertical and horizontal projection analysis is automatically switched to carry out boundary detection and result error correction; the fused layout result is converted into structured JSON data, and the structured JSON data is further mapped into a Markdown format with a reserved title level; and finally, knowledge fragment segmentation is carried out based on a Markdown hierarchical relationship, so that each fragment carries a complete title context path. According to the method, the problems of semantic deficiency, insufficient stability and context splitting in the prior art are effectively solved, and the robustness and accuracy of complex format PDF and low-quality scanning copy analysis and the interpretability of downstream knowledge retrieval are remarkably improved.
Owner:FUJIAN STAR NET WISDOM TECH CO LTD +1

Large language model long text question answering method and system based on hybrid context compression technology

The invention belongs to the field of text questioning and answering, and relates to a large language model long text questioning and answering method and system based on a hybrid context compression technology. The method comprises the following steps: analyzing and preprocessing a document, and converting an unstructured original text into a normalized text paragraph set; classifying questions of the long text question and answer scene based on a large language model; according to the question type and the preprocessed text, adaptively selecting the most suitable context compression method to compress the long text to obtain a compressed context; and generating an answer to the question based on a large language model by using the compressed context. According to the method, the advantages of two context compression technologies are integrated, the problem of context window limitation of processing a long text by a large language model is effectively solved, high-quality question and answer performance is kept while computing resource consumption and processing delay of the compression technologies are reduced, and the limitation of a single compression method on different types of problems is relieved.
Owner:HEILONGJIANG CYBERSPACE RESEARCH CENTER (HEILONGJIANG INFORMATION SECURITY EVALUATION CENTER HEILONGJIANG ACADEMY OF NATIONAL DEFENSE SCIENCE & TECHNOLOGY) +2

Archive information extraction and intelligent management system based on self-supervised learning

The invention discloses an archive information extraction and intelligent management system based on self-supervised learning, and the system comprises the following modules: a collection preprocessing module which is used for collecting original data of an archive and generating a layout sample data set and an original text mapping table; the self-supervision pre-training module is used for obtaining a three-mode unified encoder model; the document analysis module is used for text detection, character recognition and table structure recovery and generating a document analysis data set; the information extraction module is used for aligning pointer positioning and optimal transmission, outputting metadata and a relational data set and registering evidence entries; the alignment warehousing module is used for generating a standardized record library and an evidence link table based on the structure enhanced double-tower vector recall; and the strategy operation and maintenance module is used for authority control, desensitization, archiving and incremental updating. According to the invention, through combination of three-mode self-supervision and an evidence chain, accurate extraction, aligned warehousing and traceable management of archive elements are realized.
Owner:BEIJING ZHONGKE JIANYOU TECHNOLOGY CO LTD

Whole-process intelligent generation system and method for identification document

The invention discloses an identification document full-process intelligent generation system and method. The system realizes end-to-end automatic processing from data acquisition to document generation through original report screenshot sampling, file filtering, authorization document analysis, data extraction and structuring, report content generation, equipment information analysis and judicial document template filling. The system adopts a multi-stage fault-tolerant and verification mechanism, integrates a large-model intelligent analysis and extraction technology, can automatically repair a non-standard data format, and generates standardized document content according to judicial specifications. According to the method, the judicial expertise document generation efficiency, accuracy and normalization are remarkably improved, and the evidence effectiveness of the output document is ensured.
Owner:GUIZHOU XIAOQI TECHNOLOGY CO LTD

Multi-modal intelligent question-answering system and interaction method based on privately deployed large language model

The invention discloses a multi-modal intelligent question and answer system and interaction method based on a privatized deployment large language model. The system comprises a model containerization deployment module, a multi-round question and answer and intention recognition module, a document analysis and multi-modal question and answer module, a knowledge base construction and retrieval enhancement generation module and a system integration and monitoring module. A DeepSeekR1 large language model, an inference service and a dependency library are packaged into a container mirror image through a containerization operation environment Docker, and unified scheduling and dynamic capacity expansion and contraction are achieved through Kubernetes; realizing context understanding and task distribution by adopting a dialogue state management and intention recognition technology; in combination with OCR, an embedded model BGE / SimCSE and a vector database FAISS / Milvus, semantic analysis and intelligent question and answer of multi-format documents such as PDF, Word and images are supported; and a hierarchical knowledge base is constructed and a retrieval enhancement generation technology is fused to improve the answer accuracy. According to the invention, safe, efficient and intelligent management and interaction of enterprise knowledge are realized.
Owner:CHINA THREE GORGES CORPORATION