Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

667 results about "Document processing" patented technology

Document processing involves the conversion of typed and handwritten text on paper-based & electronic documents (e.g., scanned image of a document) into electronic information using one of, or a combination of, intelligent character recognition (ICR), optical character recognition (OCR) and experienced data entry clerks.

Document element extraction method based on AI large model technology

The invention discloses a document element extraction method based on an AI large model technology. The method comprises the following steps that various receipts input by a user are received; carrying out analysis and layout analysis on the document through a visual language model or a text analysis technology to generate a unified intermediate representation; a dynamic grading module is adopted to evaluate document complexity from three dimensions of format, structure and variation, and the document complexity is divided into a simple regular type, a structure complex type or a height variation type; adaptively selecting a processing flow according to a rating result, wherein the processing flow comprises high-speed template matching, multi-modal cognitive fusion or intelligent agent driving processing; the processing result is converted into an EasyEX format which is easy to process; the method comprises the following steps of: converting a natural language demand into an execution rule by extracting an Agent and utilizing a Prompt dynamic compiling technology, and positioning and extracting a target field by LLM (Logical Language Model); and finally, outputting structured data after type verification, knowledge graph verification and compliance review. According to the invention, the universality, the accuracy and the automation degree of receipt processing are improved, and the labor cost is greatly reduced.
Owner:SHENZHEN YSSTECH INFORMATION TECH CO LTD

Document format processing method and system based on large language model, terminal and medium

The invention relates to the field of document processing, and particularly provides a document format processing method, system, terminal and medium based on a large language model.The method comprises the steps that a generation instruction containing a target document type and a source material are received, and original text content is generated through the large language model integrating domain knowledge; obtaining a matched structured format template, and analyzing the style rule into a format instruction set; identifying logic elements and hierarchical relationships thereof in the original text based on a natural language processing technology; performing association mapping on the format instruction and the logic element, and performing automatic style rendering by calling a document object model interface to generate an intermediate document with a standard format; and finally, outputting after quality verification. According to the method, an automatic process of content generation and intelligent formatting is constructed, and the document writing efficiency and normalization are improved.
Owner:浪潮智慧科技有限公司 +2

Document knowledge retrieval method and system, electronic equipment and storage medium

The invention relates to the technical field of document processing, and discloses a document knowledge retrieval method and system, electronic equipment and a storage medium. Based on a language model, synchronously generating abstract contents and a semantic association query set for each text fragment, and aggregating the abstract contents of all text fragments corresponding to the same document to form a fragment abstract set; constructing a multi-path retrieval pool based on the text fragment, the abstract content and the semantic association query set, and respectively constructing a plurality of independent retrieval data sources; the method comprises the following steps: constructing a dual-granularity abstract index system, receiving a query request of a user, executing query routing processing based on dual-granularity abstract index and a hierarchical retrieval process based on a metadata structure, returning retrieval result data, performing fusion calculation and duplicate removal processing on the retrieval result data recalled by multiple paths, and obtaining a retrieval result. And meanwhile, the document abstract short sentence group is injected into a search suggestion function of a search engine. According to the method, the retrieval accuracy can be improved, and the user experience is improved.
Owner:TONGFANG KNOWLEDGE DIGITAL PUBLISHING TECH CO LTD

File processing method and device oriented to big language model retrieval enhancement generation

The embodiment of the invention provides a large language model retrieval enhancement generation-oriented file processing method and device, and the method comprises the steps: carrying out the information extraction and semantic information enhancement of a file according to different file types, and constructing a meta-information structure of the file in combination with an enterprise business scene; dividing the document content into a plurality of structured blocks based on the meta-information structure, and labeling the title and context information of each block; when the file is uploaded, according to the parent directory material quantity and content change of the directory where the file is located, performing dirty marking on the directory; when a data request is received, whether a dirty mark exists in a related directory or not is judged, if yes, the directory is requested to be locked, a summary is extracted from files extracted from bottom to top through a large language model, and the dirty mark is cleared after the summary is cached.
Owner:特赞(上海)信息科技有限公司

Power document keyword extraction method based on Prompt and knowledge graph

The invention provides an electric power document keyword extraction method based on Prompt and a knowledge graph, relates to the technical field of electric power document processing, and constructs a lightweight multi-level index knowledge graph in the electric power field by combining entity type and relation type division based on an electric power industry standard document and an electric power field corpus. The method comprises the following steps: performing vector modeling on a power document, constructing a multi-level index from an entity to a vector, realizing standardized semantic modeling and efficient hybrid retrieval of a power document field background, and obtaining a topic vector and a core paragraph of the power document in combination with power key information; according to the method, entity types are indexed in a knowledge graph by using subject vectors, similar entities are obtained to form knowledge sub-graphs, so that multilayer Prompt is obtained to guide a large language model to extract keywords, then knowledge graph similarity constraints are introduced to decode the output of the large language model, the keyword recognition capability in the power field is improved, and the keyword recognition efficiency is improved. And the accuracy of keyword type identification and the normalization of term naming are both considered.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO LTD SHAOXING POWER SUPPLY CO

Handwriting dynamic encryption and electronic signature method and system based on national secret algorithm

The invention relates to the technical field of information security, and discloses a handwriting dynamic encryption and electronic signature method based on a national secret algorithm, and the method comprises the steps: obtaining the handwriting data of a user, and carrying out the preprocessing of the data to remove noise and errors; carrying out feature extraction on the preprocessed handwriting data, extracting corresponding track features, speed features and pressure features, and generating a feature hash value of a unique identifier through a cryptographic SM3 hash algorithm; splicing and combining the feature hash value and a user private key, deriving a dynamic key through a national secret SM2 algorithm, and generating an electronic signature by using the dynamic key; an electronic signature is combined with a timestamp to form a structured signature packet, and the signature packet is hierarchically encrypted and stored in a database through a national cryptographic SM4 algorithm. According to the method and the system, the handwritten signature picture, the digital signature technology, the electronic signature technology and the biological recognition and layout file processing technology are fused, so that the safety and the reliability of electronic document signing are improved.
Owner:HENAN INFORMATIZATION GRP CO LTD

Personalized document field prediction based on learning from user feedback

Particular embodiments relate to personalized document field prediction based on user behavior and feature generation. Specifically, various embodiments have the technical effect of improved accuracy with respect to field / entity value prediction (e.g., predicting that the amount due is X via a Gradient Boosting Model) relative to document processing technologies by learning through user behavior data or feedback (e.g., through continuous reinforcement learning from human feedback (RLHF)). This is at least partially because of the technical solution of accessing or generating unique features from one or more documents previously used by a user.
Owner:BILL OPERATIONS LLC

Knowledge base construction method and system for multi-source heterogeneous files

The invention relates to a knowledge base construction method and system for multi-source heterogeneous files, and the method comprises the steps: receiving Word, PDF, Excel and other multi-source heterogeneous files, and converting the multi-source heterogeneous files into standardized representation through a file analysis function; an event boundary recognition technology fusing rule matching, entity recognition, semantic similarity and large language model verification is adopted, and a cross-paragraph complete business event is accurately extracted; generating structured information through an information extraction function; the information is converted into vectors, full-text indexes and a graph database to be stored by means of a multi-dimensional storage conversion function; and verifying and correcting an extraction result through a quality auditing function, and finally outputting a knowledge base containing an event set, an entity relationship and vector representation. According to the method, the problem of knowledge fragmentation in traditional document processing is effectively solved, and the consistency, the searchability and the reasonability of knowledge are remarkably improved.
Owner:ECCOM NETWORK SYST CO LTD

Intelligent document checking method based on large language model

The invention belongs to the technical field of artificial intelligence and natural language processing, and particularly relates to an intelligent document checking method based on a large language model. And cooperatively executing wrongly written character and term checking, standard version checking, structural integrity checking, compliance checking, consistency checking, calculation accuracy checking and common error checking in sequence. In each verification step, deep semantic extraction of a large language model and a retrieval enhancement generation technology of an external domain knowledge base are combined, and strict preprocessing regularization and consistency gating verification are supplemented. According to the method, the illusion of autoregression generation of the language model is effectively inhibited, the logic coherence of the overlength document in context and parameter characteristics is guaranteed, intelligent automatic verification with extremely high accuracy, low false alarm rate and completely traceable auditing process is realized, and the quality and efficiency of complex document processing in various industries are greatly improved.
Owner:COAL IND JINAN DESIGN & RES

Four-layer progressive mind mapping generation system

The invention discloses a four-layer progressive mind map generation system which comprises an input module which can receive input of various document formats, supports simultaneous receiving of multiple documents and adopts a multi-source heterogeneous data adapter to unify the formats; the hierarchical preprocessing module is connected with the input module, receives the converted DOCX document data, and performs hierarchical feature enhancement and hierarchical rule configuration on the data; the bimodal processing module fuses a hierarchical driving structured extraction mode and a large language model generation mode, the hierarchical driving structured extraction mode constructs a hierarchical incidence matrix and the like, and the large language model generation mode provides task instructions and the like; and the structured output module is used for providing multi-format conversion, mind map downloading and interaction functional components. The system has the advantages that all the modules work cooperatively, multi-format document processing is achieved, multiple modes are fused to generate the mind map, rich output functions are provided, and a user can conveniently and efficiently generate and use the mind map.
Owner:ZHEJIANG UNIV BINJIANG RES INST +1

Structural model generation method and device based on architectural drawing, equipment and medium

The invention relates to the technical field of structural analysis, and provides a structural model generation method and device based on an architectural drawing, equipment and a medium, which can reconstruct a line unit of a basic architectural drawing according to a line type assignment rule so as to realize visual distinguishing of component information. The drawing exchange format file processing library is called to accurately extract drawing information of the target architectural drawing and write the drawing information into the dictionary, so that hierarchical management and control of information can be realized, and component information confusion is avoided; each structure model standard layer is created according to the building drawing information file and the modeling information file, nodes are added to each structure model standard layer based on a node duplicate removal mechanism, and efficient node creation can be realized on the premise of avoiding repetition; structural component modeling is carried out on the basic framework according to the modeling information file and the pre-processing parameter file, pre-processing parameter configuration and model data engineering output processing are carried out on an initial structural model to obtain a target three-dimensional structural model, and therefore accurate modeling of the three-dimensional structural model is automatically achieved.
Owner:CHINA CONSTRUCTION SCIENCE & IND GROUP GREEN TECHNOLOGY CO LTD +1

Audit information extraction method and device based on multi-modal fusion and electronic equipment

The invention provides an audit information extraction method and device based on multi-modal fusion and electronic equipment, and can be applied to the technical field of document processing. The method comprises the steps of obtaining a to-be-queried field in response to a received extraction request for audit information; according to a modal type of data in the audit information, calling an encoder corresponding to the modal type to encode the audit information, and generating feature vectors of at least two modals; taking a vector obtained by encoding a to-be-queried field as an initial query vector, calling a stacked cross attention module, respectively performing interactive fusion with the feature vectors of at least two modalities to realize feature interaction and alignment between the modalities and in the modalities, and outputting a target modal feature vector matched with the to-be-queried field; and decoding the target modal feature vector, and extracting a target field matched with the to-be-queried field from the audit information.
Owner:STATE GRID TIANJIN ELECTRIC POWER COMPANY +1

XML semantic editing system and method based on aviation knowledge graph

The invention belongs to the field of Web application technology, XML document processing technology and online editor technology, and discloses an XML semantic editing system based on an aviation knowledge graph. The system comprises a front-end application module, an XML processing engine module, a DTD rule verification module, a visual editing module, a plug-in extension module, a collaborative editing module and an aviation knowledge graph integration module. According to the method, professional and efficient editing of XML documents in the aviation field is achieved through the full-process design of front-end interaction, XML processing, DTD verification, visual editing, semantic assistance and collaborative management, and the Web technology, the XML processing technology, the DTD verification technology, the knowledge graph technology and the collaborative editing technology are integrated, so that the XML documents in the aviation field are edited in a professional and efficient mode. The problems that an existing XML editing tool is difficult to deploy, poor in compliance, low in efficiency and weak in cooperation in the aviation maintenance field are solved, and a specialized and efficient XML semantic editing scheme is provided for the aviation maintenance industry.
Owner:WUHAN YIFAN IOT TECHNOLOGY CO LTD

Document processing method and device

PendingCN121706731ASemantic analysisVisual presentationInteractive editing
The invention provides a document processing method and device. The method comprises the steps of obtaining a to-be-processed demand document; calling an intelligent agent to perform format conversion on the to-be-processed demand document based on prompt information associated with the intelligent agent to obtain a structured tagged document in a markup language format returned by the intelligent agent; the prompt information is obtained by filling a prompt template of an intelligent agent based on a sample demand template obtained in response to the selection operation, and the prompt information is used for guiding the intelligent agent to perform structured recombination and formatted output on the content of the demand document to be processed according to structured fields and levels defined in the sample demand template; on the basis of the structured marking document, the configured online structured template is filled, and the online structured demand document supporting visual presentation and / or interactive editing is generated, so that the document processing efficiency and accuracy are improved, errors and omissions possibly occurring in manual processing are avoided, and the user experience and the working efficiency are improved.
Owner:BEIJING PACTERA JINXIN TECH LTD

Server-free document processing and real-time report method and system

The invention provides a server-free document processing and real-time reporting method and system, relates to the technical field of data processing, and starts from responding to the operation that a user uploads a source document to an object storage server, and routing an event to a pre-subscribed document processing server-free function. And after the function is triggered, obtaining a source document, performing analysis and data extraction on the source document to generate structured data, and storing the structured data in a database. The event bus then routes the event to a pre-subscribed report-generating server-less function. And automatically generating a notice containing an access link and pushing the notice to a specified user or terminal. According to the method, automation, real-time performance and flexibility of the whole process from document uploading to report generation are achieved, and through tight combination of event driving and a server-free architecture, the processing efficiency, the resource utilization rate and the system response speed are improved.
Owner:ANHUI FUXING SOFTWARE CO LTD

PDF analysis method and system based on image-text semantic alignment

The invention provides a PDF analysis method and system based on image-text semantic alignment, and belongs to the technical field of document processing. The method comprises the following steps: reading PDF or picture file byte data, and converting pictures into PDF byte data in a unified format; creating an image and a Markdown directory according to the output directory, the PDF name and an analysis method; extracting byte data of a specified page of the PDF according to starting and ending page numbers; selecting a back end to analyze PDF byte data to obtain a reasoning result, an image list and the like; generating an intermediate JSON containing PDF detailed information according to an analysis result; analyzing the positions of the text and the picture, and associating information to realize semantic alignment of the picture and the text; and generating multiple types of machine readable files according to the intermediate JSON and storing the files in a specified directory. According to the method, the PDF multi-mode content can be accurately extracted and subjected to semantic alignment, the PDF multi-mode content is efficiently converted into a machine readable format, the analysis accuracy and usability are improved, and the method is suitable for PDF analysis of multiple image-text tables such as scientific and technical literatures.
Owner:XI AN JIAOTONG UNIV +1

Virtual assembly method and system based on improved point cloud registration and precision feature extraction

The invention discloses a virtual assembly method and system based on improved point cloud registration and precision feature extraction, relates to the technical field of virtual assembly, and aims to solve the technical problem that a conventional ICP algorithm is insufficient in robustness and precision in a point cloud registration link in a current virtual assembly technology based on point cloud data. Comprising a file processing module, a preprocessing module, a point cloud registration module and a many-to-many component matching scheme solving module. According to the method, the RICP algorithm is designed, and a general adaptive robust function is introduced, so that the problem that the traditional ICP algorithm is sensitive to noise and outliers is effectively solved. The robust function can dynamically give a weight according to a point pair distance, stable and high-precision registration can still be realized even in a scene of large point cloud initial pose difference, low overlapping degree or unobvious surface features, and the problem of insufficient robustness and precision of a traditional ICP algorithm in a point cloud registration link of a current virtual assembly technology based on point cloud data is solved.
Owner:AEROSUN CORP

FastGPT-based intelligent question-answering system for construction project quality inspection and detection standards

The invention relates to the technical field of data retrieval, in particular to a FastGPT-based intelligent question-answering system for construction project quality inspection and detection standards. According to the technical scheme, the method comprises the steps of external knowledge acquisition, document processing, vectorization, vector database, retrieval and context generation and intelligent question and answer, data can be collected by a system, a semantic segmentation technology, a QA pair generation technology and a high-dimensional vectorization technology are utilized, in combination with M3E and OpenAIEmbedding models, Top-K similarity retrieval and context dynamic recombination are achieved, and structured and hierarchical professional answers are generated. According to the method, the intelligent management and real-time application capability of the construction project quality inspection and detection standard can be remarkably improved, the manual retrieval and interpretation cost is reduced, and the digital and intelligent requirements of an engineering project are met.
Owner:ZHANJIANG JIANKE ENG QUALITY TESTING CENT CO LTD

Lossless compression storage method and system for heterogeneous data based on credential environment

The invention discloses a lossless compression storage method and system based on heterogeneous data in a credential environment, and the method comprises the steps: firstly, extracting a mapping relation between a file type and a structural feature through scanning a target file, carrying out the preliminary lossless compression after classification and grouping, and further employing a sliding window matching and sequence matching algorithm to optimize compressed data for a redundancy mode, and then generating a metadata structure containing a check code and a compression parameter, packaging the metadata structure into a data packet supporting cross-platform transmission, finally transmitting the data packet to a target storage device through a secure write-in protocol, and verifying data consistency and recovering path feasibility by using the check code. According to the method, the data compression efficiency is remarkably improved, the compatibility of cross-platform transmission and the integrity of stored data are ensured, and efficient and reliable technical support is provided for complex electronic file processing.
Owner:HUNAN YUNDANG INFORMATION TECH CO LTD

Visual compression and retrieval method and device of document, equipment and storage medium

The invention discloses a document visual compression and retrieval method and device, equipment and a storage medium, and relates to the technical field of computers, the method comprises the following steps: obtaining a to-be-processed document page image, segmenting the image into a plurality of image blocks, and determining the structure category and the structure importance score of each image block; obtaining a plurality of structure regions based on structure category and spatial position aggregation, and distributing a preset number of compression tokens for each region in combination with structure category weights and importance scores; generating structure anchor point tokens corresponding to the regions by the compressed tokens to form a set; receiving a query request, converting the query request into a query vector, and performing retrieval in the anchor point token set to obtain a target structure region; and performing local decoding reconstruction based on the compressed token of the target region, and outputting a region image or a structure mask. According to the method, through structure-guided self-adaptive compression and fine-grained retrieval, the long document processing efficiency is greatly improved, and the compression effect and the retrieval accuracy are both considered.
Owner:BEIJING DIGITAL CHINA CLOUD COMPUTING CO LTD

Medical document processing method and system based on double-pipeline architecture

The invention discloses a medical document processing method and system based on a double-pipeline architecture, and relates to the technical field of document processing. According to the medical document processing method based on the double-assembly-line architecture, through a closed-loop process of document classification, preprocessing, double-assembly-line directional parallel processing and hierarchical vectorization storage, precise adaptation and efficient processing of multi-format and multi-type medical documents are achieved, the information loss rate and the key information truncation rate are greatly reduced, and the medical document processing efficiency is improved. According to the method, the document processing efficiency and the data standardization degree are improved, the warehousing success rate and the data traceability of the vector library are ensured, high-quality and structured data source support is provided for subsequent medical intelligent retrieval, clinical question and answer and retrieval enhancement generation system application, the knowledge base construction and maintenance cost is remarkably reduced, and the method is suitable for popularization and application. The problems that in existing medical document processing, medical semantics are not taken into consideration, so that key clinical information is easy to cut off, and a single processing flow cannot adapt to a structured guide and an unstructured case are solved.
Owner:SONGJIANG HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIVERSITY SCHOOL OF MEDICINE +2

File processing method, apparatus, and electronic device

Embodiments of the present disclosure provide a file processing method and apparatus, and an electronic device. The method includes: acquiring additional data written to a file end of a target file, where data in the target file is stored based on a strip, and a data end of the additional data is located in a middle of a target strip; obtaining a corresponding redundant storage policy according to an occupancy of the additional data in the target strip, and generating a restoration record of the additional data based on the redundant storage policy; and writing the restoration record into a disk.
Owner:BEIJING VOLCANO ENGINE TECH CO LTD

Localized document automatic processing method and device, equipment and storage medium

The invention relates to the technical field of document automatic processing, in particular to a localized document automatic processing method and device, equipment and a storage medium. The localized document automatic processing method comprises the following steps: acquiring a current document and a historical document associated with the current document; extracting text features of the current document and the historical document; constructing a multi-document context sensing model based on the text features of the current document and the historical document, and obtaining a document context vector containing complete context information; a document summary is generated based on the document context vector, and / or a structured task is extracted based on the document context vector. According to the method, the current working context of the user can be fully combined, the accuracy of automatically processing the document is improved, and therefore the document abstract can be automatically and accurately formed and the structured task can be accurately extracted.
Owner:HUAYIXIN (WUXI) TECHNOLOGY CO LTD

Document generation method and system

The invention discloses a document generation method and system, and belongs to the field of document processing software, the method comprises the following steps: in response to a template creation operation, creating a template structure file, a style configuration file and a component library definition file, the template structure file is used for defining a logic structure and a to-be-filled area of a target document, and the to-be-filled area is used for defining a to-be-filled area of the target document; the style configuration file is used for defining page layout of the target document, and the component library definition file is used for defining available document components of the target document; generating a first document template file based on the template structure file, the style configuration file and the component library definition file; obtaining a first content data file generated based on the first document template file, wherein the first content data file comprises filling content used for filling the to-be-filled area; and synthesizing the target document based on the template structure file, the style configuration file and the first content data file. According to the method, standardization of the generated target document can be ensured, and the document maintenance difficulty is reduced.
Owner:SUNGROW POWER SUPPLY CO LTD

Contract identification and intelligent management system and method based on MCP

The invention provides an MCP-based contract identification and intelligent management system and method. The method comprises the following steps: acquiring contract documents in various formats of a target contract, performing text processing and image processing on the contract documents, and performing multi-modal fusion to obtain intermediate data; constructing a knowledge graph, and performing clause logic conflict detection and compliance risk assessment on the target contract to obtain a contract assessment result; constructing a reinforcement learning decision model according to the knowledge graph and the contract evaluation result; based on the reinforcement learning decision model, MCP interface standard data is determined, and an adapter model is constructed; processing the contract document by using an adapter model to obtain a contract MCP data stream, reconstructing a knowledge graph and a reinforcement learning decision model for the contract MCP data stream based on MCP interface specification data, and integrating a model processing result and a model operation state; and executing the contract processing flow to obtain a contract management scheme. According to the invention, efficient and accurate contract document processing and management are realized.
Owner:BEIJING YULORE INNOVATION TECH

File processing methods, electronic device, storage medium and computer program product

The present disclosure relates to the fields of large model technology and text processing. Disclosed are file processing methods, an electronic device, a storage medium and a computer program product. A method comprises: in response to an input instruction acting on an operation interface, displaying on the operation interface a file to be processed, said file containing text to be processed of at least one modality; and, in response to a processing instruction acting on the operation interface, displaying a processing result on the operation interface, the processing result being used for representing that said file has been successfully named and stored on the basis of a target processing rule, the target processing rule being a processing rule corresponding to a target type of said file, the target type being determined by comparing said text with preset text contained in a plurality of preset files, and the types of different preset files being different. The present disclosure solves the technical problem in the prior art of relatively low file processing accuracy.
Owner:ALIBABA (CHINA) CO LTD

Power grid autonomous operation and maintenance question-answering system based on large model and RAG open source framework

The invention provides a power grid autonomous operation and maintenance question-answering system based on a large model and an RAG open source framework, and belongs to the technical field of natural language processing. Comprising a knowledge document processing module which is used for segmenting a knowledge document in a specific field into a plurality of text blocks with independent semantics and performing semantic enhancement processing on the plurality of text blocks in combination with knowledge characteristics of the specific field; the system further comprises a vector generation and storage module, and the vector generation and storage module generates high-dimensional vector representation for each text block through an embedded model. By setting a dynamic retrieval and generation mechanism, the operation and maintenance problem of the intelligent equipment of the power grid can be quickly and accurately solved, and the system not only improves the operation and maintenance efficiency and accuracy, but also reduces the development cost, and enhances the expansibility and flexibility of the system.
Owner:GUANGZHOU BUREAU CSG EHV POWER TRANSMISSION

Multi-format document analysis and verification method, system and program product

The invention provides a multi-format document analysis and verification method and system and a program product, and relates to the technical field of document processing and intelligent verification. The method adopts a layered modular architecture, and comprises the following steps: a document analysis layer selects an analysis strategy according to document types and complexity, and uniformly converts various formats of documents such as PDF, Word, Excel and the like into a Markdown intermediate format; the index extraction layer uses an abstract base class and a dynamic Prompt optimization mechanism to extract structured data from an intermediate document and integrate the structured data into a unified JSON format to realize complete decoupling of data extraction and rule verification; and the rule verification layer dynamically loads a business rule by scanning an independent file, constructs a rule-dependent directed acyclic graph, and determines an execution sequence by utilizing topological sorting. According to the method, the problems of high data and rule coupling degree, high maintenance cost and poor expansibility in an existing system are effectively solved, hot plugging of the rules, dependency on automatic management and multi-document concurrent processing are realized, and the reusability and execution efficiency of the system are remarkably improved.
Owner:HUA DATA TECH (SHANGHAI) CO LTD

Teaching courseware generation method, device and system and computer storage medium

The invention provides a teaching courseware generation method and device, and relates to the technical fields of document processing, artificial intelligence, natural language processing and the like. According to the specific scheme, the method comprises the following steps: generating a teaching design general outline comprising a unit identifier of a teaching material unit and module information of at least two teaching modules based on teaching elements of the teaching material unit input by a user; based on the teaching design general outline, teaching resources are generated for each teaching module; based on the teaching design general outline, detecting whether the learning targets of the teaching resources among different teaching modules are consistent and / or whether the teaching resources of each teaching module deviate from respective teaching targets; in response to detection that the learning targets of the teaching resources are inconsistent and / or deviate from the teaching targets, regenerating the teaching resources of the corresponding teaching modules, and continuing detection until the learning targets of the teaching resources of different teaching modules are consistent and the teaching resources of all the teaching modules do not deviate from the respective teaching targets; and generating a target teaching courseware of the teaching material unit based on the teaching resources.
Owner:HUA CHUAN INTERNATIONAL HOLDINGS GROUP CO LTD

PDF (Portable Document Format) document processing method and device, equipment and medium

The invention relates to the technical field of computers, and discloses a PDF (Portable Document Format) document processing method, device and equipment and a medium, the method comprises the following steps: carrying out global layout analysis on a PDF document to detect element information of all structural elements in the PDF document, the element information comprising bounding box coordinates, element categories and a reading sequence; based on the bounding box coordinates, cutting out a corresponding local area from the PDF document, and performing content identification on different types of structural elements corresponding to the local area; according to the element category and the reading sequence, carrying out recombination and logic division on a content recognition result obtained by carrying out content recognition to obtain a plurality of logic parts; and for each logic part, constructing a cue word, calling a large language model to carry out thinking chain reasoning so as to extract structured information corresponding to the logic part, and merging the structured information to generate a JSON file. According to the method and the device, the PDF document processing accuracy is improved.
Owner:传申弘安智能(深圳)有限公司 +1