Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

210 results about "Document structuring" patented technology

Document Structuring is a subtask of Natural language generation, which involves deciding the order and grouping (for example into paragraphs) of sentences in a generated text. It is closely related to the Content determination NLG task.

Intelligent document generation system and method for bidding document

The invention discloses an intelligent document generation system and method for bidding documents, and the system comprises a preprocessing module which is used for obtaining bidding document files in various formats, and converting the obtained bidding document files in various formats into structured bidding document files; the analysis module is used for analyzing the converted bidding document file based on the large model cue word to obtain a structured bidding document file analysis result; the directory generation module is used for generating a directory of the bidding document file according to the analysis result; the retrieval module is used for generating a corresponding knowledge content set according to the bidding document file directory; and the content generation module is used for generating the text content of the bidding document file based on the knowledge content set and the structure information of the bidding document file directory. And the content assembly module is used for performing multi-dimensional feature extraction and classification on the generated bidding document file text content, and completing content assembly according to a preset document structure tree to obtain a standard document meeting bidding document requirements.
Owner:BIAOYIZHONG DIGITAL TECHNOLOGY (ZHEJIANG) CO LTD

Context aware document augmentation and synthesis

A method includes obtaining a document structure from a data repository. The document structure includes multiple structured sections. A table is detected in a first structured section. A table representation of the table is processed by a general large language model (LLM) to generate a natural language description of the table. An image is detected in the first structured section. The image is processed by an image-processing LLM to generate a natural language description of the image. A form is detected in the first structured section. The form is processed by the general LLM to generate a natural language description of the form. The natural language descriptions of the table, image and form are inserted into the first structured section to obtain a modified first structured section. A modified document structure including the modified first structured section is outputted.
Owner:INTUIT INC

PDF (Portable Document Format) document structured extraction system based on multi-modal language model

The invention discloses a PDF (Portable Document Format) document structured extraction system based on a multi-modal language model, belongs to the technical field of document processing and optical character recognition, and aims at solving the technical problem of how to improve the existing OCR (Optical Character Recognition) technology to improve the analysis capability of a complex document structure and improve the recognition precision of handwritten forms and other non-standard fonts. According to the technical scheme, the system adopts a layered decoupling architecture and comprises an input layer, a preprocessing layer, a reasoning layer, an output layer and a monitoring and fault-tolerant module; wherein the output layer is used for multi-source data access and path management to realize a local file system or S3 cloud storage; the preprocessing layer is used for invalid document filtering and visual feature extraction; the reasoning layer is used for multi-modal model interaction and content processing; the output layer is used for outputting a content aggregation result; and the monitoring and fault-tolerant module is used for realizing real-time state monitoring, resource consumption analysis and exception handling.
Owner:JIANGSU HAIRUO INFORMATION TECHNOLOGY CO LTD

Permission configuration method and device

The invention provides a permission configuration method and device. The method comprises the following steps: acquiring a source code and relative path information of a target document edited in a current editing environment; wherein the relative path information is used for representing position information of the target document relative to a root directory of a working area of the current editing environment; analyzing the source code to obtain a document structure object of the target document; performing node traversal on the document structure object to identify a target node corresponding to the interaction component from each node of the document structure object; according to the permission attribute corresponding to the target node and the relative path information, generating an attribute value of the permission attribute in the source code; wherein the attribute value is used for performing permission configuration of the interaction component, so that the accuracy and consistency of permission configuration are improved, the workload of manual permission configuration of developers is reduced, and the development efficiency is improved.
Owner:CLP JINXIN SOFTWARE (SHANGHAI CO LTD

Document content generation method based on title structured requirement and related equipment

The invention discloses a document content generation method based on a title structuring demand and related equipment, and relates to the technical field of artificial intelligence, the method comprises the following steps: performing structured analysis on a target electronic document to extract at least one title element in the target electronic document; receiving personalized requirements input by a user for the title elements, and performing data association on the personalized requirements and the corresponding title elements; the personalized requirements are used for guiding content generation; the title elements and the personalized requirements corresponding to the data association are combined into a generation request, and the generation request is sent to an artificial intelligence model; using an artificial intelligence model to obtain generation content according to the generation request; and inserting the generated content into a position related to the corresponding title element in the target electronic document. According to the method, the document structure can be deeply understood, a user is allowed to set independent generation requirements for different title elements, a document processing scheme for batch generation and selective optimization can be realized, and the flexibility of document processing is greatly improved.
Owner:GUANGDONG ELECTRIC POWER PLANNING SURVEY & DESIGN INST

Context-aware information retrieval

Certain aspects of the disclosure provide for information retrieval that exploits context derived from document structure. Source documents can be preprocessed to identify fields and determine context attributes related to each field based on the structural layout of a source document. Resource documents can also be preprocessed to segment a resource document into passages and determine context related to the passages based on structural layout. Queries pertaining to a field can be enhanced by adding context metadata associated with the field. A query embedding can be generated and compared with previously generated passage embeddings to locate candidate matches based on similarity. A machine learning model can be provided with the top-ranked passages and tasked with re-ranking the passages based on relevancy to the original query. The highest re-ranked passage or set of passages can be output in response to the query.
Owner:INTUIT INC

Self-adaptive text extraction method and system based on artificial intelligence

The invention discloses a self-adaptive text extraction method and system based on artificial intelligence, and the method comprises the steps: carrying out the analysis of the document structure entropy of an example document set, quantifying the noise density, geometric distortion degree and background complexity of the example document set, and carrying out the self-adaptive selection of a preprocessing assembly line intensity grade according to the above; dynamically configuring image preprocessing parameters and AI recognition model parameters, and generating a recognition engine instance to output a preliminary recognition text; after regularized coarse screening extraction is carried out based on key field description, a multi-candidate generation strategy is started for low-confidence-coefficient candidate text fragments, a multi-person cooperative verification process is triggered for lower-confidence-coefficient fragments, finally all the fragments are processed through a text standardization module, and structured text extraction information is output. According to the method, accurate adaptation of processing intensity is achieved through document quality quantitative evaluation, the extraction accuracy and system robustness of complex heterogeneous documents are effectively improved through a multi-level confidence coefficient verification mechanism, and the identification error risk caused by image quality fluctuation or rule solidification is reduced.
Owner:BEIJING VOCATIONAL COLLEGE OF ECONOMICS & MANAGEMENT (BEIJING MANAGER COLLEGE)

Enterprise-level unstructured knowledge governance-oriented method and storage medium

The invention provides an enterprise-level unstructured knowledge governance-oriented method and a storage medium. The method comprises the steps that an initial document is input into a visual language model for layout analysis, and document structure information of a target document is obtained; segmenting the document structure information to generate text slices to be vectorized; metadata extraction is carried out based on the text slices, and structured knowledge content vectors are generated; and performing index configuration according to the knowledge content vectors to form enterprise-level knowledge governance standards. According to the method, the document structure is analyzed through the visual language model, the text slices are generated through the intelligent segmentation algorithm, the knowledge vector is constructed through metadata extraction, and the mixed index configuration is performed, so that the problems of inaccurate layout analysis, logic structure damage and low retrieval efficiency in the traditional technology are effectively solved, and the method has the advantage of improving the knowledge management precision and availability.
Owner:DIGITAL CHINA CHINA CO LTD +1

Zero sample template inference and document structured recognition method and device

The invention relates to the cross technical field of computer vision and natural language processing, and particularly provides a zero sample template inference and document structured recognition method and device, and the method comprises the following steps: S1, document image collection and preprocessing; s2, performing layout sensing partitioning and position coding; s3, priori or example information is constructed and injected; s4, performing cross-modal fusion and expression construction; s5, performing automatic format analysis and field slot filling; s6, generating a cue word-free extraction instruction; s7, performing field area parallel character recognition; s8, performing semantic verification and result standardization; and S9, outputting the structured field-value data. Compared with the prior art, the method has the advantages that the dependence of a traditional method on template making, cue word writing and large-scale sample training can be avoided, the flexibility, accuracy and online speed of document structured recognition are remarkably improved, and the method has good intelligence and rapid adaptation capacity and is suitable for diversified document recognition scenes.
Owner:INSPUR SOFTWARE CO LTD

Document structure extraction and model training method and device, equipment and medium

The invention discloses a document structure extraction and model training method and device, equipment and a medium, and relates to the technical field of artificial intelligence and computer vision. The method comprises the following steps: constructing a special training data set containing data of at least two document understanding tasks (including optical character recognition, layout analysis, text positioning, regional text extraction, image description and chart title generation), and a fine tuning data set for converting a document image into a machine-readable structured text format; constructing a multi-modal large model comprising a shape adaptive cutting module, a visual encoder, a visual token compression module, a modal connector and a language decoder; pre-training the model by using the special training data set to jointly learn various document understanding tasks; and performing fine tuning on the pre-trained model by using the fine tuning data set, and adapting to a document structure extraction task to obtain a document structure extraction model. By means of the technical scheme, efficient and accurate document structure extraction can be achieved.
Owner:CETC CYBERSPACE SECURITY TECH CO LTD

Word document format conversion method based on Java

The invention particularly relates to a Word document format conversion method based on Java. The Word document format conversion method based on Java comprises the following steps: analyzing the content of an original. Doc file, and extracting text paragraphs, tables, pictures, style information and document metadata; the extracted content is divided into different categories, and the XML node type corresponding to each category of elements in the target. Docx document is established; the method comprises the following steps of: constructing a pattern mapping rule base, constructing a new. Docx document structure by using an XWPF Document object model according to an Office Open XML (Extensible Markup Language) specification, sequentially inserting paragraphs, tables and pictures, and applying corresponding pattern configuration; and outputting and generating a. Docx file, detecting and processing abnormal conditions, and recording a conversion log at the same time. The Word document format conversion method based on Java is efficient and accurate, has good compatibility, expandability and cross-platform capability, is suitable for enterprise-level document management systems, cloud services and batch document processing scenes, and remarkably improves document compatibility and processing efficiency of office automation systems.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Interface verification and test method and device based on large model, medium and product

The embodiment of the invention relates to the technical field of information, and discloses an interface verification and test method and device based on a large model, a medium and a product, and the method comprises the steps: obtaining a to-be-tested interface document, analyzing the interface document through a first large model, constructing a test request parameter based on a request parameter definition, acquiring a real-time response message by calling a real interface; performing semantic comparison on a response example in the document structure information and the real-time response message by using the second large model to obtain a corrected interface document; comparing the parameter constraint condition in the corrected interface document with the verification logic in the business logic code snippet by using a third model to generate a logic verification passing result; and on the basis of the corrected interface document and a logic verification passing result, a Mock rule and an automatic test case script are generated by utilizing the fourth model, so that the maintenance cost and the false alarm rate of the automatic test script are reduced, and the automatic test efficiency is improved.
Owner:SHANGHAI SHANGHU INFORMATION TECH CO LTD

Product image-text and document association identifier generation method fused with PLM coding rule

The invention discloses a product image-text and document association identifier generation method fusing a PLM coding rule, relates to the technical field of product data management, and is used for solving the problem that an image-text and document association identifier is not clear. According to the method, document information vectors are constructed by fusing PLM coding rules and collecting document metadata of design drawings, process documents, quality management and bill-of-material systems, an image-text document structure mapping chain is generated through a structure mapping judgment model, a unique association identifier of the image-text document is generated in combination with a preset template, and the unique association identifier of the image-text document is obtained. And real-time updating is carried out during structure replacement, hierarchical reconstruction or configuration change, cross-system serial number semantic alignment is realized through a heterogeneous coding analysis rule base, heterogeneous naming conflicts are solved, identifiers are written into a product data management platform, an index relationship is established, hierarchical retrieval, version switching and historical backtracking are supported, and the product data management efficiency is improved. And the consistency and cooperation efficiency of product research and development and quality control are improved.
Owner:TIANJIN BINHAI TONGDA POWER TECH

Rich text multi-dimensional difference analysis and accurate revision method and system based on AI

The invention discloses a rich text multi-dimensional difference analysis and accurate revision method and system based on AI, and relates to the technical field of artificial intelligence and data analys.The rich text multi-dimensional difference analysis and accurate revision method comprises the steps that firstly, a rich text file is analyzed, a document structure is reconstructed, vectorization and key information labeling are conducted on semantic blocks, and structuralization and semantic enhancement of a text are achieved; the method comprises the following steps of: firstly, carrying out text surface layer, semantic deep layer and structure association dimension difference analysis in parallel, then carrying out system knowledge graph-based revision influence degree evaluation on the basis of difference analysis, tracing associated terms through penetrating analysis, calculating an influence degree score and early warning potential conflicts, and automatically generating accurate revision suggestions; visual difference preview and one-key generation of a revised draft are provided, a closed loop from problem discovery to problem solving is realized, and the efficiency and accuracy of system revision are greatly improved.
Owner:YGSOFT INC

Engineering document generation method and system based on logic structure tree and LLM and medium

The invention discloses an engineering document generation method and system based on a logic structure tree and LLM and a medium. The method comprises the steps that original data are analyzed to form an initial text segment set X; calling LLM to generate an initial technical draft based on X, extracting a logic structure, and constructing an initial logic structure tree T; maintaining a search queue, for leaf nodes in the queue, acquiring context information related to semantics from X through a retrieval enhancement mechanism, after same-level brother node deduplication and global deduplication judgment, calling LLM to expand the nodes to generate child nodes, optimizing text description, and recursively enriching a tree structure; and finally, based on the X and the optimized tree, calling LLM to generate a title, an abstract and a background, generating content, details and key point chapters through traversal, and combining into a complete engineering document. The system comprises a document analysis module, a knowledge base module, a retrieval enhancement module, a logic structure tree maintenance module and a document generation module. The method is high in automation degree, and the generated document is clear in structure, complete in content and high in specialty.
Owner:SCHOOL OF SOFTWARE ZHEJIANG UNIV (NINGBO) MANAGEMENT CENT (NINGBO SOFTWARE EDUCATION CENT)

Document analysis large model, training method and multi-task document structured analysis method

The invention relates to a large document analysis model, a training method and a multi-task document structured analysis method, which are suitable for the fields of artificial intelligence, computer vision and natural language processing. The model comprises a visual coding module used for extracting an image feature vector from a document page image, and the extraction of the image feature vector comprises the step of capturing a dependency relationship on a spatial dimension and a channel dimension by adopting a double attention module; the visual feature projection module is used for mapping the image feature vector extracted by the visual coding module from a visual feature space to a feature space of the text decoding module to generate a visual projection feature vector; and the text decoding module is a large language model based on a Transform decoder architecture, and is used for generating document space layout information or document content corresponding to the document processing task on the document page image based on the visual projection feature vector and the token sequence after word segmentation of the cue word of the document processing task.
Owner:POWERCHINA HUADONG ENG CORP LTD +1

Intelligent processing method for standard documents in ship industry

The embodiment of the invention provides a ship industry standard document intelligent processing method, which comprises the following steps: S1, document analysis: carrying out structured analysis and restoration on a ship industry standard document to generate a structured text supporting vectorization storage and semantic retrieval; s2, database construction: constructing a document vector index database based on the structured text to form a semantic retrieval task-oriented efficient data structure; and S3, content retrieval: performing user-oriented query, constructing a semantic retrieval mechanism, and completing high-precision matching from a natural language problem to a structured document content and result return. According to the embodiment of the invention, the document structure can be accurately restored, the content segmentation is more reasonable, the retrieval matching is more accurate, the retrieval result has more content depth, the technical path is clear, and the project landing is easy.
Owner:SHANGHAI WAIGAOQIAO SHIP BUILDING CO LTD

Intelligent long report writing method and system based on chapter structure control

The invention discloses an intelligent long report writing method and system based on chapter structure control, and the method comprises the steps: inputting a writing target, determining a report structure template, and carrying out the analysis and decomposition of the writing target into document structure objects with a hierarchical relation based on the report structure template; traversing each chapter node in the document structure object and distributing a chapter label to the chapter node to construct a tagged document structure, and then mapping the chapter label to a specific generation strategy; setting a group of dynamic writing path schedulers for constructing a group of nonlinear optimal writing execution paths based on the logic dependency relationship between chapter nodes and the availability of required resources; according to the optimal writing execution path, executing a generation strategy at the chapter node to generate chapter content; and integrating the chapter content into the report document. According to the invention, the structure consistency and logic preciseness of the report can be obviously improved, and the accurate control and optimization of the whole writing process can be realized.
Owner:北京京能能源技术研究有限责任公司 +1

Drug production data table processing method based on document analysis and HTML rendering

The invention discloses a medicine production data table processing method based on document analysis and HTML rendering, and relates to the technical field of document data processing. The medicine production data table processing method based on document analysis and HTML rendering comprises the steps that S1, table recognition data, position mapping data and expressive data are collected and preprocessed, and a standardized table state data set is constructed; s2, analyzing semantic association closeness between the nested table and the affiliated paragraph, and adjusting an affiliation marking strategy of the table; s3, evaluating the structuring degree of the cells, and reconstructing the logical hierarchical relationship of the table structure; s4, evaluating the information importance of the cells, and adjusting the layout priority of the cells; and S5, generating a structured record and connecting the structured record to a business process of a manufacturing execution platform. The problem that a document structure and a business field are lack of binding, and construction of a unified data main line and a business driving process is seriously hindered is solved.
Owner:CHENGDU HONGRUI TECH

Cryptographic integrity verification and adaptive artificial intelligence document extractor system for workflow automation in various domains

A system and method for secure and efficient automated workflows includes two complementary components. A digest embedding system verifies the integrity of workflow event sequences using a rolling SHA-256 digest salted with microsecond-precision timestamps. A template-caching extractor adaptively processes heterogeneous electronic documents. The digest system enables decentralized verification without querying centralized audit logs. The extractor uses a layout hash derived from document structure to route documents through either a low-latency, rule-based extraction path or a fallback artificial intelligence model path. New templates are generated for previously unseen layouts exceeding a confidence threshold. The disclosed methods improve latency, resource utilization, energy utilization and scalability in sectors including finance, healthcare, and logistics, offering advantages over existing prior art in terms of integration, specific mechanisms for timestamp salting, hardware security module utilization for workflow events, layout-based template caching, and adaptive learning.
Owner:LEGACI LABS INC

Inference enhanced RAG retrieval method and system based on semantic tree index

The invention discloses a semantic tree index-based reasoning enhanced RAG retrieval method and system. The method comprises the following steps of: S1, reading and analyzing a document; s2, structuring the semantic knowledge tree by the document; s3, carrying out tree search and retrieval; and S4, enhancing the answer by the RAG system. According to the method, a long document of a specific field task is structured into an indexable semantic knowledge tree, and a large language model is guided to perform multi-step reasoning and path search on the tree in a retrieval process, so that a logic node which is most matched with query is dynamically locked, fundamental transformation from semantic similarity matching to structured logic positioning is realized, and the search efficiency is improved. The method does not need to depend on a vector database and embedded model training.
Owner:JIANGSU HOPERUN SOFTWARE CO LTD

Iterative graph neural network-based event causal identification method, apparatus and device, and medium

The invention discloses an event causal identification method and device based on an iterative graph neural network, equipment and a medium, and relates to the technical field of artificial intelligence and machine learning, and the method comprises the steps: carrying out the sentence coding and event extraction of an input text, and obtaining a sentence embedding and event mention result; then, sentence-level embedding and document-level embedding of the event are generated using a multi-granularity context awareness mechanism. Then, constructing an initial event causal graph structure, and encoding the initial event causal graph structure to obtain graph embedding; and finally, dynamically updating an event causal graph structure by combining sentence-level embedding, document-level embedding and graph embedding through an iterative graph optimization mechanism, and realizing accurate identification of the event causal relationship. Through a multi-granularity context perception mechanism and an iterative graph optimization mechanism, local and global context information is effectively integrated, the accuracy and robustness of document-level event causal relationship recognition are improved, and the method is particularly excellent in performance when processing long texts and cross-sentence causal relationships and can better adapt to complex document structures.
Owner:NAT UNIV OF DEFENSE TECH

Structural semantic dual-driven multi-level document intelligent slicing and associating method and system

The invention discloses a structural semantic dual-driven multi-level document intelligent slicing and associating method and system, and relates to the technical field of data processing. The method comprises the following steps of: firstly, automatically constructing a document structure tree containing a father-child nesting relationship by jointly analyzing format layout features and text semantic features of a document; traversing the document structure tree to generate document slices by taking nodes as units, and generating unique hierarchical path codes based on node positions; then, semantic vectors of the slices are generated in parallel, and hierarchical path codes are mapped into structured position vectors through a neural network; and calculating a fusion weight by using an adaptive gating fusion mechanism, and carrying out deep fusion on the semantic vector and the structured position vector to obtain a final representation vector. According to the method, the problems of structural analysis distortion, homonymous title slice affiliation fuzziness and association breakage during long document processing in the prior art are effectively solved, and the retrieval precision ratio and the large model application efficiency are remarkably improved.
Owner:INFORMATION SCI RES INST OF CETC

Cryptographic integrity verification and adaptive artificial intelligence document extractor system for workflow automation in various domains

A system and method for secure and efficient automated workflows includes two complementary components. A digest embedding system verifies the integrity of workflow event sequences using a rolling SHA-256 digest salted with microsecond-precision timestamps. A template-caching extractor adaptively processes heterogeneous electronic documents. The digest system enables decentralized verification without querying centralized audit logs. The extractor uses a layout hash derived from document structure to route documents through either a low-latency, rule-based extraction path or a fallback artificial intelligence model path. New templates are generated for previously unseen layouts exceeding a confidence threshold. The disclosed methods improve latency, resource utilization, energy utilization and scalability in sectors including finance, healthcare, and logistics, offering advantages over existing prior art in terms of integration, specific mechanisms for timestamp salting, hardware security module utilization for workflow events, layout-based template caching, and adaptive learning.
Owner:LEGACI LABS INC

Multi-modal mixed document OCR (Optical Character Recognition) and structured extraction method

The invention relates to the technical field of content extraction, in particular to a multi-modal hybrid document OCR (Optical Character Recognition) and structured extraction method, which comprises the following steps of: acquiring an image text region bounding box and classifying a style, extracting a font or stroke sequence to generate a character positioning structure, dividing paragraph and sentence groups to classify semantic fields, and calculating a field matching relationship to generate structural mapping. According to the method, logic mapping is constructed through character two-dimensional coordinate sorting and paragraph contours, complete reconstruction of a page structure is enhanced, semantic field categories are extracted by using syntactic density of sentence paragraph division and inter-paragraph features, the accuracy of field classification is improved, and the method has the advantages of being simple in structure, convenient to operate and high in practicability. A field mapping relation is established through Jaccard similarity and part-of-speech consistency analysis between a head word and a standard field keyword, a field path index and a structure node link are clarified, field semantic affiliation and structure position output are unified, and the document structure reduction degree and field extraction accuracy are improved.
Owner:HANGZHOU JINGSHENG HANGXING TECH CO LTD

General document structured analysis method

The invention discloses a general document structured analysis method, and relates to the field of document intelligent analysis, and the method comprises the following steps: carrying out global layout analysis based on a thumbnail to obtain each layout area position; performing local content identification on each layout area based on the native resolution to obtain corresponding element contents; and sorting and splicing the element contents to obtain the structured document. Through decoupling global layout analysis and local content recognition, computing resources are accurately put into an information area, and unification of efficiency and precision is achieved; and multiple tasks are processed through a unified model and special elements are processed through a special algorithm, so that the integrity and accuracy of analysis are ensured.
Owner:SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT

Word document display method and system based on structured semantic analysis, terminal and medium

The invention relates to the field of document processing, and particularly provides a Word document display method and system based on structured semantic parse, a terminal and a medium. Firstly, a dynamic rule priority queue arranged in a descending order according to the number of failures is constructed by analyzing historical document samples; guiding the analysis engine to cooperate with the rule and the machine learning model to perform high-precision structured analysis on the document; constructing a document structure tree rich in semantic association, and extracting entity relationships to generate knowledge graph sub-graphs; and finally, the serialized data increment is transmitted to a front end, and catalog jump, reference tracking, term prompting and graph visualization are realized through componentization rendering. According to the method, intelligent interaction of the Word document is realized through a self-adaptive analysis strategy and deep semantic modeling, and the information acquisition efficiency is improved.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Document analysis method and system, terminal and medium

The invention belongs to the technical field of document analysis, and particularly discloses a document analysis method and system, a terminal and a medium. Comprising the steps of obtaining content and page representation of a to-be-analyzed document; performing adaptive segmentation on the document content to obtain a segmentation sequence; performing joint coding on the text features, the visual features and the layout features of the segmented sequence to generate fusion feature representation used for representing cross-modal relevance; determining a hierarchical structure relationship of the document based on the fused feature representation; inputting a text corresponding to the hierarchical structure into a language model subjected to field training to generate a semantic analysis result of the document; and according to a semantic analysis result, generating knowledge representation of the documents and establishing associated information between the documents. According to the method, a consistent processing link can be formed among multi-modal feature fusion, document structure inference and semantic understanding, and high-precision structured analysis of complex document contents is realized.
Owner:浪潮智慧科技有限公司 +2

Document segmentation method and system based on context marking and model cascading

The invention relates to the technical field of natural language processing, and provides a document segmentation method based on context marking and model cascading, which comprises the following steps: in response to a document segmentation request, loading a to-be-processed document and initializing segmentation parameters; calling a large language model to analyze a current to-be-processed text segment, identifying a logic demarcation point, and generating wedge information containing a segmentation point mark and a context; positioning absolute positions of segmentation points in the text segment according to the wedge information, and segmenting the text segment into a plurality of text sub-segments; repeating the execution until a preset recursion termination condition is reached; after recursion is completed, segmenting results of all layers are aggregated, a hierarchical document structure is constructed, and the segmenting results are output. Through a wedge mechanism, only tiny positioning marks are output, the token cost is reduced, and the overall cost is optimized in combination with a model cascading strategy. And the generated text block is highly aligned with the semantic boundary of the document, so that the context fragmentation problem is effectively solved, the context relevance is enhanced, and the model illusion is inhibited.
Owner:SHANGHAI WENYIN INTERNET INFORMATION TECHNOLOGY CO LTD

Document processing method and related equipment

The invention discloses a document processing method and related equipment, and belongs to the technical field of data processing.The method comprises the steps that in response to a document input instruction, a to-be-processed target document is obtained; performing structure analysis on a to-be-processed target document to generate a directory tree of the target document; calling a dynamic recursive slicing algorithm, traversing and analyzing the directory tree of the target document, and generating a plurality of text slices with dynamic lengths; and vectorizing the plurality of text slices with the dynamic lengths, and uploading and storing the vectorized text slices to a database of a retrieval enhancement generation system. By constructing the directory tree and traversing the directory tree by adopting the dynamic recursive slicing algorithm, the document structure can be more accurately understood, the semantic integrity of the document slices is improved, and the retrieval accuracy and the answer generation quality of the retrieval enhancement generation system are further effectively improved.
Owner:E-SURFING DIGITAL LIFE TECH CO LTD