Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

34 results about "Portable document format" patented technology

The Portable Document Format (PDF) (redundantly: PDF format) is a file format developed by Adobe in the 1990s to present documents, including text formatting and images, in a manner independent of application software, hardware, and operating systems. Based on the PostScript language, each PDF file encapsulates a complete description of a fixed-layout flat document, including the text, fonts, vector graphics, raster images and other information needed to display it. PDF was standardized as ISO 32000 in 2008, and no longer requires any royalties for its implementation.

PDF (Portable Document Format) document processing method and device, equipment and medium

The invention relates to the technical field of computers, and discloses a PDF (Portable Document Format) document processing method, device and equipment and a medium, the method comprises the following steps: carrying out global layout analysis on a PDF document to detect element information of all structural elements in the PDF document, the element information comprising bounding box coordinates, element categories and a reading sequence; based on the bounding box coordinates, cutting out a corresponding local area from the PDF document, and performing content identification on different types of structural elements corresponding to the local area; according to the element category and the reading sequence, carrying out recombination and logic division on a content recognition result obtained by carrying out content recognition to obtain a plurality of logic parts; and for each logic part, constructing a cue word, calling a large language model to carry out thinking chain reasoning so as to extract structured information corresponding to the logic part, and merging the structured information to generate a JSON file. According to the method and the device, the PDF document processing accuracy is improved.
Owner:传申弘安智能(深圳)有限公司 +1

PDF (Portable Document Format) document structured analysis method based on weighted overlap ratio

The invention discloses a PDF (Portable Document Format) document structured analysis method based on weighted overlap ratio, which comprises the following steps of: performing layout analysis, standardization and sorting processing on an original PDF document to obtain a layout frame of the original PDF document and a corresponding category label data set; carrying out content object extraction on the original PDF document and constructing to obtain a content object box set of the original PDF document; based on a weighted coincidence degree scoring method, obtaining an optimal attribution layout of the original PDF document content object, and identifying an abnormal scene; based on a configurable rule engine, performing cross validation and deviation correction on the layout frame and the content object frame, and dynamically updating the layout-document content object tree; and based on the dynamically updated layout-document content object tree, outputting an analysis result, and converting and storing the analysis result. According to the method, the problems of inaccurate layout identification, wrong content classification, document content missing and the like are solved, and high-precision, high-robustness and high-flexibility structured analysis of the complex PDF document is realized.
Owner:IOL WUHAN INFORMATION TECH CO LTD

PDF (Portable Document Format) document content identification method and device, equipment and storage medium

The invention provides a PDF (Portable Document Format) document content identification method and device, equipment and a storage medium. The method comprises the following steps: acquiring an access link of an unanalyzed document; downloading the target PDF document from the object storage service according to the access link; when the content region corresponding to the page type in the target PDF document is a text region, dividing the content region into a text region, a table region and an image region according to the page type corresponding to each content region; and when the content region is a text region, extracting a native text character sequence from the text region, and performing similarity calculation on the native text character sequence to generate a semantic coherent paragraph text. According to the method, the sentences with similar semantics in the text region are automatically divided into the same text block based on the cosine similarity, so that the text content with coherent and complete semantics is analyzed from the document, the problem that text paragraphs are broken after document recognition is solved, and the semantic coherence of the document content is effectively improved.
Owner:SHENZHEN ISSMART SCI & TECH CO LTD

PDF (Portable Document Format) file word embedding method and device and related medium

The invention discloses a PDF (Portable Document Format) file character embedding method and device and a related medium, and the method comprises the following steps: analyzing a content stream of a PDF file to establish a first data set mapped between a text object and a candidate font in the content stream; counting fonts in the first data set and carrying out combination and deduplication to generate a second data set; the measurement parameters in the second data set are analyzed to screen fonts needing to be reserved, and a third data set is obtained; performing font cutting processing on the third data set to construct subset fonts to obtain a fourth data set; generating a subset font dictionary, a font stream and font mapping based on the fourth data set, and associating with a preset document resource to obtain a fifth data set; and embedding the font binary data in the fifth data set into the PDF file to generate a target PDF file. According to the method, the font binary data in the fifth data set obtained through final calculation is embedded into the PDF file, so that the font resource volume of the PDF file is reduced, and the rendering efficiency is improved.
Owner:SHENZHEN JINNIU TECH CO LTD

Machine learning powered cloud sandbox for malware detection in portable document format (PDF) files

A cloud-based network security system (NSS) is described. The NSS uses a sandbox to safely open and extract information about a PDF file and uses machine learning algorithms to analyze the information to predict whether the PDF file contains malware. Specifically, dynamic information about the PDF file is captured while it is open in the sandbox. Static information is extracted from the PDF file as well. The dynamic and static information is input to an AI or machine learning model trained to provide an output indicating a prediction of whether the PDF file contains malware. A verdict engine uses the output from the AI or machine learning model to classify the document as malicious or clean. Security policies can then be applied based on the classification.
Owner:NETSKOPE INC

File processing method and device, computer equipment and storage medium

The invention relates to a file processing method and device, computer equipment and a storage medium. The method belongs to the technical field of file processing, and comprises the following steps: analyzing an OFD file to obtain a document object model; constructing an intermediate model based on the document object model; and performing mapping processing on the intermediate model to obtain a target portable document format (PDF) file corresponding to the OFD file. According to the method, the OFD file is analyzed to obtain the document object model, then the document object model is converted into the intermediate model, and the intermediate model is accurately mapped to obtain the target PDF file, so that the conversion efficiency between the OFD file and the PDF file is improved, and the method is suitable for high-concurrency and batch processing scenes; and the obtained target PDF file does not have the condition of content or typesetting disorder or damage, so that high-fidelity conversion from the OFD file to the PDF file is realized.
Owner:CHINA LIFE INSURANCE CO LTD

Computing system and method for extracting unstructured document data

A computing system and method are configured to receive an input in a portable document format, generate a raster image of the input, and detect a form type of the input based on the raster image. In response to detecting that the raster image of the input indicates a first form type, the computing system and method are configured to partition the raster image into sections corresponding to document regions containing document entities targeted for retrieval for a specific form type, generate sub-images based on bounding coordinates of the sections, and apply selected document data identification and retrieval computational techniques to extract the document entities from the sub-images.
Owner:VELOCITYEHS HOLDINGS INC

Training data for training artificial intelligence agents to automate multimodal software usage

A system for automating software usage includes an agent configured to automate. The agent is trained on one or more training data sets. The one or more training datasets include one or more of a first training dataset including documents containing text interleaved with images, a second training dataset including text embedded in images, a third training dataset including recorded videos of software usage, a fourth training dataset including portable document format (PDF) documents, a fifth training dataset including recorded videos of software tool usage trajectories, a sixth training dataset including images of open-domain web pages, a seventh training dataset including images of specific-domain web pages, and / or an eighth training dataset including images of agentic trajectories of the agent performing interface automation task workflows.
Owner:ANTHROPIC PBC

Machine learning powered cloud sandbox for malware detection in portable document format (PDF) files

ActiveUS20260099598A1Digital data protectionPlatform integrity maintainanceEngineeringPortable document format
A cloud-based network security system (NSS) is described. The NSS uses a sandbox to safely open and extract information about a PDF file and uses machine learning algorithms to analyze the information to predict whether the PDF file contains malware. Specifically, dynamic information about the PDF file is captured while it is open in the sandbox. Static information is extracted from the PDF file as well. The dynamic and static information is input to an AI or machine learning model trained to provide an output indicating a prediction of whether the PDF file contains malware. A verdict engine uses the output from the AI or machine learning model to classify the document as malicious or clean. Security policies can then be applied based on the classification.
Owner:NETSKOPE INC

Cooperative review management system and method based on PDF structured preview and precise annotation

The invention belongs to the technical field of computer software, and particularly relates to a collaborative review management system based on PDF (Portable Document Format) structured preview and precise annotation, which is characterized in that structural analysis is carried out on an uploaded PDF through a PDF JavaScript analysis engine, paragraph structure information is extracted, and hierarchical and paging quick preview is realized; providing a text-level precise annotation tool to firmly bind approval and original text character positions and perform reverse highlight positioning; a Web-based real-time cooperation platform is constructed, multi-user synchronous review is supported, and full-life-cycle management including states and persons in charge is carried out on review opinions; and a large PDF file is efficiently processed by adopting a paging loading and dynamic memory management strategy. According to the method, the problems of difficult cooperation, rough annotation, poor traceability, large file performance bottleneck and the like in a traditional review mode are effectively solved, and the efficiency and quality of scenes such as technical review and contract auditing are remarkably improved.
Owner:AEROSPACE SCI & IND INTELLIGENT OPERATION RES & INFORMATION SECURITY RES INST (WUHAN) CO LTD

Method, apparatus and device for information interaction and storage medium

ActiveCN119005211BNatural language data processingEngineeringPortable document format
According to embodiments of the present disclosure, methods, apparatuses, devices and storage media for information interaction and processing are provided. The information interaction method includes, in response to a preset operation, selecting a portable document format (PDF) plug-in in an interaction window of a user and a digital assistant; and using the PDF plug-in to perform a task associated with a target PDF file in an interaction process of the user and the digital assistant. Thus, the user can be supported to use the PDF plug-in to complete a task related to the PDF file in the interaction process with the digital assistant. In this way, it is beneficial to provide an efficient assistance experience to the user.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Bank flow analysis method and device, electronic equipment and storage medium

The invention provides a bank flow analysis method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a portable document format (PDF) file and a bank identifier associated with the PDF file; analyzing the PDF file, and determining a first character string set contained in the PDF file and first position information of each first character string in the first character string set; in a pre-configured candidate template library, a reference template associated with the bank identifier is obtained, and the reference template comprises a second character string corresponding to at least one table column; each second character string is matched with all the first character strings, and target character strings successfully matched with the second character strings in the first character string set and target position information of the target character strings are determined; according to all the target position information and the first position information of other first character strings, the flow data corresponding to the PDF file is obtained and stored in the database, and the flexibility and efficiency of PDF bank flow analysis are improved.
Owner:PEOPLE'S INSURANCE COMPANY OF CHINA

Data understanding and extracting method for financial data report

PendingCN122022986AHigh extraction error rateSolve the inefficiency of manual processingDigital data information retrievalFinanceData understandingEngineering
The invention belongs to the technical field of financial risk control, and particularly relates to a data understanding and extracting method for a financial data report, which comprises the following steps: acquiring an original file of the financial data report; wherein the financial data report is stored as a PDF (Portable Document Format) file of an editable version; traversing each page of table data of the original file, and constructing dynamic mapping between a header and a corresponding cell based on deep learning; identifying a corresponding header according to the to-be-analyzed target, and collecting content information recorded by the related cells in a cross-page manner by utilizing a mapping relationship between the header and the corresponding cells; and splicing all the collected content information, and generating standardized structural data and analysis quality reports according to the spliced content information, so that the large model can call the structural data and the corresponding analysis quality reports to perform financial risk control analysis.
Owner:LONGWAGE TECH (SHANGHAI) CO LTD

Artificial intelligence (AI) educational resource generation system

PendingUS20260188137A1AlgorithmDocument transformation
Embodiments receive a portable document format (PDF) from a user computing device; convert the PDF to a JPEG file by utilizing an artificial intelligence (AI) image segmentation model; convert the JPEG file to a text file by utilizing an AI vision workflow model; parse and classify the text file using a first large language model (LLM); determine a textual study guide using a second LLM; generate a question and answer exam based on the textual study guide using a third LLM; and generate a multiple choice question (MCQ) exam based on the question and answer exam using a plurality of LLMs.
Owner:EDAI SYSTEMS INC

PDF document screenshot method

PendingCN121639705AImage analysisNatural language data processingEngineeringPortable document format
The embodiment of the invention discloses a PDF (Portable Document Format) document screenshot method. The method comprises the steps that a user screenshot request is obtained, the user screenshot request comprises target keywords, a PDF document template and one or more target PDF documents, the PDF document template is configured with multiple sets of screenshot area parameters determined based on different keywords, and the target PDF documents and the PDF document template have the same format. The method comprises the following steps: positioning a plurality of target screenshot areas from one or more target PDF documents based on target keywords and a PDF document template; in addition, the method includes generating a plurality of screenshots in batches based on the plurality of target screenshot areas. In this way, according to the pre-configured document template, multi-thread parallel processing is achieved, a plurality of target screenshot areas can be located in the same PDF document, or the target screenshot areas can be located in a plurality of PDF documents at the same time, and therefore batch screenshot in the same document and cross-document batch screenshot are achieved, the screenshot efficiency is high, and screenshots are accurate, complete and uniform in format.
Owner:HANGZHOU REPUGENE TECH CO LTD

Comparison method for single PDF (Portable Document Format) file tables in power industry based on multi-modal large model retrieval enhancement

PendingCN121659924ASemantic analysisText processingEngineeringPortable document format
The invention discloses a comparison method for single PDF (Portable Document Format) file tables in the power industry based on multi-modal large model retrieval enhancement. The comparison method comprises the following steps: constructing an RAG knowledge base containing field specifications, unit conversion rules and the like in the power industry; the PDF to be compared is uploaded and converted into an image, and table information is extracted through OCR and structured reconstruction; cross-page judgment and splicing are completed through table name matching and table header checking, and a table structure is analyzed and adapted; semantic analysis is combined with RAG to achieve field alignment, units are automatically converted, and the reasonability of numerical value differences is judged; and generating a structured report containing a comparison list, difference interpretation and the like. The method solves the problems of low efficiency and poor accuracy of traditional comparison, adapts to a complex table scene, and improves the efficiency and accuracy of PDF table comparison in the power industry.
Owner:GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD

Document distribution method and device, computer equipment and readable storage medium

The invention relates to a document distribution method and device, computer equipment and a readable storage medium. The method comprises the steps of obtaining a pre-constructed initial document reading application; embedding a to-be-distributed portable document format document and an identification association rule of the portable document format document into the initial document reading application to obtain an intermediate document reading application; configuring a security access strategy of the document in the portable document format in the intermediate document reading application to obtain a target document reading application; and packaging the target document reading application into an executable file and forwarding the executable file to the target equipment, so that the target equipment runs the target document reading application based on the executable file, and positioning the portable document format document based on the identification association rule through the target document reading application. And calling and displaying the portable document format document based on the secure access policy. By adopting the method, the document distribution security can be improved.
Owner:SUZHOU ZONGWEI AUTOMATION CO LTD

Classifier for identifying suspicious PDF files to limit deep-scanning

A cloud-based network security system (NSS) is described. The NSS extracts information about a document (e.g., a portable document format (PDF) file) and uses heuristic rules to analyze the information to predict whether the document contains malicious software. Specifically, prior to detonation of the document, object features, code features, and embedded features of the document are extracted. The extracted information is input to a classification engine that applies sets of heuristic rules to groups of the features of the document to provide an output indicating a prediction of whether the document contains malware. A routing engine provides the document for further analysis (e.g., deep scanning) if the document is suspicious or bypasses the further analysis if the document is benign. Security policies can then be applied based on the classification.
Owner:NETSKOPE INC

Methods and systems for intelligently and adaptively managing and using data in a supply chain environment

PendingUS20260127546A1Ensemble learningKernel methodsEngineeringPortable document format
Disclosed herein are systems and methods for the automated ingestion and processing of orders for the food supply chain industry using artificial intelligence. An example method can comprise extracting order level information and item level information from a purchase order in the form of at least one of an email message, a text message, a voicemail audio file, an image file, a spreadsheet file, and a portable document format (PDF) file. The method can also comprise preprocessing the order level information and the item level information into a plurality of machine learning features and inputting the machine learning features into a machine learning model to obtain predictions concerning the purchase order. A sales order can then be generated based in part on the predictions outputted by the machine learning model.
Owner:GRUBMARKET INC

Document format conversion method and device, computer device and storage medium

ActiveCN114764559BNatural language data processingEngineeringPortable document format
The application relates to a document format conversion method and device, computer equipment and a storage medium. The application is applied to a main process, and the method comprises the following steps: acquiring file header information of a portable document format (PDF) document, wherein the file header information of the PDF document comprises a data amount of the PDF document; acquiring an optimal first number of fragments according to the file header information of the PDF document; creating a same number of sub-processes as the optimal first number of fragments; dividing the PDF document into a plurality of first fragments according to the optimal first number of fragments; sending each first fragment to a sub-process, converting the first fragment into a first picture through the sub-process; receiving the first picture; and generating a picture corresponding to the PDF document according to the first picture. The application can improve the efficiency of converting the PDF document into a picture.
Owner:JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD +1

Detecting object burn-in on documents in a document management system

A document management system surfaces changes to a portable document format (PDF) document to a user. The document management system converts each page of the PDF document into images, segments those images, and processes each segment of those images using computer vision and / or natural language processing. The document management system compares segments from an original copy of the PDF document with segments from a modified copy of the PDF document to identify significant changes to the PDF document.
Owner:DOCUSIGN INC

System and method for pixel perfect conversion, retrospective synthesis and migration of portable document format (PDF) file into editable design templates

Systems and methods to effectively and efficiently replace a legacy CCM product from a first Customer Communication Management (CCM) software application with a newer solution in a second CCM software application, and transition these large-scale and highly regulated Customer Communications Management documents. Embodiments of the present invention present systems and methods that minimize the need for manual template and business logic redevelopment, ensures a substantially pixel-perfect layout format and data alignment, reduces the necessity of operating both the legacy and new CCM products concurrently for an extended duration, and mitigate the risk of data integrity issues during the product transition process. Embodiments of the present invention provide for automated migration of CCM design templates and business rules when transitioning from one CCM vendor (e.g. legacy vendor) having a first CCM software application, to another CCM vendor having a second CCM software application. This was not feasible prior to this invention.
Owner:ELIXIR TECHNOLOGIES CORP

Method and system for identifying financial statement form in PDF (Portable Document Format) document

The invention discloses a method and a system for identifying a financial statement table in a PDF (Portable Document Format) document. The method comprises the following steps: analyzing the PDF document to obtain HTML (Hypertext Markup Language) data and JSON (JavaScript Object Notation) data; obtaining a prediction acquisition template output by the identification model for the table in the PDF document; based on HTML and JSON data, multi-dimensional features of each table to be recognized are extracted; inputting the multi-dimensional features into an expert rule engine, determining a service table type according to a matching result, and executing data storage; if the matching fails, marking the table as an unmatched table; for an unmatched table, if a prediction acquisition template corresponding to the model is identified, determining a service table type and executing data storage; and aiming at the matching failure of the expert rule engine and the failure of outputting the table of the effective prediction acquisition template by the identification model, transferring the table to a manual processing module, and receiving a business table mapping result input by a user for storage. According to the invention, the recognition stability and interpretability can be improved, and the computing power cost is reduced.
Owner:XIAMEN TIANJIAN CAIZHI TECH CO LTD

Data processing method and device, equipment and storage medium

The present disclosure provides a data processing method and device, equipment and storage medium, and relates to the technical field of computer. The method comprises the following steps: obtaining a to-be-processed image format page of a to-be-processed portable document format (PDF) file, wherein the to-be-processed image format page is obtained by converting a to-be-processed page format in the to-be-processed PDF file into an image; performing chart detection on the to-be-processed image format page by using a chart detection model to obtain information of a target chart region of the to-be-processed image format page; classifying the target chart region by using a chart classification model to obtain a chart category label of the target chart region, wherein the chart category label comprises a data chart category, a non-data chart category and a table category; and obtaining related content of the target chart region of the to-be-processed image format page of the to-be-processed PDF file according to the chart category label of the target chart region. The method improves the accuracy of extracting chart information in the PDF file.
Owner:泰康保险集团股份有限公司 +1

Data extraction system and method

A data extraction system and method that provides a reliable, automated alternative to the manual input of financial and other data from portable document format (PDF) documents. The solution of the present disclosure, for example, utilizes an extraction template to parse through each page of the PDF document to identify relevant data elements based on the position relative to other field names shown in the PDF document. Because the data incorporates the actual digital values from the PDF objects (as opposed to an interpreted value from an OCR analysis) the numerical values are processed with a high level of confidence and accuracy.
Owner:CDO CONSULTING LLC

A message pushing method, system and medium for instant processing of enterprise operation information

The application provides a message pushing method, system and medium for instant processing of enterprise operation information, and the method comprises the following steps: acquiring the information type of enterprise operation information; when the information type does not belong to a static type, the enterprise operation information is divided into Web chart information or Excel report information; when the information type is Web chart information, the Web chart information is converted into Web chart configuration data, and the Web chart configuration data is rendered into a first static image file; when the information type is Excel report information, the Excel chart information is converted into a portable document format file, and the portable document format file is converted into a second static image file; the first static image file or the second static image file is uploaded to a cloud storage service to obtain a corresponding image access address; and the pushing information containing visualized operation information is generated according to the image access address. The application can instantly push Web interactive charts and Excel statistical reports.
Owner:HANGZHOU XIAOMA EDUCATION TECH CO LTD

Document analysis method, device and equipment and computer readable storage medium

The invention discloses a document analysis method which comprises the following steps: receiving a to-be-analyzed document, and converting the to-be-analyzed document into a portable document format file; carrying out layout analysis on the portable document format file by utilizing the layout detection model to obtain a layout analysis result; performing content analysis on the portable document format file according to the layout analysis result to obtain a content analysis result; and performing knowledge reconstruction according to the content analysis result to obtain each structured knowledge unit. By applying the document analysis method provided by the invention, unified analysis of documents in different formats is realized, compatible analysis of multi-modal information is realized, and the accuracy of each structured knowledge unit obtained by reconstruction is improved. The invention furthermore discloses a document analysis apparatus and device, and a storage medium, which have corresponding technical effects.
Owner:MALANSHAN AUDIO & VIDEO LABORATORY

Message pushing method and system for instant processing of enterprise operation information, and medium

The invention provides a message pushing method and system for instant processing of enterprise operation information and a medium, and the method comprises the steps: obtaining the information type of the enterprise operation information, and dividing the enterprise operation information into Web chart information or Excel report information when the information type does not belong to a static type; when the Web chart information is the Web chart information, converting the Web chart information into Web chart configuration data, and rendering the Web chart configuration data into a first static image file; when the information is the Excel report information, converting the Excel chart information into a file in a portable document format, and converting the file in the portable document format into a second static image file; uploading the first static image file or the second static image file to a cloud storage service to obtain a corresponding image access address; and generating push information containing visual operation information according to the image access address. According to the method, the Web interactive chart and the Excel statistical report can be pushed in real time.
Owner:HANGZHOU XIAOMA EDUCATION TECH CO LTD