Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1507 results about "Character recognition" patented technology

Engineering document index consistency proofreading method and system based on multi-modal large model

The invention relates to an engineering document index consistency proofreading method and system based on a multi-modal large model, and the method comprises the steps: Q1. OCR detection and recognition: carrying out the optical character recognition and format analysis of a source document, converting an uploaded PDF document into a processable text message in a Markdown format, and carrying out the format discrimination of a table, a formula and a plain text; and Q2, table and formula processing: adopting a hierarchical processing strategy, intelligently selecting an optimal processing mode according to the complexity of the table, and converting table information into a descriptive long text through a language large model and cue words. According to the method, accurate, reliable and efficient document index checking service can be provided for a user, the quality and efficiency of professional document processing are remarkably improved, the efficiency and quality of knowledge graph construction are remarkably improved, a knowledge verification system capable of being evolved continuously is established, and the method is suitable for popularization and application. And a reliable technical support is provided for knowledge management and professional decision-making in a complex field.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719

Table identification reconstruction method and system, terminal and medium

The invention relates to the field of computer vision, and particularly provides a table recognition reconstruction method and system, a terminal and a medium, and the method comprises the steps: firstly decomposing a large-size table image into a plurality of overlapped sub-images, and carrying out the table structure detection and OCR character recognition of each sub-image through parallel recognition; then, sub-graph recognition results are integrated through a coordinate mapping and confidence coefficient weighted fusion algorithm, and boundary errors are eliminated; then, automatically distinguishing common cells based on an area clustering algorithm, merging the cells and a header region, and reconstructing a complete table logic structure; further understanding header semantics through a natural language model and repairing identification errors; and finally, realizing intelligent splicing and standardized output of the cross-page table. According to the method, the memory limitation of the traditional OCR technology is broken through, an oversized table can be processed, the recognition accuracy of a complex structure is improved, and the digitization efficiency of professional documents such as financial statements and engineering drawings is improved.
Owner:INSPUR YUNZHOU (SHANDONG) IND INTERNET CO LTD

RAG-based voucher classification method, medium and equipment

The invention relates to an RAG-based voucher classification method, a medium and equipment, and the method comprises the steps: receiving digital image data of a to-be-classified voucher, carrying out the multi-modal optical character recognition processing to generate a structured OCR result, extracting a text semantic feature vector and a visual layout feature vector based on the structured OCR result, carrying out the fusion of the text semantic feature vector and the visual layout feature vector to generate a multi-modal query vector, and carrying out the classification of the to-be-classified voucher. Similar samples and semantic similarity scores and category metadata thereof are obtained through approximate nearest neighbor retrieval, after an initial candidate category list is generated, key field values are extracted for each candidate category, evidence credibility scores are calculated, comprehensive confidence scores are generated by fusing the semantic similarity scores and the evidence credibility scores, reordering is conducted, and a candidate category list is obtained. And finally, selecting a classification decision path according to the score distribution, and outputting a classification result and an interpretability report. The accuracy and robustness of voucher classification are effectively improved, and complex voucher scenes with changeable formats and fuzzy semantics can be processed; and the interpretability and reliability of the classification decision are enhanced.
Owner:FUJIAN BOSS SOFTWARE

Device for time-based tracking and cost optimization in construction projects

A device for time-based tracking and cost optimization in construction projects, the device comprising the following: a robust housing suitable for use on construction sites; a processing unit located inside the housing, configured to perform real-time time-stamping, data acquisition and preprocessing tasks; a multimodal sensor unit that is operationally coupled with the processing unit, wherein the sensor unit comprises at least a motion sensor, an RFID reader, sensors for environmental conditions and a vision module with optical character recognition; a real-time clock module that is operationally connected to the processing unit to provide time synchronization for all sensor data streams; a wireless communication module that supports the Wi-Fi, LoRa and LTE protocols and is configured for transmitting time-stamped data to a central project server; a storage module that is operationally coupled with the processing unit to locally buffer time series data of construction activity during offline operation; a housing-mounted, touchscreen-based human-machine interface configured to allow site personnel to enter activity updates and confirm the status of construction tasks; a cost optimization engine running on the central server, the engine being configured to receive time-synchronized sensor data from multiple such devices and dynamically calculate time-cost trade-offs using a predictive planning technique that incorporates the principles of the critical path and the power value; furthermore, the device is configured to be integrated into a digital twin environment of the building under construction in order to provide real-time visualization of progress and to generate suggestions for resource reallocation based on a time-cost-benefit analysis.
Owner:1XL INFRA & REAL ESTATE DEVELOPMENT LLC +2

Method for identifying engineering drawing detail table and generating BOM table

The invention discloses a method for identifying an engineering drawing detail table and generating a BOM table, and belongs to the crossing field of automation technology and image processing, and the method comprises the steps: obtaining a scanning or electronic image of an engineering drawing; through preset datum line positioning, recursively detecting a nested rectangular region conforming to an area difference threshold value, and determining a title bar, a detail list region coordinate and a table image; identifying lines and cross points, analyzing the line and column boundaries of the table, and constructing a topological structure; performing character recognition by adopting multi-engine OCR integration; the characters and the cells are associated, a header is recognized through semantics, and analysis data subjected to integrity verification are generated; and automatically generating a structured BOM table in a preset standard format based on the data, and outputting an editable file. According to the method, full-process automation is achieved, manual intervention is not needed, and manual input cost and personal errors are greatly reduced.
Owner:CRRC TAIYUAN CO LTD

System and methods for document processing for data extraction and matching

System and methods are disclosed for matching extracted text data based on one or more similarity scores. The method may include receiving one or more documents from a plurality of data sources, utilizing an optical character recognition algorithm for extracting text data from the one or more documents, comparing, utilizing a fuzzy matching algorithm, the extracted text data to reference dataset(s) to determine one or more matches between the extracted text data and at least one of the reference dataset(s), wherein the one or more matches are based on at least one similarity score, inputting the determined one or more matches and the at least one similarity score into a trained machine-learning model to refine the one or more matches, and outputting a representation of the refined one or more matches and the at least one similarity score to a graphical user interface of a device.
Owner:STATE FARM MUTAL AUTOMOBILE INSURANCE COMPANY

Intelligent paper marking system based on large language model

The invention provides an intelligent paper marking system based on a large language model. The intelligent paper marking system comprises an examinee test paper character recognition module used for carrying out image processing and content recognition on scanned or shot student answer sheets or answer sheets; the subject knowledge base is used for performing systematic arrangement and representation modeling on multi-subject teaching contents; the subject knowledge retrieval module is used for performing semantic analysis and matching on test paper questions and examinee answering contents to obtain subject knowledge related to the test questions and a scoring basis; a scoring template generator module; a large language model scoring module; and the comment correction module is used for optimizing and adjusting the preliminary comments generated by the large language model and outputting final comments with more pertinence and teaching guidance significance. The technical scheme can be widely applied to automatic evaluation scenes of subjective questions in the education field.
Owner:FUZHOU UNIV

Document interpretation and report generation method and device, equipment and medium

The invention relates to the technical field of natural language processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a document interpretation and report generation method, device, equipment and medium, which comprises the following steps: receiving an original document set to generate a structured document object, executing optical character recognition on an image content set to generate a recognition text set, the recognition text set and the text content set are combined into a unified text sequence, element item extraction is executed based on the interpretation template parameter set to generate an interpretation element set, a retrieval enhancement context is retrieved and generated from the domain knowledge base, and the unified text sequence, the interpretation template parameter set and the retrieval enhancement context are input into a language model to generate an interpretation result. And generating report content based on the historical report template set. According to the method, automatic closed loop of document interpretation and report generation is realized through multi-modal unified processing and semantic enhanced reasoning, the efficiency is improved, and the manual dependence and compliance risk are reduced.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Sea cucumber growth character recognition and measurement method based on machine vision and measurement system thereof

The invention relates to a sea cucumber growth character recognition and measurement method based on machine vision and a measurement system thereof, and belongs to the field of image processing, and the method comprises the following steps: S1, image acquisition and preprocessing; s2, instance segmentation and morphological feature extraction; s3, measuring the length and width of the sea cucumber; s4, pixel-to-actual size conversion is carried out; s5, constructing a body weight prediction model; and S6, outputting a result. The method has the advantages that the sea cucumber image is obtained by using the high-definition image acquisition technology, the contour information of the sea cucumber is accurately extracted through the instance segmentation algorithm, and the problems of changeable and irregular sea cucumber shapes and the like are effectively solved by combining the optimized image processing technology, so that the measurement precision is remarkably improved. Meanwhile, a machine learning regression model is introduced, the weight of the sea cucumber is predicted based on various morphological characteristics, and the dimensionality and the utilization value of measured data are further enriched.
Owner:YELLOW SEA FISHERIES RES INST CHINESE ACAD OF FISHERIES SCI +1

Intelligent paper marking system and method for talent selection and recruitment

The invention relates to the technical field of human resource management, and provides an intelligent paper marking system and method for talent selection and recruitment. The method comprises the following steps: cutting test paper according to a preset test paper template by using an image processing technology to obtain a plurality of plates corresponding to the test paper; converting the image content of each section of the cut test paper into a text format which can be edited and processed by adopting a character recognition technology; setting a preliminary scoring standard according to question type information and knowledge point information corresponding to the test paper, and optimizing the preliminary scoring standard by using a large language model; performing semantic understanding and logic analysis on the converted test paper text answers from a plurality of scoring dimensions including semantic accuracy, logic integrity and knowledge point coverage by using two preset scoring large models to give corresponding scoring information, and marking advantage information and defect information in the answers; and obtaining a comprehensive score corresponding to the test paper according to score information printed by each score big model for each question of each test paper.
Owner:SHENZHEN TALENT GROUP CO LTD

Multi-modal information analysis and scheme reminding method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a multi-modal information analysis and scheme reminding method, device, equipment and medium, which comprises the steps of receiving input data and converting the input data into multi-modal data, executing optical character recognition and image classification recognition to generate a recognition result, and sending the recognition result to a server; a natural language processing model is used for analyzing fuzzy description to generate an analysis result, a knowledge base is inquired, a knowledge graph is combined to generate an association result, the analysis result, the association result and user feature data are fused to generate an execution scheme, the execution scheme is compared with an abnormal list, supervision confirmation is triggered, and a compliance instruction is generated. And personalized reminding contents are generated. The information analysis integrity is improved through multi-modal recognition, natural language processing and the knowledge graph, supervision confirmation and personalized reminding are introduced, and intelligent, compliant and reliable reminding management is achieved.
Owner:PING AN TECH (SHENZHEN) CO LTD

OCR character recognition method and system based on end-to-end network

The invention provides an OCR character recognition method and system based on an end-to-end network, and relates to the technical field of image recognition, and the method comprises the steps: inputting a to-be-recognized image into a preset neural network for degeneration type recognition and corresponding processing, and obtaining a preprocessed image; performing character recognition and structure recognition on the preprocessed image by using a pre-trained end-to-end comprehensive model to obtain character text information and character structure information; recognizing a character image according to the character text information, and generating a character structure constraint graph based on the character image; and constructing a feedback result based on the character structure constraint graph and the character structure information, and optimizing the end-to-end comprehensive model by using the feedback result. According to the method, the identification precision can be improved, the robustness and the adaptive capability are improved, and the problems that in the prior art, in the face of an actual complex degraded scene, the adaptive capability is lacked, errors cannot be effectively fed back and optimized, and the identification accuracy and robustness are reduced are solved.
Owner:SHENZHEN KUAITONG TECH CO LTD

Chinese ancient book character recognition method and system based on image recognition technology

The invention relates to the field of character recognition systems, and discloses a Chinese ancient book character recognition method based on an image recognition technology, which comprises the following steps: reading an ancient book image according to an open source code computer vision library to obtain an original image; performing gray processing on the original image to obtain an image after gray equalization; according to a median filtering algorithm, carrying out de-noising processing on the image after gray scale equalization to obtain a pre-processed image; according to the method, multi-level feature extraction is carried out on the image through the real-time multi-scale detection model and the convolutional neural network, and character features of different scales and angles in the image can be captured; and therefore, the character recognition accuracy is improved, especially for the common font, typesetting, inclination or blurring conditions in ancient book images.
Owner:山东齐鲁壹点传媒有限公司 +1

Medicine bottle label content identification method based on multiple cameras and YOLOv8

The invention relates to the technical field of computer vision, image recognition and intelligent medicine management, in particular to a medicine bottle label content recognition method based on multiple cameras and YOLOv8. The method at least comprises the following steps: S1, deploying a medicine bottle label generation system, and generating a label; s2, multi-camera image acquisition and preprocessing; s3, carrying out chessboard calibration and space positioning; s4, medicine bottle label detection and label character recognition and structured analysis; and S5, system integration and application. According to the invention, through combination of multi-camera and multi-angle acquisition and checkerboard calibration positioning and combination with YOLOv8 label detection and OCR identification, high precision, high efficiency, end-to-end automation and system integration of medicine bottle label identification are realized, the defects of precision, efficiency, environmental adaptability and management integration in the prior art are overcome, and the system is suitable for popularization and application. The method has obvious technical advantages and practical value.
Owner:DONGGUAN KEYAN TECHNOLOGY CO LTD

Multi-modal fusion bank receipt intelligent processing method and system based on vision and NLP

The invention discloses a multi-modal fusion bank receipt intelligent processing system and method based on vision and NLP, and the method comprises the steps: receiving a bank receipt picture or a PDF document, and completing the text detection, direction correction and character recognition through a visual processing engine; a three-level receipt independent segmentation mechanism is applied, and independent receipt records are divided through spatial clustering analysis, semantic analysis, visual verification and cross-page association processing; jointly extracting text features and layout features of each receipt through a double-flow multi-modal fusion model, and fusing the text features and the layout features; field-level data extraction is executed through a field extraction engine, and financial data verification including account number, amount, date and abnormity quadruple verification is carried out; and generating structured JSON output to obtain a bank receipt processing result. According to the invention, the innovative five-layer processing architecture realizes high-precision analysis of the bank receipts through an independently researched and developed receipt independent segmentation engine, a vision-semantic fusion model and a financial data verification system.
Owner:QINGDAO WHALE ABACUS TECHNOLOGY CO LTD

Dynamic double-layer hidden watermark and encryption binding file protection method and system based on deep learning

The invention relates to a dynamic double-layer hidden watermark and encryption binding file protection method and system based on deep learning, and belongs to the technical field of digital content security. The problems of attack resistance, traceability obstruction and key-watermark unhooking in document cross-platform circulation are solved. According to the scheme, the method comprises the following steps of: extracting semantic fingerprints by using a sentence vector model Sentence-BERT; the fuzzy extractor generates a master key and derives a time key chain; the authentication encryption algorithm AEAD encrypts and binds the source and the timestamp load; container layer structure rearrangement and document layer zero-width character double embedding are carried out; the generative adversarial network or diffusion model adversarial training improves the optical character recognition and transcoding resistance; version binding and tracing are achieved through the watermark hash chain. The technical effects cover anti-counterfeiting migration, cross-layer fault-tolerant guarantee recoverability, rearrangement attack resistance, full-period accurate traceability and post-quantum security enhancement.
Owner:SOUTHWEST UNIV

AI semantic enhanced unstructured manufacturing document data structuring method

The invention relates to the technical field of electrical digital data processing, and discloses an AI semantic enhanced unstructured manufacturing document data structuring method, which comprises the following steps: acquiring a character recognition stream of a target document to extract a semantic anchor point and topological coordinates thereof, and vectorizing a to-be-structured field to generate a to-be-processed vector; querying an engineering logic mapping table to determine association intensity, and calculating a logic gravitational field intensity value of the to-be-processed vector relative to the semantic anchor point according to the association intensity; associating the to-be-processed vector to a target semantic anchor point according to the field intensity value, starting a logic polarization program when the to-be-processed vector is identified to be interfered by the homogeneous semantic anchor point in the association stage, extracting bias characteristics to generate a polarization vector, and modulating the native gravitational field intensity; according to the method, the attribution ambiguity of similar semantic entities in a dense distribution scene is solved by introducing the logic gravitation constraint, the structured conflict caused by spatial distribution and engineering logic decoupling is eliminated, and the topology consistency in the data conversion process is ensured.
Owner:FUJIAN YOUHEKE NETWORK TECH CO LTD

PDF drawing data extraction method and system based on intelligent identification

The invention relates to the field of drawing recognition, in particular to a PDF drawing data extraction method and system based on intelligent recognition. Comprising the following steps: reading an internal structure of a PDF engineering drawing to obtain a native text stream, a vector path and a grating image; identifying the native text flow through a shunt preprocessing framework to form structured text data; rendering the vector path and the grating image to obtain a background image; analyzing the structured text data by utilizing the intelligent recognition model through the character recognition and extraction sub-model, obtaining drawing metadata and recording the position, and obtaining a character recognition result; analyzing the background image through a graphic element recognition and classification sub-model, recognizing and classifying component elements, and obtaining a graphic recognition result; and performing fusion according to the visual space corresponding relation to form a drawing analysis result. According to the method, the adaptive capacity of engineering drawings with various sources and different qualities is improved through the shunting preprocessing framework and the intelligent identification model.
Owner:TAIZHOU HUAWEI INFORMATION TECH CO LTD

Multi-dimensional automatic background investigation method and system based on Ai assistance

The invention relates to the technical field of human resources, and discloses a multi-dimensional automatic background investigation method and system based on Ai assistance. The system comprises a multi-source data acquisition module, an intelligent analysis engine module, a real-time interview processing module, a quality verification control module, a dynamic decision support module, a privacy security protection module and a visual report generation module. According to the method, public data of candidates are acquired through a multi-thread crawler technology, and the authenticity of educational background is verified; analyzing the paper certificate by adopting an optical character recognition technology; constructing a space-time relation knowledge graph to detect work experience abnormality; interview content is transcribed in real time, and emotional tendency is analyzed; dynamically adjusting the evaluation dimension weight; and generating a visual risk assessment report. According to the method, the enterprise employment risk is reduced, the comprehensiveness and depth of background investigation are improved, the accuracy and applicability of a background investigation conclusion are improved, a safe and reliable background investigation environment is constructed, and the enterprise human resource allocation effect is optimized.
Owner:CAIXING (GUANGZHOU) TECH SERVICE CO LTD

Identifying Items in Images Using Embeddings Generated from the Images and Ranking Candidates Using a Language Model

An online system applies a visual language model and an optical character recognition model to a received image to generate descriptive information about unknown items in the image. The online system prompts a generative model with the descriptive information about unknown items in the image to separate the descriptive information into different bins each corresponding to a different unknown item in the image. For each unknown item detected in the image, the online system generates a target embedding from its descriptive information and performs a nearest neighbor search on an item catalog including embeddings for various items to find a set of candidate embeddings matching the target embedding. The online system retrieves item attributes of candidate items each corresponding to a candidate embedding of the set and prompts the generative model with this information to rank candidate items for the unknown item in the image.
Owner:MAPLEBEAR INC

PDF (Portable Document Format) document structured extraction system based on multi-modal language model

The invention discloses a PDF (Portable Document Format) document structured extraction system based on a multi-modal language model, belongs to the technical field of document processing and optical character recognition, and aims at solving the technical problem of how to improve the existing OCR (Optical Character Recognition) technology to improve the analysis capability of a complex document structure and improve the recognition precision of handwritten forms and other non-standard fonts. According to the technical scheme, the system adopts a layered decoupling architecture and comprises an input layer, a preprocessing layer, a reasoning layer, an output layer and a monitoring and fault-tolerant module; wherein the output layer is used for multi-source data access and path management to realize a local file system or S3 cloud storage; the preprocessing layer is used for invalid document filtering and visual feature extraction; the reasoning layer is used for multi-modal model interaction and content processing; the output layer is used for outputting a content aggregation result; and the monitoring and fault-tolerant module is used for realizing real-time state monitoring, resource consumption analysis and exception handling.
Owner:JIANGSU HAIRUO INFORMATION TECHNOLOGY CO LTD

Automatic software testing method and system based on multi-mode AI cooperation

The invention relates to the technical field of automatic testing, in particular to an automatic software testing method and system based on multi-mode AI cooperation. According to the scheme, optical character recognition, voice recognition and natural language processing analysis tools are packaged into independent containers, and analysis results are input into a multi-mode semantic association model in real time; calculating the semantic similarity between the image description text and the corresponding voice transcription text; constructing a test demand complexity evaluation model; defining use case quality indexes, and establishing an error sample library to store manually labeled problem use cases and correction labels thereof; the intention demand is finally determined according to the screening result of the three-stage filter; and setting a test case template, optimizing a path according to the fusion quality, improving the case generation efficiency, and synchronously updating the optimized path to a knowledge base associated with the error sample library. According to the scheme, automatic software testing is achieved through multi-modal AI cooperation in combination with containerized deployment, model dynamic optimization and the like, and efficiency and recognition accuracy are improved.
Owner:RUIJIAN TECHNOLOGY (BEIJING) CO LTD

Orthopedic implant code character recognition method based on machine learning

The invention provides an orthopedic implant coded character recognition method based on machine learning, and the method comprises the steps: extracting a deformation parameter of a character region according to a surface curvature distribution diagram, determining a deformation proportionality coefficient and an angle deviation matrix, and fusing a spiral curvature radius to evaluate a path torsion coefficient; performing reverse adjustment on the angle deviation matrix by adopting an angle correction algorithm to obtain a corrected character angle parameter, and verifying the stability of the curvature peak position through a path continuity index; if the character integrity is higher than a preset threshold value, directly extracting a coding sequence to obtain preliminary coding information, and adjusting an offset correction factor according to curvature distribution density to optimize a spiral pitch parameter; and performing geometric reduction on the coding sequence through the character arrangement direction to obtain a final recognition result. According to the method, the recognition accuracy of complex curved surface characters is remarkably improved, the robustness of coding sequence extraction is optimized, and an efficient solution is provided for reliable recognition of implant surface characters.
Owner:SHANGHAI JIUXIANG DIGITAL TECH CO LTD

Information extraction from unstructured documents with hybrid retrieval augmentation using multi-modal language models

A system for extracting a number of data elements from one or more data sources. Image-based documents are indexed using optical character recognition and a text embedding model to convert the document text to vector embeddings. Relevant portions of the document are identified by comparing the vector embedding of the documents to a vector embedding of a prompt or a request to extract information. The relevant text is mapped to a corresponding page of the documents. The page may be provided to a multi-modal language model for information extraction. The multi-modal language model can process contextual information included in the layout, figures, markings, etc. of the document to extract the information. The system populates an ontological data store based on the response from the language model. Extraction accuracy is improved without significant increases in computations performed by the system.
Owner:AMERICAN INTERNATIONAL GROUP INC

Hidden character recognition and confrontation method and related equipment

The embodiment of the invention provides a hidden character recognition and confrontation method and related equipment. The hidden character recognition and confrontation method comprises the following steps: receiving a PDF file, and extracting multi-dimensional visual typesetting information of each text character; the typesetting information comprises color information, size information, position information and container range information of a container where the text characters are located; detecting each text character according to the typesetting information, and determining a target hidden character; and in response to a request for carrying out content processing on the PDF file, filtering the target hidden characters, and generating final text content. According to the technical scheme, the typesetting information such as the color, the size, the position and the container range of each character in the PDF file is extracted, the hidden characters are recognized and automatically filtered during content processing, and therefore it is guaranteed that the extracted or translated result is consistent with the visual sense of a user, and the hidden characters are prevented from being maliciously used for AI confrontation attacks.
Owner:BEIJING QIHOOD TECHNOLOGY CO LTD

UI automatic test method and test system, and storage medium

The invention provides a UI automatic test method and test system and a storage medium. The method comprises the steps that an optical character recognition result of a target text and a to-be-tested UI is obtained, and the optical character recognition result comprises a plurality of elements and element texts, coordinate frames and visual features of all the elements; obtaining the semantic similarity between the target text and each element text; obtaining attribute similarity between the target text and each element based on the visual features of each element; according to the semantic similarity and the attribute similarity, elements, corresponding to the target text, in the optical character recognition result are determined to serve as target elements; taking the position of the coordinate frame of the target element in the UI as a target position; and generating a presentation image which highlights the target position and / or the target element. Therefore, fuzzy matching based on semantics can be realized, multi-modal features including text semantics and visual features are fused, and the robustness and accuracy of UI element positioning are improved in combination with context sensing.
Owner:AIJI MICRO CONSULTING (XIAMEN) CO LTD

Zero sample template inference and document structured recognition method and device

The invention relates to the cross technical field of computer vision and natural language processing, and particularly provides a zero sample template inference and document structured recognition method and device, and the method comprises the following steps: S1, document image collection and preprocessing; s2, performing layout sensing partitioning and position coding; s3, priori or example information is constructed and injected; s4, performing cross-modal fusion and expression construction; s5, performing automatic format analysis and field slot filling; s6, generating a cue word-free extraction instruction; s7, performing field area parallel character recognition; s8, performing semantic verification and result standardization; and S9, outputting the structured field-value data. Compared with the prior art, the method has the advantages that the dependence of a traditional method on template making, cue word writing and large-scale sample training can be avoided, the flexibility, accuracy and online speed of document structured recognition are remarkably improved, and the method has good intelligence and rapid adaptation capacity and is suitable for diversified document recognition scenes.
Owner:INSPUR SOFTWARE CO LTD

Multi-modal document data processing method and system oriented to large language model training

ActiveCN121093293ANeural learning methodsBatch processingCharacter (computing)
The invention discloses a multi-modal document data processing method and system for large language model training, and the method comprises the steps: receiving a plurality of original documents in various formats, extracting the structure information of each original document, and recognizing a text region and an image region of each original document based on the structure information; performing optical character recognition on the text region and the image region by adopting a parallel OCR (Optical Character Recognition) engine based on GPU (Graphics Processing Unit) acceleration and heterogeneous calculation to generate recognition text data of the corresponding original document; performing multi-dimensional quality evaluation and cleaning on the recognition text data of each original document, and outputting normalized text data; and storing the standardized text data into a distributed knowledge base according to a predefined structure, and performing copyright and compliance test on the standardized text data. By adopting a parallel OCR recognition engine based on GPU acceleration and heterogeneous calculation, efficient and high-precision batch processing of multi-modal documents is realized, and the processing speed, the recognition accuracy and the data quality are improved.
Owner:HANGZHOU BINGTE TECH

Document structure extraction and model training method and device, equipment and medium

The invention discloses a document structure extraction and model training method and device, equipment and a medium, and relates to the technical field of artificial intelligence and computer vision. The method comprises the following steps: constructing a special training data set containing data of at least two document understanding tasks (including optical character recognition, layout analysis, text positioning, regional text extraction, image description and chart title generation), and a fine tuning data set for converting a document image into a machine-readable structured text format; constructing a multi-modal large model comprising a shape adaptive cutting module, a visual encoder, a visual token compression module, a modal connector and a language decoder; pre-training the model by using the special training data set to jointly learn various document understanding tasks; and performing fine tuning on the pre-trained model by using the fine tuning data set, and adapting to a document structure extraction task to obtain a document structure extraction model. By means of the technical scheme, efficient and accurate document structure extraction can be achieved.
Owner:CETC CYBERSPACE SECURITY TECH CO LTD

MINI chip visual inspection and accurate placement system and method

The invention relates to the technical field of semiconductor detection and automatic packaging, discloses a vision detection and accurate placement system and method for MINI chips, and aims to solve the problems that in the prior art, due to the lack of multi-source sensing fusion and closed-loop control, the chip positioning precision is insufficient, the grabbing and placement reliability is low, and the whole-process quality monitoring is lack. The method comprises the following steps: acquiring three-dimensional coordinates and attitude data of a chip through a distance measurement detection unit; the vacuum pressure of the suction nozzle is monitored in real time to confirm the grabbing state; the actual posture of the chip and the burning station reference data are fused, a six-dimensional space compensation vector is calculated, a manipulator is driven to conduct fine adjustment, and accurate placement is achieved; and visual inspection is carried out at the discharging braid end, and comprehensive judgment of position, appearance and character recognition is completed. According to the scheme, submicron positioning correction, closed-loop grabbing control and whole-process quality monitoring are achieved, and the chip placement precision, the production yield and the automation efficiency are remarkably improved.
Owner:SHENZHEN KINCOTO ELECTRONICS EQUIP CO LTD