Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

176 results about "Document format" patented technology

A document file format is a text or binary file format for storing documents on a storage media, especially for use by computers. There currently exist a multitude of incompatible document file formats.

Content editable document format conversion method and system based on visual identification

In the field of document processing, format conversion is a common demand, for example, a PDF (Portable Document Format) or a scanning copy is converted into an editable format such as Word, Excel and the like, a traditional method generally adopts an OCR (Optical Character Recognition) technology to directly recognize document content and output a target format, and the method has the defects that the recognition accuracy is insufficient, errors can only be checked and corrected item by item after the target format is converted, and the efficiency is high. In order to solve the problems of lack of visual comparison with an original manuscript, easiness in omission, incapability of adjusting an identification result in real time in a format conversion process and high subsequent editing workload, the invention provides a brand-new document format conversion method, and provides a visual interactive interface by introducing a three-column visual interface and a real-time team correction mechanism. According to the method, the user can check and adjust the recognition result in real time in the conversion process, the accuracy of document conversion is remarkably improved, the flexibility of document conversion is achieved, the user is supported to carry out secondary editing on the content before formal output, visual comparison between the recognized content and an original manuscript is achieved, and the check difficulty is reduced.
Owner:SIGNALING TECHNOLOGY (HANGZHOU) CO LTD

Method and system for automatically generating document based on template and large model

The embodiment of the invention provides a method and system for automatically generating a document based on a template and a large model, and relates to the technical field of document generation, the method comprises the following steps: step S1, a user inputs information through a browser interface and uploads a to-be-reviewed file; step S2, analyzing the uploaded multi-modal file and extracting structured data; s3, outputting a standardized template instance file according to the decision tree dynamic combination template; s4, calling a large model to generate a project background and a file content abstract; s5, the generated content is rendered and combined with the attachment to form a complete conference review file; s6, executing cross-document consistency intelligent verification and executing intelligent error correction; and S7, generating and storing a final document with version traceability information. According to the method, on the basis of the document template and the large model capacity, document formats are unified, content is accurately generated, and the document writing efficiency is remarkably improved.
Owner:TERMINUSBEIJING TECH CO LTD

Heterogeneous document set-oriented cross-modal semantic alignment and logic consistency verification system

The invention relates to document verification, in particular to a heterogeneous document set-oriented cross-modal semantic alignment and logic consistency verification system, which is used for heterogeneous document input, supports multi-format document input and comprises multi-modal elements including texts, pictures, tables and charts. The document analysis module is used for carrying out structured extraction on document contents; extracting multi-modal elements, identifying and classifying various elements in the document, and establishing position and type labels of a foundation; the knowledge graph construction module is used for uniformly modeling heterogeneous elements into a multi-modal knowledge graph; the graph neural network semantic alignment module is used for realizing accurate cross-modal semantic alignment by using a specially designed graph neural network based on the multi-modal knowledge graph; the hybrid consistency verification engine is used for performing logic consistency verification in combination with a symbol logic verification mechanism and a semantic consistency verification mechanism; according to the method, the defect that accurate cross-modal semantic alignment and logic consistency verification are difficult to carry out on professional documents with multi-modal elements can be effectively overcome.
Owner:ANHUI GAOSHAN TECH CO LTD

Document format processing method and system based on large language model, terminal and medium

The invention relates to the field of document processing, and particularly provides a document format processing method, system, terminal and medium based on a large language model.The method comprises the steps that a generation instruction containing a target document type and a source material are received, and original text content is generated through the large language model integrating domain knowledge; obtaining a matched structured format template, and analyzing the style rule into a format instruction set; identifying logic elements and hierarchical relationships thereof in the original text based on a natural language processing technology; performing association mapping on the format instruction and the logic element, and performing automatic style rendering by calling a document object model interface to generate an intermediate document with a standard format; and finally, outputting after quality verification. According to the method, an automatic process of content generation and intelligent formatting is constructed, and the document writing efficiency and normalization are improved.
Owner:浪潮智慧科技有限公司 +2

Hydrofracture question-answering system and method based on cross-language retrieval enhanced generation

The invention discloses a system and a method for enhancing generation of hydrofracture questions and answers based on cross-language retrieval. The knowledge question-answering system aiming at the hydraulic fracturing field is developed by utilizing a cross-language retrieval enhancement generation technology, so that a user can obtain an optimal answer and a corresponding source from a multi-language knowledge base by matching only by using Chinese retrieval; the complex process that a user needs to find useful information from numerous and jumbled papers and reports is changed. According to the system, a dynamic text vector library framework is adopted, multi-language and complex document formats are supported, the analysis function of tables and formulas is integrated, documents in the hydraulic fracturing field are stored in the same knowledge base, the knowledge base is continuously updated along with document expansion in the professional field, a user can inquire related problems in the field in the system, and the user experience is improved. By integrating and updating a large amount of document information, required professional knowledge in the field can be efficiently obtained.
Owner:ZHEJIANG UNIV

Big model-based receipt identification method, system and equipment

The invention belongs to the field of bill recognition, and provides a bill recognition method, system and equipment based on a large model, and the method comprises the steps: obtaining bill image data containing text information, and carrying out the preprocessing of the obtained bill image data; evaluating image quality, language distribution and layout complexity of the preprocessed image, and determining a processing path of document image data by integrating evaluation results; according to the determined processing path, calling a corresponding optical character recognition model to extract text information in the document image data; performing deep semantic analysis on the extracted text data by using a pre-trained deep learning large model, and extracting key feature information; and performing rule compliance analysis on the text information or the extracted key feature information, and generating an identification analysis report according to an analysis result. According to the invention, the processing speed and accuracy are improved, the adaptability to various document formats is enhanced, and the limitation in the prior art is effectively solved.
Owner:INSPUR GENERSOFT CO LTD

Format conversion method of PDF (Portable Document Format) document, storage medium and computer equipment

The invention discloses a PDF document format conversion method, a storage medium and computer equipment. The format conversion method of the PDF document comprises the following steps: analyzing an original PDF document based on a preset document analysis algorithm to extract a plurality of original content blocks of the original PDF document; based on the structure type of each original content block, extracting document content and layout features included in the corresponding original content block; and performing semantic conversion on the document content and the layout features based on a pre-trained large model to obtain a plurality of target content blocks in a preset target format, and arranging each target content block in the target format document based on the position information of the original content block corresponding to each target content block. By means of the method, the content of the PDF document can be efficiently and accurately analyzed and converted into the corresponding target format document, the target format document can keep the document content and layout characteristics of the original PDF document, and the integrity and readability of the converted document are effectively improved.
Owner:XIAN YOUFANG DIGITAL TECH CO LTD

Neural-symbolic hybrid system for direct binary document synthesis with integrated constraint satisfaction and hardware acceleration

A neural-symbolic hybrid system for generating binary document formats directly from natural language input comprises a binary-aware hierarchical tokenizer operating across four levels (binary bytes, structural elements, semantic content, and concepts), a constraint satisfaction engine with 64 parallel processing cores for enforcing structural integrity and mathematical consistency, format-specific processors for Excel, PowerPoint, PDF, and CAD documents, and a formal verification system generating mathematical proofs of correctness. The system includes custom AI Document Generation Processor (AIDGP) silicon spanning 600 mm2 with specialized cores providing 500 TOPS processing power. Performance characteristics include 99.7% structural accuracy, 100% format compliance, 15.3 second average generation time for complex documents, and distributed capacity of 1,000,000 documents per hour. The system eliminates intermediate conversion steps while maintaining semantic preservation through hardware-accelerated constraint satisfaction and formal verification engines ensuring structural integrity, format compliance, and security through AES-256 encryption and automated regulatory compliance across 25+ international standards.
Owner:GUPTA GAURAV +1

Document generation method and system integrating document analysis and cognitive reasoning

The invention provides an official document generation method and system fusing document analysis and cognitive reasoning, and the method comprises the steps: carrying out the multi-modal feature extraction of an input original document through employing an analysis engine combining a convolutional neural network and a graph attention network, and obtaining the multi-modal data in the original document; based on the multi-modal data, generating a classification decision result of the multi-modal data by using a hierarchical classifier enhanced by a knowledge graph; based on the multi-modal data and the classification decision result, generating a task decomposition result by combining a cognitive inference engine; and based on the classification decision result and the task decomposition result of the multi-modal data, generating an official document by combining compliance constraint with an official document format template, so that multi-modal accurate analysis of a complex document can be realized, and dynamic semantic understanding and classification are realized by utilizing a hierarchical classifier enhanced by a knowledge graph. The classification accuracy in the fuzzy semantic scene can be improved, so that the recognition error rate can be reduced, and the accuracy can be improved.
Owner:STATE GRID INFORMATION & TELECOMM BRANCH

Communication protocol automatic analysis and test data generation method and system

The invention discloses a communication protocol automatic analysis and test data generation method and system, and the method comprises the steps: analyzing an unstructured protocol document, recognizing a fixed field and a variable field in the document, and outputting structured data; automatically generating test data based on the structured data, dynamically calculating a checksum according to a specified check algorithm type in the structured data, and filling the checksum into the original message to form a complete test message; and generating a standardized test data set according to the complete test message. According to the method, manual intervention is not needed from document analysis to test data generation, the analysis efficiency is improved, and most protocol scenes are covered through the boundary value, the abnormal value and the orthogonal combination strategy. The method supports various document formats, data types and verification algorithms, adapts to protocol standards of different industries, facilitates problem positioning and regression testing through standardized output documents, and reduces the maintenance cost.
Owner:SICHUAN HONGMEI INTELLIGENT TECH CO LTD

Artificial intelligence-based system and method for automating job matching

The present artificial intelligence-based system (100) automates job matching by evaluating each candidate's resume against job requirements. It includes a user interface (102) for administrators to upload job descriptions and candidate resumes, a database directory module (104) for data storage, a document extension module (106) for file verification, and a processing module (108) with a Large Language Model (110), at least one vector database (118), and a Scoring module (112). The processing module interprets resume contents through machine-readable instructions, performs similarity searches, and scores candidates in real-time based on their suitability for each job requirement. The method (200) of the present invention involves receiving and managing job descriptions and resumes, verifying document formats, processing resumes and job descriptions, and displaying ranked results. Within the processing step (208), the method further encompasses processing and interpreting resume contents, scoring candidates for real-time suitability, and comparing qualifications and skills against job descriptions.
Owner:MYQUICKHR SDN BHD

Resume data analysis method and device, electronic equipment and storage medium

The invention relates to a resume data analysis method and device, electronic equipment and a storage medium, and the method comprises the steps: collecting a resume file of a candidate, the resume file comprising a resume file in a document format and / or a resume file in a non-document format; extracting a resume text and auxiliary feature data from the resume file; according to the resume text, identifying an industry category involved by the content of the resume file; a corresponding resume analysis model is determined according to the industry category, and the resume analysis model is a deep learning model obtained through training based on multi-mode resume data of the industry category and an industry knowledge graph; and based on the resume text and the auxiliary feature data, utilizing the resume analysis model to obtain structured resume data, wherein the structured resume data comprises a plurality of resume fields and field contents. According to the method, the accuracy and comprehensiveness of resume analysis are improved.
Owner:QIAN JIN NETWORK INFORMATION TECH SHANGHAI LTD

Word document format conversion method based on Java

The invention particularly relates to a Word document format conversion method based on Java. The Word document format conversion method based on Java comprises the following steps: analyzing the content of an original. Doc file, and extracting text paragraphs, tables, pictures, style information and document metadata; the extracted content is divided into different categories, and the XML node type corresponding to each category of elements in the target. Docx document is established; the method comprises the following steps of: constructing a pattern mapping rule base, constructing a new. Docx document structure by using an XWPF Document object model according to an Office Open XML (Extensible Markup Language) specification, sequentially inserting paragraphs, tables and pictures, and applying corresponding pattern configuration; and outputting and generating a. Docx file, detecting and processing abnormal conditions, and recording a conversion log at the same time. The Word document format conversion method based on Java is efficient and accurate, has good compatibility, expandability and cross-platform capability, is suitable for enterprise-level document management systems, cloud services and batch document processing scenes, and remarkably improves document compatibility and processing efficiency of office automation systems.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Four-layer progressive mind mapping generation system

The invention discloses a four-layer progressive mind map generation system which comprises an input module which can receive input of various document formats, supports simultaneous receiving of multiple documents and adopts a multi-source heterogeneous data adapter to unify the formats; the hierarchical preprocessing module is connected with the input module, receives the converted DOCX document data, and performs hierarchical feature enhancement and hierarchical rule configuration on the data; the bimodal processing module fuses a hierarchical driving structured extraction mode and a large language model generation mode, the hierarchical driving structured extraction mode constructs a hierarchical incidence matrix and the like, and the large language model generation mode provides task instructions and the like; and the structured output module is used for providing multi-format conversion, mind map downloading and interaction functional components. The system has the advantages that all the modules work cooperatively, multi-format document processing is achieved, multiple modes are fused to generate the mind map, rich output functions are provided, and a user can conveniently and efficiently generate and use the mind map.
Owner:ZHEJIANG UNIV BINJIANG RES INST +1

RAG question and answer optimization method and system based on fusion retrieval

The invention discloses an RAG question and answer optimization method and system based on fusion retrieval, and relates to the technical field of large model knowledge question and answer. The method comprises the steps of obtaining user query, executing keyword retrieval and semantic vector retrieval in parallel, and obtaining a keyword retrieval result list and a semantic vector retrieval result list respectively; and calculating the quality score of each text unit in each retrieval result list based on multi-dimensional quality evaluation, and filtering out the text units with the quality scores lower than a preset quality threshold value. And sorting the text units in all the filtered retrieval result lists by adopting a quality weighted reciprocal sorting fusion algorithm. And based on the length of each text unit and the document format to which the text unit belongs, performing context content enhancement and combination on each text unit after sorting, and generating an enhanced context block list. And constructing structured prompt information based on the enhanced context block list and user query, inputting structured prompt words into the large language model, and generating and outputting a final answer.
Owner:JIANGSU RED NET TECH CO LTD

Railway official document keyword extraction method and device and electronic equipment

The invention relates to a railway official document keyword extraction method and device and electronic equipment, and the method comprises the steps: based on a pre-constructed railway official document format rule base, extracting a key field of a fixed position from an input text through regular expression matching and position locking; a Jieba word segmentation device is used for loading a railway-specific term library for word segmentation, and a multi-word combination entity boundary is dynamically corrected through a dependency relationship rule; executing a TF-IDF algorithm on the text after word segmentation to generate an initial word weight, adjusting the weight according to the position area of the word in the official document and a preset coefficient, and performing position weighting; and combining the words of which the weights are greater than a set threshold value with the extracted key fields, and outputting a final keyword set after verification of a term library. According to the method, missing detection caused by low frequency of a traditional algorithm is avoided, splitting errors of a general word segmentation device are eliminated, the term recognition error rate is reduced, the core word sorting priority is improved, and the semantic weight of keywords is strengthened; the new term storage time is shortened, and the updating cost problem is solved.
Owner:INST OF COMPUTING TECH CHINA ACAD OF RAILWAY SCI +2

Intelligent printing control method based on multifunctional integration

The invention relates to the field of intelligent printing control, in particular to an intelligent printing control method based on multifunctional integration, and the method comprises the steps: collecting a user information terminal, an instruction text and a printing environment parameter text through a data collection end, and matching a control center according to the user information terminal; the printing document format optimization parameters and the reference parameters are automatically matched through the instruction text, meanwhile, the printing environment parameter text is substituted into the environment health index model to obtain the environment health index, the temperature, the pressure and the ink amount of the printing head are automatically matched, and then the printing head is started for printing. The temperature parameter, the pressure parameter, the ink quantity parameter, the pressure parameter difference coefficient and the ink quantity parameter difference coefficient of the printing head are detected in real time, and the printing risk index is obtained according to the temperature parameter difference coefficient, the pressure parameter difference coefficient, the ink quantity parameter difference coefficient and the environment health index; and finally, carrying out risk judgment on the printing risk index and outputting a risk signal to a user side.
Owner:ZHUHAI MANCHIN ELECTRONIC TECHNOLOGY CO LTD

Verification environment generation method and apparatus, and electronic device

Embodiments of the invention provide a verification environment generation method and apparatus, and an electronic device. The method comprises the steps that according to a document format of a design document corresponding to a chip to be tested, the design document is recognized, and a recognition document is obtained; the recognition document is analyzed, multiple pieces of register information corresponding to the multiple registers are obtained, and the register information comprises register identifiers, boundary information and field line indexes; based on the design document, standardizing the register information corresponding to each register to generate a standardized data structure, the standardized data structure comprising a plurality of sub-data structures corresponding to the plurality of registers; based on the standardized data structure, generating a verification code meeting a verification environment specification; and integrating the verification code into the target verification environment to complete a verification task of the to-be-tested chip. The method is used for achieving the effect of improving the chip verification efficiency.
Owner:XIAMEN UNISOC TECH CO LTD

Intelligent document writing system based on large language model

The invention discloses an intelligent document writing system based on a large language model, belongs to the technical field of natural language processing and intelligent office crossing, and aims to solve the technical problems of how to realize intellectualization, standardization and high efficiency of a whole document writing process and improve document writing efficiency and accuracy. According to the technical scheme, the system comprises an input layer, a core layer and an output layer; the input layer comprises a multi-modal input interface and a demand analysis engine; the multi-mode input interface is used for voice input and text input; the demand analysis engine is used for text recognition and automatic element providing, and only subject fields can be filled and document unit fields can be received; the core layer is used for adopting a mixed training strategy of comparative learning and curriculum learning, improving professional field content generation capability, integrating format verification and an OCR feedback mechanism, ensuring document format compliance and supporting a quick response mechanism of standard updating; and the output layer is used for realizing export and version management of documents in standard DOCX, signature PDF and HTML5 formats.
Owner:SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD

Big model-based official document processing method and device and official document generating and checking all-in-one machine

The invention discloses an official document processing method and device based on a large model and an official document generating and checking all-in-one machine. The official document processing method based on the large model can quickly respond to a user request, automatically match and call a pre-trained official document generation sub-model corresponding to a literary and sports category by automatically identifying an official document title and an official document keyword, and automatically process the official document in combination with knowledge base data of a corresponding application field. According to the method, the large language model technology is utilized to quickly and intelligently generate the official document file with high quality, official document processing which is specially refined in various official document styles and is combined with the user industry can be realized, and the official document processing efficiency is greatly improved. According to the document generating and checking all-in-one machine, software and hardware are perfectly combined, the problems of document format accurate generation, large model illusion, safety risk, computing power support and the like can be solved at the same time, safe and efficient operation of the all-in-one machine in a user intranet environment is achieved, and the difficulty and cost of building and maintaining a complex system by a unit are reduced.
Owner:NAT IND INFORMATION SECURITY DEV RES CENT

Document format conversion method and device for large model training

The invention relates to the technical field of computers, and discloses a document format conversion method and device for large model training, and the method comprises the steps: carrying out the classification of a PDF document based on a plurality of classification indexes, and obtaining a document type; when the document type is an image type, performing image conversion and preprocessing to obtain a preprocessed page image; analyzing the preprocessed page image to obtain multi-modal content, and processing the multi-modal content to obtain corresponding processing results; performing content reconstruction and optimization on the PDF document to obtain a first intermediate document; and performing content rearrangement on the first intermediate document to obtain a Markdown document. According to the method, the PDF documents are accurately classified by integrating multiple dimensions, different types of documents are subjected to differentiated format conversion, the format conversion efficiency and accuracy are improved, the generated Markdown document conforms to the original PDF document, semantic coherence and format standardization are achieved, the large model input requirement is met, and therefore the model training effect can be improved.
Owner:TIANTIANZHIYUAN (CHENGDU) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Document analysis method and system based on multiple modes

The invention belongs to the technical field of artificial intelligence and multi-modal data processing, and provides a multi-modal-based document analysis method and system. The method comprises the following steps: S1, inputting a to-be-analyzed document, and converting all pages in the to-be-analyzed document into pictures with the same size as an original page page by page; s2, performing layout analysis on the converted picture page by page by utilizing a visual model, detecting layout elements existing in the picture, and obtaining a position coordinate of each layout element; s3, inputting the current picture and the detected layout elements into a large language model for analysis, and adopting different analysis methods for different layout elements to obtain analyzed contents; and S4, displaying the analyzed content according to needs, or performing layout recovery by using the obtained position coordinates of the layout elements and the analyzed content. According to the method, the input document supports all document formats which are almost visible at present, and the analysis speed can be increased.
Owner:POWERCHINA BEIJING ENG CORP

Method and system for translating PDF (Portable Document Format) text containing complex features

The invention belongs to the technical field of text translation, and provides a method and a system for translating a PDF (Portable Document Format) text containing complex features. The method comprises the steps that a PDF analysis engine is initialized, a PDF file is read, and basic information of the PDF file is extracted; judging whether the PDF file is a complex file or not, and recording complex features; preprocessing an image in the complex document, and calling a corresponding text detection model according to the complex features to perform text region identification; and extracting layout information of the text in the text area, translating the text in the text area through a translation model, performing typesetting according to the layout information after a final translation result is obtained, and outputting a target translation file in a user-defined manner. According to the method, the complex content in the PDF document can be intelligently identified, and the original document format is completely reserved in the translation process; meanwhile, multi-language translation is supported, and the accuracy of PDF document translation is guaranteed.
Owner:AFIRSTSOFT CO LTD

Medical examination report processing method and device, electronic equipment and storage medium

The invention discloses a medical examination report processing method and device, electronic equipment and a storage medium, and the method comprises the steps: processing a medical examination report, and determining report text data corresponding to the medical examination report; processing the report text data based on a big language model, and obtaining structured data which is output by the big language model and corresponds to the medical examination report; and converting the structured data into standardized data based on a standard document format, and mapping the standardized data to a medical coding system. Based on the technical scheme, a prompt project and context reasoning mechanism is utilized to guide a large language model to automatically identify the medical entity and output the structured representation, and the structured data is mapped to a medical coding system, so that seamless joint with various medical information is realized; and the real-time adaptive analysis and structuring requirements of the multi-source medical examination report can be met.
Owner:LIANREN HEALTHCARE BIG DATA TECH CO LTD

Multi-language intelligent analysis system for medical documents

The invention provides a medical document multi-language intelligent analysis system, relates to the field of language translation, and improves the accuracy and efficiency of professional term translation. The method comprises the following steps of: firstly, performing language recognition on an original text document by a recognition module through text unitization and context vector generation, and matching a corresponding corpus; then, a translation module carries out lexical element alignment on the source lexical elements through term bank injection and an AI model, the translation process is automatically optimized, and accurate translation of the terminologies is ensured; and finally, the reconstruction module accurately replaces corresponding contents in the original text document with translation output through the mapping file, so as to ensure that the document format and typesetting are consistent. Through the automatic and optimized translation process, the quality and efficiency of professional term translation are remarkably improved, manual intervention is reduced, and the translation requirement of a high professional standard is met.
Owner:LUNAN PHARMA GROUP CORPORATION +2

Bank account bill automatic processing method and system based on OCR (Optical Character Recognition)

The invention relates to an OCR (Optical Character Recognition)-based bank account bill automatic processing method. The method comprises the following steps: inputting bank account bill data in various formats; judging whether the document format is a picture format, and converting the document format into the picture format; constructing a watermark and seal identification model; removing the watermark and the seal on the picture; character recognition is carried out, and character content in the picture is extracted to form a text; reconstructing the text according to a table plate of the original document or picture; and adjusting and confirming the reconstructed table structure to obtain formatted table data, and exporting the formatted table data into Excel and JSON formats. Manual operation can be reduced, and the data processing efficiency is improved; the OCR and NLP technologies are combined, so that high-accuracy information extraction is realized; watermarks or seals can be automatically detected and removed, and the quality of text recognition is improved; format output of Excel, JSON and the like is supported, and subsequent analysis and storage are facilitated; the method has a verification function, and ensures the accuracy and integrity of final data.
Owner:FUYANG NORMAL UNIVERSITY

Method for optimized storage and query of table data of large model knowledge base

The invention provides a method for optimally storing and querying table data of a large model knowledge base, which belongs to the field of data processing and comprises the following steps of: performing multi-step processing including data splicing, natural language processing optimization, document format conversion and vectorization storage on the table data; the logicality and coherence of the table data during storage and retrieval are enhanced, so that the accuracy and logicality of the retrieval generated content are remarkably improved, and the retrieval result is more readable and easy to understand.
Owner:INSPUR SOFTWARE TECH CO LTD

Intelligent table data escape and knowledge base construction method and system based on large language model

The invention discloses an intelligent table data escape and knowledge base construction method based on a large language model. The method comprises the steps that multiple document formats are imported; identifying the existence of the table, analyzing the row-column structure of the table, positioning the position of each cell, and extracting data in the table; the big language model understands the data significance of each cell according to the context and the data content in the table; performing escape on the original data; merging and storing the original data, logically partitioning the data, and vectorizing the data after escaping; table data in a knowledge base is obtained through a large language model knowledge base system, and a query request of a user is quickly responded through a vector index of the knowledge base. According to the method, the data in the table is intelligently converted by adopting a large language model semantic escape technology, and the data is embedded into the knowledge base through data merging operation, so that a user can conveniently query and utilize the data, and the problem of insufficient table data processing precision in the prior art is solved.
Owner:XINJIANG NORTH-WEST STAR INFORMATION TECH CO LTD

LLM-based medical translation model training method and medical document translation method

The invention discloses an LLM-based medical translation model training method and a medical document translation method. The model training method comprises the steps of constructing a Chinese-English medical corpus, pre-training a medical translation model, finely tuning the medical translation model, quantifying the medical translation model and deploying the medical translation model. The medical document translation method comprises the steps of document format conversion, document preprocessing, medical document translation, document post-processing and document format restoration. According to the medical translation model training method and the medical document translation method based on the LLM, the accuracy of medical translation is remarkably improved, professional terms, complex sentence patterns and context logic relations in medical texts are efficiently understood and translated, and translated texts which are more natural, smoother and higher in accuracy can be generated.
Owner:JINYE TIANCHENG BEIJING TECH CO LTD

Any document format annotation display method and device based on Canvas

The invention provides a Canvas-based annotation display method and a Canvas-based annotation display device for any document format, and aims to solve the problems of high file format dependency, poor annotation information interoperability, poor display consistency and the like in the prior art. The method specifically comprises the following steps: S1, acquiring a file type, and acquiring an original width and an original height of a file according to the file type; s2, constructing a first coordinate system, and representing the labeling information in the first coordinate system in a rectangular coordinate mode; s3, drawing a rectangular labeling frame and a labeling text through Canvas to form a transparent mask; s4, according to the actual display size of the current display window, calculating a scaling ratio, adjusting the size of Canvas according to the scaling ratio, and covering the file display area; and S5, displaying the file with the annotation information in the file display area. The use experience of the user is improved.
Owner:上海通办信息服务有限公司