Offline data digitization method and system based on large model and storage medium
Through the offline data digitization method based on large-models, text recognition and content extraction are automatically performed, which solves the problem of inefficient digitalization of offline data in the existing technology, and realizes a more efficient digitalization process.
Patent Information
- Application Number
- CN202510179348.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-02-18
AI Technical Summary
In the prior art, offline data digitization efficiency is low, resulting in cumbersome manual operations and reducing digital efficiency.
The offline data digitization method based on large models is adopted, and the offline data is automatically digitized through steps such as text recognition, content extraction and data verification to improve efficiency.
There is no need to manually copy and paste text, which improves the digital efficiency of offline data, and automates the content extraction process through content extraction from large models, further improving efficiency.
Smart Images

Figure CN120011619A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a large model-based offline data digitization method, system and storage medium. Background Art
[0002] In today's information age, despite the rapid development of electronic information technology, a large amount of offline data still exists widely. Offline data covers various types of paper documents, such as books, archives, contracts, manuscripts, etc. Offline data carries a wealth of knowledge, historical information and important business data. However, offline data exposes many problems in the actual use and management process. First of all, in terms of storage, paper data requires a lot of physical space. With the continuous increase in the amount of data, the cost of storage space has risen sharply. At the same time, paper data is easily affected by natural environmental factors such as moisture, fire, and insect bites, which lead to damage to the data and loss of information, seriously affecting its long-term preservation stability. Therefore, digitizing offline data has become an inevitable trend.
[0003] In the existing offline data digitization process, data digitization is generally done manually, which leads to cumbersome manual operations and reduces the efficiency of offline data digitization. Summary of the invention
[0004] The purpose of the embodiments of the present invention is to provide a method, system and storage medium for offline data digitization based on a large model to solve the problem of low efficiency of offline data digitization in the prior art.
[0005] The embodiment of the present invention is implemented as follows: a method for digitizing offline data based on a large model, the method comprising:
[0006] Acquire offline data to be digitized, and perform text recognition on the offline data to be digitized to obtain offline documents;
[0007] Obtaining content extraction requirements, and combining the content extraction requirements with the offline documents to obtain a material digitization prompt;
[0008] Inputting the digital data prompt into the pre-trained large model to extract content, obtain data extraction data, and perform data verification on the data extraction data;
[0009] An online data template is obtained, and the extracted data after data verification is filled into the online data template.
[0010] Preferably, performing text recognition on the offline data to be digitized to obtain offline documents includes:
[0011] Performing grayscale processing on the offline data to be digitized to obtain a data grayscale image, and performing normalization processing on the data grayscale image to obtain a normalized image;
[0012] According to different convolution scales, the normalized image is convolved to obtain convolution features, and the convolution features are fused according to the convolution scale to obtain a feature pyramid;
[0013] Performing text prediction according to the feature pyramid to obtain a text prediction result, and determining a target text box according to the text prediction result;
[0014] The offline document is generated according to the target text box.
[0015] Preferably, determining a target text box according to the text prediction result includes:
[0016] Obtaining a text existence probability in the text prediction result, and comparing the text existence probability with a probability threshold;
[0017] If the text existence probability is greater than the probability threshold, determining the text box corresponding to the text existence probability as a candidate text box, and calculating the overlap between different candidate text boxes;
[0018] A text box score of the candidate text box is determined according to the overlap degree and the text existence probability, and the target text box is determined according to the text box score.
[0019] Preferably, before the digital prompt of the data is input into the pre-trained large model for content extraction, the method further includes:
[0020] Acquire a digital prompt sample, and input the digital prompt sample into the large model for content extraction to obtain sample extraction data, wherein the digital prompt sample includes a sample document and a sample prompt word;
[0021] The model loss is determined according to the sample extraction data, and the parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
[0022] Preferably, after performing text recognition on the offline data to be digitized to obtain the offline document, the method further includes:
[0023] Performing entity recognition on the document paragraphs in the offline document to obtain entity recognition results, and determining the paragraph object of the document paragraph according to the entity recognition;
[0024] Classifying the document paragraphs according to the paragraph objects to obtain a paragraph set, and performing semantic recognition on the document paragraphs to obtain a semantic recognition result;
[0025] According to the semantic recognition result, the paragraph associations between different document paragraphs in the same paragraph set are determined, and the document paragraphs in the same paragraph set are sorted according to the paragraph associations.
[0026] Preferably, the content extraction requirement is combined with the offline document to obtain a material digitization prompt, including:
[0027] According to the sorting result of the document paragraphs, the document paragraphs in the same paragraph set are combined, and the first identifier is inserted into different document paragraphs to obtain a paragraph string;
[0028] Inserting the corresponding paragraph object at the beginning of the paragraph string, and combining different paragraph strings;
[0029] Inserting a second identifier into the combined different paragraph strings to obtain a string combination, and inserting the content extraction requirement at the beginning of the string combination;
[0030] A third identifier is inserted between the content extraction requirement and the beginning of the character string combination to obtain the material digitization prompt.
[0031] Preferably, filling the extracted data after data verification into the online data template includes:
[0032] Performing type recognition on the data vocabulary in the data extraction data after data verification to obtain vocabulary type, and performing type matching between the vocabulary type and the fill-in column in the online data template;
[0033] The data vocabulary is filled into the corresponding fill-in column according to the type matching result.
[0034] Another object of an embodiment of the present invention is to provide an offline data digitization system based on a large model, the system comprising:
[0035] A text recognition module is used to obtain offline data to be digitized, and perform text recognition on the offline data to be digitized to obtain offline documents;
[0036] A prompt generation module, used to obtain content extraction requirements and combine the content extraction requirements with the offline documents to obtain a material digitization prompt;
[0037] A content extraction module is used to input the digital data prompt into the pre-trained large model to extract content, obtain data extraction data, and perform data verification on the data extraction data;
[0038] The data filling module is used to obtain the online data template and fill the data extracted after data verification into the online data template.
[0039] Preferably, the text recognition module is also used for:
[0040] Performing grayscale processing on the offline data to be digitized to obtain a data grayscale image, and performing normalization processing on the data grayscale image to obtain a normalized image;
[0041] According to different convolution scales, the normalized image is convolved to obtain convolution features, and the convolution features are fused according to the convolution scale to obtain a feature pyramid;
[0042] Performing text prediction according to the feature pyramid to obtain a text prediction result, and determining a target text box according to the text prediction result;
[0043] The offline document is generated according to the target text box.
[0044] The embodiments of the present invention perform text recognition on offline data to be digitized, thereby eliminating the need for manual copying and pasting of texts, thereby improving the efficiency of offline data digitization. The digitization prompts of the data are input into a pre-trained large model for content extraction, and the content of offline documents is automatically extracted based on the powerful reasoning ability of the large model, eliminating the need for manual content extraction, thereby further improving the efficiency of offline data digitization. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 is a flow chart of a large model-based offline data digitization method provided in the first embodiment of the present invention;
[0046] Figure 2 is a structural schematic diagram of an offline data digitization system based on a large model provided by a second embodiment of the present invention;
[0047] Figure 3 It is a schematic diagram of specific implementation steps of the offline data digitization system based on a large model provided by the second embodiment of the present invention;
[0048] Figure 4 It is a schematic diagram of the structure of a terminal device provided in the third embodiment of the present invention. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0050] In order to illustrate the technical solution of the present invention, a specific embodiment is provided below for illustration.
[0051] Embodiment 1
[0052] See also Figure 1 , is a flow chart of an offline data digitization method based on a large model provided in a first embodiment of the present invention. The offline data digitization method based on a large model can be applied to any device or system. The offline data digitization method based on a large model includes the following steps:
[0053] Step S10, obtaining offline data to be digitized, and performing text recognition on the offline data to be digitized to obtain offline documents;
[0054] Among them, by performing text recognition on the digitized offline materials, there is no need to copy and paste the text manually. In this step, algorithms such as CRNN algorithm (Convolutional Recurrent Neural Networks), Attention OCR algorithm and DenseNet-OCR algorithm can be used to perform text recognition on the digitized offline materials.
[0055] Specifically, the CRNN algorithm combines convolutional neural networks (CNN) and recurrent neural networks (RNN). CNN is responsible for extracting feature sequences from the input image, RNN is used to predict the label distribution of these feature sequences, and finally the label distribution is converted into the final recognition result through the transcription layer (using the CTC algorithm). Features: It can effectively handle the recognition of indefinite-length text, without the need for explicit character cutting, and can convert text recognition into a sequence learning problem, and performs well in natural scene text recognition.
[0056] The Attention OCR algorithm is similar to CRNN. It uses the CNN+RNN network structure in the feature learning stage, but uses the attention mechanism (Attention) in the output layer. The attention mechanism allows the model to automatically focus on different parts of the input text when processing text, and dynamically allocates attention weights according to the current recognition task, so as to better capture the long sequence dependencies and contextual information in the text. Features: It can recognize more accurately when processing long texts, irregular texts, and texts with complex layouts, and has better performance than CRNN in some specific scenarios.
[0057] The DenseNet-OCR algorithm uses a densely connected convolutional neural network (DenseNet) and a sequence transducer (Transducer). The densely connected structure of DenseNet can effectively utilize feature information, improve the efficiency of feature transmission, and enable the model to better learn complex features in the image; the sequence transducer is used to model and predict text sequences. Features: It can improve text recognition accuracy and has strong robustness when facing texts of multiple fonts, sizes, styles, and complex background images.
[0058] Optionally, performing text recognition on the offline data to be digitized to obtain an offline document includes:
[0059] Performing grayscale processing on the offline data to be digitized to obtain a data grayscale image, and performing normalization processing on the data grayscale image to obtain a normalized image;
[0060] According to different convolution scales, the normalized image is convolved respectively to obtain convolution features, and the convolution features are feature fused according to the convolution scale to obtain a feature pyramid; wherein the convolution scale can be set according to demand, and the number of convolution scales is greater than or equal to two, and the normalized image is convolved respectively by different convolution scales to obtain convolution features of different feature scales, so that the feature pyramid after feature fusion can effectively have context feature information;
[0061] Text prediction is performed according to the feature pyramid to obtain a text prediction result, a target text box is determined according to the text prediction result, and the offline document is generated according to the target text box; wherein, the feature pyramid features are decoded by an encoder in the large model, and text prediction is performed based on the feature decoding result to obtain a text prediction result.
[0062] Further, determining a target text box according to the text prediction result includes:
[0063] Obtaining the text existence probability in the text prediction result, and comparing the text existence probability with a probability threshold; wherein the text existence probability is used to indicate the probability of text existing at the corresponding position, and the probability threshold can be set according to demand;
[0064] If the text existence probability is greater than the probability threshold, the text box corresponding to the text existence probability is determined as a candidate text box, and the overlap between different candidate text boxes is calculated; wherein, if the text existence probability is greater than the probability threshold, it is determined that the corresponding position exists text, the text box corresponding to the text existence probability is determined as a candidate text box, and the overlap area between different candidate text boxes is calculated, and the overlap degree is determined based on the overlap area;
[0065] The text box score of the candidate text box is determined according to the overlap and the text existence probability, and the target text box is determined according to the text box score; wherein, the text box score is obtained by performing a weighted operation on the overlap and the text existence probability, and the weighted coefficients of the overlap and the text existence probability in the weighted operation process can be set according to demand. In this step, the candidate text boxes are sorted according to the text box scores, and boxes with higher probabilities and lower overlap with other boxes are retained, and boxes with higher overlap are suppressed to obtain target text boxes, and only one optimal target text box is retained for each text area.
[0066] Furthermore, after performing text recognition on the offline data to be digitized to obtain offline documents, the method further includes:
[0067] Performing entity recognition on the document paragraphs in the offline document to obtain entity recognition results, and determining the paragraph object of the document paragraph according to the entity recognition; wherein, by performing entity recognition on the document paragraphs in the offline document, based on the entity type of each entity in the entity recognition results, the paragraph object of the document paragraph can be effectively determined;
[0068] Classifying the document paragraphs according to the paragraph objects to obtain a paragraph set, and performing semantic recognition on the document paragraphs to obtain a semantic recognition result; wherein the document paragraphs corresponding to the same paragraph object are divided into the same set to obtain a paragraph set;
[0069] According to the semantic recognition result, the paragraph association between different document paragraphs in the same paragraph set is determined, and the document paragraphs in the same paragraph set are sorted according to the paragraph association; wherein the paragraph semantics between different document paragraphs are combined to obtain a semantic combination, and the semantic combination is matched with an association query table to obtain a paragraph association.
[0070] Step S20, obtaining content extraction requirements, and combining the content extraction requirements with the offline documents to obtain a material digitization prompt;
[0071] The content extraction requirement is used to indicate the data parameters that the user needs to extract. Optionally, the content extraction requirement is combined with the offline document to obtain a data digitization prompt, including:
[0072] According to the sorting result of the document paragraphs, the document paragraphs in the same paragraph set are combined, and the first identifiers are inserted into different document paragraphs to obtain a paragraph string; wherein, by inserting the first identifiers into different document paragraphs, the large model is effectively facilitated to identify different document paragraphs;
[0073] Inserting the corresponding paragraph object at the beginning of the paragraph string, and combining different paragraph strings;
[0074] Inserting a second identifier into the different paragraph strings after combination to obtain a string combination, and inserting the content extraction requirement at the beginning of the string combination; wherein, by inserting the second identifier into the different paragraph strings after combination, the large model effectively facilitates the identification of different paragraph strings;
[0075] A third identifier is inserted between the content extraction requirement and the beginning of the string combination to obtain the data digitization prompt; wherein, by inserting the third identifier between the content extraction requirement and the beginning of the string combination, the large model is effectively facilitated to identify the content extraction requirement and the string, and the first identifier, the second identifier and the third identifier can all be set according to the requirements.
[0076] Step S30, inputting the digital data prompt into the pre-trained large model to extract content, obtain data extraction data, and perform data verification on the data extraction data;
[0077] Among them, content extraction is performed by inputting digital prompts of the data into a pre-trained large model, so as to automatically extract content from offline documents based on the powerful reasoning ability of the large model.
[0078] Optionally, before the digital prompt of the data is input into the pre-trained large model for content extraction, the method further includes:
[0079] Acquire a digital prompt sample, and input the digital prompt sample into the large model for content extraction to obtain sample extraction data, wherein the digital prompt sample includes a sample document and a sample prompt word;
[0080] Determine the model loss according to the sample extraction data, and update the parameters of the large model according to the model loss until the large model converges to obtain the pre-trained large model;
[0081] Among them, the digital prompt samples are preprocessed, including removing special characters, punctuation marks, spaces and other irrelevant content, unifying the format and encoding of the text, and at the same time performing word segmentation on the digital prompt samples, dividing the continuous text sequence into single words or phrases, and using commonly used Chinese word segmentation tools such as Jieba word segmentation, to prepare for subsequent vocabulary extraction.
[0082] In this step, large models can be set up according to needs, such as BERT model, GPT model, ERNIE model, etc., which need to be investigated and evaluated according to the characteristics of the document and the specific needs of vocabulary extraction. For example, the BERT model performs well in natural language understanding and feature extraction, and is suitable for vocabulary extraction tasks that require high semantic understanding of text; the GPT series of models have advantages in generating text and can also be used for vocabulary extraction, especially in scenarios where the meaning of words needs to be inferred based on the context.
[0083] Design appropriate prompt words according to the goal of vocabulary extraction and the type of document. The prompt words should clearly express the type, scope, and conditions of the vocabulary to be extracted. For example, if you want to extract nouns in a document, you can design the prompt words as "Please extract all nouns from the following text"; if you want to extract vocabulary related to a specific topic, such as vocabulary related to "artificial intelligence", you can design the prompt words as "Please find all vocabulary related to artificial intelligence in the text."
[0084] The pre-processed digitized prompt samples and the designed prompt words are input into the big model. The big model will analyze and process the text according to the requirements of the prompt words and output the extracted vocabulary results. The big model will identify the words that meet the conditions based on its understanding and knowledge of the language.
[0085] De-duplication and screening: De-duplication of the extracted vocabulary results is carried out to remove repeated words and improve the accuracy and practicality of the vocabulary. At the same time, the extracted vocabulary can be screened according to some rules or conditions, such as removing stop words and low-frequency words, etc., to retain more valuable words.
[0086] Step S40, obtaining an online data template, and filling the data extraction data after data verification into the online data template;
[0087] Among them, the online data template can be set according to needs. By extracting data after data verification and filling it into the online data template, the effect of digitizing offline data to be digitized online can be achieved.
[0088] Optionally, filling the extracted data after data verification into the online data template includes:
[0089] Perform type identification on the data vocabulary in the data extraction data after data verification to obtain the vocabulary type, perform type matching on the vocabulary type and the fill-in column in the online data template, and fill the data vocabulary into the corresponding fill-in column according to the type matching result; wherein, the data vocabulary is matched with the type query table to obtain the vocabulary type, and by performing type matching on the vocabulary type and the fill-in column in the online data template, the filling correspondence between the data vocabulary and the fill-in column can be effectively determined, and the data vocabulary is filled into the corresponding fill-in column based on the filling correspondence.
[0090] In this embodiment, by performing text recognition on the digitized offline data, there is no need to manually copy and paste the text, thereby improving the efficiency of offline data digitization. By inputting the data digitization prompts into a pre-trained large model for content extraction, the content of offline documents is automatically extracted based on the powerful reasoning ability of the large model, and there is no need to manually extract the content, thereby further improving the efficiency of offline data digitization.
[0091] Embodiment 2
[0092] See also Figure 2 , is a schematic diagram of the structure of an offline data digitization system 100 based on a large model provided in a second embodiment of the present invention, including:
[0093] The text recognition module 10 is used to obtain offline data to be digitized, and perform text recognition on the offline data to be digitized to obtain offline documents.
[0094] Optionally, the text recognition module 10 is further used to: perform grayscale processing on the offline data to be digitized to obtain a data grayscale image, and perform normalization processing on the data grayscale image to obtain a normalized image;
[0095] According to different convolution scales, the normalized image is convolved to obtain convolution features, and the convolution features are fused according to the convolution scale to obtain a feature pyramid;
[0096] Performing text prediction according to the feature pyramid to obtain a text prediction result, and determining a target text box according to the text prediction result;
[0097] The offline document is generated according to the target text box.
[0098] Furthermore, the text recognition module 10 is also used to: obtain the text existence probability in the text prediction result, and compare the text existence probability with a probability threshold;
[0099] If the text existence probability is greater than the probability threshold, determining the text box corresponding to the text existence probability as a candidate text box, and calculating the overlap between different candidate text boxes;
[0100] A text box score of the candidate text box is determined according to the overlap degree and the text existence probability, and the target text box is determined according to the text box score.
[0101] Furthermore, the text recognition module 10 is also used to: perform entity recognition on the document paragraph in the offline document to obtain an entity recognition result, and determine the paragraph object of the document paragraph according to the entity recognition;
[0102] Classifying the document paragraphs according to the paragraph objects to obtain a paragraph set, and performing semantic recognition on the document paragraphs to obtain a semantic recognition result;
[0103] According to the semantic recognition result, the paragraph associations between different document paragraphs in the same paragraph set are determined, and the document paragraphs in the same paragraph set are sorted according to the paragraph associations.
[0104] Preferably, the text recognition module 10 is further used to: combine the document paragraphs in the same paragraph set according to the sorting result of the document paragraphs, and insert the first identifier into different document paragraphs to obtain a paragraph string;
[0105] Inserting the corresponding paragraph object at the beginning of the paragraph string, and combining different paragraph strings;
[0106] Inserting a second identifier into the combined different paragraph strings to obtain a string combination, and inserting the content extraction requirement at the beginning of the string combination;
[0107] A third identifier is inserted between the content extraction requirement and the beginning of the character string combination to obtain the material digitization prompt.
[0108] The prompt generation module 11 is used to obtain content extraction requirements and combine the content extraction requirements with the offline documents to obtain a material digitization prompt.
[0109] The content extraction module 12 is used to input the digital data prompts into the pre-trained large model for content extraction, obtain data extraction data, and perform data verification on the data extraction data.
[0110] Optionally, the content extraction module 12 is further used to: obtain a digital prompt sample, and input the digital prompt sample into the large model for content extraction to obtain sample extraction data, wherein the digital prompt sample includes a sample document and a sample prompt word;
[0111] The model loss is determined according to the sample extraction data, and the parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
[0112] The data filling module 13 is used to obtain an online data template and fill the data extraction data after data verification into the online data template.
[0113] Optionally, the data filling module 13 is further used to: perform type recognition on the data vocabulary in the data extraction data after data verification to obtain the vocabulary type, and perform type matching between the vocabulary type and the filling column in the online data template;
[0114] The data vocabulary is filled into the corresponding fill-in column according to the type matching result.
[0115] See also Figure 3 The offline data digitization system 100 based on the big model integrates an OCR engine and a big model engine, which can perform OCR recognition on the uploaded data with one click, convert the offline data into online text, and use the text data recognized by OCR and related extraction requirements to form a data digitization prompt, which is sent to the big model service for automatic extraction, and finally outputs the data extraction data in JSON format. Finally, the data extraction data in JSON format is automatically filled into the relevant online form for saving.
[0116] The specific implementation steps of the offline data digitization system 100 based on the big model include:
[0117] 1. Integrate OCR engine and large model engine;
[0118] 2. Upload offline documents to be digitized into the system;
[0119] 3. Click the data recognition function to perform OCR recognition on offline data and convert documents in different formats into plain text data;
[0120] 4. Combine the plain text data generated in the previous step with the content extraction requirements to form a prompt;
[0121] 5. Send the assembled prompt to the big model service to automatically extract the content and output the data in JSON format;
[0122] 6. The JSON format data generated in the previous step is read by the code and automatically filled into the online form;
[0123] 7. After the business personnel checks and confirms the form data, they save it.
[0124] In this embodiment, an OCR engine is integrated to convert offline data into plain text, reducing the workload of manual copying / pasting. The engine uses a large model engine and the powerful reasoning ability of the large model to automatically extract plain text to replace the manual extraction of content, forming a closed-loop operation from offline data to online data storage, greatly improving the efficiency of digital work.
[0125] In this embodiment, by performing text recognition on the digitized offline data, there is no need to manually copy and paste the text, thereby improving the efficiency of offline data digitization. By inputting the data digitization prompts into a pre-trained large model for content extraction, the content of offline documents is automatically extracted based on the powerful reasoning ability of the large model, and there is no need to manually extract the content, thereby further improving the efficiency of offline data digitization.
[0126] Embodiment 3
[0127] Figure 4 2 is a block diagram of a terminal device 2 provided in the third embodiment of the present application. Figure 4 As shown, the terminal device 2 of this embodiment includes: a processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the processor 20, such as a program of a method for digitizing offline data based on a large model. When the processor 20 executes the computer program 22, the steps in each embodiment of the method for digitizing offline data based on a large model are implemented.
[0128] Exemplarily, the computer program 22 may be divided into one or more modules, which are stored in the memory 21 and executed by the processor 20 to complete the present application. The one or more modules may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program 22 in the terminal device 2. The terminal device may include, but is not limited to, a processor 20 and a memory 21.
[0129] The processor 20 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0130] The memory 21 may be an internal storage unit of the terminal device 2, such as a hard disk or memory of the terminal device 2. The memory 21 may also be an external storage device of the terminal device 2, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the terminal device 2. Further, the memory 21 may also include both an internal storage unit and an external storage device of the terminal device 2. The memory 21 is used to store the computer program and other programs and data required by the terminal device. The memory 21 may also be used to temporarily store data that has been output or is to be output.
[0131] In addition, each functional module in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional unit.
[0132] If the integrated module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium can be non-volatile or volatile. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include: any entity or device capable of carrying computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in computer-readable storage media can be appropriately increased or decreased according to the requirements of legislation and patent practices in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practices, computer-readable storage media do not include electrical carrier signals and telecommunication signals.
[0133] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for digitizing offline data based on a large model, characterized in that: The method comprises: Acquire offline data to be digitized, and perform text recognition on the offline data to be digitized to obtain offline documents; Obtaining content extraction requirements, and combining the content extraction requirements with the offline documents to obtain a material digitization prompt; Inputting the digital data prompt into the pre-trained large model to extract content, obtain data extraction data, and perform data verification on the data extraction data; An online data template is obtained, and the extracted data after data verification is filled into the online data template.
2. The offline data digitization method based on a large model as claimed in claim 1, characterized in that: Performing text recognition on the offline data to be digitized to obtain offline documents, including: Performing grayscale processing on the offline data to be digitized to obtain a data grayscale image, and performing normalization processing on the data grayscale image to obtain a normalized image; According to different convolution scales, the normalized image is convolved to obtain convolution features, and the convolution features are fused according to the convolution scale to obtain a feature pyramid; Performing text prediction according to the feature pyramid to obtain a text prediction result, and determining a target text box according to the text prediction result; The offline document is generated according to the target text box.
3. The offline data digitization method based on a large model as claimed in claim 2, characterized in that: Determining a target text box according to the text prediction result includes: Obtaining a text existence probability in the text prediction result, and comparing the text existence probability with a probability threshold; If the text existence probability is greater than the probability threshold, determining the text box corresponding to the text existence probability as a candidate text box, and calculating the overlap between different candidate text boxes; A text box score of the candidate text box is determined according to the overlap degree and the text existence probability, and the target text box is determined according to the text box score.
4. The offline data digitization method based on a large model as claimed in claim 1, characterized in that: Before the digital prompt of the data is input into the pre-trained large model for content extraction, it also includes: Acquire a digital prompt sample, and input the digital prompt sample into the large model for content extraction to obtain sample extraction data, wherein the digital prompt sample includes a sample document and a sample prompt word; The model loss is determined according to the sample extraction data, and the parameters of the large model are updated according to the model loss until the large model converges to obtain the pre-trained large model.
5. The offline data digitization method based on a large model as claimed in claim 1, characterized in that: After performing text recognition on the offline data to be digitized to obtain offline documents, the method further includes: Performing entity recognition on the document paragraphs in the offline document to obtain entity recognition results, and determining the paragraph object of the document paragraph according to the entity recognition; Classifying the document paragraphs according to the paragraph objects to obtain a paragraph set, and performing semantic recognition on the document paragraphs to obtain a semantic recognition result; According to the semantic recognition result, the paragraph associations between different document paragraphs in the same paragraph set are determined, and the document paragraphs in the same paragraph set are sorted according to the paragraph associations.
6. The offline data digitization method based on a large model as claimed in claim 5, characterized in that: Combining the content extraction requirements with the offline documents to obtain a data digitization prompt includes: According to the sorting result of the document paragraphs, the document paragraphs in the same paragraph set are combined, and the first identifier is inserted into different document paragraphs to obtain a paragraph string; Inserting the corresponding paragraph object at the beginning of the paragraph string, and combining different paragraph strings; Inserting a second identifier into the combined different paragraph strings to obtain a string combination, and inserting the content extraction requirement at the beginning of the string combination; A third identifier is inserted between the content extraction requirement and the beginning of the character string combination to obtain the material digitization prompt.
7. The offline data digitization method based on a large model as claimed in claim 1, characterized in that: Filling the extracted data after data verification into the online data template includes: Performing type recognition on the data vocabulary in the data extraction data after data verification to obtain vocabulary type, and performing type matching between the vocabulary type and the fill-in column in the online data template; The data vocabulary is filled into the corresponding fill-in column according to the type matching result.
8. An offline data digitization system based on a large model, characterized in that: The system comprises: A text recognition module is used to obtain offline data to be digitized, and perform text recognition on the offline data to be digitized to obtain offline documents; A prompt generation module, used to obtain content extraction requirements and combine the content extraction requirements with the offline documents to obtain a material digitization prompt; A content extraction module is used to input the digital data prompt into the pre-trained large model to extract content, obtain data extraction data, and perform data verification on the data extraction data; The data filling module is used to obtain the online data template and fill the data extracted after data verification into the online data template.
9. The offline data digitization system based on a large model as claimed in claim 8, characterized in that: The text recognition module is also used for: Performing grayscale processing on the offline data to be digitized to obtain a data grayscale image, and performing normalization processing on the data grayscale image to obtain a normalized image; According to different convolution scales, the normalized image is convolved to obtain convolution features, and the convolution features are fused according to the convolution scale to obtain a feature pyramid; Performing text prediction according to the feature pyramid to obtain a text prediction result, and determining a target text box according to the text prediction result; The offline document is generated according to the target text box.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Information extraction method and model training method and device for information extraction
CN116524523A
Method and device for digitizing paper content and electronic equipment
CN117332001A
Key information extraction method and system based on large language model
CN118210879A
Systems and Methods for Generating Task-Specific Agent Modules Based on User Requests
US20240404712A1
Cited By
Method and device for identifying and automatically merging petition documents
CN120877315A