File digital processing method, device, equipment, medium and product
By using the arrangement language model to adjust the text position in the digital processing of archives, the efficiency and accuracy problems of traditional OCR technology when dealing with complex text layout are solved, and high-quality digital archive generation is achieved.
Patent Information
- Application Number
- CN202510511974.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-23
AI Technical Summary
Traditional OCR technology is prone to text direction deviation, disordered rows and misalignment when dealing with paper archives with complex text layouts, which affects the efficiency and accuracy of digital processing of archives.
By determining multiple text blocks from pre-acquisitioned archive images and inputting them into the pre-trained arrangement language model, text-position-related information is obtained, text position is adjusted, and high-quality digital archives are finally generated.
It improves the efficiency and accuracy of digital processing of archives, reduces the need for manual verification, and ensures a high degree of consistency between digital archives and original archives.
Smart Images

Figure CN120071376A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image recognition technology, and in particular, to a method, device, equipment, medium and product for digital processing of archives. Background Art
[0002] Traditional archives are usually paper archives, which occupy a large space, are easy to be damaged and lost, have low retrieval efficiency, are difficult to update and maintain, have poor security, and do not meet the environmental protection requirements. In contrast, electronic archives save space, are easy to retrieve and update, have high security, are convenient for remote access, and are more in line with the concept of green environmental protection. Therefore, the electronic archive system is the development trend of personnel file management, which is conducive to improving the modernization level of file management and promoting the sharing and efficient utilization of information resources. In order to achieve this transformation, in efficiently converting paper archives into electronic archives, Optical Character Recognition (OCR) technology plays a key role.
[0003] However, it is found in practice that the text layout in paper archives is complex and changeable, and there are often situations such as non-horizontal arrangement, multi-directional mixed arrangement or special formats, resulting in problems such as text direction deviation, row and column disorder, and content dislocation during the recognition by traditional OCR technology, which requires manual verification and seriously affects the efficiency and accuracy of digital processing of archives. Summary of the Invention
[0004] The purpose of the present application is to provide a method, device, equipment, medium and product for digital processing of archives, which can improve the efficiency and accuracy of digital processing of archives.
[0005] To achieve the above purpose, the present application provides the following solutions: In a first aspect, the present application provides a method for digital processing of archives, including: Determining a plurality of text blocks from pre-collected archive images; Inputting the plurality of text blocks into a pre-trained arrangement language model to obtain the text-position related information of each text block output by the arrangement language model; wherein, a text-position related information includes the probability that each character included in a text block is in different positions in the text block; Based on the text-position related information, adjusting the character positions in each text block to obtain the text content of each text block; Generating a digital archive corresponding to the archive image using each text content.
[0006] Optionally, the method for digital processing of archives further includes: Performing face target detection on the archive image to obtain the face area in the archive image; Perform a portrait matting operation on the face region to obtain a target face image; Fuse the target face image with a pre-determined background image to obtain a fused face image; Determine the position of the face image in the digital archive; Fuse the fused face image to the position of the face image in the digital archive to obtain a target digital archive containing the fused face image.
[0007] Optionally, input a target text block among multiple text blocks into a pre-trained permutation language model to obtain the word-position related information of the target text block output by the permutation language model, specifically including: Determine each word contained in a target text block among multiple text blocks and the initial position of each word in the target text block; According to each word and the initial position of each word in the target text block, determine the token encoding of each word and the combined hidden state information of each word; wherein, the combined hidden state information includes the forward hidden state and the backward hidden state of the word corresponding to the combined hidden state information; Based on the token encoding of each word and the combined hidden state information of each word, calculate the word-position related information of the target text block.
[0008] Optionally, the determining each word contained in a target text block among multiple text blocks and the initial position of each word in the target text block specifically includes: Perform a zoom-in operation on an initial text block among multiple text blocks to obtain a first text block; Grayscale the first text block to obtain a second text block; Denoise the second text block to obtain a target text block; Perform character recognition on the target text block to obtain each word contained in the target text block and the initial position of each word in the target text block.
[0009] Optionally, the calculating the word-position related information of the target text block based on the token encoding of each word and the combined hidden state information of each word specifically includes: Transpose the combined hidden state information of each word to obtain the transposed combined hidden state information of each word; Multiply the token encoding of each word and the transposed combined hidden state to obtain an initial matrix; Based on a preset bias, calculate the initial matrix to obtain a target matrix; Perform a normalization operation on the target matrix to obtain the text-position related information of the target text block.
[0010] Optionally, the generating a digital archive corresponding to the archive image using each text content specifically includes: Perform type recognition on each text content to obtain the text type of each text content; wherein, the text type includes a label type and a content type; Determine the coordinate information of each text content in the archive image; According to the coordinate information of each text content, associate a text content of a content type with a text content of a label type; wherein, the distance between the text content of the content type and the text content of the label type is the closest; Obtain a pre-set digital archive template; wherein, the digital archive template contains multiple target labels; Repeat the filling operation until each blank area corresponding to each target label in the digital archive template is filled with text content, to obtain a digital archive corresponding to the archive image: Wherein, the filling content specifically includes: Determine a current target label from multiple target labels; Determine the target text content identical to the current target label from the text content of the label type; Fill the text content of the content type corresponding to the target text content into the blank area corresponding to the current target label.
[0011] In a second aspect, the present application provides a digital processing device for archives, including: A determination unit, configured to determine multiple text blocks from a pre-collected archive image; An input unit, configured to input the multiple text blocks into a pre-trained permutation language model, to obtain the text-position related information of each text block output by the permutation language model; wherein, a text-position related information includes the probability that each word included in a text block is in different positions in the text block; An adjustment unit, configured to adjust the word positions in each text block based on the text-position related information, to obtain the text content of each text block; A generation unit, configured to generate a digital archive corresponding to the archive image using each text content.
[0012] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the steps of the digital processing method of the file described in any one of the above.
[0013] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the digital processing method of the file described in any one of the above are implemented.
[0014] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the digital processing method of the file described in any one of the above are implemented.
[0015] In a sixth aspect, the present application provides a chip, the chip includes a processor and a communication interface, the communication interface is coupled to the processor, the processor is used to run a program or an instruction, and when the processor executes the program or the instruction, the steps of the digital processing method of the file described in any one of the above are implemented.
[0016] According to the specific embodiments provided by the present application, the following technical effects are disclosed in the present application: The present application provides a digital processing method, device, equipment, medium and product for files, which can accurately determine a plurality of text blocks from the pre-collected file images, laying a clear basic unit for subsequent processing; and inputting these text blocks into a pre-trained permutation language model, the word-position related information output by the model is extremely crucial. The word-position related information represents the probability that each word in each text block is in a different position, providing an accurate basis for word position adjustment; based on this, word position adjustment can greatly correct text errors caused by problems such as complex and irregular original arrangements of words in file images, making the text content of each generated text block more in line with its true semantics. Finally, using these high-quality text contents to generate digital files, compared with the traditional method, effectively reduces manual intervention and error troubleshooting time, avoids repeated work caused by text recognition errors, and thus comprehensively improves the efficiency of file digital processing. At the same time, accurate word position adjustment ensures that the content of the digital file is highly consistent with the original file, greatly improving the accuracy of file digital processing. Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a flowchart of a digital processing method for a file in an embodiment of the present application; Figure 2 It is a schematic structural diagram of an intelligent processing model for a file provided in an embodiment of the present application; Figure 3 It is a schematic diagram of a file image provided in an embodiment of the present application; Figure 4 It is a schematic diagram of a digital file template provided in an embodiment of the present application; Figure 5 It is a schematic diagram of a digital file provided in an embodiment of the present application; Figure 6 It is a schematic diagram of the functional modules of a digital processing device for a file provided in another embodiment of the present application.
[0019] Figure 7 It is a schematic structural diagram of a computer device provided in an embodiment of the present application. Detailed implementation manners
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0021] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the drawings and specific implementation manners.
[0022] In an exemplary embodiment, as Figure 1 shown, a digital processing method for a file is provided. This method is executed by a computer device, and specifically can be executed independently by a computer device such as a terminal or a server, or can be jointly executed by a terminal and a server. In the embodiments of the present application, it includes the following steps 101 to 104. Among them: Step 101, determine a plurality of text blocks from the pre-collected file images.
[0023] In the embodiments of the present application, text blocks can be determined by segmenting according to the grid included in the archive image; the distance between words can also be obtained, and then multiple text blocks can be determined according to the distance between words, and the distance between two adjacent words in each text block is less than a preset distance threshold.
[0024] In the embodiments of the present application, an archive image can be collected by a collection device such as a camera, a video camera, a scanner, etc.
[0025] Step 102: Input the multiple text blocks into a pre-trained permutation language model to obtain the word-position related information of each text block output by the permutation language model.
[0026] In the embodiments of the present application, a word-position related information includes the probability that each word included in a text block is in different positions in the text block.
[0027] In the embodiments of the present application, the permutation language model may include BERT and BiLSTM. The permutation language model task is a natural language processing task, aiming to rearrange the scrambled text into the correct order. In this task, a combined model of BERT and BiLSTM is used to predict the position of each word in the correct order, and then the words are rearranged according to these positions to obtain the correct order.
[0028] First, in order to train this model, a dataset needs to be prepared.
[0029] The construction of the dataset can be completed through the following steps: 1. Collect a large number of ordered texts related to the archive.
[0030] 2. Perform jieba word segmentation on the ordered texts, and randomly select the words to be scrambled. At this time, these words are called whole words.
[0031] 3. On the basis of the whole words, further select consecutive n-grams; an n-gram is a sequence composed of n consecutive tokens. Among them, a token represents a character.
[0032] 4. Randomly exchange the tokens inside the selected whole words and n-grams to scramble the original order of the ordered texts.
[0033] 5. While scrambling the ordered texts, mark the position of each character in the original order. In this task, the maximum text length is limited to , so the predicted values that the model needs to output are these 50 positions.
[0034] The present invention uses the bert - base - chinese version of the BERT model, which can effectively generate text encodings. In the BERT model, Chinese text tokenization is carried out with individual Chinese characters as the basic unit. Each Chinese character will be converted into a token and assigned word embeddings, position embeddings, and segment embeddings. These three embeddings are added together and then fed into a network composed of 12 layers of Transformer encoders. In the output of the last layer of the Transformer encoder, the present invention obtains the token encodings of all input tokens. These token encodings are rich in the semantic information of each token in the current context and are an embodiment of the model's comprehensive understanding of the input text.
[0035] The parameters of the BiLSTM can be updated through gradient descent during the training process to adapt to specific task requirements.
[0036] BiLSTM is a recurrent neural network that can capture long - distance dependencies in text. It is composed of two LSTMs combined. Its calculation process is shown by the following formula:
[0037] Among them, represents the input at time step t, that is, the vector corresponding to the token position output by BERT; represents the processing of the LSTM model from the beginning to the end of the sequence to capture the previous information; is the forward hidden state of the forward LSTM at time step t; represents the processing of the LSTM model from the end to the beginning of the sequence to capture the subsequent information; is the backward hidden state of the backward LSTM at time step t; represents the concatenation operation; is the output of BiLSTM at time step t (word - position related information), which is the combination of the forward and backward output results. In this way, the representation of each time step can contain both previous and subsequent information The specific calculation formula of LSTM is shown as follows:
[0038] Among them, represents the result of the input gate, and are the first weight parameter and the second weight parameter, is the first bias parameter, represents the activation function sigmoid, is the current time step is the input of is the hidden state of the previous time step; represents the result of the forget gate, and are the third weight parameter and the fourth weight parameter, is the second bias parameter; represents the result of the output gate, and are the fifth weight parameter and the sixth weight parameter, is the third bias parameter; represents the candidate memory cell, and are the seventh weight parameter and the eighth weight parameter, is the fourth bias parameter; represents the time step when the memory cell, is the memory cell of the previous time step, represents the Hadamard product, and this step of calculation determines how much of the memory cell features of the previous time step are retained through the forget gate and how much of the candidate memory cell is used through the input gate features; represents the forward / backward hidden state at the current time step t, and tanh is the hyperbolic tangent activation function. This step of calculation uses the output gate to control the transmission of memory information.
[0039] It should be emphasized here that the output dimension of BERT is 768 dimensions, and the BiLSTM used in the present invention has a hidden layer of 384 dimensions, so that the result after bidirectional splicing can also be 768 dimensions, which is convenient for subsequent calculations.
[0040] To calculate the position of the original token, the present invention performs matrix multiplication on the output of BiLSTM and the output of the BERT token embedding module, and through the combination of all hidden states of BiLSTM and the encoding combination of all tokens of the BERT token embedding module transpose ( ) for matrix multiplication, and the correlation between each token and each position can be obtained.
[0041] The dimension of the calculation result is k×k, and the bias b is added to this calculation result. Subsequently, the softmax operation is performed row by row, and the present invention obtains the probability that each token belongs to each position. The softmax function is used to convert the output of the model into a probability distribution.
[0042] Finally, based on these probabilities, the position of each token in the correct order can be predicted, and the words can be rearranged to obtain the correct order. The formula for probability calculation is:
[0043] The loss function adopts the cross-entropy loss L, as shown in the following formula:
[0044] Among them, y represents the true label, which is a one-hot encoding of length M; p is the predicted value, also of length M, which is the probability prediction value for each class; M represents the number of classes. Here, M = k, and 50 positions mean 50 classes.
[0045] According to the loss function, the BiLSTM can be trained to obtain a BiLSTM model with more accurate output prediction probabilities.
[0046] As an alternative implementation manner, the way of inputting a target text block in multiple text blocks into a pre-trained permutation language model in step 102 to obtain the word-position related information of the target text block output by the permutation language model may specifically include: Determine each word included in a target text block in multiple text blocks and the initial position of each word in the target text block; According to each word and the initial position of each word in the target text block, determine the token encoding of each word and the combined hidden state information of each word; wherein, the combined hidden state information includes the forward hidden state and the backward hidden state of the word corresponding to the combined hidden state information; Based on the token encoding of each word and the combined hidden state information of each word, calculate the word-position related information of the target text block.
[0047] Among them, in implementing this implementation manner, by accurately obtaining the initial position information of the text, a basic coordinate system is constructed for the model to analyze the text spatial layout; innovatively combining the semantic representation of token encoding with the combined information of bidirectional hidden states, not only preserves the semantic features of the text itself, but also through the interaction modeling of the forward hidden state and the backward hidden state, enables the model to dynamically capture the logical order and spatial association between texts; finally, the text position relationship is deduced through probability calculation, enabling the system to have the ability to understand the arrangement rules of complex layouts. This deep context awareness mechanism effectively breaks through the recognition limitations of traditional OCR for artistic fonts, handwritten texts, and multi-directional mixed-layout texts, and significantly enhances the adaptability to special-layout archives.
[0048] Optionally, the method for determining each character included in a target text block among multiple text blocks and the initial position of each character in the target text block may specifically include: Performing a zoom-in operation on an initial text block among multiple text blocks to obtain a first text block; Performing grayscale processing on the first text block to obtain a second text block; Performing denoising on the second text block to obtain a target text block; Performing character recognition on the target text block to obtain each character included in the target text block and the initial position of each character in the target text block.
[0049] Among them, in implementing this implementation manner, first, an adaptive zoom-in strategy is adopted to enhance the detailed features of tiny characters or blurred areas, providing a higher-definition text base for subsequent processing; then, color interference is eliminated through grayscale processing, compressing the three-dimensional color space into a two-dimensional grayscale gradient, reducing the computational complexity while retaining the key information of the character contours; then, an intelligent denoising algorithm is used to accurately filter out interference factors such as image noise and scratches, constructing a text recognition environment with a high signal-to-noise ratio; finally, through refined character positioning technology, the initial coordinates of each character are accurately obtained, providing an accurate spatial layout basis for the subsequent arrangement model. This full-process automated optimization mechanism effectively solves common problems such as blurring, noise, and color interference in low-quality archival images, and significantly enhances the reliability of character recognition in complex scenarios.
[0050] Optionally, the method for calculating the text-position related information of the target text block based on the token encoding of each character and the combined information of the hidden states of each character may specifically include: Transposing the combined information of the hidden states of each character to obtain the transposed combined information of the hidden states of each character; Multiplying the token encoding of each character and the transposed combined information of the hidden states to obtain an initial matrix; Calculating the initial matrix based on a preset bias to obtain a target matrix; Perform a normalization operation on the target matrix to obtain the text-position related information of the target text block.
[0051] Among them, in implementing this implementation manner, first, the hidden state transposition strategy is adopted to enable the forward and backward semantic information to form complementary interaction in matrix operations, enhancing the model's bidirectional perception ability of the spatial relationship of characters; secondly, through the matrix fusion operation of token encoding and hidden state, while retaining the deep semantic features of the text, context clues required for position distribution prediction are dynamically injected; then a learnable bias parameter is introduced to calibrate the initial matrix, enabling the model to adaptively adjust the position prediction weights in different layout scenarios; finally, the original prediction value is converted into a probability distribution through a normalization operation to ensure that the text position inference result conforms to the mathematical statistical law. This multi-level joint optimization mechanism effectively solves the recognition problem of traditional OCR for special layouts such as multi-directional mixed arrangement and curved text, and significantly enhances the full-process automation level of archival digitization processing.
[0052] Step 103: Based on the text-position related information, adjust the text positions in each text block to obtain the text content of each text block.
[0053] In the embodiment of the present application, each character can be placed at the position with the highest probability, so as to realize the adjustment of the text positions in each text block.
[0054] Step 104: Generate a digital archive corresponding to the archival image using each text content.
[0055] The embodiment of the present application can be applied to Figure 2 the shown archival intelligent processing model. Among them, the archival intelligent processing model includes a template customization and management module and an OCR-based document structured warehousing module. Specifically: A. The template customization and management module includes template information filling, region segmentation setting, data configuration setting, and document template management. This template customization and management module uses the website interface as an interaction platform, integrates the core functions of Excel, and is implemented through three main steps: template information filling, region segmentation setting, and data configuration setting. The design of this module aims to provide a standardized framework for the digitization processing of cadre archives, ensuring the consistency and accuracy of subsequent content filling and display, and laying a solid foundation for the digitization processing of cadre archives.
[0056] To enhance the display effect of the information recognized and extracted by the model subsequently, users can use the template customization and management module to design templates personalized. In this way, users can flexibly customize the display format and content of information according to their specific needs, thereby improving the readability and practicality of the information and bringing a better experience to users. This module mainly consists of four parts: template information filling, area segmentation setting, data configuration setting, and document template management.
[0057] The template information filling interface provides users with convenient functions to input and define relevant information of the template. On this interface, users are required to fill in three main parts: template name, template type, and template description.
[0058] First of all, the template name is a required item, which provides a unique identifier for the template created by the user. This name should be concise and clear, able to reflect the content or purpose of the template, and facilitate users to search for and use it in the future.
[0059] Secondly, the template type is also a required item, which helps users classify and identify the purpose of the template. For example, the template type can be full text recognition and full table recognition, depending on how the user wants to use this template. Selecting the correct template type can ensure that the template is used correctly in the right occasion.
[0060] Finally, the template description is optional. The template description provides a more detailed overview of the template, including its design purpose, applicable scenarios, special functions, etc. This not only helps users themselves remember the detailed information of the template, but also helps other users who may use this template understand its background and purpose.
[0061] After completing the entry of the basic template information, users will enter the area segmentation setting session, which allows users to define the structure and layout of the template. This page integrates the functions of Excel, enabling users to flexibly plan each part of the template in a familiar interface, just like operating in Excel. This design method greatly simplifies the template creation process, especially for users who are already familiar with Excel operations, and they can design templates more intuitively and efficiently. This customization ability enables each user to design a practical and beautiful template according to their own needs and preferences.
[0062] After the template format customization is completed, users should perform data configuration settings to ensure that each area of the template can correctly recognize and process different types of data. This step involves classifying each area in the template, designating it as a title area, atomic label area, combined label area, content area, or image area, and making detailed configurations for the specific attributes of each area.
[0063] For the title area, the user needs to set a title name and the corresponding title alias so that the system can accurately identify and classify the main parts in the document.
[0064] In the atomic tag area, the user needs to define the name of the tag, the tag recognition rule, and the text layout direction. These settings help the system accurately extract and classify basic data elements.
[0065] The configuration of the combined tag area is more complex. The user needs to specify the tag name, the tag recognition rule, the level of the tag, the parent tag, and the text direction. These detailed settings help build the hierarchical structure of the data and ensure the correct parsing and extraction of complex information.
[0066] The settings for the content area require the user to specify the mapping fields, the associated tags, and decide whether to extract the content of this area from the document.
[0067] The configuration of the image area is similar to that of the content area. The user needs to specify the mapping fields, the associated tags, and select whether to extract the image content. This is particularly important when the document contains visual elements.
[0068] It should be noted that meaningful text content needs to be filled in the title area, the atomic tag area, and the combined tag area. These text contents will be stored in the locally built tag library, and the subsequent OCR recognition results will be matched with the contents in the tag library.
[0069] Through the carefully designed tag system, a solid foundation is laid for the structured warehousing part of the present invention, and it can greatly improve the search efficiency. Under such a tag system, the data is classified and stored according to the established rules and hierarchical structure, so that in the subsequent search process, faster and more accurate positioning can be achieved.
[0070] For example, when managing a database with a large number of documents, by establishing indexes for the titles, basic tags, and composite tags of each document, the search time can be significantly reduced. When the user conducts a search, the system can immediately locate the documents containing the relevant tags, avoiding the need to review the document library one by one. The tag-based search method can reduce the search time from several minutes to several milliseconds, achieving instant response. In addition, the structured storage method also helps to improve the accuracy of the search. Since the data has been pre-classified and labeled, the system can more accurately understand and interpret the user's query intention, thus providing more accurate search results.
[0071] During the data configuration phase, the model provides an interactive preview mechanism, including preview of configuration items, service parameter preview, and calibration page preview, to achieve instant feedback on the design. This mechanism allows users to observe the layout and style of the template in real time during the configuration process, ensuring that the final presented template design meets the expected requirements. Once the template design is completed and meets the user's needs, the user can choose to publish the template to the system. The publishing process integrates the template into the system, making it available for users to reuse in future tasks, improving the reusability of the template and work efficiency. If the settings are not completed, the user can also choose to save, and the saved template will also appear in the list of the template management module.
[0072] The document template management module provides a centralized platform for users. This part is used to configure, delete, preview, and publish the templates they have created. The design of this module aims to simplify the management and maintenance work of the user for the template library, ensuring the timely update and optimization of the templates. Through the template management module, users can easily browse and retrieve their template collections and quickly find the templates they need for editing.
[0073] B. The OCR-based document structured warehousing module includes task creation, task management, precise information extraction, structured storage, and intelligent content display. This module has the core functions of precise information extraction, structured storage, and intelligent content display. At the information extraction level, this module integrates the PaddleOCR algorithm and innovatively proposes a permutation language model that combines BERT and BiLSTM, which can achieve precise recognition and extraction of the text of archival materials, thus ensuring the accuracy and integrity of the information. At the same time, the module uses the advanced PaddleX library, face_recognition technology, and PaddleSeg suite to perform refined segmentation processing on the picture content, further enriching the dimension and depth of information extraction. After the information extraction is completed, the module further develops a table recognition mechanism that uses distance calculation and visual hierarchy analysis based on the recognition results and the rectangular coordinate arrangement rules of their outputs. This mechanism can not only achieve the structured storage of the table content but also perform intelligent content display according to the specific format and layout of the table, greatly improving the readability and usability of the information.
[0074] To efficiently extract text and image information from digital copies, users can utilize the OCR-based document structured warehousing module. This module can not only identify and extract information but also display the extracted data in an intelligent manner according to the templates customized by users, ensuring the accuracy of the information and the consistency of the presentation. This module mainly consists of five parts: task creation, task management, precise information extraction, structured storage, and intelligent content display.
[0075] First, task creation is required. Task creation is the process of setting up a new task, including the task name, acquisition method, and file selection. The task name and acquisition method are mandatory fields. The acquisition method can be selected as local upload or file synchronization.
[0076] Once the task is successfully created, it will be displayed on the task management page. This page allows for editing, deletion, calibration viewing, unrecognized document viewing, and archiving operations on the task.
[0077] Steps 101 to 104 of this application can be implemented through the precise information extraction module. The system receives the digital copy uploaded by the user. Then, the system will analyze the document to extract the information therein. Since the OCR model cannot directly perform image segmentation, the information extraction process includes two main aspects: text information extraction and image segmentation extraction.
[0078] Text information extraction involves using OCR technology to identify and convert the text content in the document into an editable and searchable text format. This process involves precise detection, efficient recognition, and accurate conversion of the text in the document, aiming to ensure that the extracted text information is correct and meets various application requirements.
[0079] In the present invention, the PaddleOCR OCR technology is adopted. PaddleOCR has high-precision recognition capabilities, supports multi-language recognition and multi-task processing, and is an open-source OCR tool. It also provides annotation tools for fine-tuning and optimizing the application scenario model. By using PaddleOCR, the present invention can more accurately and efficiently extract the text information in the document, providing strong support for subsequent text processing and analysis.
[0080] When performing text recognition, PaddleOCR first performs preprocessing operations such as scaling, grayscaling, and denoising to improve the accuracy of text recognition. These steps help improve the image quality, especially for images with noise or skew. The output results include text position information and text content recognition results. Since the PaddleOCR tool already has very good recognition effects on text position and text content, the present invention focuses on using the output of PaddleOCR and designs a deep learning model to restore the problem of non-horizontal text layout.
[0081] Specifically, the present invention first uses the position information of the text output by PaddleOCR to splice text fragments. This process helps to recombine scattered text fragments into a complete text sequence, thereby restoring the original meaning of the text to the greatest extent. Subsequently, by combining pre-trained permutation language models (BERT and BiLSTM), the spliced text is processed to restore the correct order and meaning of the text. The processed text will be further classified. If the recognized text belongs to the title, atomic label, or combined label area, the system will match it with the locally constructed label library and associate similar labels for subsequent template filling. If the recognized part is content text, it will be associated and bound with the coordinates of adjacent labels.
[0082] As an alternative implementation, step 104 generates a digital file corresponding to the file image using each text content, specifically including: Perform type recognition on each text content to obtain the text type of each text content; wherein, the text type includes a label type and a content type; Determine the coordinate information of each text content in the file image; According to the coordinate information of each text content, associate a text content of the content type with a text content of the label type; wherein, the distance between the text content of the content type and the text content of the label type is the closest; Obtain a pre-set digital file template; wherein, the digital file template contains multiple target labels; Repeat the filling operation until the blank area corresponding to each target label in the digital file template is filled with text content, obtaining a digital file corresponding to the file image: Wherein, the filling content specifically includes: Determine a current target label from multiple target labels; Determine the target text content identical to the current target label from the text content of the label type; Fill the text content of the content type corresponding to the target text content into the blank area corresponding to the current target label.
[0083] Among them, to implement this implementation method, first, type recognition technology is adopted to automatically distinguish between tag-type and content-type texts, and a structured information classification system is constructed; secondly, based on spatial coordinates, the spatial proximity between tags and content is accurately calculated to achieve intelligent matching of semantic associations; then, a preset digital archive template is used to standardize the archive data structure, and through a dynamic loop filling mechanism, the information integrity and format consistency are ensured; finally, an end-to-end automated process from the original image to the standardized digital archive is formed. This multi-dimensional information fusion and rule-driven processing mode effectively avoids the mismatch risk in manual association, eliminates information loss in format conversion, and significantly improves the processing efficiency and quality standardization level of complex archives.
[0084] As an alternative implementation method, after step 104, the following steps can also be executed: Perform face target detection on the archive image to obtain the face area in the archive image; Perform portrait matting operation on the face area to obtain a target face image; Fuse the target face image with a pre-determined background image to obtain a fused face image; Determine the position of the face image from the digital archive; Fuse the fused face image to the position of the face image in the digital archive to obtain a target digital archive containing the fused face image.
[0085] Among them, to implement this implementation method, first, accurate face detection and adaptive matting algorithms are adopted to effectively extract facial features in the archive image and strip complex backgrounds, solving quality problems such as blurred and yellowed faces caused by the storage years of traditional archives; secondly, through intelligent image fusion technology, the optimized face is seamlessly combined with the standardized background, not only retaining the historical features of the original archive but also giving it a modern visual presentation effect; then, coordinate positioning technology is used to accurately map the enhanced face to the corresponding area of the digital archive to achieve non-destructive content upgrade; finally, a new type of archive carrier with both historical authenticity and digital readability is formed, significantly improving the integrity and recognition of archive content.
[0086] In the embodiment of the present application, first, PaddleX can be used for target detection to roughly identify and locate the face image area. After narrowing down the range where the face image is located, the advantages of the face_recognition library can be utilized to accurately locate and intercept the face area in the ID photo.
[0087] In the embodiment of the present application, the target face image obtained by matting has a transparent background. This hierarchical processing method not only ensures the accurate separation of the portrait and the background but also provides high-quality image materials for subsequent background merging and image reshaping.
[0088] Finally, the present invention combines the matte result with the required background and reshapes the background to form the final fused face image. This process not only preserves the clarity and details of the portrait but also makes the background more coordinated with the overall effect. In this process, the present invention not only improves the accuracy of image segmentation but also provides more possibilities for subsequent processing and applications of the image.
[0089] After the information extraction is completed, the system will perform structured processing on the extracted text content and image content for easy storage and management.
[0090] Specifically, the present invention uses the n-gram method for the matching between the sequentially restored text and the locally constructed tag library. In natural language processing, shorter n-grams tend to capture the surface form of the vocabulary, while longer n-grams can capture more context information. The present invention combines 1-gram, 2-gram, 3-gram, and 4-gram to capture the features of the text at different granularities for matching, and gives higher weights to larger n-grams. This method can identify the tag categories, and those that do not match the local tag library are marked as content. The present invention combines distance calculation and visual hierarchy analysis, optimizes the binding relationship by calculating the center point distance between the text and the tag and considering their relative positions in the visual layout.
[0091] First, calculate the distances between the text blocks and the center points of each tag, and select the closest one for preliminary binding.
[0092] Then, adjust the binding according to the visual hierarchy analysis to ensure compliance with the visual logic.
[0093] Finally, accurate and reasonable binding between the text and the tag is achieved.
[0094] Classifying and tagging the extracted information can facilitate retrieval and analysis.
[0095] During the structuring process, the system can use this information classification technology to more precisely organize the information in the document according to a specific data structure.
[0096] For example, for a resume document, the system may separately mark and classify the information in different parts such as personal information, educational background, work experience, etc. for subsequent query and application. Subsequently, the system stores the structured information in the database to complete the information extraction and structured storage process of the document. In this way, the present invention not only optimizes the accuracy of information extraction but also improves the overall efficiency of data processing.
[0097] After being structured, the present invention fills the stored content into a template designed by the user to display information in a clear and orderly manner, thereby helping the user easily browse and understand the extracted data.
[0098] Please refer to Figures 3 to 5 , Figure 3 which is a schematic diagram of an archive image provided by an embodiment of the present application; Figure 4 which is a schematic diagram of a digital archive template provided by an embodiment of the present application; Figure 5 which is a schematic diagram of a digital archive provided by an embodiment of the present application, Figure 5 The confirmation button in can be pressed when the user believes that the generation of the archive is completed, indicating that the archive has been completed. Figure 5 The * mark in indicates that the content corresponding to the label "Resume" is a required item. The relevant experiments are based on Python 3.9 and PaddlePaddle, requiring cuda 11.2 or above in cooperation with cudnn 8.2.4. The main data packets include PaddleOCR, PaddleX, PaddleSeg, and face_recognition, etc.
[0099] Implementing the above steps 101 to 104 to generate digital archives using these high-quality text contents effectively reduces the manual intervention and error troubleshooting time compared with the traditional method, avoids the repetitive work caused by text recognition errors, and thus comprehensively improves the efficiency of archive digitization. At the same time, the precise adjustment of the text position ensures that the content of the digital archive is highly consistent with the original archive, greatly improving the accuracy of archive digitization. In addition, the present application can also enhance the adaptability to archives with special formats. In addition, the present application can also enhance the reliability of text recognition in complex scenarios. In addition, the present application can also enhance the full-process automation level of archive digitization. In addition, the present application can also improve the processing efficiency and quality standardization level of complex archives. In addition, the present application can also improve the integrity and recognition rate of archive content.
[0100] Based on the same inventive concept, the embodiment of the present application also provides a digital processing device for archives for implementing the above-mentioned digital processing method for archives. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the digital processing device for archives provided below can refer to the limitations on the digital processing method for archives in the above text and will not be repeated here.
[0101] In an exemplary embodiment, as Figure 6 shown, a digital processing device for archives is provided, including: A determination unit 601, configured to determine a plurality of text blocks from pre-collected archive images; An input unit 602, configured to input the plurality of text blocks into a pre-trained permutation language model to obtain word-position related information of each text block output by the permutation language model; wherein, one piece of word-position related information includes the probabilities of each word included in a text block at different positions in the text block; An adjustment unit 603, configured to adjust the word positions in each text block based on the word-position related information to obtain the text content of each text block; A generation unit 604, configured to generate a digital archive corresponding to the archive image using each text content.
[0102] As an optional implementation manner, the generation unit 604 is further configured to: Perform face target detection on the archive image to obtain a face region in the archive image; Perform portrait matting operation on the face region to obtain a target face image; Fuse the target face image with a pre-determined background image to obtain a fused face image; Determine a face image position from the digital archive; Fuse the fused face image to the face image position in the digital archive to obtain a target digital archive including the fused face image.
[0103] Wherein, when implementing this implementation manner, first, a precise face detection and adaptive matting algorithm is adopted to effectively extract facial features in the archive image and strip complex backgrounds, solving quality problems such as facial blurring and yellowing caused by the storage years of traditional archives; second, through intelligent image fusion technology, the optimized face is seamlessly combined with a standardized background, not only retaining the historical features of the original archive but also giving it a modern visual presentation effect; then, coordinate positioning technology is used to accurately map the enhanced face to the corresponding area of the digital archive, realizing non-destructive content upgrade; finally, a new type of archive carrier with both historical authenticity and digital readability is formed, significantly improving the integrity and recognition of archive content.
[0104] As an optional implementation manner, the manner in which the input unit 602 inputs a target text block among the plurality of text blocks into a pre-trained permutation language model to obtain the word-position related information of the target text block output by the permutation language model may specifically be: Determine each word included in a target text block among the plurality of text blocks and the initial position of each word in the target text block; Determine the token encoding of each character and the hidden state combination information of each character according to each character and the initial position of each character in the target text block; wherein, the hidden state combination information includes the forward hidden state and the backward hidden state of the character corresponding to the hidden state combination information. Calculate the character-position related information of the target text block based on the token encoding of each character and the hidden state combination information of each character.
[0105] Among them, implementing this implementation method, by accurately obtaining the initial position information of the characters, a basic coordinate system is constructed for the model to analyze the text space layout; innovatively combining the semantic representation of token encoding with the bidirectional hidden state combination information, not only preserves the semantic features of the characters themselves, but also through the interaction modeling of the forward hidden state and the backward hidden state, enables the model to dynamically capture the logical order and spatial correlation between characters; finally, the position relationship of the characters is deduced through probability calculation, enabling the system to have the ability to understand the arrangement rules of complex layouts. This deep context awareness mechanism effectively breaks through the recognition limitations of traditional OCR for artistic fonts, handwritten fonts, and multi-directional mixed text, and significantly enhances the adaptability to special layout archives.
[0106] As an optional implementation method, the way for the input unit 602 to determine each character included in a target text block among multiple text blocks and the initial position of each character in the target text block can specifically be: Perform a magnification operation on an initial text block among multiple text blocks to obtain a first text block; Perform grayscale processing on the first text block to obtain a second text block; Perform a denoising operation on the second text block to obtain a target text block; Perform character recognition on the target text block to obtain each character included in the target text block and the initial position of each character in the target text block.
[0107] Among them, implementing this implementation method, first, an adaptive magnification strategy is adopted to enhance the detail features of tiny characters or blurred areas, providing a higher-definition text base for subsequent processing; then, color interference is eliminated through grayscale processing, compressing the three-dimensional color space into a two-dimensional grayscale gradient, reducing the computational complexity while retaining the key information of the character outline; then, an intelligent denoising algorithm is used to accurately filter out interference factors such as image noise and scratches, constructing a text recognition environment with high signal-to-noise ratio; finally, through refined character positioning technology, the initial coordinates of each character are accurately obtained, providing an accurate spatial layout basis for the subsequent arrangement model. This full-process automatic optimization mechanism effectively solves the common problems such as blur, noise, and color interference in low-quality archival images, and significantly enhances the reliability of character recognition in complex scenarios.
[0108] As an alternative implementation, the way for the input unit 602 to calculate the character-position related information of the target text block based on the token encoding of each character and the combined hidden state information of each character can be specifically as follows: Transpose the combined hidden state information of each character to obtain the transposed combined hidden state information of each character; Multiply the token encoding of each character by the transposed combined hidden state to obtain an initial matrix; Based on a preset bias, calculate the initial matrix to obtain a target matrix; Perform a normalization operation on the target matrix to obtain the character-position related information of the target text block.
[0109] Among them, when implementing this implementation, first, the hidden state transposition strategy is adopted to enable the forward and backward semantic information to form complementary interaction in matrix operations, enhancing the model's bidirectional perception ability of the spatial relationship of characters; second, through the matrix fusion operation of token encoding and hidden state, while retaining the deep semantic features of characters, context clues required for position distribution prediction are dynamically injected; then, a learnable bias parameter is introduced to calibrate the initial matrix, enabling the model to adaptively adjust the position prediction weights in different layout scenarios; finally, the original prediction values are converted into probability distributions through normalization operations, ensuring that the character position inference results conform to mathematical statistical laws. This multi-level joint optimization mechanism effectively solves the recognition problems of traditional OCR for special layouts such as multi-directional mixed arrangements and curved texts, and significantly enhances the full-process automation level of archival digitization processing.
[0110] As an alternative implementation, the way for the generation unit 604 to generate a digital archive corresponding to the archival image using each text content can be specifically as follows: Perform type recognition on each text content to obtain the text type of each text content; among them, the text type includes a label type and a content type; Determine the coordinate information of each text content in the archival image; According to the coordinate information of each text content, associate a text content of a content type with a text content of a label type; among them, the distance between the text content of a content type and the text content of a label type is the closest; Obtain a pre-set digital archive template; among them, the digital archive template contains multiple target labels; Repeat the filling operation until the blank area corresponding to each target label in the digital archive template is filled with text content, obtaining a digital archive corresponding to the archival image: Among them, the filling content specifically includes: Determine a current target label from multiple target labels; Determine the target text content that is the same as the current target tag from the text content of the tag type; Fill the text content of the content type corresponding to the target text content into the blank area corresponding to the current target tag.
[0111] Among them, to implement this implementation method, first, use type recognition technology to automatically distinguish tag-type and content-type texts and construct a structured information classification system; second, accurately calculate the spatial proximity between tags and content based on spatial coordinates to achieve intelligent matching of semantic associations; then use a preset digital archive template to standardize the archive data structure, and ensure information integrity and format consistency through a dynamic loop filling mechanism; finally, form an end-to-end automated process from the original image to the standardized digital archive. This multi-dimensional information fusion and rule-driven processing mode effectively avoids the mismatch risk in manual association, eliminates information loss in format conversion, and significantly improves the processing efficiency and quality standardization level of complex archives.
[0112] Implementing the above implementation method, using these high-quality text contents to generate digital archives, compared with the traditional method, effectively reduces the time of manual intervention and error troubleshooting, avoids repeated work caused by text recognition errors, and thus comprehensively improves the efficiency of archive digitization processing. At the same time, the accurate adjustment of the text position ensures that the content of the digital archive is highly consistent with the original archive, greatly improving the accuracy of archive digitization processing. In addition, this application can also enhance the adaptability to special format archives. In addition, this application can also enhance the reliability of text recognition in complex scenarios. In addition, this application can also enhance the end-to-end automation level of archive digitization processing. In addition, this application can also improve the processing efficiency and quality standardization level of complex archives. In addition, this application can also improve the integrity and recognition rate of archive content.
[0113] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 7As shown in the figure. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the digital processing data of the files. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it realizes a method for digital processing of files.
[0114] Those skilled in the art can understand that Figure 7 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0115] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are realized.
[0116] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are realized.
[0117] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are realized.
[0118] In an exemplary embodiment, a chip is provided. The chip includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to realize the steps in the above method embodiments, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0119] It should be understood that the chip mentioned in the embodiments of this application can also be referred to as a system-on-chip, system chip, chip system, or system-on-chip, etc.
[0120] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0121] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0122] The databases involved in the various embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the various embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0123] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0124] In this article, specific examples are used to elaborate on the principles and implementation manners of the present application. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for digital processing of archives, characterized in that: The digital processing method of the archives includes: determining a plurality of text blocks from pre-captured archival images; Input the plurality of text blocks into a pre-trained permutation language model to obtain character-position related information of each text block output by the permutation language model; wherein a character-position related information includes the probability of each character contained in a text block being located at a different position in the text block; Based on the text-position related information, the text position in each text block is adjusted to obtain the text content of each text block; A digital archive corresponding to the archive image is generated using each text content.
2. The method for digitalizing archives according to claim 1, characterized in that: The digital processing method of the archives also includes: Performing face target detection on the archive image to obtain a face region in the archive image; Performing a portrait cutout operation in the face area to obtain a target face image; Merging the target face image with a predetermined background image to obtain a fused face image; Determining a facial image location from the digital file; The fused face image is fused to the face image position in the digital file to obtain a target digital file containing the fused face image.
3. The method for digitalizing archives according to claim 1, characterized in that: Inputting a target text block among the multiple text blocks into a pre-trained permutation language model to obtain text-position related information of the target text block output by the permutation language model, specifically including: Determine each character contained in a target text block among the plurality of text blocks and an initial position of each character in the target text block; Determine the word unit encoding of each character and the hidden state combination information of each character according to each character and the initial position of each character in the target text block; wherein the hidden state combination information includes the forward hidden state and the backward hidden state of the character corresponding to the hidden state combination information; Based on the word-element encoding of each character and the hidden state combination information of each character, the character-position related information of the target text block is calculated.
4. The method for digitalizing archives according to claim 3, characterized in that: The step of determining each character contained in a target text block among the plurality of text blocks and the initial position of each character in the target text block specifically includes: Performing an enlargement operation on an initial text block among the multiple text blocks to obtain a first text block; Gray-scaling the first text block to obtain a second text block; Performing a denoising operation on the second text block to obtain a target text block; Character recognition is performed on the target text block to obtain each character contained in the target text block and the initial position of each character in the target text block.
5. The method for digitalizing archives according to claim 3 or 4, characterized in that: The word-unit encoding of each word and the hidden state combination information of each word are used to calculate the word-position related information of the target text block, specifically including: Transpose the hidden state combination information of each character to obtain the transposed hidden state combination information of each character; Multiply the word unit encoding and transposed hidden state combination of each word to get the initial matrix; Based on a preset bias, the initial matrix is calculated to obtain a target matrix; The target matrix is normalized to obtain the text-position related information of the target text block.
6. The method for digitalizing archives according to claim 1, characterized in that: The step of using each text content to generate a digital archive corresponding to the archive image specifically includes: Performing type identification on each text content to obtain the text type of each text content; wherein the text type includes a tag type and a content type; Determine the coordinate information of each text content in the archive image; According to the coordinate information of each text content, a text content of a content type is associated with a text content of a tag type; wherein the distance between the text content of the content type and the text content of the tag type is the shortest; Obtaining a preset digital file template; wherein the digital file template includes a plurality of target tags; Repeat the filling operation until the blank area corresponding to each target label in the digital archive template is filled with text content, and obtain the digital archive corresponding to the archive image: The content to be filled in specifically includes: Determine a current target label from multiple target labels; Determine, from the text content of the tag type, the target text content that is the same as the current target tag; Fill the text content of the content type corresponding to the target text content into the blank area corresponding to the current target tag.
7. A digital file processing device, characterized in that: The digital processing device of the archives comprises: A determination unit, used for determining a plurality of text blocks from a pre-collected archive image; An input unit, used to input the plurality of text blocks into a pre-trained permutation language model, and obtain character-position related information of each text block output by the permutation language model; wherein a character-position related information includes the probability of each character contained in a text block being located at a different position in the text block; An adjustment unit, configured to adjust the position of the characters in each text block based on the character-position related information to obtain the text content of each text block; A generating unit is used to generate a digital archive corresponding to the archive image using each text content.
8. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method for digitalizing archives described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the digital processing method of archives described in any one of claims 1 to 6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the digital processing method of archives described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Resume analysis method and system based on deep learning
CN111737969A
Slice document key information single model extraction method and system
CN113536797A
Method and system for automatically structuring key information of document image
CN114328845A
License key information extraction method based on small sample data
CN116229494A
Document image automatic classification and cleaning method, device and system and storage medium
CN118116016A