Key information extraction model training method for set pattern and related product
By converting document data into pictures and extracting setting pattern information, combining the annotation and association of text information, training the setting pattern extraction model, solving the problem of difficulty in extracting non-text information in the prior art, and improving the accuracy of information extraction.
Patent Information
- Application Number
- CN202510130667.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-30
AI Technical Summary
It is difficult for the prior art to effectively extract non-text information in documents, such as check box status, check sign or cross status, resulting in a decrease in the accuracy of information extraction.
By obtaining key information document data containing the setting pattern, converting it into a picture, extracting the setting pattern position information and category information in the image, and marking the text position information and meaning information, forming a relationship to train the setting pattern extraction model.
The semantics of the setting pattern are extracted, the accuracy of information extraction is improved, and the depth and breadth of the model's understanding of the setting pattern is enhanced.
Smart Images

Figure CN120071375A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of key information extraction from documents, and particularly relates to a method for training a key information extraction model for a set pattern and related products. Background Art
[0002] The technology of key information extraction from documents is a key technology in the field of natural language processing (NLP). Its development background is rooted in the urgent need for effective information in a vast amount of text data. With the development of the Internet and digital technologies, we have entered the era of information explosion, and the technology of key information extraction from documents has emerged, aiming to help people quickly identify and extract valuable information from complex texts, improving the efficiency and accuracy of information processing.
[0003] The application scenarios of the technology of key information extraction from documents are extremely extensive. In the legal field, key information such as contract terms, case facts, and legal bases are identified from legal documents to assist legal professionals in case analysis and document review; in the field of medical and health, key medical information such as symptoms, diagnoses, and treatment suggestions are extracted from medical records and medical literature to support clinical decision-making and medical research. In the field of financial analysis, key financial indicators, market trends, etc. are extracted from financial reports and market analysis reports to assist investment decision-making and risk assessment, and so on.
[0004] The technology of key information extraction has been relatively mature, but there are still some problems. The technology of key information extraction is based on text semantics as the basic feature and cannot extract non-text information, such as the status of checkboxes next to text in a document; the status of ticks or crosses after text; other semantic patterns that cannot be extracted by OCR technology, or semantics or status that cannot be characterized by characters. These kinds of information often play a crucial role in the accuracy of the extracted information, and there is no mature technology to solve this problem currently, thus reducing the accuracy of information extraction. Summary of the Invention
[0005] Based on the above problems, the embodiments of the present application provide a method for training a key information extraction model for a set pattern and related products to solve the problems existing in the above-mentioned prior art.
[0006] The present application discloses a method for training a key information extraction model for a set pattern, including: obtaining key information document data containing the set pattern and converting the key information document data into pictures; extracting the set pattern position information and set pattern category information in the pictures to obtain a first data set; annotating the text position information and text meaning information in the pictures to obtain a second data set; based on the set pattern position information and text position information in the second data set, associating the set pattern category information in the second data set with the text meaning information to obtain a third data set; and training the set pattern extraction model to be trained based on the third data set until the training of the set pattern extraction model is completed.
[0007] Optionally, the obtaining key information document data containing the set pattern and converting the key information document data into pictures includes: obtaining key information document data containing the set pattern and converting the document data into pictures based on a format conversion tool.
[0008] Optionally, the extracting the set pattern position information and set pattern category information in the pictures to obtain a first data set includes: based on an annotation tool, extracting the set pattern position information and set pattern category information in the pictures according to the annotation of the set pattern on the pictures to obtain a first data set.
[0009] Optionally, the annotating the text position information and text meaning information in the pictures to obtain a second data set includes: based on an OCR tool, extracting the corresponding text position information and text meaning information according to the annotation of the text on the pictures to obtain a second data set.
[0010] Optionally, the associating the set pattern category information in the second data set with the text meaning information based on the set pattern position information and text position information in the second data set to obtain a third data set includes: associating the set pattern category information in the second data set with the text meaning information based on the coordinates of the set pattern position information and text position information in the second data set to obtain a third data set.
[0011] Optionally, the associating the set pattern category information in the second data set with the text meaning information based on the coordinates of the set pattern position information and text position information in the second data set to obtain a third data set includes: mapping the text meaning information to the pattern category information based on the coordinates of the set pattern position information and text position information in the second data set for association to obtain a third data set.
[0012] Optionally, the method includes: training a to-be-trained set pattern detection model based on the first data set, so as to extract set pattern position information and set pattern category information in the picture based on the trained set pattern detection model.
[0013] The present application further provides a training device for a key information extraction model of a set pattern, including: a first information collection unit: obtaining key information document data containing the set pattern, and converting the key information document data into a picture; extracting the set pattern position information and the set pattern category information in the picture to obtain a first data set; a second information collection unit: annotating the text position information and the text meaning information in the picture to obtain a second data set; an information fusion unit: associating the set pattern category information in the second data set with the text meaning information based on the set pattern position information and the text position information in the second data set to obtain a third data set; a model training unit: training a to-be-trained set pattern extraction model based on the third data set until the training of the set pattern extraction model is completed.
[0014] The present application further provides a training method for a key information extraction model for a set target object, including: Obtaining key information carrier data containing the set target object, and extracting set target object position information and set target object category information therefrom; Annotating the relevant identifier position information and the relevant identifier meaning information in the specific visualization data form; Associating the set target object category information with the relevant identifier meaning information based on the set target object position information and the relevant identifier position information to obtain a data set for training the model; Training a to-be-trained set target object extraction model based on the data set for training the model until the training of the set target object extraction model is completed.
[0015] The present application further provides a key information extraction method, which includes: Invoking a target object extraction model trained based on any embodiment of the present application; Inputting key information carrier data containing a target object to be extracted into the target object extraction model to extract the target object to be extracted therefrom.
[0016] The present application provides a method for training a key information extraction model for a set pattern and related products. Among them, a method for training a key information extraction model for a set pattern includes: obtaining key information document data containing the set pattern and converting the key information document data into pictures; extracting the position information and category information of the set pattern in the pictures to obtain a first data set; annotating the position information and semantic information of the text in the pictures to obtain a second data set; based on the position information of the set pattern and the position information of the text in the second data set, associating the category information of the set pattern in the second data set with the semantic information of the text to obtain a third data set; training the set pattern extraction model to be trained based on the third data set until the training of the set pattern extraction model is completed. By converting the key information document data containing the set pattern into pictures, advanced technologies and algorithms in the field of image processing can be fully utilized to efficiently and accurately identify and extract the key information of the set pattern. This conversion makes the originally complex text processing tasks intuitive and easy to handle. In this embodiment, by simultaneously extracting the position information and category information of the set pattern in the picture, and annotating the position information and semantic information of the text, this training method realizes the fusion of image information and text information. The fusion of such multi-source information can provide a richer context environment, which helps to improve the depth and breadth of the model's understanding of the set pattern. Moreover, based on the position information of the set pattern and the position information of the text, the category information of the set pattern is associated with the semantic information of the text to form a third data set. This intelligent association mechanism not only enhances the semantic expression ability of the data, but also provides strong support for the model to learn the potential relationship between the set pattern and the text, enabling the model to more accurately understand and extract the key information of the set pattern. In addition, training the model based on the third data set that integrates the category information of the set pattern and the semantic information of the text can significantly improve the recognition accuracy and extraction efficiency of the model for the set pattern. Since the training data contains rich context information and semantic associations, the model can learn more comprehensive and in-depth feature representations, thus showing stronger generalization ability and robustness in practical applications. Finally, this training method has strong versatility and flexibility, and can be applied to various scenarios that require extracting specific pattern information from documents or pictures, such as extracting specific identifiers in financial documents, identifying official seals in legal documents, etc. By adjusting the model structure and training parameters, it can adapt to changes in different fields and specific requirements. In summary, by establishing a third data set through training in the present application, a key information extraction model for a set pattern is obtained, which can extract the semantics of the set pattern and improve the accuracy of information extraction. Brief Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0018] Figure 1 Flowchart of a method for training a key information extraction model for a set pattern in an embodiment of the present application; Figure 2 Example diagram of key information document data containing a set pattern in an embodiment of the present application; Figure 3 Example diagram of set pattern category information in an embodiment of the present application; Figure 4 Example diagram of the first dataset in an embodiment of the present application; Figure 5 Example diagram of the second dataset in an embodiment of the present application; Figure 6 Example diagram of the third dataset in an embodiment of the present application; Figure 7 Recognition effect diagram of the set pattern extraction model in an embodiment of the present application. Detailed implementation manners
[0019] Implementing any of the technical solutions in the embodiments of the present application does not necessarily require achieving all of the above advantages simultaneously.
[0020] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0021] Figure 1 Flowchart of a method for training a key information extraction model for a set pattern in an embodiment of the present application; Figure 2 Example diagram of key information document data containing a set pattern in an embodiment of the present application; Figure 3 Example diagram of set pattern category information in an embodiment of the present application; Figure 4 Example diagram of the first dataset in an embodiment of the present application; Figure 5 Example diagram of the second dataset in an embodiment of the present application; Figure 6 Example diagram of the third dataset in an embodiment of the present application; Figure 7This is the recognition effect diagram of the set pattern extraction model in the embodiment of the present application. As Figures 1-7 shown: A method for training a key information extraction model for a set pattern includes: obtaining key information document data containing the set pattern and converting the key information document data into pictures; extracting the position information and category information of the set pattern in the pictures to obtain a first data set; annotating the position information and semantic information of the text in the pictures to obtain a second data set; based on the position information of the set pattern and the position information of the text in the second data set, associating the category information of the set pattern in the second data set with the semantic information of the text to obtain a third data set; training the set pattern extraction model to be trained based on the third data set until the training of the set pattern extraction model is completed. In this embodiment, by converting the key information document data containing the set pattern into pictures, advanced technologies and algorithms in the field of image processing can be fully utilized to efficiently and accurately identify and extract the key information of the set pattern. This conversion makes the originally complex text processing task intuitive and easy to handle. In this embodiment, by simultaneously extracting the position information and category information of the set pattern in the picture, as well as annotating the position information and semantic information of the text, this training method realizes the fusion of image information and text information. The fusion of this multi-source information can provide a richer context environment, which helps to improve the depth and breadth of the model's understanding of the set pattern. Moreover, based on the position information of the set pattern and the position information of the text, the category information of the set pattern is associated with the semantic information of the text to form a third data set. This intelligent association mechanism not only enhances the semantic expression ability of the data, but also provides strong support for the model to learn the potential relationship between the set pattern and the text, enabling the model to more accurately understand and extract the key information of the set pattern. In addition, training the model based on the third data set that combines the category information of the set pattern and the semantic information of the text can significantly improve the recognition accuracy and extraction efficiency of the model for the set pattern. Since the training data contains rich context information and semantic associations, the model can learn more comprehensive and in-depth feature representations, thus showing stronger generalization ability and robustness in practical applications. Finally, this training method has strong versatility and flexibility, and can be applied to various scenarios that require extracting specific pattern information from documents or pictures, such as extracting specific identifiers in financial documents and identifying official seals in legal documents. By adjusting the model structure and training parameters, it can adapt to changes in different fields and specific requirements. In summary, through the present application, a key information extraction model for a set pattern is obtained by establishing a third data set for training, which can extract the semantics of the set pattern and improve the accuracy of information extraction Optionally, obtaining the key information document data containing the set pattern and converting the key information document data into a picture includes: obtaining the key information document data containing the set pattern and converting the document data into a picture based on a format conversion tool. In this embodiment, using a format conversion tool can automatically convert document data into a picture, greatly reducing the steps and time costs of manual operations. This automated processing not only improves work efficiency but also reduces the possibility of human errors. Moreover, a format conversion tool can usually convert document data into a picture according to preset rules and standards, ensuring the consistency of the converted picture in terms of format, resolution, color, etc. This consistency is very important for subsequent picture processing and model training and can avoid performance degradation caused by inconsistent data formats. In addition, during the process of converting document data into a picture, a format conversion tool usually tries to retain as much information as possible in the original document, including text, patterns, layout, etc. This characteristic of retaining the original information helps to more accurately extract the key information and text information of the set pattern in subsequent steps. Finally, compared with methods such as manual screenshotting or copy-pasting, using a format conversion tool for document-to-picture conversion can more effectively reduce the risk of data loss and distortion. This is because the tool usually preprocesses and optimizes the document to ensure that the quality of the converted picture meets the requirements of subsequent processing.
[0022] Optionally, extracting the set pattern position information and the set pattern category information in the picture to obtain a first data set includes: based on an annotation tool, extracting the set pattern position information and the set pattern category information in the picture according to the annotation of the set pattern on the picture to obtain a first data set. In this embodiment, using an annotation tool for annotation can ensure the accuracy of the set pattern position information and category information. In addition, annotation tools usually provide various annotation options and tools, such as rectangular boxes, polygon boxes, labels, etc., to adapt to set patterns of different shapes and sizes, and appropriate annotation methods can be selected according to the actual situation to ensure the accuracy and effectiveness of the annotation results. Moreover, the annotation information generated by the annotation tool is usually stored in a structured form, such as JSON, XML, etc. This structured data format is convenient for subsequent data management and reuse. For example, the annotation information can be easily modified or extended, or the annotation data set can be used for other related image processing or machine learning tasks.
[0023] Optionally, annotating the text position information and text meaning information in the picture to obtain a second data set, including: based on an OCR tool, extracting the corresponding text position information and text meaning information according to the annotation of the text in the picture to obtain a second data set. In this embodiment, the OCR tool can automatically scan the text in the picture and recognize its content and position, greatly reducing the time and cost of manual annotation. This automated processing not only improves work efficiency but also can process a large number of pictures in a short time to meet the needs of large-scale data sets. In addition, with the continuous development of OCR technology, modern OCR tools have been able to achieve a relatively high recognition accuracy. Through advanced image processing and pattern recognition algorithms, they can accurately recognize the text in the picture, including text of various fonts, sizes, and layouts. This high accuracy ensures the reliability of the annotation results and provides strong support for subsequent data processing and model training. In addition to this, the OCR tool can not only recognize the text content but also provide the position information of the text. This enables the simultaneous acquisition of the text information and spatial layout information of the text during the annotation process, providing more comprehensive data support for subsequent text analysis and image processing. Finally, OCR tools usually have good interfaces and scalability and can be easily integrated with other image processing or machine learning tools. This means that during subsequent data processing or model training, the OCR tool can be conveniently called to extract and annotate text information, thus constructing a more complex and efficient information processing system.
[0024] Optionally, associating the set pattern category information in the second data set with the literal meaning information based on the set pattern position information and literal position information in the second data set to obtain a third data set includes: associating the set pattern category information in the second data set with the literal meaning information based on the coordinates of the set pattern position information and literal position information in the second data set to obtain a third data set. In this embodiment, by directly using the coordinates of the position information for association, it can be ensured that the set pattern and its corresponding literal description maintain an accurate correspondence in space. This accuracy is crucial for understanding the structure and content of the document, especially when dealing with documents having a complex layout and multiple set patterns. Additionally, associating the category information of the set pattern with the literal meaning information through spatial positions not only enriches the semantic information of the data but also makes the relationship between the data clearer and more intuitive. This enhanced data semantics helps the model better understand the internal connection between the set pattern and its description during the training process, thereby improving the model's recognition and understanding capabilities. Moreover, during the model training process, using the third data set containing accurate spatial correspondence relationships for training can help the model learn the accurate matching rules between the set pattern and its description. Such accurate matching rules help the model more accurately identify and understand the set pattern in practical applications, thereby improving the accuracy and reliability of the model. Finally, when dealing with complex documents containing multiple set patterns and a large amount of text, associating based on the coordinates of the position information can effectively organize and manage the data. This method enables the model to handle more complex and diverse document structures, improving the versatility and scalability of the model.
[0025] Optionally, associating the set pattern category information in the second data set with the literal meaning information based on the coordinates of the set pattern position information and the literal position information in the second data set to obtain a third data set, including: mapping the literal meaning information onto the pattern category information based on the coordinates of the set pattern position information and the literal position information in the second data set for association to obtain a third data set. In this embodiment, by directly mapping the literal meaning information onto the pattern category information, the logical consistency and accuracy of the data are ensured. This mapping relationship helps to reduce ambiguity and misunderstanding in the process of data processing and analysis, making subsequent operations more reliable. In addition, after mapping, the data structure and content of the third data set are clearer and more explicit. This structured data form is convenient for subsequent data mining, analysis, and application. Whether it is used for training machine learning models or directly applied to actual business processes, it can significantly improve the availability of data. Additionally, in many application scenarios, there are often complex context relationships between the set pattern and the literal description. By associating and mapping the two, this context relationship can be better understood, thereby supporting more complex document parsing, image understanding, and natural language processing tasks. Moreover, in the fields of machine learning and artificial intelligence, cross-modal learning is an important research direction. By associating and mapping the set pattern (visual modality) and the literal description (text modality), the development of cross-modal learning can be promoted, enabling the model to simultaneously understand and process information from different modalities, thus possessing stronger comprehensive analysis and decision-making capabilities.
[0026] Optionally, the method includes: training a to-be-trained set pattern detection model based on the first data set to extract set pattern position information and set pattern category information in the picture based on the trained set pattern detection model. In this embodiment, by using the first data set containing rich samples and accurate annotations for training, the set pattern detection model can learn various features and variation rules of the set pattern. This learning method enables the model to more accurately identify the set pattern in the picture during the detection process, including its position and category information, thereby improving the detection accuracy. In addition, the first data set usually contains diverse samples, which may vary in terms of shape, size, color, background, etc. By learning these diverse samples, the set pattern detection model can accumulate rich experience and gradually form a general understanding of the set pattern. Additionally, using the trained set pattern detection model to extract the key information of the set pattern in the picture can achieve automated processing. Compared with manual detection and annotation, automated processing not only greatly improves the processing efficiency but also reduces the influence of human errors and subjectivity. This enables the detection and extraction of the set pattern to be completed more quickly in large-scale picture processing scenarios. Moreover, the trained set pattern detection model can be further optimized and adjusted as needed. For example, the recognition range of the model can be expanded by adding new sample data, or its detection performance can be optimized by adjusting the model parameters. This scalability and flexibility enable the model to adapt to different application scenarios and requirements. Finally, the position information and category information of the set pattern are important bases for subsequent data processing and analysis. By accurately extracting this information, it can provide strong support for subsequent tasks such as image recognition, natural language processing, and data mining. This support makes the entire data processing process more complete and efficient.
[0027] Optionally, training the to-be-trained set pattern detection model based on the first data set includes: training the to-be-trained set pattern detection model based on the first data set, and adjusting the model parameters of the to-be-trained set pattern detection model by minimizing the loss function by adjusting the iteration parameters. In this embodiment, by minimizing the loss function, the model continuously optimizes its parameters during training, making the detection of set patterns more accurate. This helps the model better identify and distinguish different patterns in practical applications. Also, by adjusting the iteration parameters (such as learning rate, batch size, etc.), the convergence speed of the model during training can be controlled. Appropriate selection can make the model accelerate the training speed and reduce the training time while maintaining performance. Additionally, during the training process, the model learns the commonalities and differences between different samples. By minimizing the loss function, the model can better handle various complex situations and improve its robustness in different scenarios. Finally, the iteration parameters are the parts that can be flexibly adjusted during the training process. According to actual requirements and data characteristics, different parameter combinations can be configured to optimize the model performance.
[0028] Optionally, training the to-be-trained set pattern extraction model based on the third data set until the training of the set pattern extraction model is completed includes: preprocessing the third data set to obtain a fourth data set; dividing the fourth data set into a training set, a validation set, and a test set according to a set ratio; inputting the training set into the to-be-trained set pattern extraction model for training to obtain a to-be-validated set pattern extraction model; verifying the accuracy of the to-be-validated set pattern extraction model based on the validation set to complete the verification after reaching the set accuracy rate and obtain a to-be-set pattern extraction model; testing the to-be-tested set pattern extraction model based on the test set to obtain a set pattern extraction model after meeting the set test requirements. In this embodiment, by preprocessing the third data set, noise can be removed, the data format can be standardized, data features can be enhanced, etc., thereby improving the quality and consistency of the data. This helps the model better learn and understand the data during training and improve the training efficiency and effect. Also, verifying the accuracy of the to-be-validated set pattern extraction model obtained during the training process based on the validation set can timely discover the problems and deficiencies of the model and make corresponding adjustments and optimizations. This verification process can ensure that the model completes training after reaching the set accuracy rate, thereby improving the accuracy and reliability of the model. Additionally, through the verification and testing processes, feedback information about the model performance can be collected. This feedback information can be used to guide the further optimization and iteration of the model to improve the performance of the model. Finally, through a series of standardized training, verification, and testing processes, it can be ensured that the model follows unified specifications and standards during development. This helps the rapid deployment and stable operation of the model in practical applications.
[0029] Optionally, preprocessing the third data set to obtain a fourth data set includes: cleaning the third data set to remove noise data and obtaining a third data set after impurity removal; performing format conversion on the data set after impurity removal to enable the to-be-trained set pattern extraction model to recognize it and obtaining a fourth data set. In this embodiment, the data cleaning process can identify and remove noise data in the third data set, that is, those incomplete, incorrect, or abnormal data points. These noise data may have a negative impact on the training of the model, causing the model to learn incorrect information or patterns. By removing these noise data, the overall quality of the data set can be significantly improved, thereby improving the effect and accuracy of model training. Additionally, data format conversion is an important step to ensure that the to-be-trained set pattern extraction model can correctly recognize and process the data. Different models may require different input data formats, including data types, data ranges, data structures, etc. By performing format conversion on the data set after impurity removal, it can be ensured that the input data meets the requirements of the model, thereby improving the recognition efficiency and accuracy of the model. In addition, format conversion can also standardize and normalize the data, further reducing the differences and interferences between data. Moreover, noise data and data that do not meet the format requirements will increase the computational burden of model training. By removing these data through the preprocessing step, the unnecessary computational amount during the model training process can be reduced, thereby accelerating the training speed and saving computational resources. Finally, the data set after removing noise data is more pure and consistent, which enables the model to learn more accurate and useful information during the training process. Therefore, the generalization ability of the trained model on unknown data will also be improved.
[0030] The present application also provides a training device for a key information extraction model of a set pattern, including: a first information collection unit: obtaining key information document data containing the set pattern and converting the key information document data into pictures; extracting the set pattern position information and the set pattern category information in the pictures to obtain a first data set; a second information collection unit: annotating the text position information and the text meaning information in the pictures to obtain a second data set; an information fusion unit: based on the set pattern position information and the text position information in the second data set, associating the set pattern category information in the second data set with the text meaning information to obtain a third data set; a model training unit: training the to-be-trained set pattern extraction model based on the third data set until the training of the set pattern extraction model is completed.
[0031] The present application also provides a method for training a key information extraction model for a set target object, including: Obtaining key information carrier data containing the set target object and extracting the set target object position information and the set target object category information therefrom; Annotate the relevant identification position information and relevant identification meaning information in the specific visual data form; Based on the set target object position information and relevant identification position information, associate the set target object category information with the relevant identification meaning information to obtain a data set for training the model; Train the extraction model of the set target object to be trained based on the data set for training the model until the training of the extraction model of the set target object is completed.
[0032] In this embodiment, exemplary descriptions of each step can be found in the above embodiments.
[0033] This application also provides a key information extraction method, which includes: Invoke the extraction model of the target object trained based on any embodiment of this application; Input the key information carrier data containing the target object to be extracted into the extraction model of the target object to extract the target object to be extracted therefrom.
[0034] In this embodiment, exemplary descriptions of each step can be found in the above embodiments.
[0035] Among them, "target object" includes but is not limited to "pattern", "key information carrier" includes but is not limited to various information-carrying carriers such as "document", "visual data form" includes but is not limited to the presentation form of pictures, and "relevant identification" includes but is not limited to "text".
[0036] This application also provides an electronic device, including: a memory and a processor, where a computer-executable program is stored on the memory, and the processor is used to execute the computer-executable program to implement the method described in any one of this embodiment.
[0037] This application also provides a computer storage medium, on which a computer-executable program is stored, and the computer-executable program implements the method described in any one of this embodiment when it runs.
[0038] It should be noted that for the various embodiments in this specification, the same or similar parts can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, they are described relatively simply, and the relevant parts can be referred to the partial description of the method embodiments. The device and system embodiments described above are only illustrative. The modules described as separate components may or may not be physically separated, and the components indicated as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0039] As described above, this is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A key information extraction model training method for a set pattern, characterized in that: include: Acquire key information document data containing a set pattern, and convert the key information document data into a picture; Extracting the set pattern position information and the set pattern category information in the picture to obtain a first data set; Annotating the text position information and text meaning information in the image to obtain a second data set; Based on the pattern position information and the text position information set in the second data set, associating the pattern category information set in the second data set with the text meaning information to obtain a third data set; The set pattern extraction model to be trained is trained based on the third data set until the training of the set pattern extraction model is completed.
2. The key information extraction model training method for a set pattern according to claim 1, characterized in that: The step of obtaining key information document data containing a set pattern and converting the key information document data into a picture includes: Acquire key information document data containing a set pattern, and convert the document data into a picture based on a format conversion tool.
3. The key information extraction model training method for a set pattern according to claim 1, characterized in that: The extracting the set pattern position information and the set pattern category information in the picture to obtain a first data set includes: Based on a marking tool, according to the marking of the set pattern on the picture, the set pattern position information and the set pattern category information in the picture are extracted to obtain a first data set.
4. The key information extraction model training method for a set pattern according to claim 1, characterized in that: The step of marking the text position information and the text meaning information in the picture to obtain a second data set includes: Based on the OCR tool, the corresponding text position information and text meaning information are extracted according to the annotations of the text in the image to obtain the second data set.
5. The key information extraction model training method for a set pattern according to claim 1, characterized in that: The step of associating the pattern category information set in the second data set with the text meaning information based on the pattern position information and the text position information set in the second data set to obtain a third data set includes: Based on the coordinates of the set pattern position information and the text position information in the second data set, the set pattern category information in the second data set is associated with the text meaning information to obtain a third data set.
6. The key information extraction model training method for a set pattern according to claim 5, characterized in that: The step of associating the pattern category information set in the second data set with the text meaning information based on the coordinates of the pattern position information and the text position information set in the second data set to obtain a third data set includes: The coordinates of the pattern position information and the text position information are set based on the second data set, and the text meaning information is mapped to the pattern category information for association to obtain a third data set.
7. The key information extraction model training method for a set pattern according to claim 1, characterized in that: The method comprises: The to-be-trained setting pattern detection model is trained based on the first data set, so as to extract setting pattern position information and setting pattern category information in the image based on the trained setting pattern detection model.
8. A key information extraction model training device for setting a pattern, characterized in that: include: The first information collection unit is used to obtain key information document data containing a set pattern, and convert the key information document data into a picture; Extracting the set pattern position information and the set pattern category information in the picture to obtain a first data set; A second information collection unit: annotating the text position information and text meaning information in the image to obtain a second data set; An information fusion unit: based on the pattern position information and the text position information set in the second data set, associating the pattern category information set in the second data set with the text meaning information to obtain a third data set; Model training unit: training the set pattern extraction model to be trained based on the third data set until the training of the set pattern extraction model is completed.
9. A key information extraction model training method for a set target object, characterized in that: include: Acquire key information carrier data containing a set target object, and extract the set target object location information and the set target object category information therefrom; Marking relevant identification position information and relevant identification meaning information in the specific visual data form; Based on the set target object position information and the related identification position information, the set target object category information is associated with the related identification meaning information to obtain a data set for training a model; The set target object extraction model to be trained is trained based on the data set used for the training model until the training of the set target object extraction model is completed.
10. A key information extraction method, characterized in that: include: Calling a target object extraction model trained based on claim 9; The key information carrier data containing the target object to be extracted is input into the target object extraction model to extract the target object to be extracted therefrom.