A text information processing method, device and equipment and storage medium

By extracting contextual information from financial contract images to generate new text information and training a text recognition model, the problem of poor recognition results caused by insufficient training samples is solved, achieving efficient key element detection and improved model accuracy.

CN115393870BActive Publication Date: 2026-06-02SHANGHAI PUDONG DEVELOPMENT BANK

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI PUDONG DEVELOPMENT BANK
Filing Date
2022-08-30
Publication Date
2026-06-02

Smart Images

  • Figure CN115393870B_ABST
    Figure CN115393870B_ABST
Patent Text Reader

Abstract

The application discloses a kind of text information processing method, device, equipment and storage medium.The text information processing method comprises: extracting original context information including target contract element from original contract image;According to the original context information, generate new text information including target contract element;According to the new text information and the original contract image, generate new contract image;Using the new contract image, train the character recognition model.The embodiment of the application generates new contract image according to new text information and original contract image, trains the character recognition model, improves the effect of key element detection and identification when the number of training samples is less, and improves the accuracy of the character recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of semi-supervised learning technology, and in particular to a text information processing method, apparatus, device, and storage medium. Background Technology

[0002] In the lending business of the financial industry, business personnel will repeatedly process a large number of contracts with fixed formats and diverse types every day, while ensuring the accurate identification of various key elements, which will consume a lot of human resources.

[0003] Meanwhile, due to the constraints of the financial system's business processes, real data is not disclosed to the public due to privacy concerns. As a result, image text recognition models lack a large number of real and effective samples, leading to poor performance in text image recognition. Summary of the Invention

[0004] This invention provides a text information processing method, apparatus, device, and storage medium, which can improve the effect of key element detection and recognition when the number of training samples is small, and can improve the accuracy of text recognition models.

[0005] According to one aspect of the present invention, a text information processing method is provided, the method comprising:

[0006] Extract raw contextual information, including the target contract elements, from the original contract image;

[0007] Generate new text information including the target contract elements based on the original context information;

[0008] A new contract image is generated based on the new text information and the original contract image;

[0009] The new contract image is used to train the text recognition model.

[0010] According to another aspect of the present invention, a text information processing apparatus is provided, the apparatus comprising:

[0011] The information extraction module is used to extract raw contextual information, including the target contract elements, from the original contract image;

[0012] The information generation module is used to generate new text information including the target contract elements based on the original context information;

[0013] An image generation module is used to generate a new contract image based on the new text information and the original contract image;

[0014] The model training module is used to train the character recognition model using the new contract image.

[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the text information processing method according to any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the text information processing method according to any embodiment of the present invention.

[0020] The technical solution of this invention extracts original contextual information including target contract elements from the original contract image, automatically generates new text information including target contract elements based on the original contextual information, generates a new contract image based on the new text information and the original contract image, and trains the text recognition model using the new contract image. This solves the problem of poor text image recognition performance due to a small number of training samples in the prior art, improves the key element detection and recognition performance when the number of training samples is small, and improves the accuracy of the text recognition model.

[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a flowchart of a text information processing method provided in Embodiment 1 of the present invention;

[0024] Figure 2 This is a flowchart of a text information processing method provided in Embodiment 2 of the present invention;

[0025] Figure 3A flowchart of a text information processing method provided in an embodiment of the present invention;

[0026] Figure 4 This is a schematic diagram of the structure of a text information processing device according to Embodiment 3 of the present invention;

[0027] Figure 5 This is a schematic diagram of the structure of an electronic device that implements a text information processing method according to an embodiment of the present invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] Example 1

[0031] Figure 1 This is a flowchart illustrating a text information processing method according to Embodiment 1 of the present invention. This embodiment is applicable to processing a small amount of text. The method can be executed by a text information processing device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0032] S110. Extract the original contextual information, including the target contract elements, from the original contract image.

[0033] The original contract image can be a scanned copy of a contract document with a small number of samples to be optimized, uploaded by the user through the system platform. Contract elements are important components of the contract, while target contract elements are key contents to be identified within the contract. Contract elements are determined based on the contract's usage scenario. For example, contract elements differ in scenarios such as leasing, finance, and insurance. In a leasing scenario, contract elements might include the lessor, lessee, lease term, lease recipient, and lease method. In a financial scenario, taking a loan contract as an example, contract elements might include the lender and borrower, loan date, loan amount, and loan limit. In an insurance scenario, taking a car insurance contract as an example, contract elements might include the vehicle model, insured person, and insurance start and end dates. Contract elements with a small number of training samples or whose text recognition model's performance needs optimization can be used as target contract elements.

[0034] The original contextual information refers to the contextual textual information of the target contract element within the original contract image. It can be determined as follows: OCR (Optical Character Recognition) is performed on the original contract image to obtain the corresponding contract text information; the original contextual information of the target contract element is then extracted from the contract text information. The original contextual information can be the content preceding, following, or both preceding and following the target contract element. Furthermore, the text at the target contract element can be truncated according to a set original contextual length to obtain the original contextual information.

[0035] S120. Generate new text information including the target contract elements based on the original context information.

[0036] Both the new text information and the original context information include the target contract elements. The new text information also includes new contextual information about the target contract elements. It should be noted that the new elements of the target contract elements in the new text information differ from the original elements of the target contract elements in the original context information. Taking the time period as an example, the original context information could be "Party A's loan term starts from a certain year, month, and day," while the new context information could be "Party B's loan term starts from a certain year, month, and day." This embodiment does not specifically limit the generation method of the new text information; for example, it can be generated based on semantic rules or based on a text generation network model. The text generation network model can be a Seq2Seq text generation network model. Seq2Seq is a variant of a recurrent neural network, consisting of an encoder and a decoder. Seq2Seq is an important model in natural language processing and can be used for machine translation, dialogue systems, and automatic summarization. By automatically generating new text information based on the original context information, a foundation is laid for generating new contract images based on the new text information, thereby increasing the diversity of contract images to which contract elements belong, and thus enhancing the quality of processing contract elements based on contract images.

[0037] S130. Generate a new contract image based on the new text information and the original contract image.

[0038] The new text information includes textual information that incorporates the elements of the target contract. The new contract image is used to train the character recognition model.

[0039] Optionally, generating a new contract image based on the new text information and the original contract image includes: replacing the original context information in the original contract image with the new text information to obtain a replaced image; and performing illumination transformation and / or projection transformation on the replaced image to obtain a new contract image.

[0040] Specifically, by replacing the original context information in the original contract image with new text information, a large number of training sample images based on the original context information can be generated, increasing the number of training sample images. Performing illumination transformation and / or projection transformation on the replaced image to obtain a new contract image can increase the variety of training sample images, laying the foundation for training the accuracy of the text recognition model.

[0041] Optionally, before generating a new contract image based on the new text information and the original contract image, the method further includes: determining the text length deviation between the new text information and the original context information; determining the fluency confidence level of the new text information; and filtering the new text information based on the text length deviation and / or the fluency confidence level.

[0042] Specifically, fluency confidence score represents the fluency of the new text information; text length deviation can be determined based on the length of the new text information and the length of the original context information. A fluency confidence score threshold and a text length deviation threshold can be preset. If the text length deviation between the new text information and the original context information is greater than the text length deviation threshold, the new text information is discarded; if the fluency confidence score of the new text information is less than the fluency confidence score threshold, the new text information is also discarded. By filtering new text information based on text length deviation and fluency confidence score, the quality of the retained new text information is improved, thereby enhancing the performance of the text recognition model.

[0043] S140. Using the new contract image, train the text recognition model.

[0044] Specifically, the text recognition model is used to perform text recognition on contract elements to be identified in a contract image. It can identify the image position of the contract element in the contract image and the text content of the contract element. The image position can be displayed through the coordinates of the bounding box. This embodiment of the disclosure does not specifically limit the network structure of the text recognition model; for example, a convolutional neural network can be used. The text generation model is trained by leveraging the characteristics of the original context information, and new text information with semantically similar meaning is generated based on the text generation model for the original context information. A new contract image is constructed based on the new text information and the original contract image as incremental training data for the text recognition model, which is used to optimize the text recognition model. This greatly improves the limitations of the text recognition model in the data domain and increases the upper limit of the OCR model's performance.

[0045] Optionally, training the text recognition model using the new contract image includes: determining a new text label for the target contract element in the new contract image based on the original text label of the target contract element in the original contract image; performing text recognition on the new contract image using the text recognition model to obtain a new text recognition result; extracting predicted text information of the target contract element from the new text recognition result, and generating a loss function based on the predicted text information and the new text label; and training the text recognition model based on the loss function.

[0046] The new text label can include the new image position of the target contract element in the new contract image and the new text content of the contract element. Specifically, the new text content can be extracted from the new text information, and the original image position of the target contract element in the original contract image can be adaptively adjusted to obtain the new image position. It should be noted that the adaptive adjustment method corresponds to the way the new text information replaces the original context information in the original contract image. For example, the original image position of the target contract element in the original contract image can be adaptively translated and rotated.

[0047] Optionally, after training the text recognition model using the new contract image, the method further includes: inputting an unlabeled auxiliary contract image into the trained text recognition model to obtain auxiliary text recognition results; determining auxiliary labels for auxiliary contract elements in the auxiliary contract image based on the auxiliary text recognition results; obtaining verification results for the auxiliary labels; and incrementally training the text recognition model based on the auxiliary contract image and the verification results.

[0048] Specifically, by inputting unlabeled auxiliary contract images into a trained text recognition model, determining auxiliary labels based on the auxiliary text recognition results, and performing incremental training on the text recognition model based on unsupervised learning according to the verification results of the auxiliary contract images and auxiliary labels, the training efficiency of the text recognition model can be improved, the training samples of the text recognition model can be enhanced, and the performance of the text recognition model can be further improved.

[0049] The technical solution of this invention extracts original contextual information including target contract elements from the original contract image, automatically generates new text information including the target contract elements based on the original contextual information, generates a new contract image based on the new text information and the original contract image, and trains the text recognition model using the new contract image, thus realizing incremental learning technology for text recognition models based on semi-supervised learning. By using a small number of labeled original contract images to guide a large amount of unlabeled training data, combined with text generation technology, incremental training of the text recognition model can be achieved, and the customized optimization function of the text recognition model for specific contract elements can also be realized.

[0050] Example 2

[0051] Figure 2 This is a flowchart of a text information processing method provided in Embodiment 2 of the present invention. Based on the above embodiments, this embodiment further optimizes and expands the generation of new text information including target contract elements from the original context information, and can be combined with the above optional implementation methods. The step of generating new text information including target contract elements from the original context information includes: inputting the original context information into the embedding layer of a text generation network model to obtain an original context representation; performing feature encoding on the original context representation through the encoding layer of the text generation network model to obtain original context features; processing the original context features through the perceptron in the text generation network model to obtain processed original context features; and decoding the processed original context features through the decoder in the text generation network model to obtain new text information including target contract elements. Figure 2 As shown, the method includes:

[0052] S210. Extract the original contextual information, including the target contract elements, from the original contract image.

[0053] S220. Input the original context information into the embedding layer of the text generation network model to obtain the original context representation.

[0054] The input text length for text generation network models is limited, so the original context information needs to be truncated to a suitable length before being used as input to the text generation network model.

[0055] S230. The original context representation is feature-encoded through the encoding layer in the text generation network model to obtain the original context features.

[0056] S240. The original context features are processed by the perceptron in the text generation network model to obtain the processed original context features;

[0057] S250. The original context features that have been processed are decoded by the decoder in the text generation network model to obtain new text information including the target contract elements.

[0058] S260. Generate a new contract image based on the new text information and the original contract image.

[0059] S270. Using the new contract image, train the text recognition model.

[0060] The text generation network model can utilize massive corpora and be pre-trained on Seq2Seq, providing it with a vast corpus foundation. This allows the model to generate any number of semantically similar new text elements from the original contract image for various contract application scenarios, serving as incremental training data. The text generation network model includes an encoder, a perceptron, and a decoder. The perceptron can employ a fully connected neural network structure. The encoder encodes the original context representation to obtain original context features, which are then decoded by the decoder to obtain new text information. Furthermore, before decoding, the perceptron processes the original context features to better reflect the semantic features of the original context information, thereby improving the matching degree between the new text information and the original context information.

[0061] The following specific embodiment will be used to illustrate a text information processing method according to the above embodiment. See details below. Figure 3 .

[0062] refer to Figure 3The original contextual information, including the target contract elements, is extracted from the scanned contract. This extracted contextual information is then input into a Seq2Seq text generation network model to generate new text information containing the target contract elements. Based on the generated new text information and the original contract image, a new contract image is generated. The new contract image is preprocessed and used to train the text recognition model. Subsequently, the text recognition model is optimized based on semi-supervised training, and sample data is continuously supplemented during the training process.

[0063] The technical solution of this embodiment extracts original contextual information including target contract elements from the original contract image, inputs the original contextual information into the embedding layer of the text generation network model to obtain the original contextual representation, encodes the original contextual representation through the encoding layer of the text generation network model to obtain the original contextual features, processes the original contextual features through the perceptron of the text generation network model to obtain the processed original contextual features, decodes the processed original contextual features through the decoder of the text generation network model to obtain a large amount of new textual information including the target contract elements, generates a new contract image based on the new textual information and the original contract image, and uses the new contract image to train the text recognition model to optimize the effect of the target contract elements, thereby improving the key element detection and recognition effect when the number of training samples is small, and improving the accuracy of the text recognition model.

[0064] Example 3

[0065] Figure 4 This is a schematic diagram of the structure of a text information processing device provided in Embodiment 3 of the present invention. Figure 4 As shown, the device includes:

[0066] Information extraction module 410 is used to extract original contextual information, including target contract elements, from the original contract image;

[0067] Information generation module 420 is used to generate new text information including target contract elements based on the original context information;

[0068] Image generation module 430 is used to generate a new contract image based on the new text information and the original contract image;

[0069] The model training module 440 is used to train the character recognition model using the new contract image.

[0070] The technical solution of this invention extracts original contextual information including target contract elements from the original contract image, automatically generates new text information including the target contract elements based on the original contextual information, generates a new contract image based on the new text information and the original contract image, and trains the text recognition model using the new contract image, thus realizing incremental learning technology for text recognition models based on semi-supervised learning. By using a small number of labeled original contract images to guide a large amount of unlabeled training data, combined with text generation technology, incremental training of the text recognition model can be achieved, and the customized optimization function of the text recognition model for specific contract elements can also be realized.

[0071] Furthermore, the information generation module 420 includes:

[0072] The original context representation unit is used to input the original context information into the embedding layer in the text generation network model to obtain the original context representation;

[0073] The original context feature unit is used to encode the original context representation through the encoding layer in the text generation network model to obtain the original context features;

[0074] The feature processing unit is used to process the original context features through the perceptron in the text generation network model to obtain the processed original context features;

[0075] The new text information unit is used to decode the processed original context features through the decoder in the text generation network model to obtain new text information including the target contract elements.

[0076] Furthermore, the image generation module 430 includes:

[0077] An information replacement unit is used to replace the original context information in the original contract image with the new text information to obtain a replaced image;

[0078] The image transformation unit is used to perform illumination transformation and / or projection transformation on the replaced image to obtain a new contract image.

[0079] Furthermore, the model training module 440 includes:

[0080] The new text label determination unit is used to determine the new text label of the target contract element in the new contract image based on the original text label of the target contract element in the original contract image;

[0081] The text recognition unit is used to perform text recognition on the new contract image using the text recognition model to obtain a new text recognition result.

[0082] The loss function generation unit is used to extract the predicted text information of the target contract element from the new text recognition result, and generate a loss function based on the predicted text information and the new text label;

[0083] The model training unit is used to train the character recognition model according to the loss function.

[0084] Furthermore, before generating a new contract image based on the new text information and the original contract image, the process also includes:

[0085] Determine the text length deviation between the new text information and the original context information;

[0086] Determine the fluency confidence level of the new text information;

[0087] The new text information is filtered based on the text length deviation and / or fluency confidence level.

[0088] Furthermore, after training the character recognition model using the new contract image, the process also includes:

[0089] Input the unlabeled auxiliary contract image into the trained text recognition model to obtain the auxiliary text recognition result;

[0090] Based on the auxiliary text recognition results, auxiliary labels for auxiliary contract elements in the auxiliary contract image are determined;

[0091] Obtain the verification result of the auxiliary label, and perform incremental training on the text recognition model based on the auxiliary contract image and the verification result.

[0092] The text information processing device provided in the embodiments of the present invention can execute the text information processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0093] Example 4

[0094] Figure 5 A schematic diagram of an electronic device 50 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0095] like Figure 5 As shown, the electronic device 50 includes at least one processor 51 and a memory, such as a read-only memory (ROM) 52 or a random access memory (RAM) 53, communicatively connected to the at least one processor 51. The memory stores computer programs executable by the at least one processor. The processor 51 can perform various appropriate actions and processes based on the computer program stored in the ROM 52 or loaded from storage unit 58 into the RAM 53. The RAM 53 can also store various programs and data required for the operation of the electronic device 50. The processor 51, ROM 52, and RAM 53 are interconnected via a bus 54. An input / output (I / O) interface 55 is also connected to the bus 54.

[0096] Multiple components in electronic device 50 are connected to I / O interface 55, including: input unit 56, such as keyboard, mouse, etc.; output unit 57, such as various types of monitors, speakers, etc.; storage unit 58, such as disk, optical disk, etc.; and communication unit 59, such as network card, modem, wireless transceiver, etc. Communication unit 59 allows electronic device 50 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0097] Processor 51 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 51 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 51 performs the various methods and processes described above, such as a text information processing method.

[0098] In some embodiments, text information processing may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 58. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 50 via ROM 52 and / or communication unit 59. When the computer program is loaded into RAM 53 and executed by processor 51, one or more steps of the text information processing described above may be performed. Alternatively, in other embodiments, processor 51 may be configured to perform the text information processing method by any other suitable means (e.g., by means of firmware).

[0099] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems-on-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0100] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0101] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0102] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (Cathode Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0103] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0104] A computing system can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability.

[0105] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0106] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A text information processing method, characterized in that, include: Extract raw contextual information, including the target contract elements, from the original contract image; The original context information is input into the embedding layer of the text generation network model to obtain the original context representation; the original context representation is then encoded by the encoding layer of the text generation network model to obtain the original context features; the original context features are then processed by the perceptron of the text generation network model to obtain the processed original context features; the processed original context features are then decoded by the decoder of the text generation network model to obtain new text information including the target contract elements; wherein, the perceptron adopts a fully connected neural network structure; A new contract image is generated based on the new text information and the original contract image; The new contract image is used to train the text recognition model.

2. The method according to claim 1, characterized in that, The step of generating a new contract image based on the new text information and the original contract image includes: The original context information in the original contract image is replaced with the new text information to obtain a replaced image; Perform illumination transformation and / or projection transformation on the replaced image to obtain a new contract image.

3. The method according to claim 1, characterized in that, The step of training the text recognition model using the new contract image includes: Based on the original text labels of the target contract element in the original contract image, determine the new text labels of the target contract element in the new contract image; The new contract image is subjected to text recognition using the text recognition model to obtain the new text recognition result; The predicted text information of the target contract elements is extracted from the new text recognition results, and a loss function is generated based on the predicted text information and the new text label. The character recognition model is trained based on the loss function.

4. The method according to claim 1, characterized in that, Before generating the new contract image based on the new text information and the original contract image, the process further includes: Determine the text length deviation between the new text information and the original context information; Determine the fluency confidence level of the new text information; The new text information is filtered based on the text length deviation and / or fluency confidence level.

5. The method according to claim 1, characterized in that, After training the text recognition model using the new contract image, the process further includes: Input the unlabeled auxiliary contract image into the trained text recognition model to obtain the auxiliary text recognition result; Based on the auxiliary text recognition results, auxiliary labels for auxiliary contract elements in the auxiliary contract image are determined; Obtain the verification result of the auxiliary label, and perform incremental training on the text recognition model based on the auxiliary contract image and the verification result.

6. A text information processing device, characterized in that, include: The information extraction module is used to extract raw contextual information, including the target contract elements, from the original contract image; The information generation module includes: an original context representation unit, an original context feature unit, a feature processing unit, and a new text information unit; The original context representation unit is used to input the original context information into the embedding layer in the text generation network model to obtain the original context representation. The original context feature unit is used to encode the original context representation through the encoding layer in the text generation network model to obtain the original context features; The feature processing unit is used to process the original context features through the perceptron in the text generation network model to obtain the processed original context features; The new text information unit is used to decode the processed original context features through the decoder in the text generation network model to obtain new text information including the target contract elements; An image generation module is used to generate a new contract image based on the new text information and the original contract image; The model training module is used to train the character recognition model using the new contract image.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the text information processing method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the text information processing method according to any one of claims 1-5.