A method, apparatus and related device for processing text data
By segmenting text information and using recursive attention network and memory network for iterative training, the long-term forgetting problem in the text recognition method based on deep learning is solved, and the accuracy of text recognition is improved.
Patent Information
- Application Number
- CN202110400491.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-04-14
AI Technical Summary
Text recognition methods based on deep learning, especially recurrent neural networks, have long-term forgetting problems, resulting in reduced text recognition accuracy, especially when text content is long, information loss is serious.
A text data processing method is proposed, by obtaining sample pictures carrying sample training tags, text segmentation of training text information, and obtaining the first sample fragment and the second sample fragment of the initial network model. Then, the sample text features and image features are input into the recursive attention network, and iteratively trained through the memory network to generate a target network model for text recognition.
Through the combination of text segmentation and memory network, the problem of long-term forgetting is effectively solved, and the accuracy of text recognition is improved, especially when processing long texts, reducing information loss.
Smart Images

Figure CN113705552B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method, apparatus, and related device for processing text data. Background Art
[0002] Currently, in the process of text recognition for certain pictures, text recognition methods based on deep learning are often used. However, in the process of using a deep learning-based recognition method (for example, using a recurrent neural network) to recognize text, the accuracy of text recognition is often limited by the length of the text content in the text area.
[0003] For example, once the text content in the text area of the entire picture is too long, due to the inherent property problems of the recurrent neural network itself, information loss will occur in the text content in the second half of the text area, and the information loss will become more and more serious as the length of the text content increases. It can be seen that in the process of using a deep learning-based recognition method (for example, using a recurrent neural network) to recognize text, the problem of long-term forgetting will occur, which will lead to the phenomenon of text recognition errors, thus reducing the accuracy of text recognition in pictures. Summary of the Invention
[0004] This application provides a method, apparatus, and related device for processing text data, which can improve the accuracy of text recognition for pictures.
[0005] On the one hand, an embodiment of this application provides a method for processing text data, including:
[0006] Obtain a sample picture carrying a sample training label, perform text segmentation on the training text information indicated by the training sample label to obtain a first sample fragment and a second sample fragment for training an initial network model; the second sample fragment is the next sample fragment of the first sample fragment;
[0007] Input the first sample text feature of the first sample fragment and the sample image feature of the sample picture into the recurrent attention network in the initial network model, determine the first text-image relationship feature between the first sample fragment and the sample picture through the recurrent attention network, output the first predicted text of the first sample fragment based on the first text-image relationship feature, and add the first predicted text to the memory network in the initial network model;
[0008] Use the first predicted text stored in the memory network as the training auxiliary text for the second sample fragment. Input the first sample text feature, the second sample text feature, and the sample image feature corresponding to the training auxiliary text into the recurrent attention network. Determine the second text-image relationship feature between the second sample fragment and the sample picture through the recurrent attention network. Output the second predicted text of the second sample fragment based on the second text-image relationship feature, and add the second predicted text to the memory network;
[0009] Based on the first predicted text and the second predicted text in the memory network, determine the sample prediction label of the training text information. Based on the sample training label and the sample prediction label, perform iterative training on the initial network model, and use the iteratively trained initial network model as the target network model for text recognition of the target picture.
[0010] One aspect of the embodiments of the present application provides a text data processing device, including:
[0011] A sample picture acquisition module, configured to acquire a sample picture carrying a sample training label, perform text segmentation on the training text information indicated by the training sample label to obtain a first sample fragment and a second sample fragment for training the initial network model; the second sample fragment is the next sample fragment of the first sample fragment;
[0012] A first relationship determination module, configured to input the first sample text feature of the first sample fragment and the sample image feature of the sample picture into the recurrent attention network in the initial network model, determine the first text-image relationship feature between the first sample fragment and the sample picture through the recurrent attention network, output the first predicted text of the first sample fragment based on the first text-image relationship feature, and add the first predicted text to the memory network in the initial network model;
[0013] A second relationship determination module, configured to use the first predicted text stored in the memory network as the training auxiliary text for the second sample fragment, input the first sample text feature, the second sample text feature, and the sample image feature corresponding to the training auxiliary text into the recurrent attention network, determine the second text-image relationship feature between the second sample fragment and the sample picture through the recurrent attention network, output the second predicted text of the second sample fragment based on the second text-image relationship feature, and add the second predicted text to the memory network;
[0014] A model training module, configured to determine the sample prediction label of the training text information based on the first predicted text and the second predicted text in the memory network, perform iterative training on the initial network model based on the sample training label and the sample prediction label, and use the iteratively trained initial network model as the target network model for text recognition of the target picture.
[0015] Among them, the sample picture acquisition module includes:
[0016] A text segmentation unit, configured to obtain a sample picture carrying a sample training label, segment the training text information indicated by the training sample label based on text segmentation parameters, and obtain a sample fragment set of the training text information; the sample fragment set is used to store all sample fragments obtained by segmenting the training text information; one sample fragment corresponds to one sample segmentation identifier; one sample segmentation identifier is used to represent the fragment position of a sample fragment in the training text information;
[0017] A fragment acquisition unit, configured to obtain a first sample fragment for training an initial network model from the sample fragment set based on the sample segmentation identifier of each sample fragment, and use the next sample fragment of the first sample fragment as a second sample fragment based on the fragment position of the first sample fragment in the training text information.
[0018] Wherein, the initial network model includes a first sample branch model and a second sample branch model; the second sample branch model includes a semantic extraction network; the apparatus further includes:
[0019] An image feature extraction module, configured to extract sample image features of the sample picture through the first sample branch model;
[0020] A text feature determination module, configured to determine a first sample text feature of the first sample fragment through the semantic extraction network, and determine a second sample text feature of the second sample fragment through the semantic extraction network.
[0021] Wherein, the text feature determination module includes:
[0022] A training feature extraction unit, configured to extract training text features of the training text information through the semantic extraction network;
[0023] A fragment feature determination unit, configured to determine a fragment text feature of the first sample fragment in the training text features based on the fragment position of the first sample fragment in the training text information, and determine a fragment text feature of the second sample fragment in the training text features based on the fragment position of the second sample fragment in the training text information;
[0024] A first feature determination unit, configured to use the fragment position of the first sample fragment and the element position of the fragment element in the first sample fragment as the first relative coding position information of the first sample fragment, perform relative coding on the first sample fragment based on the first relative coding position information to obtain a relative position feature of the first sample fragment, and obtain a first sample text feature of the first sample fragment based on the fragment text feature of the first sample fragment and the relative position feature of the first sample fragment;
[0025] A second feature determination unit, configured to use the fragment position of the second sample fragment and the element position of the fragment element in the second sample fragment as the second relative coding position information of the first sample fragment, perform relative coding on the second sample fragment based on the second relative coding position information to obtain the relative position feature of the second sample fragment, and obtain the second sample text feature of the second sample fragment based on the fragment text feature and the relative position feature of the second sample fragment.
[0026] Wherein, the initial network model includes a first sample branch model and a second sample branch model; the second sample branch model includes a semantic extraction network, a recursive attention network, and a memory network; the first sample branch model is used to extract the sample image feature of the sample picture; the semantic extraction network is used to extract the first sample text feature of the first sample fragment; the fragment elements of the first sample fragment include a first fragment element and a second fragment element; the second fragment element is the next fragment element of the first fragment element.
[0027] The first relationship determination module includes:
[0028] A first auxiliary feature determination unit, configured to obtain a first element auxiliary feature associated with the first fragment element from the first sample text feature of the first sample fragment based on the element position of the first fragment element in the first sample fragment.
[0029] A first element output unit, configured to determine a first weight relationship feature between the first fragment element and the sample picture based on the first element auxiliary feature, the sample image feature of the sample picture, the first attention layer in the recursive attention network, and the first selection gate corresponding to the first attention layer, input the first weight relationship feature into the first language layer in the recursive attention network, and output a first predicted element corresponding to the first fragment element by the first language layer.
[0030] A second element output unit, configured to recursively use the first element feature of the first predicted element as a second element auxiliary feature associated with the second fragment element, determine a second weight relationship feature between the second fragment element and the sample picture based on the second element auxiliary feature, the sample image feature, the second attention layer in the recursive attention network, and the second selection gate corresponding to the second attention layer, input the second weight relationship feature into the second language layer in the recursive attention network, and output a second predicted element corresponding to the second fragment element by the second language layer.
[0031] A predicted text output unit, configured to determine a first text-image relationship feature between the first sample fragment and the sample picture based on the first weight relationship feature and the second weight relationship feature, and output a first predicted text of the first sample fragment based on the first text-image relationship feature; the first predicted text includes the first predicted element and the second predicted element.
[0032] A first text addition unit for adding a first predicted text including a first predicted element and a second predicted element to a memory network in an initial network model.
[0033] Among them, the first element output unit includes:
[0034] A first input subunit for using a first element auxiliary feature and a sample image feature of a sample picture as a first input feature associated with a first sample fragment;
[0035] A first attention determination subunit for inputting the first input feature into a first attention layer in a recursive attention network, and determining, by the first attention layer, a first attention map feature of a first fragment element in the sample picture;
[0036] A first weight selection subunit for inputting the first attention map feature and the first element auxiliary feature into a first selection gate corresponding to the first attention layer, and selecting and outputting, by the first selection gate, a first weight relationship feature between the first fragment element and the sample picture;
[0037] A first recognition subunit for inputting the first weight relationship feature into a first language layer in the recursive attention network, and performing element recognition on the first fragment element by the first language layer to obtain a first predicted element corresponding to the first fragment element.
[0038] Among them, the first weight selection subunit includes:
[0039] A first feature extraction subunit for inputting the first attention map feature and the first element auxiliary feature into a first selection gate corresponding to the first attention layer, extracting a first inner product similarity between the first attention map feature and the first element auxiliary feature through a first selection rule in the first selection gate, and taking the product of a first weight coefficient indicated by the first selection rule and the first inner product similarity as a first similarity feature;
[0040] A second feature extraction subunit for extracting a first Gaussian similarity between the first attention map feature and the first element auxiliary feature through a second selection rule in the first selection gate, and taking the product of a second weight coefficient indicated by the second selection rule and the first Gaussian similarity as a second similarity feature;
[0041] A third feature extraction subunit for extracting a first string similarity between the first attention map feature and the first element auxiliary feature through a third selection rule in the first selection gate, and taking the product of a third weight coefficient indicated by the third selection rule and the first string similarity as a third similarity feature;
[0042] A first matrix determination subunit, configured to determine a first weight matrix between a first fragment element and a sample picture based on a first similarity feature, a second similarity feature, and a third similarity feature, and perform masked multiplication on the first weight matrix and a first attention map feature to obtain a first weight relationship feature between the first fragment element and the sample picture.
[0043] Wherein, the first matrix determination subunit includes:
[0044] A feature fusion subunit, configured to fuse the first similarity feature, the second similarity feature, and the third similarity feature to obtain a first fusion feature;
[0045] A feature normalization subunit, configured to perform normalization processing on the first fusion feature to obtain a first weight matrix between the first fragment element and the sample picture;
[0046] A masked multiplication subunit, configured to perform masked multiplication on the first weight matrix and the first attention map feature to obtain a first weight relationship feature between the first fragment element and the sample picture.
[0047] Wherein, the second element output unit includes:
[0048] A second input subunit, configured to recursively use the first element feature of the first predicted element as a second element auxiliary feature associated with a second fragment element, and use the second element auxiliary feature and a sample image feature as second input features associated with a first sample fragment;
[0049] A second attention determination subunit, configured to input the second input features into a second attention layer in a recursive attention network, and the second attention layer determines a second attention map feature of the second fragment element in the sample picture;
[0050] A second weight selection subunit, configured to input the second attention map feature and the second element auxiliary feature into a second selection gate corresponding to the second attention layer, and the second selection gate selects and outputs a second weight relationship feature between the second fragment element and the sample picture;
[0051] A second recognition subunit, configured to input the second weight relationship feature into a second language layer in the recursive attention network, and the second language layer performs element recognition on the second fragment element to obtain a second predicted element corresponding to the second fragment element.
[0052] Wherein, the second weight selection subunit includes:
[0053] The fourth feature extraction subunit is configured to input the second attention map feature and the second element auxiliary feature into the second selection gate corresponding to the second attention layer, extract the second inner product similarity between the second attention map feature and the second element auxiliary feature through the first selection rule in the second selection gate, and use the product between the first weight coefficient indicated by the first selection rule and the second inner product similarity as the fourth similarity feature;
[0054] The fifth feature extraction subunit is configured to extract the second Gaussian similarity between the second attention map feature and the second element auxiliary feature through the second selection rule in the second selection gate, and use the product between the second weight coefficient indicated by the second selection rule and the second Gaussian similarity as the fifth similarity feature;
[0055] The sixth feature extraction subunit is configured to extract the second string similarity between the second attention map feature and the second element auxiliary feature through the third selection rule in the second selection gate, and use the product between the third weight coefficient indicated by the third selection rule and the second string similarity as the sixth similarity feature;
[0056] The second matrix determination subunit is configured to determine the second weight matrix between the second fragment element and the sample picture based on the fourth similarity feature, the fifth similarity feature, and the sixth similarity feature, and perform masked multiplication on the second weight matrix and the second attention map feature to obtain the second weight relationship feature between the second fragment element and the sample picture.
[0057] Wherein, the fragment lengths of the first sample fragment and the second sample fragment are both determined by the text segmentation parameters of the initial network model; the first predicted text is determined when the second sample branch model counts that the cumulative number of predicted elements output by the language layer in the recursive attention network reaches the text segmentation parameters; the predicted element is determined after the recursive attention network associates the fragment elements in the first sample fragment with the sample picture;
[0058] The second relationship determination module includes:
[0059] The auxiliary text determination unit is configured to use the first predicted text stored in the memory network as the training auxiliary text of the second sample fragment, and use the first sample text feature corresponding to the training auxiliary text as the training auxiliary text feature;
[0060] The feature association unit is configured to input the training auxiliary text feature, the second sample text feature, and the sample image feature into the recursive attention network, and associate the second sample fragment with the sample picture through the recursive attention network to obtain the second text-image relationship feature for characterizing the association relationship between the second sample fragment and the sample picture;
[0061] A prediction element output unit, configured to input the second graphic-text relationship feature into a language layer in a recursive attention network, and output, by the language layer in the recursive attention network, a prediction element associated with a fragment element in the second sample fragment;
[0062] A second text addition unit, configured to, when detecting that the number of prediction elements output by the language layer in the recursive attention network reaches a text segmentation parameter, output a second predicted text of the second sample fragment based on the prediction elements output by the language layer in the recursive attention network, and add the second predicted text to the memory network.
[0063] An embodiment of the present application provides a text data processing method on the one hand, including:
[0064] Obtain a target picture carrying target text information, and extract a target image feature of the target picture;
[0065] Determine a first association relationship feature between a first target fragment in the target text information and the target picture based on the target image feature, and output a first target text of the first target fragment based on the first association relationship feature;
[0066] Use the first target text as a target auxiliary text of a second target fragment in the target text information, determine a second association relationship feature between the second target fragment and the target picture based on the first target text feature corresponding to the target auxiliary text and the target image feature, and output a second target text of the second target fragment based on the second association relationship feature;
[0067] Determine the target text information recognized from the target picture based on the first target text and the second target text.
[0068] Wherein, obtaining a target picture carrying target text information and extracting a target image feature of the target picture includes:
[0069] Obtain a target picture carrying target text information, and extract a target image feature of the target picture through a first target branch model in the target network model; the target network model is obtained by training an initial network model with sample pictures carrying sample training labels.
[0070] Wherein, the target network model further includes a second target branch model juxtaposed to the first target branch model, and the second target branch model includes a recursive attention network;
[0071] Determining a first association relationship feature between a first target fragment in the target text information and the target picture based on the target image feature, and outputting a first target text of the first target fragment based on the first association relationship feature includes:
[0072] Input the target image features into the recursive attention network, determine the first correlation relationship feature between the first target fragment in the target text information and the target picture through the recursive attention network, and output the first target text of the first target fragment based on the first correlation relationship feature.
[0073] Optionally, the second target branch model includes a memory network; the method further includes: adding the first target text to the memory network in the target network model.
[0074] Wherein, taking the first target text as the target auxiliary text of the second target fragment in the target text information, determining the second correlation relationship feature between the second target fragment and the target picture based on the first target text feature corresponding to the target auxiliary text and the target image features, and outputting the second target text of the second target fragment based on the second correlation relationship feature includes:
[0075] Taking the first target text stored in the memory network as the target auxiliary text of the second target fragment in the target text information, inputting the first target text feature corresponding to the target auxiliary text and the target image features into the recursive attention network, determining the second correlation relationship feature between the second target fragment and the target picture through the recursive attention network, and outputting the second target text of the second target fragment based on the second correlation relationship feature.
[0076] Optionally, the method further includes: adding the second target text to the memory network.
[0077] An embodiment of the present application provides a text data processing device on the one hand, including:
[0078] A target picture acquisition module, configured to acquire a target picture carrying target text information and extract the target image features of the target picture;
[0079] A first text output module, configured to determine the first correlation relationship feature between the first target fragment in the target text information and the target picture based on the target image features, and output the first target text of the first target fragment based on the first correlation relationship feature;
[0080] A second text output module, configured to take the first target text as the target auxiliary text of the second target fragment in the target text information, determine the second correlation relationship feature between the second target fragment and the target picture based on the first target text feature corresponding to the target auxiliary text and the target image features, and output the second target text of the second target fragment based on the second correlation relationship feature;
[0081] A target text determination module, configured to determine the target text information recognized from the target picture based on the first target text and the second target text.
[0082] Among them, the target image acquisition module is specifically configured to acquire a target image carrying target text information, and extract target image features of the target image through a first target branch model in the target network model; the target network model is obtained by training an initial network model with sample images carrying sample training labels.
[0083] Among them, the target network model further includes a second target branch model juxtaposed to the first target branch model, and the second target branch model includes a recursive attention network;
[0084] The first text output module is specifically configured to input the target image features into the recursive attention network, determine a first association relationship feature between a first target fragment in the target text information and the target image through the recursive attention network, and output a first target text of the first target fragment based on the first association relationship feature.
[0085] Optionally, the second target branch model includes a memory network;
[0086] The first text output module is further specifically configured to add the first target text to the memory network in the target network model.
[0087] The second text output module is specifically configured to use the first target text stored in the memory network as a target auxiliary text for a second target fragment in the target text information, input a first target text feature corresponding to the target auxiliary text and the target image features into the recursive attention network, determine a second association relationship feature between the second target fragment and the target image through the recursive attention network, and output a second target text of the second target fragment based on the second association relationship feature.
[0088] Optionally, the second text output module is further specifically configured to add the second target text to the memory network.
[0089] Among them, the target text determination module is specifically configured to determine the target text information recognized from the target image based on the first target text and the second target text in the memory network.
[0090] On the one hand, an embodiment of the present application provides a computer device, which includes: a processor and a memory;
[0091] The processor is connected to the memory. Among them, the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method in any aspect of the embodiment of the present application.
[0092] On the one hand, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. The computer program is suitable for being loaded and executed by a processor so that a computer device with a processor executes the method in any aspect of the embodiment of the present application.
[0093] On the one hand, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method in any aspect of the embodiment of the present application.
[0094] The computer device in the embodiments of the present application can obtain a sample picture carrying a sample training label, and then can perform text segmentation on the training text information indicated by the sample training label to obtain a first sample fragment and a second sample fragment for training the initial network model. It should be understood that since the training text information in the sample picture is known when the embodiments of the present application obtain the sample picture, text segmentation can be performed on the training text information during the model training stage to achieve fragmentation processing of training text information of any length, so as to improve the generalization ability of the trained target network model when using the trained initial network model (i.e., the target network model) to perform text recognition on the target text information of any length in the target picture during the subsequent model application stage. It should be understood that the first sample fragment and the second sample fragment here are any two sample fragments with adjacent fragment positions obtained after text segmentation of the training text information. For example, the second sample fragment is the next sample fragment of the first sample fragment. Further, after the computer device gives the first sample text feature of the first sample fragment and the sample picture feature of the sample picture to the recursive attention network in the initial network model, the recursive attention network can associate the first sample fragment with the sample picture to obtain a first text-picture relationship feature between the first sample fragment and the sample picture. At this time, the computer device can output a first predicted text of the first sample fragment through the first text-picture relationship feature, and then the output first predicted text can participate in the text prediction of the next sample fragment (i.e., the second sample fragment). For example, the computer device can use the first predicted text in the memory network as an auxiliary text, and then input the first sample feature corresponding to the auxiliary text, the second sample text feature of the second sample fragment, and the aforementioned sample image feature into the recursive attention network together, so that the recursive attention network can associate the second sample fragment with the sample picture to obtain a second text-picture relationship feature between the second sample fragment and the sample picture. Further, the computer device can output a second predicted text of the second sample fragment based on the second text-picture relationship feature, and then add the second predicted text to the aforementioned memory network, so that the memory network can subsequently determine a sample prediction label of the training text information based on the added first predicted text and second predicted text. Finally, the computer device can perform iterative training on the initial network model based on the sample training label and the sample prediction label, and use the iteratively trained initial network model as the target network model for text recognition of the target picture. It should be understood that the embodiments of the present application can extract the text-picture dependence relationship between each sample fragment and the sample picture through the recursive attention network, and the text-picture dependence relationship at least includes the above first text-picture relationship feature and second text-picture relationship feature.Therefore, when outputting each predicted text through the image-text dependency relationship, not only can the image-text features be associated through the recursive attention network, but the problem of long-term forgetting can also be solved from the root through the introduced memory network. For example, the current fragment text output by the computer device can be used as the training auxiliary text for the next fragment text through the memory network, and then the contextual relationship between the fragments can be fully utilized through the memory network to improve the accuracy of text recognition in images. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0096] Figure 1 It is a structural diagram of a network architecture provided by an embodiment of the present application;
[0097] Figure 2 This is a schematic diagram of a scenario for data interaction provided by an embodiment of the present application;
[0098] Figure 3 It is a flowchart of a text data processing method provided in an embodiment of the present application;
[0099] Figure 4 This is a schematic diagram of a scenario for text segmentation provided in an embodiment of the present application;
[0100] Figure 5 is a schematic diagram of a scenario for determining sample text features of sample fragments provided by an embodiment of the present application;
[0101] Figure 6 is a schematic diagram of a scenario of outputting a first predicted text of a first sample fragment provided by an embodiment of the present application;
[0102] Figure 7 It is a schematic diagram of a scenario for determining weight relationship features through a recursive attention network provided by an embodiment of the present application;
[0103] Figure 8 It is a schematic diagram of a scenario of recursively outputting predicted elements of each fragment element in training text information provided by an embodiment of the present application;
[0104] Fig. 9 It is a flowchart of a text data processing method provided in an embodiment of the present application;
[0105] Fig.10It is a schematic diagram of a scenario for outputting weight relationship features through a selection gate provided by this application;
[0106] Fig.11 It is a schematic diagram of a scenario for decoding and outputting predicted text through a recursive attention network provided by an embodiment of this application;
[0107] Fig.12 It is a text data processing method provided by an embodiment of this application;
[0108] Fig.13 It is a schematic diagram of the structure of a text data processing device provided by an embodiment of this application;
[0109] Fig.14 It is a schematic diagram of the structure of a text data processing device provided by an embodiment of this application;
[0110] Fig.15 It is a schematic diagram of a computer device provided by an embodiment of this application. Detailed implementation manners
[0111] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0112] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0113] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0114] The solution provided by the embodiments of this application belongs to machine learning (ML) in the field of artificial intelligence. It can be understood that machine learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning usually includes technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0115] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of a network architecture provided by the embodiments of this application. As Figure 1 shown, the network architecture may include a first user terminal cluster, a second user terminal cluster, and a service server 2000. It can be understood that both the first user terminal cluster and the second user terminal cluster may include one or more user terminals, and the number of user terminals in each user terminal cluster will not be limited here. Among them, as Figure 1 shown, the first user terminal cluster may include multiple user terminals, specifically including Figure 1 the user terminals 3000a, 3000b, 3000c,..., 3000n shown. As Figure 1 shown, the user terminals 3000a, 3000b, 3000c,..., 3000n can be network-connected to the service server 2000 so that each user terminal in the first user terminal cluster can perform data interaction with the service server 2000 through this network connection.
[0116] For example, in the scenario of document recognition, a user terminal (e.g., user terminal 3000a) in the first user terminal cluster can locate a document picture that needs to be recognized for text and graphics in the document information (e.g., document A) of a certain service through a front-end text detection method. Then, the document picture that needs to be recognized for text and graphics can be used as the target picture and sent to the service server 2000, so that the service server 2000 can recognize the target text information (i.e., the text information composed of characters in the target picture) in the target picture through the trained target network model. Then, the service server 2000 can return the recognized target text information and the text box coordinates associated with the target text information to the aforementioned user terminal 3000a, so that the user terminal 3000a can output the target text information based on the text box coordinates in the area where the target picture is located in the text A.
[0117] It can be understood that the document information (e.g., document A) here can include one or more document pictures that need to be recognized for text and graphics. The number of document pictures included in the document information will not be limited here. In addition, it can be understood that the document picture can specifically include a document retrieval picture corresponding to the document retrieval service, an electronic bill picture corresponding to the bill recognition service, an electronic insurance policy picture corresponding to the insurance policy reimbursement service, etc.
[0118] It can be understood that when the service server 2000 receives the document pictures (i.e., the aforementioned target pictures) sent by one or more user terminals in the first user terminal cluster, it can obtain the positioning position identifiers corresponding to the picture display areas of these target pictures in the corresponding document information at the same time. Then, it can perform batch text recognition on these document pictures with positioning position identifiers through the trained target network model (e.g., a target network model based on a recursive attention network). In this way, when the service server 2000 accurately recognizes the target text information in each document picture, it can directly return the target text information to the corresponding user terminal quickly based on the positioning position identifier.
[0119] In addition, it can be understood that the business server 2000 can also add these target images and the target text information identified in these target images to the prior knowledge base, so that the business processing terminal (for example, the user terminal in the second user terminal cluster) associated with the above-mentioned business (for example, document retrieval business, bill recognition business and insurance policy reimbursement business) can obtain the added target images and the target text information in the target images from the prior knowledge base, and then the obtained target images and the target text information in the target images can be used as training sample information for further training the target network model, so that a new target network model can be obtained. The training sample information here can include sample images with sample training labels, and the text information indicated by the training sample labels can be training text information.
[0120] For ease of understanding, the embodiments of the present application may collectively refer to the network model before training as the initial network model, and the new network model obtained after training the initial network model as the aforementioned target network model. In addition, the embodiments of the present application may also collectively refer to the user terminal sending the target image as the first terminal in the first user terminal cluster, and the user corresponding to the first terminal as the first user.
[0121] It is understood that the first terminal here may include: a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart TV, and other smart terminals with document loading and display functions. Figure 1 The user terminal 3000a shown is used as the first terminal, and one or more application clients can be run in the first terminal, and these application clients can provide interfaces for image and text recognition of document information for the aforementioned services (for example, document retrieval services, bill recognition services, insurance policy reimbursement services, etc.). In this way, when the first user corresponding to the first terminal uses a certain application client (for example, client A), the image and text recognition interface provided by the client A can be triggered to switch the display interface of the first terminal to the image and text recognition interface, so as to display the target text information after image and text recognition for the aforementioned target images (i.e., the document retrieval images corresponding to the document retrieval services, the electronic bill images corresponding to the bill recognition services, and the electronic insurance policy images corresponding to the insurance policy reimbursement services) in the image and text recognition interface. Among them, the application clients here can include social clients, shopping clients, office clients, education clients, electronic reading clients, and other clients with text information loading and display functions.
[0122] Among them, Figure 1 The second terminal cluster shown may also include multiple user terminals, specifically including Figure 1The user terminals 4000a, 4000b, 4000c, …, 4000n shown. As Figure 1 shown, the user terminals 4000a, 4000b, 4000c, …, 4000n can be network-connected to the service server 2000, so that each user terminal in the second user terminal cluster can perform data interaction with the service server 2000 through this network connection. For example, in the above text and image recognition scenario, each user terminal in the second terminal cluster can serve as a service processing terminal (i.e., the second terminal) to receive the target image and the target text information in the target image sent after the service server 2000 performs text and image recognition.
[0123] For example, the service server 2000 can use the target text information recognized through the above target network model and the target image to which the target text information belongs as training sample information, and add it to the above prior knowledge base, so that the user terminals (i.e., the aforementioned second terminals) in the second terminal cluster can obtain the new training sample information from the prior knowledge base, and use the new training sample information to perform model training on the current network model (i.e., the aforementioned target network model) to obtain a new target network model, and then can give the newly trained target network model to the service server 2000, so that the service server 2000 can, when obtaining a new target image, intelligently recognize the text information in the target image.
[0124] It should be understood that the second terminal here may include: intelligent terminals with model training functions such as smart phones, tablet computers, laptop computers, and desktop computers. For example, in the embodiments of the present application, Figure 1 the user terminal 4000a in the second user terminal cluster shown can be used as the second terminal. When the second terminal obtains the above training sample information from the prior knowledge base, it can use the training sample information to perform iterative training on the current network model (for example, the initial network model), and use the network model obtained by the iterative training (i.e., the initial network model after iterative training) as the target network model, and return the target network model to the above service server 2000. In this way, when the service server 2000 obtains the target image sent by the first terminal in the first terminal cluster, it can recognize the target text information in the target image through the trained target network model.
[0125] Among them, as Figure 1The business server 2000 shown may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. It will not be limited here.
[0126] Optionally, it should be understood that in the blockchain scenario, each second terminal in the second terminal cluster can be used as a blockchain node. In this way, when these blockchain nodes train a new target network model based on the foregoing new training sample information, it is necessary to reach a consensus on the model parameters of the newly generated target network model among these blockchain nodes. Then, when a consensus is reached, the model parameters of the newly generated target network model can be uploaded to the blockchain, so that when the business server 2000 obtains a new target picture, it can obtain the model parameters of the newly generated target network model from the blockchain in real time. Then, the model parameters of the target network model in the business server 2000 can be updated based on the model parameters of the newly generated target network model, and then the text information in the new target picture can be recognized based on the target network model with updated parameters.
[0127] For ease of understanding, further, please refer to Figure 2 , Figure 2 is a schematic diagram of a scenario for data interaction provided by an embodiment of the present application. As Figure 2 shown, the first terminal may be the user terminal 3000a shown above. As Figure 1 shown, when the first terminal outputs document information (for example, Article A) on the display interface, it may allow the first user to trigger the above-mentioned graphic recognition interface, so that the first terminal can detect the Figure 2 shown target picture (for example, Figure 2 shown picture 20a carrying text information) in the document information (for example, Article A) through a text detection method, and can convert the image display area of the target picture in the document information (for example, Article A) into a positioning position identifier, and then can send the positioning position identifier and the target picture to the Figure 2 shown business server. At this time, the business server can recognize the text information in the target picture through the trained target network model, and then can use the recognized text information as the target text information, so that the target text information can be returned to the Figure 2 shown first terminal based on the positioning position identifier. Figure 2 shown first terminal.
[0128] Among them, Figure 2The target network model shown may include multiple branch models. Specifically, the multiple branch models here may include Figure 2 the branch model 1 and the branch model 2 shown. It should be understood that in the model application stage, the embodiments of the present application may Figure 2 collectively refer to the branch model 1 shown as the first target branch model, and may Figure 2 collectively refer to the branch model 2 shown as the second target branch model. As Figure 2 shown, the first target branch model may include a convolutional neural network (CNN) model or a neural architecture search (NAS) model, etc. The NAS model can be used to determine a candidate network structure through a search strategy in a defined high-dimensional hand space, and search for model parameters based on the candidate network structure to find the most robust network structure. The second target branch model may be a language model. As Figure 2 shown, specifically, the language model may include Figure 2 the network 201a and the network 201b shown. Among them, the network 201a is a recurrent attention network introduced in the language model, and the network 201b is a memory network newly added in the language model.
[0129] Among them, it can be understood that when Figure 2 the business server shown obtains the target picture sent by the first terminal (for example, Figure 2 the picture 20a shown), it can give the target picture (for example, Figure 2 the picture 20a shown) to Figure 2 the target network model shown. Furthermore, the first target branch model (for example, Figure 2 the branch model 1 shown) in the target network model can be used to extract image features, and then the extracted image features can be collectively referred to as target image features. The target image features can be Figure 2 the image feature 21a shown.
[0130] It should be noted that since the business server does not know the specific semantic information of each character in the picture 20a before performing text and image recognition on the text information in the picture 20a. Therefore, when the business server extracts Figure 2 the image feature 21a shown through Figure 2 the branch model 1 shown, it can input the image feature 21a into the one located at Figure 2The network 210a within the shown branch model 2 decodes the image features 21a globally through the recursive attention mechanism of the network 210a, so as to quickly learn the association relationship features between each target element in the text information of the picture 20a and the picture 20a, and then can identify and output the predicted elements of each target element in the text information recognized from the picture 20a based on these association relationship features.
[0131] As Figure 2 shown, the service server can directly accumulate the number of elements of the predicted elements corresponding to each target element of the output text information according to the text segmentation parameter of the target network model (for example, the text segmentation length is 6). Once the accumulated number of elements reaches the text segmentation parameter, the elements when reaching the text segmentation parameter can be used as the first target text of the first target fragment. Here, the first target fragment is determined by the service server for the text information in the picture 20a based on the text segmentation parameter. Before recognizing the content in the first target fragment, the semantic information of each target element in the first target fragment is not known until Figure 2 the network 201a shown (for example, the recursive attention network 1) outputs at time T1 Figure 2 the text 203a shown (that is Figure 2 the " <sos>When it is “ABACD”), the text 203a output in segments by the network 201a (for example, the recursive attention network 1) can be used as the first target text of the first target fragment. Among them, here " <sos>”The starting identifier used to represent the target text information. Similarly, Figure 2 the text 203d shown (i.e., Figure 2 "GHXYZ shown <eos>”) in the <eos>”An end identifier used to indicate the end of the target text information. It should be understood that the start identifier and the end identifier here are invisible to the user (e.g., the first user mentioned above).
[0132] As Figure 2 shown, the service server can also add text 203a to Figure 2 the network 201b shown (i.e., the aforementioned memory network), so as to use the text 203a added to the memory network (i.e., the first target text) as the target auxiliary text of the next text fragment (i.e., the second target fragment) of the first target fragment, so that the text features of the target auxiliary text (i.e., the first target text features of the aforementioned first target text) and the image features 21a can be input into Figure 2 the network 201a shown (e.g., the recursive attention network 2) at time T2, so that the network 201a outputs at time T2 Figure 2 the text 203b shown (i.e., Figure 2 the "EF GHI" shown), the text 203b output in segments by the network 201a (e.g., the recursive attention network 2) can be used as the second target text of the second target fragment.
[0133] As Figure 2 shown, the embodiment of the present application can record the context information between any two adjacent fragments through the network 201b (i.e., the memory network) between the two networks 201a (i.e., the two recursive attention networks in the recursive attention network), so as to solve the long-term forgetting problem when the target network model recognizes ultra-long text in a picture through the recursive attention network. It should be understood that by analogy,[[]] Figure 2 when the network 201a shown (e.g., the recursive attention network 3) outputs at time T3 Figure 2 the text 203c shown (i.e., Figure 2 the "JKL MM" shown), the text 203c output in segments by the network 201a (e.g., the recursive attention network 3) can be used as the second target text of the new second target fragment. Wherein, the specific implementation manner of the service server outputting the second target text (i.e., text 203c) of the new second target fragment through the network 201a (e.g., the recursive attention network 3) can refer to the description of outputting text 203b above, and will not be elaborated here.
[0134] It can be understood that Figure 2 the target network model shown is a text model obtained after model training for the above initial network model, which means that the initial network model constructed in the model training stage can also include Figure 2 Branch model 1 and branch model 2 in the target network model shown, for example, the branch model 1 before training (i.e., the first sample branch model) can specifically include a convolutional neural network before training; for another example, the branch model 2 before training (i.e., the second sample branch model) can specifically include a recurrent attention network before training and a memory network before training. It can be understood that the model parameters of the target network model here (for example, model parameter 2) are different from the model parameters of the initial network model (for example, model parameter 1), because during the training process of the initial network model, the model parameters of the networks under multiple sample branch models in the initial network model will be continuously adjusted, so that the adjusted initial network model can be iteratively trained later. Until the trained initial network model has the minimum loss function, it is considered that the trained initial network model meets the model convergence condition, and then the trained initial network model with the minimum loss function can be determined as the target network model.
[0135] Among them, the specific implementation manner of the service server 2000 to obtain the target network model by training the initial network model and to perform text recognition on the target text information in the target picture through the target network model can be seen in the following Figure 3-Figure 12 description of the corresponding embodiment.
[0136] Furthermore, please refer to Figure 3 , Figure 3 which is a schematic flowchart of a text data processing method provided by an embodiment of the present application. Among them, it can be understood that the method provided by the embodiment of the present application can be executed by a computer device, and the computer device here includes but is not limited to a user terminal (for example, the first terminal and the second terminal in the above Figure 2 corresponding embodiments) or a service server (for example, the service server in the above Figure 2 corresponding embodiments). For the convenience of understanding, the embodiment of the present application takes this computer device as a service server as an example to elaborate on the specific process of training the initial network model in the service server to obtain a target network model for identifying the target text information in the target picture. As Figure 3 shown, the method can at least include the following steps S101 - step S104:
[0137] Step S101, obtain a sample picture carrying a sample training label, perform text segmentation on the training text information indicated by the training sample label to obtain a first sample fragment and a second sample fragment for training the initial network model;
[0138] Among them, it can be understood that the second sample fragment here is the next sample fragment of the first sample fragment.
[0139] Specifically, the computer device can obtain sample pictures carrying sample training labels, and can perform text segmentation on the training text information indicated by the training sample labels based on text segmentation parameters (where the text segmentation parameters here refer to the minimum segmentation unit for text segmentation of training text information), so as to obtain a sample fragment set of the training text information; among them, it can be understood that the sample fragment set here can be used to store all sample fragments obtained by text segmentation of the training text information; one sample fragment corresponds to one sample segmentation identifier; one sample segmentation identifier is used to represent the fragment position of a sample fragment in the training text information; further, the computer device can obtain the first sample fragment for training the initial network model from the sample fragment set based on the sample segmentation identifier of each sample fragment, and can use the next sample fragment of the first sample fragment as the second sample fragment based on the fragment position of the first sample fragment in the training text information.
[0140] For ease of understanding, further, please refer to Figure 4 , Figure 4 which is a schematic diagram of a scenario for text segmentation provided by an embodiment of the present application. As Figure 4 shown, the training text information 401a is the text information in the sample picture obtained by the computer device from the above-mentioned prior knowledge base. In the model training stage, the computer device can obtain picture information with training sample labels from the prior knowledge base, and can use the obtained picture as a sample picture. It can be understood that the training text information indicated by the training sample label can be Figure 4 the training text information 401a shown (i.e., "ABCDEF GHIJKLMMI……GHXYZ"). As Figure 4 shown, the computer device can read the start identifier (for example, the above-mentioned "sos") and the end position identifier (for example, the above-mentioned "eos") from the training text information 401a, and then can perform text segmentation on the text information corresponding to the string between the start identifier and the end identifier through text segmentation parameters for text segmentation (for example, the text segmentation parameter is used to indicate that L (for example, 6) text elements in the training text information are used as the text segmentation unit), and then can obtain N sample fragments with a text segmentation length of 6 text elements, and can add these N sample fragments to the sample fragment set associated with the training text information 401a for storage.
[0141] It should be understood that if the training text information is an English sentence, the text elements here can be word elements in the English sentence. Optionally, if the training text information is a Chinese sentence, the text elements here can be character elements in the Chinese sentence. The specific language and specific presentation form of the strings in the training text information in the sample image will not be limited here. It should be noted that the presentation form of the training text information in the sample image can include but is not limited to printed text or handwritten text.
[0142] For ease of understanding, in the embodiments of the present application, when splitting to obtain N sample fragments, a sample splitting identifier can be configured for each sample fragment, and the sample splitting identifier here can be used to represent the fragment position of the corresponding sample fragment in the training text information 401a. For example, as Figure 4 shown, these N sample fragments in the sample fragment set can specifically include sample fragment 1, sample fragment 2, sample fragment 3, …, and sample fragment N. Among them, as Figure 4 shown, sample fragment 1 contains the above-mentioned start identifier. Based on this, the sample splitting identifier of sample fragment 1 can be used to represent the first fragment position of sample fragment 1 in the training text information 401a, and so on. Figure 4 shown, sample fragment N contains the above-mentioned end identifier. Therefore, the sample splitting identifier of sample fragment N can be used to represent the last fragment position of sample fragment N in the training text information 401a. As Figure 4 shown, based on the sample splitting identifiers of these sample fragments, the computer device can determine that sample fragment 2 is the next sample fragment of sample fragment 1. And so on, sample fragment 3 is the next sample fragment of sample fragment 2.
[0143] For ease of understanding, in the embodiments of the present application, sample fragment 1 as shown in Figure 4 can be obtained from the sample fragment set as the first sample fragment for training the initial network model, and then the next sample fragment of the first sample fragment can be used as the second sample fragment, so as to further perform the following steps S102 - step S1104.
[0144] Optionally, it is understood that the initial network model may include a first sample branch model and a second sample branch model; the second sample branch model here may include a semantic extraction network; based on this, before executing step S102, the computer device may also perform the following steps in collaboration through multiple sample branch models to improve the efficiency of model training. For example, the computer device may extract sample image features of the sample image through the first sample branch model; in addition, the computer device may also determine the first sample text feature of the first sample fragment through the semantic extraction network in the second sample branch model, and determine the second sample text feature of the second sample fragment through the semantic extraction network.
[0145] For further understanding, please refer to Figure 5 , Figure 5 is a schematic diagram of a scenario for determining sample text features of sample fragments provided by an embodiment of the present application. Figure 5 The training text information 501a shown may be the above Figure 4 The training text information 401a that needs to be segmented in the corresponding embodiment. Figure 5 The fragment A1 involved can be described above Figure 4 The sample fragment 1 in the corresponding embodiment is similarly Figure 5 The fragment A2 involved can be described above Figure 4 The corresponding sample fragment 2 in the embodiment. By analogy, Figure 5 The fragments involved can be described above Figure 4 The sample fragment N in the corresponding embodiment.
[0146] like Figure 5 As shown, the computer device gives the training text information 501 to Figure 5 When the semantic extraction network shown in FIG. 5 is used, the training text features of the training text information 501a can be extracted through the semantic extraction network. Then, the computer device can extract the training text features of the training text information 501a based on the fragment position of the fragment A1 in the training text information 501a (for example, Figure 5 The fragment position 1 shown in the figure, the fragment text feature of the fragment A1 is determined in the training text feature, and the fragment text feature of the fragment A1 can be Figure 5 Similarly, the computer device can be based on the fragment position of the fragment A2 in the training text information 501a (for example, Figure 5 The fragment position 2 shown in the figure, the fragment text feature of the fragment A2 is determined in the training text feature, and the fragment text feature of the fragment A2 can be Figure 5 The fragment text feature C2 shown. Similarly, the computer device can be based on the fragment position of the fragment An in the training text information 501a (for example, Figure 5 the fragment position n) shown, determine the fragment text feature of the fragment An in this training text feature, and the fragment text feature of the fragment An can be Figure 5 the fragment text feature Cn shown.
[0147] It should be understood that during the model training phase, since each text fragment can correspond to a part of the sample image. Therefore, to ensure the accuracy of subsequent feature decoding of the sample image features, the computer device needs to record the fragment position of each fragment itself, and also needs to record the element positions of the fragment elements in each fragment. Therefore, as Figure 5 shown, the computer device can take the fragment position of fragment A1 (for example, Figure 5 the fragment position 1 shown) and the element positions of the fragment elements in the fragment A1 (for example, element position 1-1, element position 1-2,..., element position 1-6) and collectively call them the relative position encoding information of the fragment A, and then can perform relative encoding on the relative position encoding information of the fragment A to obtain the relative position feature of the fragment A, as Figure 5 shown. At this time, the computer device can comprehensively consider the fragment text feature C1 of the fragment A1 and the relative position feature of the fragment A1 to obtain the sample text feature of the fragment A1. The sample text feature of the fragment A1 can be Figure 5 the sample text feature B1 shown.
[0148] Similarly, as Figure 5 shown, the computer device can comprehensively consider the fragment text feature C2 of the fragment A2 and the relative position feature of the fragment A2 to obtain the sample text feature of the fragment A2. The sample text feature of the fragment A2 can be Figure 5 the sample text feature B2 shown. Among them, the specific implementation method for the computer device to calculate the relative position feature of the fragment A2 can be referred to the description of calculating the relative position feature of the fragment A1 for details. By analogy, the computer device can comprehensively consider the fragment text feature C3 of the fragment A3 and the relative position feature of the fragment A3 to obtain the sample text feature of the fragment A3. The sample text feature of the fragment A3 can be Figure 5 the sample text feature B3 shown. Among them, the specific implementation method for the computer device to calculate the relative position feature of the fragment A3 can be referred to the description of calculating the relative position feature of the fragment A1 for details. In addition, the computer device can also comprehensively consider the fragment text feature Cn of the fragment An and the relative position feature of the fragment An to obtain the sample text feature of the fragment An. The sample text feature of the fragment An can be Figure 5 The sample text feature Bn shown. Among them, for the specific implementation method of the computer device to calculate the relative position feature of the fragment An, reference can be made to the description of calculating the relative position feature of the fragment A1.
[0149] It can be seen from this that for the above-mentioned first sample fragment (for example, Figure 5 the fragment A1 shown) and the second sample fragment (for example, the fragment A2), the computer device can be based on the fragment position of the first sample fragment in the training text information (for example, Figure 5 the fragment position 1 shown), determine the fragment text feature of the first sample fragment in the training text feature (for example, Figure 5 the fragment text feature C1 shown), and based on the fragment position of the second sample fragment in the training text information (for example, Figure 5 the fragment position 2 shown), determine the fragment text feature of the second sample fragment in the training text feature (for example, Figure 5 the fragment text feature C2 shown). Further, as Figure 5 shown, the computer device can take the fragment position of the first sample fragment (for example, Figure 5 the fragment position 1 shown) and the element position of the fragment element in the first sample fragment as the first relative coding position information of the first sample fragment, and then can perform relative coding on the first sample fragment based on the first relative coding position information to obtain the relative position feature of the first sample fragment. Then, the computer device can obtain the first sample text feature of the first sample fragment (for example, Figure 5 the sample text feature B1 shown) based on the fragment text feature of the first sample fragment (for example, Figure 5 the fragment text feature C1 shown) and the relative position feature of the first sample fragment.
[0150] Similarly, as Figure 5 shown, the computer device can take the fragment position of the second sample fragment (for example, Figure 5 the fragment position 2 shown) and the element position of the fragment element in the second sample fragment as the second relative coding position information of the first sample fragment, and then can perform relative coding on the second sample fragment based on the second relative coding position information to obtain the relative position feature of the second sample fragment. Then, the computer device can obtain the second sample text feature of the second sample fragment (for example, Figure 5 the fragment text feature C2 shown) and the relative position feature of the second sample fragment. Figure 5 the sample text feature B2 shown).
[0151] Step S102: Input the first sample text feature of the first sample fragment and the sample image feature of the sample picture into the recurrent attention network in the initial network model. Determine the first text-image relationship feature between the first sample fragment and the sample picture through the recurrent attention network, output the first predicted text of the first sample fragment based on the first text-image relationship feature, and add the first predicted text to the memory network in the initial network model.
[0152] Among them, the initial network model includes a first sample branch model and a second sample branch model; the second sample branch model includes a semantic extraction network, a recurrent attention network, and a memory network; the first sample branch model is used to extract the sample image feature of the sample picture; the semantic extraction network is used to extract the first sample text feature of the first sample fragment; the fragment elements of the first sample fragment include a first fragment element and a second fragment element; the second fragment element is the next fragment element of the first fragment element; specifically, the computer device can obtain the first element auxiliary feature associated with the first fragment element from the first sample text feature of the first sample fragment based on the element position of the first fragment element in the first sample fragment; further, the computer device can determine the first weight relationship feature between the first fragment element and the sample picture based on the first element auxiliary feature, the sample image feature of the sample picture, the first attention layer in the recurrent attention network, and the first selection gate corresponding to the first attention layer, input the first weight relationship feature into the first language layer in the recurrent attention network, and output the first predicted element corresponding to the first fragment element by the first language layer; further, the computer device can recursively use the first element feature of the first predicted element as the second element auxiliary feature associated with the second fragment element, determine the second weight relationship feature between the second fragment element and the sample picture based on the second element auxiliary feature, the sample image feature, the second attention layer in the recurrent attention network, and the second selection gate corresponding to the second attention layer, input the second weight relationship feature into the second language layer in the recurrent attention network, and output the second predicted element corresponding to the second fragment element by the second language layer; further, the computer device can determine the first text-image relationship feature between the first sample fragment and the sample picture based on the first weight relationship feature and the second weight relationship feature, output the first predicted text of the first sample fragment based on the first text-image relationship feature; the first predicted text includes the first predicted element and the second predicted element; further, the computer device can add the first predicted text including the first predicted element and the second predicted element to the memory network in the initial network model.
[0153] For ease of understanding, further, please refer to Figure 6 , Figure 6 is a schematic diagram of a scenario for outputting the first predicted text of the first sample fragment provided by an embodiment of the present application. As Figure 6 The shown branch model 62a is the above-mentioned second sample branch model. Based on this, the network 601a located in the branch model 62a can be the above-mentioned recursive attention network, and the network 601b located in the branch model 62a can be the above-mentioned memory network.
[0154] It can be understood that in the model training stage of this application embodiment, the memory network (for example, Figure 6 the shown network 601b) can be used to store the information interaction between fragments, and the context relationship between fragments can be modeled by the relative encoding method used in the above-mentioned second sample branch model. Furthermore, different sample fragments can be made to have different relative encoding position information. In this way, it can effectively avoid the problem of encoding confusion caused by using the same encoding method for the same characters appearing at the same fragment position in different fragments when using absolute position encoding.
[0155] For example, as Figure 6 shown, the relative position feature of the fragment A1 is obtained by the computer device performing relative encoding on the relative encoding position information of the fragment A1. The relative position feature of the fragment A2 is obtained by the computer device performing relative encoding on the relative encoding position information of the fragment A2. The relative position feature of the fragment A3 is obtained by the computer device performing relative encoding on the relative encoding position information of the fragment A3. And so on, the relative position feature of the fragment An is obtained by the computer device performing relative encoding on the relative encoding position information of the fragment A1. Therefore, even if the same characters appear at the same fragment position in different fragments, the relative encoding method can use different encoding methods for these same characters in different fragments, and then the relative element position features of these fragment elements in different fragments can be mapped. It should be understood that at this time, the computer device can comprehensively obtain Figure 6 the relative position feature of each sample fragment shown.
[0156] As Figure 6 shown, the computer device can also establish an association relationship between the fragment elements in different fragments and the image characters in the sample picture through the above-mentioned recursive attention network. For example, as Figure 6 shown, the computer device can, through Figure 6 the shown network 601a (for example, the recursive attention network), determine the graphic-text relationship feature D1 between the fragment A1 and Figure 6 the shown sample picture at the moment T1', and then can output Figure 6 the text 603a shown based on the graphic-text relationship feature D1. Similarly, as Figure 6 shown, the computer device can, through Figure 6 The network 601a shown (e.g., a recursive attention network) determines the graphic - text relationship feature D2 between the fragment A2 and Figure 6 the sample picture shown at time T2', and then can output Figure 6 the text 603b shown based on the graphic - text relationship feature D2. Similarly, as Figure 6 shown, the computer device can, through Figure 6 the network 601a shown (e.g., a recursive attention network), determine the graphic - text relationship feature D3 between the fragment A3 and Figure 6 the sample picture shown at time T3', and then can output Figure 6 the text 603c shown based on the graphic - text relationship feature D3. And so on, as Figure 6 shown, the computer device can, through Figure 6 the network 601a shown (e.g., a recursive attention network), determine the graphic - text relationship feature Dn between the fragment An and Figure 6 the sample picture shown at time Tn', and then can output Figure 6 the text 603d shown based on the graphic - text relationship feature Dn.
[0157] Among them, for the sake of understanding, in the embodiments of this application, Figure 6 taking the fragment A1 shown as the above - mentioned first sample fragment as an example, to elaborate on the specific process of outputting the first predicted text of the first sample fragment through Figure 6 the network 601a shown. It should be understood that the computer device can collectively refer to the graphic - text relationship feature D1 between the fragment A1 and Figure 6 the sample picture determined by the recursive attention network at the above - mentioned time T1' as the first graphic - text relationship feature between the first sample fragment and the sample picture. It should be understood that the first graphic - text relationship feature here is determined by the relationship features between each fragment element in the first sample fragment and the sample picture, and Figure 6 the text 603a shown is the first predicted text of the first sample fragment, and the prediction elements in the first predicted text can specifically include Figure 6 the 6 prediction elements shown. These 6 prediction elements can specifically include: Figure 6 the prediction element " <sos>”, Figure 6 the predicted element "A" shown in Figure 6 the predicted element "B" shown in Figure 6 the predicted element "A" shown in Figure 6 the predicted element "C" shown in, and Figure 6 the predicted element "D" shown in. It should be understood that when the computer device outputs the predicted elements of each fragment element in the first sample fragment through Figure 6 the network 601a shown (i.e., the recursive attention network), it can output the predicted text composed of these 6 predicted elements in segments according to the above text segmentation unit (for example, L = 6) (for example, Figure 6 the text 603a shown), and then can continue to execute the following step S103 to use the first predicted text of the first sample fragment as the auxiliary text of the second sample fragment, so as to subsequently predict the second predicted text of the second sample fragment, and the second predicted text can be Figure 6 the text 603b of. Similarly, for the output manner of the predicted texts of other sample fragments (such as fragment A2, fragment A3,..., fragment An) of the training text information, reference can be made to the description of the specific process of outputting the second predicted text in the embodiments of the present application, and details will not be elaborated here.
[0158] In addition, it can be understood that the embodiments of the present application can also take a sample picture containing written text as an example, and then can elaborate on how to associate the first sample fragment in the training text information in the sample picture with the sample picture through the recursive attention network, so as to subsequently perform text recognition on the target picture containing written text. For ease of understanding, the embodiments of the present application take any two adjacent fragment elements in the first sample fragment as an example to elaborate on the specific process of determining the weight relationship characteristics between any two adjacent fragment elements and the sample picture through the recursive attention network. The first sample fragment here can include a first fragment element and a second fragment element. Note that the second fragment element here is the next fragment element of the second fragment element.
[0159] For ease of understanding, further, please refer to Figure 7 , Figure 7 which is a schematic diagram of a scenario for determining the attention map feature through the recursive attention network provided by the embodiments of the present application. As Figure 7 shown, the sample picture can be input into Figure 7 the branch model 71a, and the branch model 71a can be the above-mentioned first sample branch model. Through the convolutional kernel in the first sample branch model, convolution processing can be performed on Figure 7 the sample picture shown, and then Figure 7 the image feature 701a after convolution processing shown can be obtained. As Figure 7 As shown, the computer device can provide the training text information corresponding to the sample picture to Figure 7 the branch model 72a shown. The branch model 72a can be the second sample branch model above, so as to real-time process the Figure 7 sample fragment S1 shown (for example, " <sos>But o”), sample fragment S2 (e.g., "nline"), …, sample fragment Sn (e.g., "ng <eos>”) Perform semantic feature extraction and relative position encoding to obtain in real-time encoding Figure 7 The sample text feature P1 of the sample fragment S1 shown, the sample text feature P2 of the sample fragment S2,..., the sample text feature Pn of the sample fragment Sn.
[0160] Such as Figure 7 As shown, the computer device can establish Figure 7 The graphic-text alignment relationship 1 between the sample text feature P1 of and the image feature 701a, establish Figure 7 The graphic-text alignment relationship 2 between the sample text feature P2 of and the image feature 701a, and establish Figure 7 The graphic-text alignment relationship n between the sample text feature Pn of and the image feature 701a. For ease of understanding, in the embodiments of the present application, Figure 7 The sample fragment S1 shown (for example, " <sos>Taking “Buto” as an example to illustrate the establishment of the weight relationship features between any two adjacent fragment elements in the sample fragment S1 and the sample picture in the recursive attention network shown in Figure 7
[0161] It should be noted that the fragment elements in the sample fragment S1 (such as " <sos>”) is the starting identifier of the training text information (the fragmented element corresponding to this starting identifier is a placeholder element and has no actual meaning. Therefore, the recursive attention network can directly output the placeholder element corresponding to this starting identifier as the element for starting element prediction). Thus, the element feature corresponding to this starting identifier can be directly used as the auxiliary feature of the first element in the training text information. At this time, the embodiment of the present application can use the first element in the training text information (i.e., the element adjacent to the starting identifier) as the first fragmented element, and can use the element feature corresponding to this starting identifier as the first element auxiliary feature of this first fragmented element. At this time, the computer device can determine the first weight relationship feature between the first fragmented element and the sample picture based on this first element auxiliary feature, the sample image feature of the sample picture, the first attention layer in the recursive attention network, and the first selection gate corresponding to the first attention layer. The first weight relationship feature between the first fragmented element and the sample picture is determined by Figure 7 the attention map feature W1 (i.e., the first attention map feature) in the recursive attention network shown. In Figure 7 the attention map corresponding to the attention map feature W1 shown, the position area where the element position of this first fragmented element (for example, Figure 7 the element "B" shown) is highlighted, which means that this highlighted first fragmented element (i.e., Figure 7 the element "B" shown) has a higher weight in the attention map corresponding to the attention map feature W1. Then, the computer device can output the first predicted element of the first fragmented element (for example, Figure 7 the character "B" in the recognition result shown) based on the aforementioned first text-image relationship feature.
[0162] Similarly, the computer device can use the next fragmented element of this first fragmented element ( Figure 7 the element "u" shown) as the second fragmented element, and use the first predicted element corresponding to this first fragmented element as the auxiliary element of this second fragmented element. Furthermore, the first element feature of this first predicted element can be recursively used as the second element auxiliary feature associated with the second fragmented element. Then, based on the second element auxiliary feature, the sample image feature, the second attention layer in the recursive attention network, and the second selection gate corresponding to the second attention layer, the second weight relationship feature between the second fragmented element and the sample picture can be determined. It can be understood that the second weight relationship feature between the second fragmented element and the sample picture is determined by Figure 7 the attention map feature W2 (i.e., the second attention map feature) in the recursive attention network shown. In Figure 7 the attention map corresponding to the attention map feature W2 shown, this second fragmented element (for example, Figure 7 The position area where the element position of the element "u" shown is highlighted brightly, which represents that the second fragmented element with such bright highlighting (i.e., Figure 7 the element "u" shown has a relatively high weight in the attention map corresponding to the attention map feature W2. Similarly, the computer device can output a second predicted element of the second fragmented element based on the foregoing second text-image relationship feature (for example, Figure 7 the character "u" in the recognition result shown).
[0163] And so on, in the embodiment of the present application, the second predicted element output by the recursive attention network at the current moment can be used as a new auxiliary element, and then the original feature of the new auxiliary element can be recursively used as the auxiliary feature of the new second fragmented element (for example, Figure 7 the element "t" shown), that is, the new second element auxiliary feature. Furthermore, based on the new second element auxiliary feature, the sample image feature (for example, Figure 7 the image feature 701a), the second attention layer in the recursive attention network, and the second selection gate corresponding to the second attention layer, the weight relationship feature (i.e., the new second weight relationship feature) between the new second fragmented element and the sample picture can be determined, so as to facilitate the subsequent output of the predicted element of the new second fragmented element (for example, the new second predicted element can be Figure 7 the element "t" in the recognition result shown). The specific implementation manner for the computer device to output the new second predicted element can refer to the description of the specific process of outputting the second predicted element above, and will not be elaborated here.
[0164] As Figure 7 shown, the computer device can, based on Figure 7 each attention map feature shown, traverse and output the weight relationship feature between each element in the training text information and the sample picture, and then these traversed and output weight relationship features can be collectively referred to as the above-mentioned text-image relationship features. Furthermore, based on these text-image relationship features, Figure 7 m (for example, m = 19) predicted elements in the training text information shown can be obtained. These m predicted elements can form Figure 7 the recognition result shown, that is, the m predicted elements can be "But online shopping".
[0165] It should be understood that during the model training phase, the computer device can determine in real time whether the number of elements of the predicted elements output by the recursive attention network reaches the above-mentioned text segmentation parameter (for example, the above L = 6). If it reaches, the predicted text composed of L predicted elements when the text segmentation parameter is reached can be used as the first predicted text of the first sample fragment, so that the following step S103 can be executed subsequently. Otherwise, the computer device can recursively execute the process of outputting the second predicted element.
[0166] For ease of understanding, further, please refer to Figure 8 , Figure 8 is a schematic diagram of a scenario for recursively outputting predicted elements of each fragment element in the training text information provided by an embodiment of the present application. As Figure 8 shown, the training sample picture can be the sample picture in the corresponding embodiment of the above Figure 7 . As Figure 8 shown, the computer device can decode each character in the sample picture through the above-mentioned recursive attention network, and each character here is the Figure 8 predicted element of each fragment element shown. As Figure 8 shown, the auxiliary feature U1 is the element feature corresponding to the element of the start identifier. As Figure 8 shown, the computer device can use the auxiliary feature U1 as the first element auxiliary feature of the first fragment element, and then can use the first element auxiliary feature and the Figure 8 sample image feature of the sample picture as the first input feature of the first sample fragment, and then can input the first input feature into the Figure 8 shown attention layer 1 (i.e., the first attention layer) to determine the first attention map feature of the first fragment element in the sample picture through the attention layer 1. The first attention map feature can be the Figure 8 attention map feature W1 shown. As Figure 8 shown, the computer device can input the attention map feature and the Figure 8 shown auxiliary feature (i.e., the first element auxiliary feature) into the Figure 8 shown selection gate 1 (i.e., the first selection gate. It should be understood that the selection gate 1 here corresponds one-to-one with the attention layer 1), so that the selection gate 1 can select and output the weight relationship feature 1 (i.e., the above-mentioned first weight relationship feature) between the first fragment element (for example, the above fragment element "B") and the sample picture (i.e., the Figure 8 shown training sample picture).
[0167] As Figure 8 shown, the computer device takes the graphic-text relationship feature 1 as the first graphic-text relationship feature to give the first graphic-text relationship feature to Figure 8 The shown language layer 1 can further perform element recognition on the first fragment element by this language layer 1 (i.e., the first language layer) to recognize the first predicted element corresponding to the first fragment element, and the first predicted element can be Figure 8 the predicted element "B" shown. As Figure 8 shown, the embodiment of the present application can utilize the recursive characteristic of the recursive attention network and use the predicted element "B" (i.e., the first predicted element) output by the language layer 1 as the auxiliary element of the next fragment element. The element feature of the predicted element "B" (i.e., Figure 8 the auxiliary feature U2 shown) can be used as the second element auxiliary feature of the next fragment element. As Figure 8 shown, the computer device can use the auxiliary feature U2 and the sample image feature as the second input feature of the first sample fragment, and then can input the second input feature into Figure 8 the shown attention layer 2 (i.e., the second attention layer) so that the attention layer 2 determines the second attention map feature of the second fragment element in the sample picture, and the second attention map feature can be Figure 8 the attention map feature W2 shown. Then, the computer device can input the attention map feature W2 (i.e., the second attention map feature) and the auxiliary feature U2 (i.e., the second element auxiliary feature) into Figure 8 the shown selection gate 2 (i.e., the second selection gate), and the selection gate 2 selects and outputs the second weight relationship feature between the second fragment element and the sample picture, and the second weight relationship feature can be Figure 8 the weight relationship feature 2 shown. As Figure 8 shown, the computer device can use the weight relationship feature 2 as the first text-image relationship feature to Figure 8 the shown language layer 2 so that the language layer 2 performs element recognition on the second fragment element to obtain the second predicted element corresponding to the second fragment element, and the second predicted element can be Figure 8 the predicted element "u" shown.
[0168] As Figure 8 shown, the computer device can recursively use the predicted element "u" as the new auxiliary element to recursively output the predicted elements of each fragment element through these recursive attention networks. For example, when the computer device predicts and outputs the predicted element corresponding to the element position of the end identifier in the last fragment, it can obtain the predicted elements of each fragment element in the training text information, and the predicted text composed of these predicted elements can be the recognition result in the corresponding embodiment of the above Figure 7 .
[0169] Step S103: Use the first predicted text stored in the memory network as the training auxiliary text for the second sample fragment. Input the first sample text features, second sample text features, and sample image features corresponding to the training auxiliary text into the recursive attention network. Determine the second text-image relationship feature between the second sample fragment and the sample picture through the recursive attention network. Output the second predicted text of the second sample fragment based on the second text-image relationship feature, and add the second predicted text to the memory network;
[0170] Among them, the fragment lengths of the first sample fragment and the second sample fragment are both determined by the text segmentation parameters of the initial network model; the first predicted text is determined by the second sample branch model when the cumulative number of predicted elements output by the language layer in the recursive attention network reaches the text segmentation parameter; the predicted element is determined after the recursive attention network associates the fragment elements in the first sample fragment with the sample picture.
[0171] Specifically, as described above Figure 6 shown, the computer device can add the first predicted text of the first sample fragment obtained by prediction (for example, the text 603a shown above Figure 6 shown) to the memory network, and then can use the first predicted text in the memory network as the training auxiliary text for the second sample fragment. Further, as described above Figure 6 shown, the computer device can use the first sample text features corresponding to the training auxiliary text as the training auxiliary text features, and can input the training auxiliary text features, second sample text features (not shown above Figure 6 shown), and sample image features into the recursive attention network, and can associate the second sample fragment with the sample picture through the recursive attention network to obtain the second text-image relationship feature used to represent the association relationship between the second sample fragment and the sample picture. It can be understood that the specific implementation manner for the computer device to determine the second text-image relationship feature through the recursive attention network can refer to the description of the specific process of determining the first text-image relationship feature through the recursive attention network above, and will not be elaborated here.
[0172] Further, the computer device inputs the second text-image relationship feature into the language layer in the recursive attention network, and the language layer in the recursive attention network can output the predicted elements associated with the fragment elements in the second sample fragment; it can be understood that the specific implementation manner for the computer device to output the predicted elements of each fragment element in the second sample fragment can refer to the description of the specific processes of the first predicted element of the first fragment element and the second predicted element of the second fragment element in the corresponding embodiment above, and will not be elaborated here. Figure 8 described, and will not be elaborated here.
[0173] Further, when the computer device detects that the number of predicted elements output by the language layer in the recursive attention network reaches the text segmentation parameter, it outputs the second predicted text of the second sample fragment based on the predicted elements output by the language layer in the recursive attention network (for example, the text 603a in the corresponding embodiment described above), and can add the second predicted text to the memory network, so that the second predicted text can be used as the training auxiliary text for the next sample fragment, and then the new training auxiliary text can participate in the prediction of the predicted text of the next sample fragment to ensure that the recursive attention network can make full use of the information interaction between the fragments. Figure 6 The text 603a) in the corresponding embodiment, and can add the second predicted text to the memory network, so that the second predicted text can be used as the training auxiliary text for the next sample fragment, and then the new training auxiliary text can participate in the prediction of the predicted text of the next sample fragment to ensure that the recursive attention network can make full use of the information interaction between the fragments.
[0174] Step S104, determine the sample prediction label of the training text information based on the first predicted text and the second predicted text in the memory network, and perform iterative training on the initial network model based on the sample training label and the sample prediction label, and use the iteratively trained initial network model as the target network model for text recognition of the target picture.
[0175] Specifically, the computer device can determine the initial loss function for adjusting the model parameters in the initial network model based on the predicted probability value corresponding to the sample prediction label and the true probability value corresponding to the sample training label; further, if the computer device determines that the value of the initial loss function does not meet the model convergence condition, it adjusts the model parameters of each branch network in the initial network model based on the value of the initial loss function, and then can perform iterative training on the adjusted initial network model through the training sample information composed of the training text information and the sample picture to obtain the target loss function of the iteratively trained initial network model; further, if the value of the target loss function meets the model convergence condition, the computer device can determine the iteratively trained initial network model that meets the model convergence condition as the target network model. Among them, the target network model is used for text recognition of the target text information of the acquired target picture.
[0176] Among them, it can be understood that after the computer device predicts the prediction label, it can compare it with the true label of the sample picture (for example, the above Figure 8 Compare with the sample training label of the training sample picture in the corresponding embodiment to obtain the initial loss function value of the model loss function. It can be understood that if the initial loss function value does not meet the above model convergence condition (for example, the initial loss function value is not the minimum loss function value in the model training stage), the computer device can reversely adjust the model parameters in the initial network model based on the initial loss function value, so as to perform iterative training on the adjusted initial network model through the new training sample information until the target loss function value of the initial network model after iterative training meets the model convergence condition, and then determine the initial network model after iterative training that meets the model convergence condition as the target network model.
[0177] It should be understood that optionally, when the above computer device trains the target network model, the trained target network model can also be integrated in the above first terminal (the first terminal here can be any user terminal in the above first user terminal cluster). In this way, after the first terminal determines the target picture through text detection at the front end, it can directly perform text recognition on the target text information in the target picture based on the pre-trained network model (that is, the above target network model) in the first terminal, thereby reducing the signaling interaction with the server. In this way, the text in the target picture can be recognized locally on the first terminal, thereby improving the efficiency of text recognition.
[0178] The computer device in the embodiment of the present application can obtain a sample picture carrying a sample training label in the model training stage, and then can perform text segmentation on the training text information indicated by the sample training label to obtain a first sample fragment and a second sample fragment for training the initial network model. Among them, it should be understood that since the training text information in the sample picture is known when the embodiment of the present application obtains the sample picture, text segmentation can be performed on the training text information in the model training stage to realize the fragmentation processing of training text information of any length, so as to improve the generalization ability of the trained model when using the trained initial network model (that is, the target network model) for text recognition in the future. In addition, the embodiment of the present application can also extract the graphic-text dependency relationship between each sample fragment and the sample picture through a recursive attention network, and the graphic-text dependency relationship at least includes the above first graphic-text relationship feature and the second graphic-text relationship feature. Therefore, when outputting each predicted text through the graphic-text dependency relationship, not only can the graphic and text features be associated through the recursive attention network, but also the problem of long-term forgetting can be solved fundamentally by introducing the memory network. For example, through the memory network, the fragment text currently output by the computer device can be used as the training auxiliary text for the next fragment text, and then the context relationship between the fragments can be fully utilized through the memory network, thereby improving the accuracy of text recognition of the picture.
[0179] Further, please refer to Fig. 9 , Fig. 9 which is a schematic flowchart of a text data processing method provided by an embodiment of the present application. It can be understood that the method provided by the embodiment of the present application can be executed by a computer device, and the computer device herein includes but is not limited to a user terminal or a server. As Fig. 9 shown, the method may at least include the following steps S201 - step S208;
[0180] Step S201, obtain a sample picture carrying a sample training label, perform text segmentation on the training text information indicated by the training sample label based on text segmentation parameters, and obtain a sample fragment set of the training text information;
[0181] Among them, the sample fragment set is used to store all sample fragments obtained by performing text segmentation on the training text information; one sample fragment corresponds to one sample segmentation identifier; one sample segmentation identifier is used to represent the fragment position of a sample fragment in the training text information.
[0182] Step S202, based on the sample segmentation identifier of each sample fragment, obtain a first sample fragment for training an initial network model from the sample fragment set, and use the next sample fragment of the first sample fragment as a second sample fragment based on the fragment position of the first sample fragment in the training text information;
[0183] Among them, for the specific implementation manners of steps S201 - step S202, reference can be made to the description of the specific process of performing text segmentation on the training text information in the corresponding embodiment above, and details will not be elaborated here. Figure 3
[0184] Step S203, extract sample image features of the sample picture through a first sample branch model;
[0185] Step S204, determine a first sample text feature of the first sample fragment through a semantic extraction network, and determine a second sample text feature of the second sample fragment through the semantic extraction network.
[0186] Among them, for the specific implementation manners of steps S203 - step S204, reference can be made to the description of the specific process of obtaining the first sample text feature and the second sample text feature in the corresponding embodiment above, and details will not be elaborated here. Figure 3
[0187] Step S205: Input the first sample text feature of the first sample fragment and the sample image feature of the sample picture into the recurrent attention network in the initial network model. Determine the first text-image relationship feature between the first sample fragment and the sample picture through the recurrent attention network. Output the first predicted text of the first sample fragment based on the first text-image relationship feature, and add the first predicted text to the memory network in the initial network model.
[0188] Among them, the initial network model includes a first sample branch model and a second sample branch model; the second sample branch model includes a semantic extraction network, a recurrent attention network, and a memory network; the first sample branch model is used to extract the sample image feature of the sample picture; the semantic extraction network is used to extract the first sample text feature of the first sample fragment; the fragment elements of the first sample fragment include a first fragment element and a second fragment element; the second fragment element is the next fragment element of the first fragment element.
[0189] Specifically, the computer device can obtain the first element auxiliary feature associated with the first fragment element from the first sample text feature of the first sample fragment based on the element position of the first fragment element in the first sample fragment; further, the computer device can determine the first weight relationship feature between the first fragment element and the sample picture based on the first element auxiliary feature, the sample image feature of the sample picture, the first attention layer in the recurrent attention network, and the first selection gate corresponding to the first attention layer, and input the first weight relationship feature into the first language layer in the recurrent attention network, and the first language layer outputs the first predicted element corresponding to the first fragment element; further, the computer device can recursively use the first element feature of the first predicted element as the second element auxiliary feature associated with the second fragment element, determine the second weight relationship feature between the second fragment element and the sample picture based on the second element auxiliary feature, the sample image feature, the second attention layer in the recurrent attention network, and the second selection gate corresponding to the second attention layer, and input the second weight relationship feature into the second language layer in the recurrent attention network, and the second language layer outputs the second predicted element corresponding to the second fragment element; further, the computer device can determine the first text-image relationship feature between the first sample fragment and the sample picture based on the first weight relationship feature and the second weight relationship feature, and output the first predicted text of the first sample fragment based on the first text-image relationship feature; the first predicted text includes the first predicted element and the second predicted element; further, the computer device can add the first predicted text including the first predicted element and the second predicted element to the memory network in the initial network model.
[0190] Among them, the specific process of the computer device outputting the first predicted element can be described as follows: The computer device can use the first element auxiliary feature and the sample image feature of the sample picture as the first input feature associated with the first sample fragment; further, the computer device can input the first input feature into the first attention layer in the recursive attention network, and the first attention layer determines the first attention map feature of the first fragment element in the sample picture; further, the computer device can input the first attention map feature and the first element auxiliary feature into the first selection gate corresponding to the first attention layer, and the first selection gate selects and outputs the first weight relationship feature between the first fragment element and the sample picture; further, the computer device can input the first weight relationship feature into the first language layer in the recursive attention network, and the first language layer performs element recognition on the first fragment element to obtain the first predicted element corresponding to the first fragment element.
[0191] For ease of understanding, based on the above Figure 8 selection gate in the corresponding embodiment (for example, the above selection gate 1), the scenario schematic diagram of how to output the first weight relationship feature through the selection gate 1 is described. Further, please refer to Fig.10 , Fig.10 which is a scenario schematic diagram of outputting the weight relationship feature through the selection gate provided by this application. As Fig.10 shown, the attention image feature can be the attention map feature W1 in the corresponding embodiment above Figure 8 . Similarly, Fig.10 the element auxiliary feature shown can be the auxiliary feature U1 in the corresponding embodiment above Figure 8 . As Fig.10 shown, the computer device can send the auxiliary feature U1 (i.e., the above first element auxiliary feature) and the attention map feature W1 (i.e., the first attention map feature) into Fig.10 the multi-similarity calculation module shown. It can be understood that the multi-similarity calculation module in the embodiments of this application can include various selection rules, and the various selection rules here can specifically include the first selection rule, the second selection rule, and the third selection rule. Fig.10 Among them, through the first selection rule in the selection gate, the first inner product similarity between the attention map feature W1 (i.e., the first attention map feature) and the auxiliary feature U1 (i.e., the above first element auxiliary feature) can be extracted, and then the product of the first weight coefficient (for example, β1) indicated by the first selection rule and the first inner product similarity can be used as the first similarity feature.
[0192]
[0193] Among them, the first Gaussian similarity between the attention map feature W1 (i.e., the first attention map feature) and the auxiliary feature U1 (i.e., the above-mentioned first element auxiliary feature) can be extracted through the second selection rule in the selection gate. Furthermore, the product of the second weight coefficient (e.g., β2) indicated by the second selection rule and the first Gaussian similarity can be used as the second similarity feature.
[0194] Similarly, the first string similarity between the attention map feature W1 (i.e., the first attention map feature) and the auxiliary feature U1 (i.e., the above-mentioned first element auxiliary feature) can be extracted through the third selection rule in the selection gate. Furthermore, the product of the third weight coefficient (e.g., β3) indicated by the third selection rule and the first string similarity can be used as the third similarity feature.
[0195] Then, as Fig.10 shown, the computer device can perform feature fusion on the similarity features calculated by the multi-pixel point calculation module according to the following formula (1):
[0196] K = β1K inner + β2K gaussian + β3K JS Formula (1);
[0197] Among them, K inner is used to represent the inner product similarity (e.g., the above-mentioned first inner product similarity). Therefore, the product of β1 and K inner (i.e., β1K inner ) can be used as the above-mentioned first similarity feature. Similarly, K gaussian is used to represent the Gaussian similarity (e.g., the above-mentioned first Gaussian similarity). Therefore, the product of β2 and K gaussian (i.e., β2K gaussian ) can be used as the above-mentioned second similarity feature. By analogy, K JS is used to represent the string similarity (e.g., the above-mentioned first string similarity). Therefore, the product of β3 and K JS (i.e., β3K JS ) can be used as the above-mentioned third similarity feature. Among them, K can be used to represent the fusion feature. Among them, β1 + β2 + β3 = 1.
[0198] As Fig.10 shown, after the computer device performs feature fusion on these three features, the first fusion feature (i.e., the aforementioned K) can be obtained. At this time, the computer device can also perform normalization processing on the first fusion feature through the softmax function to obtain Fig.10 the attention weight matrix 113a shown. The embodiment of the present application can collectively refer to the attention weight matrix 113a as the first weight matrix between the first fragment element and the sample picture. As Fig.10 As shown, the computer device can also multiply the attention weight matrix (i.e., the first weight matrix) with Fig.10 the attention map feature W1 (i.e., the first attention map feature) shown to perform masked multiplication, so as to select and obtain the first weight relationship feature between the first fragment element and the sample picture from the attention map feature W1 (i.e., the first attention map feature). The first weight relationship feature can be Fig.10 the weight relationship feature 114a shown. Then, the computer device can input the first weight relationship feature into the first language layer in the recurrent attention network (for example, the Figure 8 language layer 1 shown above), and the first language layer performs element recognition on the first fragment element to obtain the first predicted element corresponding to the first fragment element.
[0199] It should be understood that the specific implementation manner for the computer device to determine the second weight relationship feature between the second fragment element and the sample picture through the selection gate 2 shown above can refer to the description of the output weight relationship feature 114 in the corresponding Figure 8 embodiment, and details will not be elaborated here. Fig.10
[0200] Step S206: Use the first predicted text stored in the memory network as the training auxiliary text for the second sample fragment, input the first sample text feature, the second sample text feature, and the sample image feature corresponding to the training auxiliary text into the recurrent attention network, determine the second text-image relationship feature between the second sample fragment and the sample picture through the recurrent attention network, output the second predicted text of the second sample fragment based on the second text-image relationship feature, and add the second predicted text to the memory network;
[0201] Specifically, the semantic extraction network is further used to extract the second sample text feature of the second sample fragment; the fragment elements of the second sample fragment include a third fragment element and a fourth fragment element; the fourth fragment element is the next fragment element of the third fragment element. Specifically, the computer device can obtain the third element auxiliary feature associated with the third fragment element from the second sample text feature of the second sample fragment based on the element position of the third fragment element in the second sample fragment; note that in the embodiments of the present application, the predicted element of the last fragment element of the first sample fragment can be used as the auxiliary element associated with the third fragment element in the second sample fragment (for example, the Figure 8 fragment element "n" in the corresponding embodiment above, and the fragment position of the fragment element "n" in the second sample fragment of the training text information is the above 1-6), and then the auxiliary element (for example, the Figure 8 fragment element "o" in the corresponding embodiment above, and the fragment position of the fragment element "o" in the second sample fragment of the training text information is the above 2-1) can be used, and then the auxiliary element (for example, the Figure 8 The element feature of the fragment element "o") in the corresponding embodiment is used as the third element auxiliary feature of the third fragment element. It should be understood that the embodiments of the present application can determine the third weight relationship feature (i.e., the new first weight relationship feature) between the third fragment element and the sample picture based on the third element auxiliary feature, the sample image feature of the sample picture (e.g., Figure 8 the training sample picture shown), the third attention layer (i.e., the new first attention layer) in the recursive attention network, and the third selection gate (i.e., the new first selection gate) corresponding to the third attention layer. Furthermore, the third weight relationship feature (i.e., the new first weight relationship feature) can be input into the third language layer (i.e., the new first language layer) in the recursive attention network, and the third prediction element corresponding to the third fragment element can be output by the third language layer (i.e., the new first language layer). Similarly, during the process of decoding the above sample image feature, the computer device can continue to send the predicted element decoded in the previous step and the sample image feature into the fourth attention layer (i.e., the new second attention layer). Furthermore, after obtaining the fourth attention map feature between the fourth fragment element and the sample picture, the element feature (i.e., the fourth auxiliary element feature) of the predicted element decoded most recently in the previous step and the fourth attention map feature are given to the fourth selection gate (i.e., the new second selection gate), so that when the fourth selection gate (i.e., the new second selection gate) can obtain the fourth weight relationship feature of the fourth fragment element, the fourth fragment element can be further element-recognized through the fourth language layer in the recursive attention network, and thus the fourth prediction element corresponding to the fourth fragment element can be obtained.
[0202] For ease of understanding, further, please refer to Fig.11 , Fig.11 which is a schematic diagram of a scenario for decoding and outputting predicted text through a recursive attention network provided by an embodiment of the present application. As Fig.11 shown, the computer device can generate an attention map as shown in Fig.11 through the recursive attention network. The attention map is determined by the attention map features between each fragment element in the training text information and the sample picture. For the fragment element itself corresponding to the above start identifier, the predicted element of the start identifier can be directly decoded and output through the recursive attention network, and the predicted element of the start identifier can be directly used as the element auxiliary feature of the next fragment element to participate in the element prediction of the next fragment element. The specific prediction process can refer to the process of obtaining the first prediction element and the second prediction element associated with the first sample fragment given in the corresponding embodiment of the above Figure 8 . As Fig.11 shown, the predicted text 1 obtained by the computer device can be the first predicted text of the above first sample fragment.
[0203] Similarly, for the second sample fragment, when the computer device detects that the number of predicted elements output by the language layer in the recursive attention network (e.g., the above-mentioned third and fourth language layers) reaches the text segmentation parameter (e.g., the above L = 6), it can traverse and output the predicted elements associated with the second sample fragment based on the language layer in the recursive attention network (e.g., Fig.11 "n", "l", "i", "n", "e", " <space>”), so as to obtain the second predicted text of the second sample fragment (for example, Fig.11 The predicted text formed by these predicted elements shown in Fig.11 is the predicted text 2 shown), and then the second predicted text can be added to the memory network, so that the second predicted text can be used as the auxiliary text of the next sample fragment to participate in the prediction process of the predicted text of the next sample fragment.
[0204] Step S207: Based on the first predicted text and the second predicted text in the memory network, determine the sample predicted label of the training text information. Based on the sample training label and the sample predicted label, perform iterative training on the initial network model, and use the iteratively trained initial network model as the target network model for text recognition of the target picture.
[0205] Optionally, after the computer device executes the above step S207, it can also use the target network model to execute the following steps:
[0206] The computer device can obtain a target picture carrying target text information, and extract the target image features of the target picture through the first target branch model in the target network model; wherein, the target network model is obtained by training the initial network model based on the sample pictures carrying sample training labels; the target network model further includes a second target branch model juxtaposed to the first target branch model, and the second target branch model includes a recursive attention network and a memory network; further, the computer device can input the target image features into the recursive attention network, determine the first correlation relationship feature between the first target fragment in the target text information and the target picture through the recursive attention network, output the first target text of the first target fragment based on the first correlation relationship feature, and add the first target text to the memory network in the target network model; further, the computer device can use the first target text stored in the memory network as the target auxiliary text of the second target fragment in the target text information, input the first target text feature corresponding to the target auxiliary text and the target image features into the recursive attention network, determine the second correlation relationship feature between the second target fragment and the target picture through the recursive attention network, output the second target text of the second target fragment based on the second correlation relationship feature, and add the second target text to the memory network; further, the computer device can determine the target text information recognized from the target picture based on the first target text and the second target text in the memory network.
[0207] Optionally, when the computer device receives a target picture sent by a user terminal (for example, the above-mentioned first terminal), it can also recognize the target text information in the target picture through the trained target network model.
[0208] It can be seen that the embodiments of the present application can extract the graphic-text dependence relationship between each sample fragment and the sample picture through a recursive attention network. The graphic-text dependence relationship here at least includes the above-mentioned first graphic-text relationship feature, second graphic-text relationship feature, third graphic-text relationship feature, and fourth graphic-text relationship feature. Therefore, when outputting each predicted text through these graphic-text dependence relationships, not only can the sample fragment features of the sample fragment and the sample image features of the sample picture be feature-associated through the recursive attention network, but also the problem of long-term forgetting can be fundamentally solved by introducing a memory network. For example, the memory network can be used to take the fragment text currently output by the computer device as the training auxiliary text for the next fragment text, and then the memory network can make full use of the context relationship between the fragments to improve the accuracy of text recognition for pictures.
[0209] Further, please refer to Fig.12 , Fig.12 which is a text data processing method provided by the embodiments of the present application. This method can be executed by the above computer device. Among them, this method can include the following steps S301 - step S304;
[0210] Step S301: Obtain a target picture carrying target text information, and extract the target image features of the target picture.
[0211] Specifically, the computer device can obtain a target picture carrying target text information, and can extract the target image features of the target picture through the first target branch model in the target network model.
[0212] Among them, the target network model is obtained by training the initial network model with sample pictures carrying sample training labels; the target network model also includes a second target branch model juxtaposed to the first target branch model, and the second target branch model includes a recursive attention network and a memory network.
[0213] Step S302: Based on the target image features, determine the first association relationship feature between the first target fragment in the target text information and the target picture, and output the first target text of the first target fragment based on the first association relationship feature.
[0214] Specifically, the computer device can input the target image features into the recursive attention network, and can determine the first association relationship feature between the first target fragment in the target text information and the target picture through the recursive attention network, and then can output the first target text of the first target fragment based on the first association relationship feature. Further, the computer device can also add the first target text to the memory network in the target network model, so that during the entire process of recognizing the target text information in the target picture, the first target text can be stored in the memory network for a long time.
[0215] In step S303, use the first target text as the target auxiliary text for the second target fragment in the target text information. Based on the first target text features corresponding to the target auxiliary text and the target image features, determine the second correlation relationship features between the second target fragment and the target picture, and output the second target text of the second target fragment based on the second correlation relationship features.
[0216] Specifically, the computer device can use the first target text stored in the memory network as the target auxiliary text for the second target fragment in the target text information, and can input the first target text features corresponding to the target auxiliary text and the target image features into the recursive attention network. Furthermore, the second correlation relationship features between the second target fragment and the target picture can be determined through the recursive attention network. Further, the computer device can output the second target text of the second target fragment based on the second correlation relationship features, and can add the second target text to the above-mentioned memory network.
[0217] In step S304, determine the target text information recognized from the target picture based on the first target text and the second target text.
[0218] Specifically, the computer device can determine the target text information recognized from the target picture based on the first target text and the second target text in the memory network.
[0219] Among them, for the specific implementation manners of steps S301 - S304, reference can be made to the description of the specific process of recognizing the target text information through the target network model in the corresponding embodiments above. Details will not be elaborated here. Additionally, it can be understood that before obtaining the target network model, the computer device can also use the above Figure 2 or Figure 3 or Fig. 9 The text data processing method in the corresponding embodiment trains the initial network model. Furthermore, during the model training phase, the initial network model can segment the training text information in the sample pictures participating in the training. This means that the embodiments of the present application can use training text information of any length to train the initial network model. Additionally, through the recurrent attention network in the initial network model, the context information between the fragmented segments after text segmentation can be fully utilized, thereby improving the modeling ability for long text sequences such as ultra-long texts from the root. That is, the embodiments of the present application can store the semantic information of the context through the recurrent attention model and can process text information of any length. Moreover, by incorporating the selection gate mechanism, the ability of the target network model to integrate the target image features of the target picture and the semantic information of the sample fragments corresponding to the above sample pictures can be further improved, which is not available in previous text recognition methods. It should be noted that in the above text recognition scenario, when the embodiments of the present application train the initial network model through a large number of document-type test sets (including test pictures of any length such as documents, papers, news, etc.) to obtain the target network model, the recall performance of text recognition for these test pictures can reach 94.2, the precision performance can reach 93.1, and the F-measure (i.e., the comprehensive evaluation index of recall performance and precision performance) index can reach 93.6. This represents a significant improvement in recall performance, precision performance, and the F-measure index compared to current mainstream deep learning-based text recognition methods.
[0220] In this way, when using the trained target network model to recognize a target picture carrying ultra-long text, the context relationship of the ultra-long text in the target picture can be fully utilized through the recurrent attention network in the target network model. For example, the embodiments of the present application can use the text segmentation parameter of the target network model (e.g., the above L = 6) as the text segmentation unit to determine the predicted text belonging to each target fragment from the elements of the ultra-long text output by traversal. It should be understood that the embodiments of the present application can adaptively adjust the text segmentation parameter according to actual business requirements.
[0221] It can be seen that for some extremely long texts (i.e., texts where the text length of the target text information in the target picture is greater than the text length of the training text information in the sample picture), when obtaining a predicted text for a target fragment (i.e., the first target text of the first target fragment) based on the above text segmentation parameters, the first target text is further added to the memory network, so that the first target text added to this memory network (here referring to a network with long-term memory function) can participate in the text prediction of the next target text. In this way, the problem of long-term forgetting during text recognition of extremely long texts can be solved at the root. At the same time, in the embodiments of the present application, the first target text can also be used as the target auxiliary text for the next target fragment (i.e., the second target fragment), so that when the second target text of the second target fragment is output later, the predicted texts of each target fragment stored in the memory network can be integrated to quickly and accurately identify the target text information in the target picture.
[0222] Further, please refer to Fig.13 , Fig.13 which is a schematic structural diagram of a text data processing device provided by an embodiment of the present application. The above text data processing device 1 may be a computer program (including program code) running in a computer device. For example, the text data processing device 1 may be an application software; the device may be used to execute the corresponding steps in the method provided by the embodiments of the present application. Among them, the text data processing device 1 may include: a sample picture acquisition module 11, a first relationship determination module 12, a second relationship determination module 13, a model training module 14; optionally, the text data processing device 1 may further include: an image feature extraction module 15, a text feature determination module 16.
[0223] The sample picture acquisition module 11 is used to acquire a sample picture carrying a sample training label, perform text segmentation on the training text information indicated by the training sample label, and obtain a first sample fragment and a second sample fragment for training an initial network model; the second sample fragment is the next sample fragment of the first sample fragment;
[0224] Among them, the sample picture acquisition module 11 includes: a text segmentation unit 111 and a fragment acquisition unit 112;
[0225] The text segmentation unit 111 is used to acquire a sample picture carrying a sample training label, perform text segmentation on the training text information indicated by the training sample label based on text segmentation parameters, and obtain a sample fragment set of the training text information; the sample fragment set is used to store all sample fragments obtained by performing text segmentation on the training text information; one sample fragment corresponds to one sample segmentation identifier; one sample segmentation identifier is used to represent the fragment position of a sample fragment in the training text information;
[0226] The fragment acquisition unit 112 is used to acquire the first sample fragment for training the initial network model from the sample fragment set based on the sample segmentation identifier of each sample fragment, and based on the fragment position of the first sample fragment in the training text information, use the next sample fragment of the first sample fragment as the second sample fragment.
[0227] The specific implementation of the text segmentation unit 111 and the fragment acquisition unit 112 can be found in the above Figure 3 The description of step S101 in the corresponding embodiment will not be repeated here.
[0228] The initial network model includes a first sample branch model and a second sample branch model; the second sample branch model includes a semantic extraction network;
[0229] Optionally, an image feature extraction module 15 is used to extract sample image characteristics of the sample picture through the first sample branch model;
[0230] Optionally, the text feature determination module 16 is used to determine a first sample text feature of the first sample fragment through a semantic extraction network, and to determine a second sample text feature of the second sample fragment through a semantic extraction network.
[0231] The text feature determination module 16 includes: a training feature extraction unit 161, a fragment feature determination unit 162, a first feature determination unit 163 and a second feature determination unit 164;
[0232] A training feature extraction unit 161, used to extract training text features of training text information through a semantic extraction network;
[0233] A fragment feature determination unit 162, configured to determine a fragment text feature of the first sample fragment in the training text feature based on a fragment position of the first sample fragment in the training text information, and to determine a fragment text feature of the second sample fragment in the training text feature based on a fragment position of the second sample fragment in the training text information;
[0234] A first feature determination unit 163 is configured to use the fragment position of the first sample fragment and the element position of the fragment element in the first sample fragment as first relative coding position information of the first sample fragment, relatively encode the first sample fragment based on the first relative coding position information to obtain a relative position feature of the first sample fragment, and obtain a first sample text feature of the first sample fragment based on the fragment text feature of the first sample fragment and the relative position feature of the first sample fragment;
[0235] A second feature determination unit 164, configured to use the fragment position of the second sample fragment and the element position of the fragment element in the second sample fragment as the second relative coding position information of the first sample fragment, perform relative coding on the second sample fragment based on the second relative coding position information to obtain the relative position feature of the second sample fragment, and obtain the second sample text feature of the second sample fragment based on the fragment text feature and the relative position feature of the second sample fragment.
[0236] Among them, for the specific implementation manners of the training feature extraction unit 161, the fragment feature determination unit 162, the first feature determination unit 163, and the second feature determination unit 164, reference may be made to the description of the specific process of relative position coding in the corresponding embodiments above, and details will not be elaborated here. Figure 3 The description of the specific process of relative position coding in the corresponding embodiments above will not be repeated here.
[0237] A first relationship determination module 12 is configured to input the first sample text feature of the first sample fragment and the sample image feature of the sample picture into the recurrent attention network in the initial network model, determine the first text-image relationship feature between the first sample fragment and the sample picture through the recurrent attention network, output the first predicted text of the first sample fragment based on the first text-image relationship feature, and add the first predicted text to the memory network in the initial network model;
[0238] Among them, the initial network model includes a first sample branch model and a second sample branch model; the second sample branch model includes a semantic extraction network, a recurrent attention network, and a memory network; the first sample branch model is configured to extract the sample image feature of the sample picture; the semantic extraction network is configured to extract the first sample text feature of the first sample fragment; the fragment element of the first sample fragment includes a first fragment element and a second fragment element; the second fragment element is the next fragment element of the first fragment element;
[0239] The first relationship determination module 12 includes: a first auxiliary feature determination unit 121, a first element output unit 122, a second element output unit 123, a predicted text output unit 124, and a first text addition unit 125;
[0240] The first auxiliary feature determination unit 121 is configured to obtain the first element auxiliary feature associated with the first fragment element from the first sample text feature of the first sample fragment based on the element position of the first fragment element in the first sample fragment;
[0241] The first element output unit 122 is configured to determine a first weight relationship feature between the first fragmented element and the sample picture based on the first element auxiliary feature, the sample image feature of the sample picture, the first attention layer in the recursive attention network, and the first selection gate corresponding to the first attention layer, and input the first weight relationship feature into the first language layer in the recursive attention network, and the first language layer outputs a first predicted element corresponding to the first fragmented element;
[0242] Among them, the first element output unit 122 includes: a first input subunit 1221, a first attention determination subunit 1222, a first weight selection subunit 1223, and a first recognition subunit 1224;
[0243] The first input subunit 1221 is configured to use the first element auxiliary feature and the sample image feature of the sample picture as first input features associated with the first sample fragment;
[0244] The first attention determination subunit 1222 is configured to input the first input feature into the first attention layer in the recursive attention network, and the first attention layer determines a first attention map feature of the first fragmented element in the sample picture;
[0245] The first weight selection subunit 1223 is configured to input the first attention map feature and the first element auxiliary feature into the first selection gate corresponding to the first attention layer, and the first selection gate selects and outputs a first weight relationship feature between the first fragmented element and the sample picture;
[0246] Among them, the first weight selection subunit 1223 includes: a first feature extraction subunit 12231, a second feature extraction subunit 12232, a third feature extraction subunit 12233, and a first matrix determination subunit 12234;
[0247] The first feature extraction subunit 12231 is configured to input the first attention map feature and the first element auxiliary feature into the first selection gate corresponding to the first attention layer, extract a first inner product similarity between the first attention map feature and the first element auxiliary feature through the first selection rule in the first selection gate, and use the product of the first weight coefficient indicated by the first selection rule and the first inner product similarity as a first similarity feature;
[0248] The second feature extraction subunit 12232 is configured to extract a first Gaussian similarity between the first attention map feature and the first element auxiliary feature through the second selection rule in the first selection gate, and use the product of the second weight coefficient indicated by the second selection rule and the first Gaussian similarity as a second similarity feature;
[0249] The third feature extraction subunit 12233 is configured to extract the first string similarity between the first attention map feature and the first element auxiliary feature through the third selection rule in the first selection gate, and use the product of the third weight coefficient indicated by the third selection rule and the first string similarity as the third similarity feature;
[0250] The first matrix determination subunit 12234 is configured to determine the first weight matrix between the first fragment element and the sample image based on the first similarity feature, the second similarity feature, and the third similarity feature, and perform masked multiplication of the first weight matrix and the first attention map feature to obtain the first weight relationship feature between the first fragment element and the sample image.
[0251] Among them, the first matrix determination subunit 12234 includes: a feature fusion subunit 122341, a feature normalization subunit 122342, and a masked multiplication subunit 122343;
[0252] The feature fusion subunit 122341 is configured to perform feature fusion on the first similarity feature, the second similarity feature, and the third similarity feature to obtain the first fusion feature;
[0253] The feature normalization subunit 122342 is configured to perform normalization processing on the first fusion feature to obtain the first weight matrix between the first fragment element and the sample image;
[0254] The masked multiplication subunit 122343 is configured to perform masked multiplication of the first weight matrix and the first attention map feature to obtain the first weight relationship feature between the first fragment element and the sample image.
[0255] Among them, for the specific implementation manners of the feature fusion subunit 122341, the feature normalization subunit 122342, and the masked multiplication subunit 122343, reference can be made to the description of the specific process of selecting the weight feature in the corresponding embodiment above. Fig. 9 Details will not be elaborated here.
[0256] Among them, for the specific implementation manners of the first feature extraction subunit 12231, the second feature extraction subunit 12232, the third feature extraction subunit 12233, and the first matrix determination subunit 12234, reference can be made to the description of the specific process of determining the first weight relationship feature in the corresponding embodiment above. Fig. 9 Details will not be elaborated here.
[0257] The first recognition subunit 1224 is configured to input the first weight relationship feature into the first language layer in the recursive attention network, and the first language layer performs element recognition on the first fragment element to obtain the first predicted element corresponding to the first fragment element.
[0258] Among them, for the specific implementation manners of the first input subunit 1221, the first attention determination subunit 1222, the first weight selection subunit 1223, and the first recognition subunit 1224, reference may be made to the description of the specific process of outputting the first predicted element in the corresponding embodiment above, and details will not be elaborated here. Fig. 9 For the description of the specific process of outputting the first predicted element in the corresponding embodiment above, details will not be elaborated here.
[0259] The second element output unit 123 is configured to recursively use the first element feature of the first predicted element as the second element auxiliary feature associated with the second fragmented element, determine the second weight relationship feature between the second fragmented element and the sample picture based on the second element auxiliary feature, the sample image feature, the second attention layer in the recursive attention network, and the second selection gate corresponding to the second attention layer, and input the second weight relationship feature into the second language layer in the recursive attention network, and the second language layer outputs the second predicted element corresponding to the second fragmented element;
[0260] Among them, the second element output unit 123 includes: a second input subunit 1231, a second attention determination subunit 1232, a second weight selection subunit 1233, and a second recognition subunit 1234;
[0261] The second input subunit 1231 is configured to recursively use the first element feature of the first predicted element as the second element auxiliary feature associated with the second fragmented element, and use the second element auxiliary feature and the sample image feature as the second input feature associated with the first sample fragment;
[0262] The second attention determination subunit 1232 is configured to input the second input feature into the second attention layer in the recursive attention network, and the second attention layer determines the second attention map feature of the second fragmented element in the sample picture;
[0263] The second weight selection subunit 1233 is configured to input the second attention map feature and the second element auxiliary feature into the second selection gate corresponding to the second attention layer, and the second selection gate selects and outputs the second weight relationship feature between the second fragmented element and the sample picture;
[0264] Among them, the second weight selection subunit 1233 includes: a fourth feature extraction subunit 12331, a fifth feature extraction subunit 12332, a sixth feature extraction subunit 12333, and a second matrix determination subunit 12334;
[0265] The fourth feature extraction subunit 12331 is configured to input the second attention map feature and the second element auxiliary feature into the second selection gate corresponding to the second attention layer, extract the second inner product similarity between the second attention map feature and the second element auxiliary feature through the first selection rule in the second selection gate, and use the product between the first weight coefficient indicated by the first selection rule and the second inner product similarity as the fourth similarity feature;
[0266] The fifth feature extraction subunit 12332 is configured to extract the second Gaussian similarity between the second attention map feature and the second element auxiliary feature through the second selection rule in the second selection gate, and use the product between the second weight coefficient indicated by the second selection rule and the second Gaussian similarity as the fifth similarity feature;
[0267] The sixth feature extraction subunit 12333 is configured to extract the second string similarity between the second attention map feature and the second element auxiliary feature through the third selection rule in the second selection gate, and use the product between the third weight coefficient indicated by the third selection rule and the second string similarity as the sixth similarity feature;
[0268] The second matrix determination subunit 12334 is configured to determine the second weight matrix between the second fragment element and the sample picture based on the fourth similarity feature, the fifth similarity feature, and the sixth similarity feature, and perform masked multiplication on the second weight matrix and the second attention map feature to obtain the second weight relationship feature between the second fragment element and the sample picture.
[0269] Among them, for the specific implementation manners of the fourth feature extraction subunit 12331, the fifth feature extraction subunit 12332, the sixth feature extraction subunit 12333, and the second matrix determination subunit 12334, reference can be made to the description of the output of the second weight relationship feature in the corresponding embodiments above Fig. 9 and details will not be elaborated here.
[0270] The second recognition subunit 1234 is configured to input the second weight relationship feature into the second language layer in the recursive attention network, and the second language layer performs element recognition on the second fragment element to obtain the second predicted element corresponding to the second fragment element.
[0271] Among them, for the specific implementation manners of the second input subunit 1231, the second attention determination subunit 1232, the second weight selection subunit 1233, and the second recognition subunit 1234, reference can be made to the description of the specific process of outputting the second predicted element in the corresponding embodiments above Fig. 9 and details will not be elaborated here.
[0272] A predicted text output unit 124 is configured to determine a first text-image relationship feature between a first sample fragment and a sample picture based on a first weight relationship feature and a second weight relationship feature, and output a first predicted text of the first sample fragment based on the first text-image relationship feature; the first predicted text includes a first predicted element and a second predicted element;
[0273] A first text addition unit 125 is configured to add the first predicted text including the first predicted element and the second predicted element to a memory network in an initial network model.
[0274] Among them, for the specific implementation manners of the first auxiliary feature determination unit 121, the first element output unit 122, the second element output unit 123, the predicted text output unit 124, and the first text addition unit 125, reference may be made to the description of the specific process of outputting the first predicted text in the corresponding embodiments above, and details will not be elaborated here. Figure 3 The description of the specific process of outputting the first predicted text in the corresponding embodiments above will not be repeated here.
[0275] A second relationship determination module 13 is configured to use the first predicted text stored in the memory network as a training auxiliary text for a second sample fragment, input the first sample text feature, the second sample text feature, and the sample image feature corresponding to the training auxiliary text into a recursive attention network, determine a second text-image relationship feature between the second sample fragment and the sample picture through the recursive attention network, output a second predicted text of the second sample fragment based on the second text-image relationship feature, and add the second predicted text to the memory network;
[0276] Among them, the fragment length of the first sample fragment and the fragment length of the second sample fragment are both determined by the text segmentation parameter of the initial network model; the first predicted text is determined when the second sample branch model counts that the cumulative number of predicted elements output by the language layer in the recursive attention network reaches the text segmentation parameter; the predicted element is determined after the recursive attention network associates the fragment elements in the first sample fragment with the sample picture;
[0277] The second relationship determination module 13 includes: an auxiliary text determination unit 131, a feature association unit 132, a predicted element output unit 133, and a second text addition unit 134;
[0278] The auxiliary text determination unit 131 is configured to use the first predicted text stored in the memory network as a training auxiliary text for the second sample fragment, and use the first sample text feature corresponding to the training auxiliary text as a training auxiliary text feature;
[0279] A feature association unit 132, configured to input the training auxiliary text feature, the second sample text feature, and the sample image feature into a recurrent attention network, and perform feature association between the second sample fragment and the sample picture through the recurrent attention network to obtain a second text-image relationship feature for characterizing the association relationship between the second sample fragment and the sample picture;
[0280] A prediction element output unit 133, configured to input the second text-image relationship feature into the language layer in the recurrent attention network, and output a prediction element associated with the fragment element in the second sample fragment by the language layer in the recurrent attention network;
[0281] A second text addition unit 134, configured to, when detecting that the number of prediction elements output by the language layer in the recurrent attention network reaches the text segmentation parameter, output a second prediction text of the second sample fragment based on the prediction elements output by the language layer in the recurrent attention network, and add the second prediction text to the memory network.
[0282] Wherein, for the specific implementation manners of the auxiliary text determination unit 131, the feature association unit 132, the prediction element output unit 133, and the second text addition unit 134, reference may be made to the description of the specific process of outputting the second prediction text in the corresponding embodiment above. Details will not be elaborated here. Figure 3 For the description of the beneficial effects of adopting the same method, details will not be elaborated here either.
[0283] A model training module 14, configured to determine a sample prediction label of training text information based on the first prediction text and the second prediction text in the memory network, and perform iterative training on the initial network model based on the sample training label and the sample prediction label, and use the iteratively trained initial network model as a target network model for text recognition of a target picture.
[0284] Wherein, for the specific implementation manners of the sample picture acquisition module 11, the first relationship determination module 12, the second relationship determination module 13, and the model training module 14, reference may be made to the description of steps S101 - S104 in the corresponding embodiment above. Details will not be elaborated here. Additionally, for the description of the beneficial effects of adopting the same method, details will not be elaborated here either. Figure 3 For the description of the beneficial effects of adopting the same method, details will not be elaborated here either.
[0285] Furthermore, please refer to Fig.14 , Fig.14 It is a schematic structural diagram of a text data processing device provided by an embodiment of the present application. The text data processing device 2 may be a computer program (including program code) running on a computer device. For example, the text data processing device 2 may be an application software; the device may be used to execute corresponding steps in the method provided by an embodiment of the present application. The text data processing device 2 may include: a target picture acquisition module 21, a first text output module 22, a second text output module 23, and a target text determination module 24;
[0286] The target picture acquisition module 21 is used to acquire a target picture carrying target text information and extract target image features of the target picture;
[0287] Among them, the target picture acquisition module 21 is specifically used to acquire a target picture carrying target text information and extract target image features of the target picture through a first target branch model in the target network model; the target network model is obtained by training an initial network model with sample pictures carrying sample training labels.
[0288] The first text output module 22 is used to determine a first association relationship feature between a first target fragment in the target text information and the target picture based on the target image features, and output a first target text of the first target fragment based on the first association relationship feature;
[0289] Among them, the target network model further includes a second target branch model parallel to the first target branch model, and the second target branch model includes a recursive attention network and a memory network;
[0290] The first text output module 22 is specifically used to input the target image features into the recursive attention network, determine a first association relationship feature between a first target fragment in the target text information and the target picture through the recursive attention network, and output a first target text of the first target fragment based on the first association relationship feature.
[0291] Among them, optionally, the second target branch model further includes a memory network, and the first text output module 22 is further specifically used to add the first target text to the memory network in the target network model.
[0292] The second text output module 23 is used to use the first target text as a target auxiliary text of a second target fragment in the target text information, determine a second association relationship feature between the second target fragment and the target picture based on the first target text feature corresponding to the target auxiliary text and the target image features, and output a second target text of the second target fragment based on the second association relationship feature.
[0293] Among them, the second text output module 23 is specifically configured to use the first target text stored in the memory network as the target auxiliary text of the second target fragment in the target text information, input the first target text feature corresponding to the target auxiliary text and the target image feature into the recursive attention network, determine the second association relationship feature between the second target fragment and the target picture through the recursive attention network, and output the second target text of the second target fragment based on the second association relationship feature.
[0294] Optionally, the second text output module 23 is further specifically configured to add the second target text to the memory network.
[0295] The target text determination module 24 is configured to determine the target text information recognized from the target picture based on the first target text and the second target text.
[0296] Among them, the target text determination module 24 is specifically configured to determine the target text information recognized from the target picture based on the first target text and the second target text in the memory network.
[0297] Among them, for the specific implementation manners of the target picture acquisition module 21, the first text output module 22, the second text output module 23, and the target text determination module 24, reference can be made to the description of the specific processes of steps S301 - S304 in the corresponding embodiments above, and details will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. Fig.12 For the specific processes of steps S301 - S304 in the corresponding embodiments above, details will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either.
[0298] Furthermore, please refer to Fig.15 , Fig.15 which is a schematic diagram of a computer device provided by an embodiment of the present application. As shown in Fig.15 , the computer device 1000 may include: at least one processor 1001, such as a CPU, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002. Among them, the communication bus 1002 is used to implement connection communication between these components. Among them, the network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The memory 1005 may optionally also be at least one storage device located far from the aforementioned processor 1001. As shown in Fig.15 , in the memory 1005 as a computer storage medium, there may be included an operating system, a network communication module, a user interface module, and a device control application program.
[0299] In Fig.15 In the computer device 1000 shown, the network interface 1004 is mainly used to provide network communication functions; the user interface 1003 is mainly used to provide an interface for user input; and the processor 1001 can be used to call the device control application program stored in the memory 1005 to execute the description of the text data processing method in the corresponding embodiment mentioned above, and can also execute the description of the text data processing device 1 in the corresponding embodiment mentioned above, and can also execute the description of the text data processing device 2 in the corresponding embodiment mentioned above, which will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. Figure 3 or Fig. 9 or Fig.12 the description of the text data processing method in the corresponding embodiment mentioned above, and can also execute the description of the text data processing device 1 in the corresponding embodiment mentioned above, which will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. Fig.13 the description of the text data processing device 1 in the corresponding embodiment mentioned above, and can also execute the description of the text data processing device 2 in the corresponding embodiment mentioned above, which will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. Fig.14 the description of the text data processing device 2 in the corresponding embodiment mentioned above, which will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either.
[0300] In addition, it should be noted here that: The embodiments of the present application also provide a computer-readable storage medium, and the computer-readable storage medium stores the computer program executed by the computer device 1000 mentioned above, and the computer program includes program instructions. When the above-mentioned processor executes the above-mentioned program instructions, it can execute the description of the above-mentioned text data processing method in the corresponding embodiment mentioned above, and therefore, it will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application. Figure 3 or Fig. 9 or Fig.12 the description of the above-mentioned text data processing method in the corresponding embodiment mentioned above, and therefore, it will not be elaborated here. In addition, the description of the beneficial effects of using the same method will not be elaborated either. For the technical details not disclosed in the embodiments of the computer-readable storage medium involved in the present application, please refer to the description of the method embodiments of the present application.
[0301] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0302] The above-disclosed are only the preferred embodiments of the present application. Of course, the scope of the rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.< / space> < / sos> < / sos> < / eos> < / sos> < / sos> < / eos> < / eos> < / sos> < / sos>
Claims
1. A method for processing text data, characterized in that, Including: Obtain a sample picture carrying a sample training label, perform text segmentation on the training text information indicated by the sample training label to obtain a first sample fragment and a second sample fragment for training an initial network model; the second sample fragment is the next sample fragment of the first sample fragment; in the text and image recognition scenario, the presentation form of the training text information in the sample picture includes printed text or handwritten text; Input the first sample text feature of the first sample fragment and the sample image feature of the sample picture into the recurrent attention network in the initial network model, determine the first text and image relationship feature between the first sample fragment and the sample picture through the recurrent attention network, output the first predicted text of the first sample fragment based on the first text and image relationship feature, and add the first predicted text to the memory network in the initial network model; Use the first predicted text stored in the memory network as the training auxiliary text for the second sample fragment, input the first sample text feature corresponding to the training auxiliary text, the second sample text feature of the second sample fragment, and the sample image feature into the recurrent attention network, determine the second text and image relationship feature between the second sample fragment and the sample picture through the recurrent attention network, output the second predicted text of the second sample fragment based on the second text and image relationship feature, and add the second predicted text to the memory network; Determine the sample prediction label of the training text information based on the first predicted text and the second predicted text in the memory network, perform iterative training on the initial network model based on the sample training label and the sample prediction label, and use the iteratively trained initial network model as the target network model for text recognition of the target text information in the target picture; the target text information is the text information composed of the characters in the target picture.
2. The method according to claim 1, wherein The obtaining of the sample picture carrying the sample training label and performing text segmentation on the training text information indicated by the sample training label to obtain the first sample fragment and the second sample fragment for training the initial network model includes: Obtain a sample picture carrying a sample training label, perform text segmentation on the training text information indicated by the sample training label based on text segmentation parameters to obtain a sample fragment set of the training text information; the sample fragment set is used to store all the sample fragments obtained by performing text segmentation on the training text information; one sample fragment corresponds to one sample segmentation identifier; one sample segmentation identifier is used to represent the fragment position of a sample fragment in the training text information; Based on the sample segmentation identifier of each sample fragment, obtain the first sample fragment for training the initial network model from the sample fragment set, and use the next sample fragment of the first sample fragment as the second sample fragment based on the fragment position of the first sample fragment in the training text information.
3. The method according to claim 1, characterized in that, The initial network model includes a first sample branch model and a second sample branch model; the second sample branch model includes a semantic extraction network; the method further includes: Extracting sample image characteristics of the sample picture through the first sample branch model; Determining first sample text features of the first sample fragment through the semantic extraction network, and determining second sample text features of the second sample fragment through the semantic extraction network.
4. The method according to claim 3, wherein The determining first sample text features of the first sample fragment through the semantic extraction network, and determining second sample text features of the second sample fragment through the semantic extraction network includes: Extracting training text features of the training text information through the semantic extraction network; Based on the fragment position of the first sample fragment in the training text information, determining the fragment text features of the first sample fragment in the training text features, and based on the fragment position of the second sample fragment in the training text information, determining the fragment text features of the second sample fragment in the training text features; Taking the fragment position of the first sample fragment and the element position of the fragment element in the first sample fragment as the first relative coding position information of the first sample fragment, performing relative coding on the first sample fragment based on the first relative coding position information to obtain the relative position features of the first sample fragment, and based on the fragment text features and the relative position features of the first sample fragment, obtaining the first sample text features of the first sample fragment; Taking the fragment position of the second sample fragment and the element position of the fragment element in the second sample fragment as the second relative coding position information of the first sample fragment, performing relative coding on the second sample fragment based on the second relative coding position information to obtain the relative position features of the second sample fragment, and based on the fragment text features and the relative position features of the second sample fragment, obtaining the second sample text features of the second sample fragment.
5. The method according to claim 1, characterized in that, The initial network model includes a first sample branch model and a second sample branch model; the second sample branch model includes a semantic extraction network, a recursive attention network, and a memory network; the first sample branch model is used to extract sample image features of the sample picture; The semantic extraction network is used to extract first sample text features of the first sample fragment; the fragment elements of the first sample fragment include a first fragment element and a second fragment element; the second fragment element is the next fragment element of the first fragment element; Inputting the first sample text features of the first sample fragment and the sample image features of the sample picture into the recursive attention network in the initial network model, determining first graphic-text relationship features between the first sample fragment and the sample picture through the recursive attention network, outputting a first predicted text of the first sample fragment based on the first graphic-text relationship features, and adding the first predicted text to the memory network in the initial network model includes: Based on the element position of the first fragment element in the first sample fragment, obtain a first element auxiliary feature associated with the first fragment element from the first sample text feature of the first sample fragment; Based on the first element auxiliary feature, the sample image feature of the sample picture, the first attention layer in the recursive attention network, and the first selection gate corresponding to the first attention layer, determine a first weight relationship feature between the first fragment element and the sample picture, and input the first weight relationship feature into the first language layer in the recursive attention network, and output a first predicted element corresponding to the first fragment element by the first language layer; Recursively use the first element feature of the first predicted element as a second element auxiliary feature associated with the second fragment element. Based on the second element auxiliary feature, the sample image feature, the second attention layer in the recursive attention network, and the second selection gate corresponding to the second attention layer, determine a second weight relationship feature between the second fragment element and the sample picture, and input the second weight relationship feature into the second language layer in the recursive attention network, and output a second predicted element corresponding to the second fragment element by the second language layer; Based on the first weight relationship feature and the second weight relationship feature, determine a first text-image relationship feature between the first sample fragment and the sample picture, and output a first predicted text of the first sample fragment based on the first text-image relationship feature; the first predicted text includes the first predicted element and the second predicted element; Add the first predicted text including the first predicted element and the second predicted element to the memory network in the initial network model.
6. The method according to claim 5, characterized in that The method of determining a first weight relationship feature between the first fragment element and the sample picture based on the first element auxiliary feature, the sample image feature of the sample picture, the first attention layer in the recursive attention network, and the first selection gate corresponding to the first attention layer, and inputting the first weight relationship feature into the first language layer in the recursive attention network, and outputting a first predicted element corresponding to the first fragment element by the first language layer includes: Use the first element auxiliary feature and the sample image feature of the sample picture as a first input feature associated with the first sample fragment; Input the first input feature into the first attention layer in the recursive attention network, and the first attention layer determines a first attention map feature of the first fragment element in the sample picture; Input the first attention map feature and the first element auxiliary feature into the first selection gate corresponding to the first attention layer, and the first selection gate selects and outputs a first weight relationship feature between the first fragment element and the sample picture; Input the first weight relationship feature into the first language layer in the recursive attention network, and the first language layer performs element recognition on the first fragment element to obtain a first predicted element corresponding to the first fragment element.
7. The method according to claim 6, wherein Inputting the first attention map feature and the first element auxiliary feature into the first selection gate corresponding to the first attention layer, and selecting and outputting the first weight relationship feature between the first fragment element and the sample image by the first selection gate, includes: Inputting the first attention map feature and the first element auxiliary feature into the first selection gate corresponding to the first attention layer, extracting the first inner product similarity between the first attention map feature and the first element auxiliary feature through the first selection rule in the first selection gate, and taking the product between the first weight coefficient indicated by the first selection rule and the first inner product similarity as the first similarity feature; Extracting the first Gaussian similarity between the first attention map feature and the first element auxiliary feature through the second selection rule in the first selection gate, and taking the product between the second weight coefficient indicated by the second selection rule and the first Gaussian similarity as the second similarity feature; Extracting the first string similarity between the first attention map feature and the first element auxiliary feature through the third selection rule in the first selection gate, and taking the product between the third weight coefficient indicated by the third selection rule and the first string similarity as the third similarity feature; Based on the first similarity feature, the second similarity feature, and the third similarity feature, determining the first weight matrix between the first fragment element and the sample image, and performing masked multiplication on the first weight matrix and the first attention map feature to obtain the first weight relationship feature between the first fragment element and the sample image.
8. The method according to claim 7, wherein The step of determining the first weight matrix between the first fragment element and the sample image based on the first similarity feature, the second similarity feature, and the third similarity feature, and performing masked multiplication on the first weight matrix and the first attention map feature to obtain the first weight relationship feature between the first fragment element and the sample image, includes: Performing feature fusion on the first similarity feature, the second similarity feature, and the third similarity feature to obtain a first fusion feature; Performing normalization processing on the first fusion feature to obtain the first weight matrix between the first fragment element and the sample image; Performing masked multiplication on the first weight matrix and the first attention map feature to obtain the first weight relationship feature between the first fragment element and the sample image.
9. The method according to claim 5, characterized in that, Recursively using the first element feature of the first predicted element as the second element auxiliary feature associated with the second fragment element, and based on the second element auxiliary feature, the sample image feature, the second attention layer in the recursive attention network, and the second selection gate corresponding to the second attention layer, determining the second weight relationship feature between the second fragment element and the sample image, and inputting the second weight relationship feature into the second language layer in the recursive attention network, and outputting the second predicted element corresponding to the second fragment element by the second language layer, includes: Recursively use the first element feature of the first prediction element as the second element auxiliary feature associated with the second fragment element, and use the second element auxiliary feature and the sample image feature as the second input feature associated with the first sample fragment; Input the second input feature into the second attention layer in the recursive attention network, and the second attention layer determines the second attention map feature of the second fragment element in the sample picture; Input the second attention map feature and the second element auxiliary feature into the second selection gate corresponding to the second attention layer, and the second selection gate selects and outputs the second weight relationship feature between the second fragment element and the sample picture; Input the second weight relationship feature into the second language layer in the recursive attention network, and the second language layer performs element recognition on the second fragment element to obtain the second prediction element corresponding to the second fragment element.
10. The method according to claim 9, wherein The step of inputting the second attention map feature and the second element auxiliary feature into the second selection gate corresponding to the second attention layer, and the second selection gate selects and outputs the second weight relationship feature between the second fragment element and the sample picture includes: Input the second attention map feature and the second element auxiliary feature into the second selection gate corresponding to the second attention layer, extract the second inner product similarity between the second attention map feature and the second element auxiliary feature through the first selection rule in the second selection gate, and use the product of the first weight coefficient indicated by the first selection rule and the second inner product similarity as the fourth similarity feature; Extract the second Gaussian similarity between the second attention map feature and the second element auxiliary feature through the second selection rule in the second selection gate, and use the product of the second weight coefficient indicated by the second selection rule and the second Gaussian similarity as the fifth similarity feature; Extract the second string similarity between the second attention map feature and the second element auxiliary feature through the third selection rule in the second selection gate, and use the product of the third weight coefficient indicated by the third selection rule and the second string similarity as the sixth similarity feature; Based on the fourth similarity feature, the fifth similarity feature, and the sixth similarity feature, determine the second weight matrix between the second fragment element and the sample picture, and perform masked multiplication on the second weight matrix and the second attention map feature to obtain the second weight relationship feature between the second fragment element and the sample picture.
11. The method according to claim 1, characterized in that, The fragment lengths of the first sample fragment and the second sample fragment are both determined by the text segmentation parameters of the initial network model; the first predicted text is determined by the second sample branch model in the initial network model when the cumulative number of predicted elements output by the language layer in the recursive attention network reaches the text segmentation parameter; the predicted element is determined by the recursive attention network after feature association of the fragment elements in the first sample fragment and the sample picture; Using the first predicted text stored in the memory network as the training auxiliary text for the second sample fragment, and inputting the first sample text feature, the second sample text feature, and the sample image feature corresponding to the training auxiliary text into the recursive attention network, determining the second text-picture relationship feature between the second sample fragment and the sample picture through the recursive attention network, outputting the second predicted text of the second sample fragment based on the second text-picture relationship feature, and adding the second predicted text to the memory network, includes: Using the first predicted text stored in the memory network as the training auxiliary text for the second sample fragment, and using the first sample text feature corresponding to the training auxiliary text as the training auxiliary text feature; Inputting the training auxiliary text feature, the second sample text feature, and the sample image feature into the recursive attention network, and performing feature association between the second sample fragment and the sample picture through the recursive attention network to obtain a second text-picture relationship feature for characterizing the association relationship between the second sample fragment and the sample picture; Inputting the second text-picture relationship feature into the language layer in the recursive attention network, and outputting, by the language layer in the recursive attention network, predicted elements associated with the fragment elements in the second sample fragment; When it is detected that the number of predicted elements output by the language layer in the recursive attention network reaches the text segmentation parameter, outputting the second predicted text of the second sample fragment based on the predicted elements output by the language layer in the recursive attention network, and adding the second predicted text to the memory network.
12. A method for processing text data, characterized in that, Includes: Obtaining a target picture carrying target text information, and extracting the target image feature of the target picture through a target network model; The target text information is the text information formed by the characters in the target picture, and the target network model is obtained by training an initial network model based on the method according to any one of the above claims 1-11; Determining a first association relationship feature between a first target fragment in the target text information and the target picture based on the target image feature, and outputting a first target text of the first target fragment based on the first association relationship feature; Use the first target text as the target auxiliary text for the second target fragment in the target text information. Based on the first target text feature corresponding to the target auxiliary text and the target image feature, determine the second correlation relationship feature between the second target fragment and the target picture. Based on the second correlation relationship feature, output the second target text of the second target fragment. Based on the first target text and the second target text, determine the target text information recognized from the target picture.
13. A text data processing device, characterized in that, Including, A sample picture acquisition module, configured to acquire a sample picture carrying a sample training label, perform text segmentation on the training text information indicated by the sample training label to obtain a first sample fragment and a second sample fragment for training an initial network model; the second sample fragment is the next sample fragment of the first sample fragment; in the text and image recognition scenario, the presentation form of the training text information in the sample picture includes printed text or handwritten text. A first relationship determination module, configured to input the first sample text feature of the first sample fragment and the sample image feature of the sample picture into the recurrent attention network in the initial network model, determine the first text and image relationship feature between the first sample fragment and the sample picture through the recurrent attention network, output the first predicted text of the first sample fragment based on the first text and image relationship feature, and add the first predicted text to the memory network in the initial network model. A second relationship determination module, configured to use the first predicted text stored in the memory network as the training auxiliary text for the second sample fragment, input the first sample text feature corresponding to the training auxiliary text, the second sample text feature of the second sample fragment, and the sample image feature into the recurrent attention network, determine the second text and image relationship feature between the second sample fragment and the sample picture through the recurrent attention network, output the second predicted text of the second sample fragment based on the second text and image relationship feature, and add the second predicted text to the memory network. A model training module, configured to determine the sample prediction label of the training text information based on the first predicted text and the second predicted text in the memory network, perform iterative training on the initial network model based on the sample training label and the sample prediction label, and use the iteratively trained initial network model as the target network model for text recognition of the target text information in the target picture; the target text information is the text information composed of characters in the target picture.
14. A text data processing device, characterized in that, Including, A target picture acquisition module, configured to acquire a target picture carrying target text information, and extract the target image feature of the target picture through the target network model. The target text information is the text information composed of characters in the target picture, and the target network model is obtained by training the initial network model based on the method according to any one of the above claims 1-11. A first text output module, configured to determine a first association relationship feature between a first target fragment in the target text information and the target picture based on the target image feature, and output a first target text of the first target fragment based on the first association relationship feature; A second text output module, configured to use the first target text as a target auxiliary text of a second target fragment in the target text information, determine a second association relationship feature between the second target fragment and the target picture based on a first target text feature corresponding to the target auxiliary text and the target image feature, and output a second target text of the second target fragment based on the second association relationship feature; A target text determination module, configured to determine the target text information recognized from the target picture based on the first target text and the second target text.
15. A computer device, characterized in that, Comprising: A processor and a memory; The processor is connected to the memory, wherein the memory is used to store a computer program, and the processor is used to call the computer program so that the computer device executes the method according to any one of claims 1-12.
16. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and the computer program is suitable for being loaded and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1-12.
17. A computer program product, characterized in that, The computer program product includes computer instructions, and the computer instructions are stored in the computer-readable storage medium. The computer instructions are suitable for being read and executed by a processor so that a computer device having the processor executes the method according to any one of claims 1-12.
Citation Information
Patent Citations
Semantic representation model pre-training method and device and storage medium
CN112232089A