Text recognition method and device, equipment and storage medium

By inserting feature fusion layer into the multi-task recognition model and adopting a multi-stage learning method, the problems of mutual influence and conflict between tasks in multi-task learning are solved, and prediction accuracy and task processing efficiency are improved.

CN120032379APending Publication Date: 2025-05-23TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311561883.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In multitask learning, since the underlying network is shared in the model, some tasks may have a negative effect on other tasks, resulting in mutual "influence and conflict" between tasks, affecting the prediction effect of the model, and reducing the prediction accuracy.

Method used

By inserting a feature fusion layer into the target multi-task recognition model, fusing the output results of the feedforward neural network layer and the feedforward neural network adjustment layer corresponding to each text recognition task, a second multi-task recognition model is formed, and through a multi-stage gradual learning method, the model can gradually learn single-task and multi-task knowledge.

Benefits of technology

The prediction accuracy of the ultimate target multi-task recognition model is improved, thereby improving task processing efficiency and reducing computing resource consumption during task processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032379A_ABST
    Figure CN120032379A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and particularly relates to a text recognition method and device, equipment and a storage medium. The method comprises the steps of obtaining a to-be-recognized text corresponding to a recognition task identifier; inputting the to-be-recognized text and the recognition task identifier corresponding to the to-be-recognized text into the target multi-task recognition model for text recognition processing to obtain a target text recognition result; the target multi-task recognition model training method comprises the following steps: inserting a feature fusion layer into a feature extraction module in a first multi-task recognition model to obtain a second multi-task recognition model; and training the second multi-task recognition model based on the sample recognition text and the corresponding sample recognition task identifier and the sample recognition tag to obtain a target multi-task recognition model. The method can be applied to various scenes such as the map field, the traffic field, the automatic driving field, the vehicle-mounted scene, the cloud technology, artificial intelligence, intelligent traffic and auxiliary driving, and the text recognition task processing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of artificial intelligence technology, and specifically relates to a text recognition method, device, equipment and storage medium. Background Art

[0002] With the development of artificial intelligence technology, multi-task learning (MTL) is increasingly used in the process of machine learning. Multi-task learning, as the name implies, is the training of multiple tasks at the same time. Traditional machine learning is for single tasks (single-task learning), such as sentiment classification and named entity recognition (NER) in natural language processing, which often trains the corresponding model independently for a single task. Multi-task learning, on the other hand, puts all the training data together to train a unified model.

[0003] In related technologies, when implementing multi-task learning, multiple tasks usually share a base model, and then specific layers for different tasks are added on top of this base model, so that independent predictions can be made for each task on these task-specific layers. However, since the underlying network in this multi-task learning implementation is shared, some tasks may have a negative effect on other tasks, which may lead to problems of "influence and conflict" between tasks. This affects the prediction effect of the model, resulting in a decrease in prediction accuracy, which in turn affects the inefficiency of task processing and increases the consumption of computing resources during task processing. Summary of the invention

[0004] In order to solve the above technical problems, the present application provides a text recognition method, device, equipment and storage medium.

[0005] On the one hand, an embodiment of the present application provides a text recognition method, the method comprising:

[0006] Obtain the text to be recognized corresponding to the recognition task identifier;

[0007] Input the text to be recognized and its corresponding recognition task identifier into the target multi-task recognition model for text recognition processing, and obtain the target text recognition result corresponding to the text to be recognized;

[0008] Among them, the training method of the target multi-task recognition model includes:

[0009] A sample text data set and a first multi-task recognition model are obtained; the sample recognition text in the sample text data set corresponds to a sample recognition task identifier and a sample recognition label; the feature extraction module in the first multi-task recognition model includes a feedforward neural network layer and a feedforward neural network adjustment layer corresponding to each text recognition task; the first multi-task recognition model is obtained by performing multi-task recognition training and single-task recognition adjustment training on the initial multi-task recognition model;

[0010] Inserting a feature fusion layer into the feature extraction module in the first multi-task recognition model to obtain a second multi-task recognition model; the feature fusion layer is used to fuse the output results of the feedforward neural network layer and each feedforward neural network adjustment layer;

[0011] The second multi-task recognition model is trained based on the sample recognition text and its corresponding sample recognition task identifier and the sample recognition label to obtain a target multi-task recognition model.

[0012] On the other hand, an embodiment of the present application further provides a text recognition device, the device comprising:

[0013] The module for obtaining text to be recognized is used to obtain the text to be recognized corresponding to the recognition task identifier;

[0014] The target text recognition result determination module is used to input the text to be recognized and its corresponding recognition task identifier into the target multi-task recognition model for text recognition processing to obtain the target text recognition result corresponding to the text to be recognized;

[0015] The device further comprises a target multi-task recognition model training device; the target multi-task recognition model training device comprises:

[0016] The data set and the first multi-task recognition model acquisition submodule are used to acquire a sample text data set and a first multi-task recognition model; the sample recognition text in the sample text data set corresponds to a sample recognition task identifier and a sample recognition label; the feature extraction module in the first multi-task recognition model includes a feedforward neural network layer and a feedforward neural network adjustment layer corresponding to each text recognition task; the first multi-task recognition model is obtained by performing multi-task recognition training and single-task recognition adjustment training on the initial multi-task recognition model;

[0017] A second multi-task recognition model determination submodule is used to insert a feature fusion layer into the feature extraction module in the first multi-task recognition model to obtain a second multi-task recognition model; the feature fusion layer is used to fuse the output results of the feedforward neural network layer and each feedforward neural network adjustment layer;

[0018] The target multi-task recognition model determination submodule is used to train the second multi-task recognition model based on the sample recognition text and its corresponding sample recognition task identifier and sample recognition label to obtain the target multi-task recognition model.

[0019] On the other hand, an embodiment of the present application also provides an electronic device for text recognition, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the text recognition method as described above.

[0020] On the other hand, an embodiment of the present application also provides a computer-readable storage medium, in which at least one instruction or at least one program is stored, and the at least one instruction or the at least one program is loaded and executed by a processor to implement the text recognition method as described above.

[0021] On the other hand, an embodiment of the present application further provides a computer program product, which implements the text recognition method as described above when the computer program is executed by a processor.

[0022] The text recognition method, device, electronic device and storage medium proposed in the embodiment of the present application use the target multi-task recognition model to perform text recognition on the text to be recognized. Since the target multi-task recognition model is obtained by training the second multi-task recognition model based on the sample recognition text corresponding to a plurality of text recognition tasks, the trained target multi-task recognition model has multi-task processing capability. At the same time, since the second multi-task recognition model is obtained by inserting a feature fusion layer into the feature extraction module in the first multi-task recognition model, and the first multi-task recognition model is obtained by performing multi-task recognition training and single-task recognition adjustment training on the initial multi-task recognition model, through a multi-stage progressive learning method, the model gradually learns single-task and multi-task knowledge, so that the tasks promote each other as much as possible instead of conflicting with each other. In addition, by setting a feedforward neural network layer and a feedforward neural network adjustment layer corresponding to each text recognition task in the feature extraction module in the first multi-task recognition model, and using a feature fusion layer to fuse the output results of the feedforward neural network layer and each feedforward neural network adjustment layer, the prediction accuracy of the target multi-task recognition model finally obtained can be improved, thereby improving the task processing efficiency, and thus reducing the computing power resource consumption during the task processing process. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0024] Figure 1 The figure is a schematic diagram of an implementation environment of a text recognition method according to an exemplary embodiment.

[0025] Figure 2 The figure is a flowchart of a text recognition method according to an exemplary embodiment.

[0026] Figure 3 It is a schematic diagram of a partial structure of a feature extraction module of a second multi-task recognition model according to an exemplary embodiment.

[0027] Figure 4 It is a structural diagram of a feature extraction module of an initial multi-task recognition model according to an exemplary embodiment.

[0028] Figure 5 It is a schematic diagram of a partial structure of a feature extraction module of a fourth multi-task recognition model according to an exemplary embodiment.

[0029] Figure 6 The present invention is a schematic diagram of a target multi-task recognition model determination process according to an exemplary embodiment.

[0030] Figure 7 The figure is a block diagram of a text recognition device according to an exemplary embodiment.

[0031] Figure 8 The present invention is a hardware structure block diagram of a server of a text recognition method provided according to an exemplary embodiment. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0033] It should be noted that the terms "first", "second", etc. in the description and claims of the embodiments of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features.

[0034] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0035] In the embodiments of the present application, unless otherwise specified, "plurality" means two or more than two. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or server including a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0036] In order to make the purpose, technical solution and advantages disclosed in the embodiments of the present application more clearly understood, the embodiments of the present application are further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the embodiments of the present application and are not used to limit the embodiments of the present application.

[0037] The embodiments of the present application relate to artificial intelligence (AI), machine learning (ML) technology and natural language processing (NLP).

[0038] Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0039] Artificial intelligence is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation interaction systems, mechatronics and other technologies; the software technologies of artificial intelligence generally include computer vision technology, natural language processing technology, and machine learning / deep learning and other major directions. With the development and progress of artificial intelligence, artificial intelligence is being studied and applied in many fields, such as common smart homes, smart customer service, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, robots, smart medical care, etc. It is believed that with the further development of future technology, artificial intelligence will be applied in more fields and play an increasingly important role.

[0040] Machine learning is a multi-disciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance.

[0041] Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Deep learning is the core of machine learning and a technology to achieve machine learning. Machine learning usually includes deep learning, reinforcement learning, transfer learning, inductive learning and other technologies. Deep learning includes mobile visual neural network (Mobilenet), convolutional neural network (CNN), deep belief network, recursive neural network, autoencoder, generative adversarial network and other technologies.

[0042] Natural language processing is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.

[0043] In the process of single-task machine learning, the training data sets are independent. However, in practical applications, the training data sets may be related to each other. For example, Task A is a general NER task, and Task B is a NER task in the gaming industry. Task A will recognize hundreds of NER types such as names of people and places, while Task B needs to recognize game names, game characters, props, etc. If Task A is used in conjunction with Task B, it will definitely be of great help to Task B. Therefore, multi-task learning is an efficient training paradigm in the field of NLP.

[0044] In related technologies, solutions for implementing multi-task learning all have a common problem, which is that since the underlying network is shared, it is difficult to avoid the problem of "influence and conflict" between tasks. In other words, some tasks may have a negative effect on other tasks, which will reduce the task processing effect of the model.

[0045] In view of this, the embodiment of the present application proposes a text recognition method, device, electronic device and storage medium, which uses a target multi-task recognition model to perform text recognition on the text to be recognized. Since the target multi-task recognition model is obtained by training the second multi-task recognition model based on the sample recognition text corresponding to a plurality of text recognition tasks, the trained target multi-task recognition model has multi-task processing capability. At the same time, since the second multi-task recognition model is obtained by inserting a feature fusion layer into the feature extraction module in the first multi-task recognition model, and the first multi-task recognition model is obtained by performing multi-task recognition training and single-task recognition adjustment training on the initial multi-task recognition model, through a multi-stage progressive learning method, the model gradually learns single-task and multi-task knowledge, so that the tasks promote each other as much as possible instead of conflicting with each other. In addition, by setting a feedforward neural network layer and a feedforward neural network adjustment layer corresponding to each text recognition task in the feature extraction module in the first multi-task recognition model, and using a feature fusion layer to fuse the output results of the feedforward neural network layer and each feedforward neural network adjustment layer, the prediction accuracy of the target multi-task recognition model finally obtained can be improved, thereby improving the task processing efficiency, and thus reducing the computing power resource consumption during the task processing process.

[0046] See also Figure 1 , Figure 1 FIG. 1 is a schematic diagram showing an implementation environment of a text recognition method according to an exemplary embodiment. Figure 1 As shown, the application environment may include at least a server 01 and a terminal 02. In actual applications, the server 01 and the terminal 02 may be directly or indirectly connected via wired or wireless communication, which is not limited in the present embodiment.

[0047] In the embodiment of the present application, server 01 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. Specifically, the servers involved above may include physical devices, which may specifically include network communication submodules, processors and memories, etc., and may also include software running in physical devices, which may specifically include applications, etc.

[0048] In the embodiment of the present application, terminal 02 may include physical devices such as smart phones, desktop computers, tablet computers, laptops, digital assistants, augmented reality (AR) / virtual reality (VR) devices, intelligent voice interaction devices, smart home appliances, smart wearable devices, and vehicle-mounted terminal devices, and may also include software running in physical devices, such as applications.

[0049] In an embodiment of the present application, server 01 may be a background server that provides services for the target application. The background server may be used to manage the version of the target application, process the text obtained by the target application and return the processing results to terminal 02, or perform background training on the machine learning model developed by the developer, etc.

[0050] In the embodiment of the present application, terminal 02 may include a terminal used by a developer or a user. When terminal 02 is a terminal used by a developer, the developer may develop a machine learning model for performing a specified text recognition task on a text through terminal 02, and deploy the machine learning model to server 01 or a terminal used by a user. When terminal 02 is a terminal used by a user, a target application for obtaining and presenting processing results of text may be installed in terminal 02. After terminal 02 obtains the text, the processing results obtained by performing a specified text recognition task on the text may be obtained through the above target application, and the processing results may be presented.

[0051] It should be noted that Figure 1 This is just an example. In other scenarios, other implementation environments may also be included.

[0052] Figure 2 is a flowchart of a text recognition method according to an exemplary embodiment. The text recognition method can be used for Figure 1 In the implementation environment of the embodiment. This specification provides the method operation steps as described in the embodiment or flowchart, but may include more or fewer operation steps based on routine or non-creative work. The order of steps listed in the embodiment is only one way of executing the steps among many steps, and does not represent the only execution order. When the actual system or server product is executed, it can be executed in sequence or in parallel (for example, in a parallel processor or multi-threaded processing environment) according to the method shown in the embodiment or the drawings. Specifically, Figure 2 As shown, the method may include:

[0053] S101: Acquire a text to be recognized corresponding to a recognition task identifier.

[0054] In an embodiment of the present application, the text to be recognized may be a field, a sentence, a document, an email, a web page, etc. When recognizing the text to be recognized, it is necessary to simultaneously obtain the recognition task identifier corresponding to the text to be recognized. The recognition task identifier is used to identify the type of text recognition task to be performed on the text to be recognized, such as part-of-speech tagging, named entity recognition, sentiment analysis, text classification, entity relationship extraction, topic modeling, etc. Optionally, the recognition task identifier may be the same natural language as the text to be recognized, or it may be a predefined character or string. For example, if the recognition task identifier is "NER", the recognition task represented is "named entity recognition".

[0055] S103: Input the text to be recognized and its corresponding recognition task identifier into the target multi-task recognition model for text recognition processing to obtain the target text recognition result corresponding to the text to be recognized; wherein the target multi-task recognition model is obtained by training the second multi-task recognition model based on the sample text data set; the sample text data set includes sample recognition texts corresponding to at least two text recognition tasks; the second multi-task recognition model is obtained by inserting a feature fusion layer into the feature extraction module in the first multi-task recognition model; the first multi-task recognition model is obtained by performing multi-task recognition training and single-task recognition adjustment training on the initial multi-task recognition model; the feature extraction module in the first multi-task recognition model includes a feedforward neural network layer and a feedforward neural network adjustment layer corresponding to each text recognition task; the feature fusion layer is used to fuse the output results of the feedforward neural network layer and each feedforward neural network adjustment layer.

[0056] In an embodiment of the present application, a target multi-task recognition model can be used to recognize the text to be recognized. The target multi-task recognition model is used to perform a variety of text recognition tasks, such as part-of-speech tagging, named entity recognition, sentiment analysis, text classification, entity relationship extraction, topic modeling, etc. When recognizing the text to be recognized, the recognition task corresponding to the text to be recognized needs to be input into the target multi-task recognition model together, so that the target multi-task recognition model can output the corresponding task recognition result.

[0057] In an embodiment of the present application, the target multi-task recognition model is obtained by training the second multi-task recognition model, and the second multi-task recognition model is obtained by inserting a feature fusion layer into the first multi-task recognition model.

[0058] Specifically, when training a target multi-task recognition model, a sample text dataset and a first multi-task recognition model may be first obtained.

[0059] The sample text data set may include sample recognition texts corresponding to two or more text recognition tasks. The types of sample recognition texts in the sample text data set can be determined according to the types of text recognition tasks that the target multi-task recognition model needs to process. For example, the text recognition tasks that the target multi-task recognition model needs to process include named entity recognition, sentiment analysis, and text classification. Then the sample text data set needs to include sample recognition texts corresponding to the three text recognition tasks of named entity recognition, sentiment analysis, and text classification. Optionally, the number of sample recognition texts corresponding to each text recognition task can be multiple, so as to improve the accuracy of the target multi-task recognition model finally trained. For each sample recognition text in the sample text data set, there is a sample recognition task identifier and a sample recognition label. The sample recognition task identifier is used to identify the text recognition task corresponding to the sample recognition text, and the sample recognition label is used to identify the text recognition result corresponding to the sample recognition text. The text recognition result corresponding to the sample recognition text can be obtained by pre-marking, or it can be determined by the recognition result of the historical recognition text.

[0060] The first multi-task recognition model includes a feature extraction module. Optionally, the feature extraction module in the first multi-task recognition model can use an attention mechanism to improve the natural language processing model of the model training speed, such as a Transformer model. In some embodiments, the feature extraction module in the first multi-task recognition model may only include the encoder structure or decoder structure of the Transformer model, or other models obtained based on the Transformer model transformation, such as a bidirectional encoder representation (Bidirectional Encoder Representation from Transformers, Bert) model based on a transformer. In other embodiments, the feature extraction module in the first multi-task recognition model may also be a long short-term memory network (Long Short-Term Memory, LSTM), a gated recurrent unit (Gated Recurrent Unit, GRU) model, etc.

[0061] The feature extraction module of the first multi-task recognition model is provided with a feedforward neural network layer and a feedforward neural network adjustment layer corresponding to each text recognition task. Taking the Transformer model as an example, the encoder and decoder of the traditional Transformer model are both provided with a feedforward neural network layer. Different from the traditional Transformer model, in the feature extraction module of the first multi-task recognition model, in addition to the feedforward neural network layer in the encoder and decoder, a feedforward neural network adjustment layer corresponding to each text recognition task is also provided next to the feedforward neural network layer.

[0062] It should be noted that the above-mentioned first multi-task recognition model is a pre-trained model, and the adjustable parameters in each module in the model, such as the parameters in the feedforward neural network layer and the feedforward neural network adjustment layer in the feature extraction module, are all obtained through model training.

[0063] In the embodiment of the present application, after obtaining the first multi-task recognition model, a second multi-task recognition model can be obtained by inserting a feature fusion layer into the feature extraction module in the first multi-task recognition model. Specifically, the feature fusion layer is inserted above the feedforward neural network layer and the feedforward neural network adjustment layer. The feature fusion layer can fuse the output results of the feedforward neural network layer and each feedforward neural network adjustment layer. As an example, Figure 3 is a schematic diagram of a partial structure of a feature extraction module of a second multi-task recognition model according to an exemplary embodiment. Figure 3As shown, the semantic feature data obtained before the feedforward neural network layer and the feedforward neural network adjustment layer can be respectively input into the feedforward neural network layer and each feedforward neural network adjustment layer for processing. The processing results of the feedforward neural network layer and each feedforward neural network adjustment layer will be input into the feature fusion layer for fusion, and then the feature fusion layer will output a feature fusion data. The feature fusion layer processes the input data by first calculating the similarity of the processing results output by the feedforward neural network layer and each feedforward neural network adjustment layer and a preset parameter embedding M. Optionally, the similarity can be calculated using the inner product. Then, the calculated similarity data is converted into a probability distribution through a normalization function (such as softmax). Then, the obtained probability distribution is weighted and summed. Finally, the final feature fusion data is obtained through a nonlinear transformation. The above processing process can be expressed by the following formula:

[0064] S i =dot(V delta_i , M)

[0065]

[0066] V fusion_token =a 1 V main_fnn +a 2 V delta_2 +…+a n V delta_n

[0067] V new_token_i =relu(WV fusion_token )

[0068] Among them, M is a preset parameter, which is a parameter vector introduced by the feature fusion layer. It can be randomly initialized and adjusted in subsequent training. dot is the inner product. V delta It is the semantic feature data output by each feedforward neural network adjustment layer. 1 、a 2 ...a n is the weight parameter learned through model training. W is a parameter matrix in the feature fusion layer, which can be obtained through model training.

[0069] After obtaining the second multi-task recognition model, the sample text data set can be used to train the second multi-task recognition model. When training the second multi-task recognition model, for any sample recognition text in the sample text data set, the sample recognition text and its corresponding sample recognition task identifier can be input into the second multi-task recognition model for text recognition processing, and the second multi-task recognition model can output the first text recognition result corresponding to the sample recognition text. After obtaining the first text recognition result. The parameters in the feature fusion layer can be adjusted based on the first text recognition result and the sample recognition label until the difference between the first text recognition result and the sample recognition label meets the preset conditions to obtain the target multi-task recognition model. That is to say, when training the second multi-task recognition model, since the second multi-task recognition model is obtained by inserting the feature fusion layer into the feature extraction module of the first multi-task recognition model, and the first multi-task recognition model has been trained. Therefore, when training the second multi-task recognition model, only the newly inserted feature fusion layer can be trained, that is, only M and a in the above formulas are trained. i , W. All other parameters in the model are fixed and do not participate in training. The advantage of doing this is to maximize the use of the knowledge learned in the first multi-task recognition model. In this process, only the output results of the feedforward neural network layer and each feedforward neural network adjustment layer are fused, so that the target multi-task recognition model can be applied to a variety of text recognition tasks, output accurate feature extraction results, improve the training efficiency of the model, and reduce resource consumption during model training.

[0070] In an embodiment of the present application, the first multi-task recognition model can be obtained based on the initial multi-task recognition model. Specifically, when determining the first multi-task recognition model, the initial multi-task recognition model can be obtained first, and then the initial multi-task recognition model can be trained for multi-task recognition based on a sample text data set to obtain a third multi-task recognition model, and the third multi-task recognition model can be trained for single-task recognition adjustment based on the sample text data set to obtain the first multi-task recognition model. By first training the initial multi-task recognition model, the third multi-task recognition model obtained can be made to have multi-text task recognition capabilities, and then single-task adjustment training can be performed on the third multi-task recognition model to improve the accuracy of the second multi-task model obtained on each text recognition task.

[0071] Specifically, when the initial multi-task recognition model is trained, the sample recognition text and its corresponding sample recognition task identifier can be respectively input into the initial multi-task recognition model for text recognition processing to obtain a second text recognition result corresponding to the sample recognition text. Then, based on the second text recognition result and the sample recognition label, the parameters in the initial multi-task recognition model are adjusted until the difference between the second text recognition result and the sample recognition label meets the preset conditions, thereby obtaining a third multi-task recognition model. By performing multi-task recognition training on the initial multi-task recognition model based on the sample text data set, the model can learn multi-task knowledge, so that the obtained third multi-task recognition model has multi-task recognition capabilities, thereby improving the work efficiency of the model application and the resource utilization of the system.

[0072] In practical applications, training the initial multi-task recognition model can be regarded as the process of training a general multi-task recognition model. The purpose is to enable the model to learn multi-task knowledge and enable it to have multi-task recognition capabilities.

[0073] The initial multi-task recognition model may include a feature extraction module and a recognition result output module. The feature extraction module may be a structure including an encoder and / or a decoder of a Transformer model. In some embodiments, the feature extraction module of the initial multi-task recognition model may also be a model structure such as Bert, LSTM, GRU, etc. The semantic feature data extracted by the feature extraction module may be input into the recognition result output module for processing, thereby inputting the text recognition result.

[0074] It should be noted that the recognition result output module in the embodiment of the present application can output the recognition results of multiple text tasks. In this way, there is no need to set up a recognition result output module corresponding to each task, thereby simplifying the model structure and reducing resource overhead during the model training process.

[0075] As an example, Figure 4 is a structural diagram of a feature extraction module of an initial multi-task recognition model according to an exemplary embodiment. Figure 4 As shown, the feature extraction module of the initial multi-task recognition model can be a transformer model, which can be composed of one or more feature extraction submodules stacked. Each feature extraction submodule includes an attention calculation layer and a feedforward neural network layer arranged on the attention calculation layer. The attention calculation layer can perform attention calculation, and optionally, the attention calculation performed by the attention calculation layer includes but is not limited to self-attention calculation, multi-head attention calculation, etc.

[0076] When using a sample text dataset to train an initial multi-task recognition model, the sample recognition text in the sample text dataset may be preprocessed first. Specifically, preprocessing the sample recognition text may include organizing the sample recognition text into a sequence-sequence format. The sequence-sequence format refers to integrating the sample recognition task identifier, sample recognition text, and sample recognition label corresponding to the sample recognition text into a question-answer format. For example, for an article corresponding to a text classification task, it can be converted into the following format:

[0077] Please categorize the following articles:

[0078] {article}

[0079] Answer: This article is classified as {category}".

[0080] Optionally, when converting the sample recognition text into the above sequence-sequence format, the conversion may be performed based on a preset template corresponding to each recognition task.

[0081] In addition, in order to ensure the training effect of the initial multi-task recognition model, when determining the sample data for model training of the initial multi-task recognition model, it is necessary to ensure the uniformity of the sample data. Specifically, when determining the sample data for model training of the initial multi-task recognition model, the sample data for model training of the initial multi-task recognition model can be obtained from the sample text data set according to a certain sampling strategy. Specifically, by setting a sampling lower limit (M) and upper limit (N). If the number of sample recognition texts corresponding to a recognition task in the sample text data set is less than M, it is upsampled to M. If the number of sample recognition texts corresponding to a recognition task exceeds N, it is downsampled to N. In this way, the distribution of sample recognition texts corresponding to each recognition task in the sample data obtained by sampling can be more evenly distributed, thereby ensuring the training effect of the initial multi-task recognition model.

[0082] After sampling to obtain sample data for model training of the initial multi-task recognition model, the initial multi-task recognition model can be trained based on the sample data obtained by sampling. During the model training process, the model loss data can be determined based on the difference between the second text recognition result and the sample recognition label. Optionally, the loss data of the model can be determined based on a cross entropy loss function. The formula of the cross entropy loss function is as follows:

[0083]

[0084] Where n is the number of tokens in the sample recognition text input into the initial multi-task recognition model. labelIt is the probability value that the sample recognition text output by the initial multi-task recognition model belongs to the target label.

[0085] In an embodiment of the present application, a third multi-task recognition model can be obtained by performing model training on the initial multi-task recognition model. On the basis of the third multi-task recognition model, single-task recognition adjustment training can be performed to obtain a first multi-task recognition model. Specifically, when single-task recognition adjustment training is performed on the third multi-task recognition model, a preset number of third multi-task recognition models can be obtained, wherein the number of third multi-task recognition models is the same as the number of text recognition tasks. In the feature extraction module of each third multi-task recognition model, a feedforward neural network adjustment layer corresponding to the text recognition task is inserted to obtain a fourth multi-task recognition model corresponding to each text recognition task. Then, based on the sample recognition text corresponding to each text recognition task, single-task recognition adjustment training is performed on the fourth multi-task recognition model corresponding to each text recognition task to obtain a fifth multi-task recognition model corresponding to each text recognition task, and each fifth multi-task recognition model is fused to obtain the first multi-task recognition model. By inserting a feedforward neural network adjustment layer corresponding to the text recognition task into the third multi-task recognition model and training the model using a sample text data set, the performance of the model on each text recognition task can be enhanced. On the one hand, the recognition accuracy of the model on a single task can be improved. On the other hand, since each feedforward neural network adjustment layer is trained separately, the negative effects between different tasks can be avoided, thereby improving the accuracy of the final target multi-task recognition model on different recognition tasks. Moreover, during the model training process, parameters other than the feedforward neural network adjustment layer are adjusted, thereby improving the model training efficiency.

[0086] Specifically, by performing model training on the initial multi-task recognition model, the obtained third multi-task recognition model has the recognition capability of general multi-task. Then, for each single text recognition task, the performance of the model on each single text recognition task can be improved through single task recognition adjustment training. Before performing single task recognition adjustment training, it is necessary to insert a feedforward neural network adjustment layer into the third multi-task recognition model. The feedforward neural network adjustment layer corresponds to the single text recognition task, and only one feedforward neural network adjustment layer is inserted into each third multi-task recognition model to ensure that each text recognition task will not be affected by other tasks during the single task recognition adjustment training process.

[0087] After inserting a feedforward neural network adjustment layer into the third multi-task recognition model, a fourth multi-task recognition model can be obtained. It should be understood that the fourth multi-task recognition model actually corresponds to a single text recognition task, that is, each single text recognition task corresponds to a fourth multi-task recognition model. As an example, Figure 5 is a schematic diagram of a partial structure of a feature extraction module of a fourth multi-task recognition model according to an exemplary embodiment. Figure 5 As shown, in the feature extraction module of the fourth multi-task recognition model, in addition to the feedforward neural network layer that processes the main path, a branch of a feedforward neural network adjustment layer is added.

[0088] In the NLP model, the role of the feedforward neural network layer is usually to further extract local features. The feedforward neural network layer can capture the dependencies between words, thereby enhancing the semantic expression ability of the model. At the same time, the feedforward neural network layer can also reduce or increase the dimension of the input vector, thereby achieving feature compression or expansion, further enhancing the expression ability of the model. Moreover, before adding the feedforward neural network layer, the attention mechanism in the attention calculation can already capture the global dependencies, and the feedforward neural network layer can further extract local features based on the attention mechanism. Therefore, the embodiment of the present application can further extract the local features corresponding to a single recognition task by adding a branch of the feedforward neural network adjustment layer to the side of the feedforward neural network layer, thereby improving the performance of the model on a single recognition task.

[0089] After obtaining the fourth multi-task recognition model, the fourth multi-task recognition model can be trained. For any multi-task recognition model, when training the model, all model parameters will be frozen, that is, they will not participate in the single-task recognition adjustment training, and only the parameters in the newly added feedforward neural network adjustment layer will participate in the single-task recognition adjustment training. Each recognition task corresponds to independent feedforward neural network adjustment layer parameters.

[0090] In the fourth multi-task recognition model, the semantic feature data (token embedding) output by the attention calculation layer will be input into the feedforward neural network layer and the feedforward neural network adjustment layer for calculation respectively. For the newly added feedforward neural network adjustment layer, the calculation performed on the input semantic feature data is to first perform a linear transformation, and perform a nonlinear transformation on the result of the linear transformation through an activation function, and then obtain the final semantic feature data through a linear transformation. Among them, the linear transformation can be implemented using a fully connected layer, and the activation function can use a relu function or a gelu function, etc. As an example, the specific calculation formula for the input semantic feature data processed by the feedforward neural network adjustment layer is as follows:

[0091] V′ token_i = relu(W 1 V token_i )

[0092] V″ token_i =W 2 V′ token_i

[0093] Among them, W 1 and W 2 are two parameter matrices in the feedforward neural network adjustment layer. W 1 and W 2 It can be randomly initialized and trained together with the fourth multi-task recognition model. During the training of the fourth multi-task recognition model, other parameters in the fourth multi-task recognition model are fixed during the training. token_i is the semantic feature data input into the adjustment layer of the feedforward neural network, V′ token_i It is the result obtained after nonlinear transformation by activation function, V″ token_i It is the output semantic feature data of the feedforward neural network adjustment layer.

[0094] Finally, the semantic feature data output by the feedforward neural network adjustment layer is added to the semantic feature data output by the feedforward neural network layer to obtain the semantic feature data output by the current feature extraction submodule. If a feature extraction submodule is also provided above the current feature extraction submodule, the semantic feature data will be input into the feature extraction submodule thereon for processing. If a recognition result output module is provided above the current feature extraction submodule, the semantic feature data will be input into the recognition result output module for processing.

[0095] In an embodiment of the present application, when single-task recognition adjustment training is performed on the fourth multi-task recognition model corresponding to each text recognition task, the current sample recognition text corresponding to the current text recognition task and the current sample recognition task identifier corresponding to the current sample recognition text can be input into the current fourth multi-task recognition model for text recognition processing to obtain a third text recognition result corresponding to the current sample recognition text. Among them, the current text recognition task is any one of the at least two text recognition tasks, and the current fourth multi-task recognition model is the fourth multi-task recognition model corresponding to the current text recognition task. Then, based on the third text recognition result and the current sample recognition label corresponding to the current sample recognition text, the parameters in the feedforward neural network adjustment layer in the current fourth multi-task recognition model are adjusted until the difference between the third text recognition result and the current sample recognition label meets the preset conditions, and the fifth multi-task recognition model corresponding to the current text recognition task is obtained. Then, any recognition task other than the current text recognition task among the at least two text recognition tasks is re-used as the current text recognition task, and the steps of identifying the current sample recognition text corresponding to the current text recognition task and the current sample recognition task corresponding to the current sample recognition text are repeated until the fifth multi-task recognition model corresponding to the current text recognition task is obtained, and the fifth multi-task recognition model corresponding to each text recognition task is obtained. By using the sample text data corresponding to each text recognition task, the fourth multi-task recognition model corresponding to each text recognition task is respectively trained for single-task recognition adjustment, so as to ensure that the parameters in the adjustment layer of the feedforward neural network do not affect each other during the training process, and each independently learns its own parameters, thereby avoiding mutual influence between different recognition tasks and improving the accuracy of the model on different recognition tasks.

[0096] The process of training different fourth multi-task recognition models can be a continuous training process or an intermittent training process.

[0097] In an embodiment of the present application, after training all the fourth multi-task recognition models, the obtained fifth multi-task recognition models can be fused to obtain a first multi-task recognition model. Specifically, for any fifth multi-task recognition model, the feature extraction module in the fifth multi-task recognition model includes an attention calculation layer, a feedforward neural network layer arranged on the attention calculation layer, and a feedforward neural network adjustment layer. When all the fifth multi-task recognition models are fused, the attention calculation layer in each fifth multi-task recognition model can be fused to obtain a fused attention calculation layer, and the feedforward neural network layer in each fifth multi-task recognition model can be fused to obtain a fused feedforward neural network layer arranged on the fused attention calculation layer. Then the feedforward neural network adjustment layer in each fifth multi-task recognition model is inserted into the fused attention calculation layer to obtain the first multi-task recognition model.

[0098] In practical applications, all parameters of all the fifth multi-task recognition models are the same except for the feedforward neural network adjustment layer. Therefore, when the fifth multi-task recognition models are fused, only the feedforward neural network adjustment layers in different fifth multi-task recognition models can be inserted into the same fifth multi-task recognition model to obtain the first multi-task recognition model.

[0099] In an embodiment of the present application, by fusing the fifth multi-task recognition model, the first multi-task recognition model can contain a feedforward neural network adjustment layer corresponding to all recognition tasks, thereby ensuring the recognition accuracy of the target multi-task recognition model subsequently obtained based on the first multi-task recognition model on different recognition tasks.

[0100] Figure 6 FIG. 1 is a schematic diagram of a target multi-task recognition model determination process according to an exemplary embodiment. Figure 6 As shown, in practical applications, when determining the target multi-task recognition model, it is first necessary to obtain an initial multi-task recognition model, and then use the sample text data set to perform model training on the initial multi-task recognition model to obtain a third multi-task recognition model. Then, a feedforward neural network adjustment layer corresponding to the text recognition task is inserted into the feature recognition module of the third multi-task recognition model to obtain a fourth multi-task recognition model. The number of fourth multi-task recognition models is the same as the number of recognition tasks compatible with the target multi-task recognition model. For each fourth multi-task recognition model, a corresponding fifth multi-task recognition model can be obtained after model training. Then all the fifth multi-task recognition models are fused into one model to obtain a first multi-task recognition model. Then, a feature fusion layer is inserted into the feature extraction module of the first multi-task recognition model to obtain a second multi-task recognition model. Finally, the target multi-task recognition model is obtained by model training the second multi-task recognition model.

[0101] In the above process of determining the target multi-task recognition model, by allowing all tasks to share the backbone model parameters, the resource consumption in the model training process can be reduced and the model training efficiency can be improved. Moreover, by adding a feedforward neural network adjustment layer to each recognition task, the learning of specific tasks can be achieved, so that different tasks will not affect each other. In the above process, the model training process can be divided into general multi-task recognition training, single-task recognition adjustment training, and multi-task fusion adjustment training. Through this progressive learning method, the model gradually learns single-task and multi-task knowledge, so that the tasks can promote each other as much as possible instead of conflicting with each other, thereby greatly improving the effect of multi-task learning.

[0102] In the embodiment of the present application, the target multi-task model finally obtained includes a feature extraction module and a recognition result output module, wherein the feature extraction module of the target multi-task recognition model is provided with a target attention calculation layer, a target feedforward neural network layer, a target feedforward neural network adjustment layer corresponding to at least two text recognition tasks, and a target feature fusion layer, the target feedforward neural network layer and the target feedforward neural network adjustment layer corresponding to at least two text recognition tasks are arranged on the target attention calculation layer, and the target feature fusion layer is arranged on the target feedforward neural network layer and the target feedforward neural network adjustment layer corresponding to at least two text recognition tasks. The recognition result output module is used to output the task recognition results corresponding to the multiple text recognition tasks. Optionally, the recognition result output module may include a normalization layer (such as a softmax layer), and the normalization layer may normalize the input semantic feature data to obtain the text recognition result.

[0103] When the text to be recognized and its corresponding recognition task identifier are input into the target multi-task recognition model for text recognition processing, the text to be recognized and its corresponding recognition task identifier are first input into the target attention calculation layer for attention calculation to obtain attention data. Then the attention data are respectively input into the target feedforward neural network layer and each target feedforward neural network adjustment layer for local feature extraction processing to obtain feature data output by the target feedforward neural network layer and each target feedforward neural network adjustment layer. Then the feature data output by the target feedforward neural network layer and each target feedforward neural network adjustment layer are respectively input into the target feature fusion layer for feature fusion processing to obtain feature fusion data. Finally, according to the feature fusion data, the target text recognition result corresponding to the text to be recognized is determined. The above-mentioned target multi-task model can improve the accuracy of recognition of various text recognition tasks, thereby reducing resource consumption in the task recognition process and improving task processing efficiency.

[0104] The present application also provides a text recognition device. Figure 7 FIG. 1 is a block diagram of a text recognition device according to an exemplary embodiment. Figure 7 As shown, the device may include at least:

[0105] The to-be-recognized text acquisition module 201 is used to acquire the to-be-recognized text corresponding to the recognition task identifier;

[0106] The target text recognition result determination module 203 is used to input the to-be-recognized text and its corresponding recognition task identifier into the target multi-task recognition model for text recognition processing, and obtain the target text recognition result corresponding to the to-be-recognized text;

[0107] The device further includes a target multi-task recognition model training module 205, which includes:

[0108] The data set and the first multi-task recognition model acquisition submodule are used to acquire a sample text data set and a first multi-task recognition model; the sample recognition text in the sample text data set corresponds to a sample recognition task identifier and a sample recognition label; the feature extraction module in the first multi-task recognition model includes a feedforward neural network layer and a feedforward neural network adjustment layer corresponding to each text recognition task; the first multi-task recognition model is obtained by performing multi-task recognition training and single-task recognition adjustment training on the initial multi-task recognition model;

[0109] A second multi-task recognition model determination submodule is used to insert a feature fusion layer into the feature extraction module in the first multi-task recognition model to obtain a second multi-task recognition model; the feature fusion layer is used to fuse the output results of the feedforward neural network layer and each feedforward neural network adjustment layer;

[0110] The target multi-task recognition model determination submodule is used to train the second multi-task recognition model based on the sample recognition text and its corresponding sample recognition task identifier and sample recognition label to obtain the target multi-task recognition model.

[0111] In some exemplary embodiments, the apparatus further includes a first multi-task recognition model determination module, and the first multi-task recognition model determination module includes:

[0112] An initial multi-task recognition model acquisition submodule is used to acquire an initial multi-task recognition model;

[0113] A third multi-task recognition model determination submodule is used to perform multi-task recognition training on the initial multi-task recognition model based on the sample text data set to obtain a third multi-task recognition model;

[0114] The first multi-task recognition model determination submodule is used to perform single-task recognition adjustment training on the third multi-task recognition model based on the sample text data set to obtain a first multi-task recognition model.

[0115] In some exemplary embodiments, the third multi-task recognition model determination submodule includes:

[0116] A second text recognition result determination unit is used to input the sample recognition text and its corresponding sample recognition task identifier into the initial multi-task recognition model for text recognition processing to obtain a second text recognition result corresponding to the sample recognition text;

[0117] The third multi-task recognition model determination unit is used to adjust the parameters in the initial multi-task recognition model based on the second text recognition result and the sample recognition label until the difference between the second text recognition result and the sample recognition label meets a preset condition, thereby obtaining a third multi-task recognition model.

[0118] In some exemplary embodiments, the first multi-task recognition model determination submodule includes:

[0119] A third multi-task recognition model acquisition unit, used to acquire a preset number of third multi-task recognition models; the number of the third multi-task recognition models is the same as the number of text recognition tasks;

[0120] a fourth multi-task recognition model determination unit, configured to insert a feedforward neural network adjustment layer corresponding to the text recognition task into each feature extraction module of the third multi-task recognition model, to obtain a fourth multi-task recognition model corresponding to each text recognition task;

[0121] a fifth multi-task recognition model determination unit, configured to perform single-task recognition adjustment training on the fourth multi-task recognition model corresponding to each text recognition task based on the sample recognition text corresponding to each text recognition task, so as to obtain the fifth multi-task recognition model corresponding to each text recognition task;

[0122] The first multi-task recognition model determining unit is used to fuse each fifth multi-task recognition model to obtain a first multi-task recognition model.

[0123] In some exemplary embodiments, the fifth multi-task recognition model determination unit includes:

[0124] The third text recognition result determination subunit is used to input the current sample recognition text corresponding to the current text recognition task and the current sample recognition task identifier corresponding to the current sample recognition text into the current fourth multi-task recognition model for text recognition processing to obtain a third text recognition result corresponding to the current sample recognition text; the current text recognition task is any one of the at least two text recognition tasks; the current fourth multi-task recognition model is the fourth multi-task recognition model corresponding to the current text recognition task;

[0125] a parameter adjustment subunit, configured to adjust parameters in a feedforward neural network adjustment layer in a current fourth multi-task recognition model based on the third text recognition result and a current sample recognition label corresponding to the current sample recognition text, until a difference between the third text recognition result and the current sample recognition label satisfies a preset condition, thereby obtaining a fifth multi-task recognition model corresponding to the current text recognition task;

[0126] The fifth multi-task recognition model determination subunit is used to take any recognition task among at least two text recognition tasks except the current text recognition task as the current text recognition task, and repeat the steps of identifying the current sample recognition text corresponding to the current text recognition task and the current sample recognition task corresponding to the current sample recognition text until the fifth multi-task recognition model corresponding to the current text recognition task is obtained, thereby obtaining the fifth multi-task recognition model corresponding to each text recognition task.

[0127] In some exemplary embodiments, the feature extraction module in the fifth multi-task recognition model includes an attention calculation layer, a feedforward neural network layer and a feedforward neural network adjustment layer arranged on the attention calculation layer; the first multi-task recognition model determination unit includes:

[0128] The fused attention calculation layer determines a subunit, which is used to fuse the attention calculation layers in each fifth multi-task recognition model to obtain a fused attention calculation layer;

[0129] A fused feedforward neural network layer determining subunit, used to fuse the feedforward neural network layers in each fifth multi-task recognition model to obtain a fused feedforward neural network layer set on the fused attention calculation layer;

[0130] The first multi-task recognition model determination subunit is used to insert the feedforward neural network adjustment layer in each fifth multi-task recognition model into the fusion attention calculation layer to obtain the first multi-task recognition model.

[0131] In some exemplary embodiments, a target attention calculation layer, a target feedforward neural network layer, a target feedforward neural network adjustment layer corresponding to at least two text recognition tasks, and a target feature fusion layer are provided in a feature extraction module of a target multi-task recognition model. The target feedforward neural network layer and the target feedforward neural network adjustment layer corresponding to at least two text recognition tasks are provided on the target attention calculation layer, and the target feature fusion layer is provided on the target feedforward neural network layer and the target feedforward neural network adjustment layer corresponding to at least two text recognition tasks. The target text recognition result determination module includes:

[0132] The attention data determination submodule is used to input the text to be recognized and its corresponding recognition task identifier into the target attention calculation layer to perform attention calculation and obtain attention data;

[0133] A feature data determination submodule is used to input the attention data into the target feedforward neural network layer and each target feedforward neural network adjustment layer for local feature extraction processing, and obtain feature data output by the target feedforward neural network layer and each target feedforward neural network adjustment layer;

[0134] A feature fusion data determination submodule is used to input the feature data outputted by the target feedforward neural network layer and each target feedforward neural network adjustment layer into the target feature fusion layer for feature fusion processing to obtain feature fusion data;

[0135] The target text recognition result determination submodule is used to determine the target text recognition result corresponding to the text to be recognized based on the feature fusion data.

[0136] It should be noted that the text recognition device embodiment provided in the embodiments of the present application and the above-mentioned text recognition method embodiment are based on the same inventive concept.

[0137] An embodiment of the present application also provides an electronic device for text recognition, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement a text recognition method provided in any of the above embodiments.

[0138] An embodiment of the present application also provides a computer-readable storage medium, which can be set in a terminal to store at least one instruction or at least one program for implementing a text recognition method in an embodiment of the text recognition method, and the at least one instruction or at least one program is loaded and executed by a processor to implement the text recognition method provided in the above method embodiment.

[0139] Optionally, in the embodiment of this specification, the storage medium may be located in at least one of the multiple network servers of the computer network. Optionally, in this embodiment, the above storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store program codes.

[0140] The memory of the embodiment of this specification can be used to store software programs and modules, and the processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, applications required for functions, etc.; the data storage area may store data created according to the use of the device, etc. In addition, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory may also include a memory controller to provide the processor with access to the memory.

[0141] The embodiment of the present application also provides a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the text recognition method provided by the above method embodiment.

[0142] The method embodiments provided in the embodiments of the present application can be executed in a terminal, a computer terminal, a server or a similar computing device. Taking running on a server as an example, Figure 8 FIG. 1 is a hardware structure diagram of a server of a text recognition method provided according to an exemplary embodiment. Figure 8 As shown, the server 300 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 310 (the central processing unit 310 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 330 for storing data, and one or more storage media 320 (such as one or more mass storage devices) for storing application programs 323 or data 322. Among them, the memory 330 and the storage medium 320 can be short-term storage or permanent storage. The program stored in the storage medium 320 may include one or more modules, each of which may include a series of instruction operations on the server. Furthermore, the central processing unit 310 can be configured to communicate with the storage medium 320 and execute a series of instruction operations in the storage medium 320 on the server 300. The server 300 may also include one or more power supplies 360, one or more wired or wireless network interfaces 350, one or more input and output interfaces 340, and / or one or more operating systems 321, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0143] The input / output interface 340 may be used to receive or send data via a network. The specific example of the network may include a wireless network provided by a communication provider of the server 300. In one example, the input / output interface 340 includes a network adapter (Network Interface Controller, NIC), which may be connected to other network devices via a base station so as to communicate with the Internet. In one example, the input / output interface 340 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0144] It can be understood by those skilled in the art that Figure 8 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 8 More or fewer components as shown, or with Figure 8 Different configurations are shown.

[0145] It should be noted that the above-mentioned sequence of the embodiments of the present application is for description only and does not represent the advantages and disadvantages of the embodiments. The above-mentioned specific embodiments of this specification are described. Other embodiments are within the scope of the attached claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0146] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and server embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0147] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0148] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A text recognition method, It is characterized in that The method comprises: Obtain the text to be recognized corresponding to the recognition task identifier; Inputting the text to be recognized and its corresponding recognition task identifier into the target multi-task recognition model for text recognition processing to obtain a target text recognition result corresponding to the text to be recognized; The training method of the target multi-task recognition model includes: Acquire a sample text data set and a first multi-task recognition model; the sample recognition text in the sample text data set corresponds to a sample recognition task identifier and a sample recognition label; the feature extraction module in the first multi-task recognition model includes a feedforward neural network layer and a feedforward neural network adjustment layer corresponding to each text recognition task; the first multi-task recognition model is obtained by performing multi-task recognition training and single-task recognition adjustment training on an initial multi-task recognition model; Inserting a feature fusion layer into the feature extraction module in the first multi-task recognition model to obtain a second multi-task recognition model; the feature fusion layer is used to fuse the output results of the feedforward neural network layer and each of the feedforward neural network adjustment layers; The second multi-task recognition model is trained based on the sample recognition text and its corresponding sample recognition task identifier and the sample recognition label to obtain the target multi-task recognition model.

2. The method according to claim 1, It is characterized in that The method for determining the first multi-task recognition model includes: Obtain an initial multi-task recognition model; Performing multi-task recognition training on the initial multi-task recognition model based on the sample text data set to obtain a third multi-task recognition model; The third multi-task recognition model is trained for single-task recognition adjustment based on the sample text data set to obtain the first multi-task recognition model.

3. The method according to claim 2, It is characterized in that The performing multi-task recognition training on the initial multi-task recognition model based on the sample text data set to obtain a third multi-task recognition model includes: Inputting the sample recognition text and its corresponding sample recognition task identifier into the initial multi-task recognition model for text recognition processing, respectively, to obtain a second text recognition result corresponding to the sample recognition text; The parameters in the initial multi-task recognition model are adjusted based on the second text recognition result and the sample recognition label until the difference between the second text recognition result and the sample recognition label meets a preset condition, thereby obtaining the third multi-task recognition model.

4. The method according to claim 2 or 3, It is characterized in that The step of performing single-task recognition adjustment training on the third multi-task recognition model based on the sample text data set to obtain the first multi-task recognition model includes: Acquire a preset number of the third multi-task recognition models; the number of the third multi-task recognition models is the same as the number of text recognition tasks; Inserting a feedforward neural network adjustment layer corresponding to the text recognition task into each feature extraction module of the third multi-task recognition model to obtain a fourth multi-task recognition model corresponding to each text recognition task; Based on the sample recognition texts corresponding to each text recognition task, separately perform single-task recognition adjustment training on the fourth multi-task recognition model corresponding to each text recognition task to obtain the fifth multi-task recognition model corresponding to each text recognition task; Fuse each fifth multi-task recognition model to obtain the first multi-task recognition model.

5. The method according to claim 4, wherein, the separately performing single-task recognition adjustment training on the fourth multi-task recognition model corresponding to each text recognition task based on the sample recognition texts corresponding to each text recognition task to obtain the fifth multi-task recognition model corresponding to each text recognition task includes: Input the current sample recognition text corresponding to the current text recognition task and the current sample recognition task identifier corresponding to the current sample recognition text into the current fourth multi-task recognition model for text recognition processing to obtain the third text recognition result corresponding to the current sample recognition text; the current text recognition task is any one of at least two text recognition tasks; the current fourth multi-task recognition model is the fourth multi-task recognition model corresponding to the current text recognition task; Based on the third text recognition result and the current sample recognition label corresponding to the current sample recognition text, adjust the parameters in the feed-forward neural network adjustment layer in the current fourth multi-task recognition model until the difference between the third text recognition result and the current sample recognition label meets a preset condition to obtain the fifth multi-task recognition model corresponding to the current text recognition task; Take any recognition task other than the current text recognition task among at least two text recognition tasks as the current text recognition task, and repeat the steps from inputting the current sample recognition text corresponding to the current text recognition task and the current sample recognition task identifier corresponding to the current sample recognition text until obtaining the fifth multi-task recognition model corresponding to the current text recognition task to obtain the fifth multi-task recognition model corresponding to each text recognition task.

6. The method according to claim 4, wherein, the feature extraction module in the fifth multi-task recognition model includes an attention calculation layer, a feed-forward neural network layer and a feed-forward neural network adjustment layer arranged on the attention calculation layer; the fusing each fifth multi-task recognition model to obtain the first multi-task recognition model includes: Fuse the attention calculation layers in each fifth multi-task recognition model to obtain a fused attention calculation layer; Fuse the feed-forward neural network layers in each fifth multi-task recognition model to obtain a fused feed-forward neural network layer arranged on the fused attention calculation layer; Insert the feed-forward neural network adjustment layer in each fifth multi-task recognition model onto the fused attention calculation layer to obtain the first multi-task recognition model.

7. The method according to claim 1, wherein, The feature extraction module of the target multi-task recognition model is provided with a target attention calculation layer, a target feedforward neural network layer, a target feedforward neural network adjustment layer corresponding to at least two text recognition tasks, and a target feature fusion layer. The target feedforward neural network layer and the target feedforward neural network adjustment layer corresponding to at least two text recognition tasks are arranged on the target attention calculation layer, and the target feature fusion layer is arranged on the target feedforward neural network layer and the target feedforward neural network adjustment layer corresponding to at least two text recognition tasks; the text to be recognized and its corresponding recognition task identifier are input into the target multi-task recognition model for text recognition processing to obtain a target text recognition result corresponding to the text to be recognized, including: Input the to-be-recognized text and its corresponding recognition task identifier into the target attention calculation layer to perform attention calculation to obtain attention data; Inputting the attention data into the target feedforward neural network layer and each target feedforward neural network adjustment layer for local feature extraction processing, respectively, to obtain feature data output by the target feedforward neural network layer and each target feedforward neural network adjustment layer; Inputting the feature data outputted by the target feedforward neural network layer and each of the target feedforward neural network adjustment layers into the target feature fusion layer for feature fusion processing to obtain feature fusion data; According to the feature fusion data, a target text recognition result corresponding to the text to be recognized is determined.

8. A text recognition device, It is characterized in that The device comprises: The module for obtaining text to be recognized is used to obtain the text to be recognized corresponding to the recognition task identifier; A target text recognition result determination module is used to input the text to be recognized and its corresponding recognition task identifier into a target multi-task recognition model for text recognition processing to obtain a target text recognition result corresponding to the text to be recognized; Wherein, the device also includes a target multi-task recognition model training device; the target multi-task recognition model training device includes: The data set and the first multi-task recognition model acquisition submodule are used to acquire a sample text data set and a first multi-task recognition model; the sample recognition text in the sample text data set corresponds to a sample recognition task identifier and a sample recognition label; the feature extraction module in the first multi-task recognition model includes a feedforward neural network layer and a feedforward neural network adjustment layer corresponding to each text recognition task; the first multi-task recognition model is obtained by performing multi-task recognition training and single-task recognition adjustment training on the initial multi-task recognition model; A second multi-task recognition model determination submodule is used to insert a feature fusion layer into the feature extraction module in the first multi-task recognition model to obtain the second multi-task recognition model; the feature fusion layer is used to fuse the output results of the feedforward neural network layer and each of the feedforward neural network adjustment layers; The target multi-task recognition model determination submodule is used to train the second multi-task recognition model based on the sample recognition text and its corresponding sample recognition task identifier and the sample recognition label to obtain the target multi-task recognition model.

9. An electronic device for text recognition, It is characterized in that The device includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded by the processor and executes the text recognition method according to any one of claims 1 to 7.

10. A computer-readable storage medium, It is characterized in that The storage medium stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement the text recognition method according to any one of claims 1 to 7.

11. A computer program product comprising a computer program, It is characterized in that When the computer program is executed by a processor, the text recognition method according to any one of claims 1 to 7 is implemented.