Text recognition method and device, recognition model, electronic equipment and storage medium

CN117892728BActive Publication Date: 2026-09-04VIVO MOBILE COMM CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311565053.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-22
Publication Date
2026-09-04
Estimated Expiration
2043-11-22

AI Technical Summary

Technical Problem

[0004]本申请实施例的目的是提供一种文本识别方法及装置、识别模型、电子设备和存储介质,能够解决文本特征识别的计算开销大的问题

Benefits of technology

[0012] In this embodiment, text is identified by a fused independent model. The independent model includes an overall backbone network and lightweight modules connected to the backbone network. The feature vector of the original text is extracted by the backbone network. This vector is used as the input of the named entity recognition module and the classification module, respectively. Therefore, the named entity recognition module and the classification module no longer need to perform feature extraction processing independently. The independent model only needs to perform feature extraction calculation once, which effectively reduces the computational overhead of text feature recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117892728B_ABST
    Figure CN117892728B_ABST
Patent Text Reader

Abstract

The application discloses a text recognition method and device, a recognition model, an electronic device and a storage medium, and belongs to the technical field of feature extraction. The text recognition method comprises the following steps: inputting text data into a recognition model, extracting a feature vector of the text data through a backbone network of the recognition model; inputting the feature vector into a named entity recognition module of the recognition model, recognizing a keyword in the text data based on the feature vector through the named entity recognition module; inputting a type embedding vector corresponding to the keyword and the feature vector into a classification module of the recognition model, determining recognition information corresponding to the keyword according to the type embedding vector and the feature vector through the classification module; and outputting the keyword and the recognition information corresponding to the keyword.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of feature extraction technology, specifically relating to a text recognition method and apparatus, recognition model, electronic device and storage medium. Background Technology

[0002] In related technologies, identifying the emotions and intentions contained in a text requires extracting keywords from the text and then matching them with the corresponding emotions or intentions to form a multi-dimensional information set. For example, in the application scenario of public opinion analysis, it is necessary to extract scene words, attribute words, and opinion words from the text, determine whether they match and determine their emotions, and form a four-tuple to identify public opinion information such as positive public opinion and negative public opinion.

[0003] The extraction of the aforementioned multivariate information groups generally adopts a pipeline model, which includes multiple stages such as feature extraction, keyword recognition, relationship matching, and sentiment classification. Each stage uses an independent algorithm model, and the output of the previous model and the original text are used as the input of the next model, resulting in repetitive reasoning processes and high computational costs. Summary of the Invention

[0004] The purpose of this application is to provide a text recognition method, apparatus, recognition model, electronic device, and storage medium that can solve the problem of high computational overhead in text feature recognition.

[0005] In a first aspect, embodiments of this application provide a text recognition method, the method comprising: The text data is input into the recognition model, and the feature vector of the text data is extracted through the backbone network of the recognition model; The feature vector is input into the named entity recognition module of the recognition model, and the named entity recognition module identifies keywords in the text data based on the feature vector; The type embedding vector and feature vector corresponding to the keyword are input into the classification module of the recognition model. The classification module determines the recognition information corresponding to the keyword based on the type embedding vector and feature vector. Output keywords and their corresponding recognition information.

[0006] Secondly, embodiments of this application provide a text recognition device, including: Input module, used for: The text data is input into the recognition model, and the feature vector of the text data is extracted through the backbone network of the recognition model; The feature vector is input into the named entity recognition module of the recognition model, and the named entity recognition module identifies keywords in the text data based on the feature vector; The type embedding vector and feature vector corresponding to the keyword are input into the classification module of the recognition model. The classification module determines the recognition information corresponding to the keyword based on the type embedding vector and feature vector. The output module is used to output keywords and their corresponding recognition information.

[0007] Thirdly, embodiments of this application provide an identification model, including: Backbone network: The input data of the backbone network is text data. The backbone network is used to extract feature vectors from the text data. The output data of the backbone network is the feature vector. The Named Entity Recognition (NAME) module takes a feature vector as its input data and extracts keywords from the text data based on the feature vector. The output data of the Named Entity Recognition (NAME) module is the keywords. The classification module takes feature vectors and type embedding vectors corresponding to keywords as input. It determines the recognition information corresponding to the keywords based on the type embedding vectors and feature vectors. The output data of the recognition model consists of keywords and the corresponding recognition information.

[0008] Fourthly, embodiments of this application provide an electronic device, including a processor and a memory, wherein the memory stores a program that can run on the processor, and the program, when executed by the processor, implements the steps of the method as described in the first aspect.

[0009] Fifthly, embodiments of this application provide a readable storage medium on which a program is stored, which, when executed by a processor, implements the steps of the method as described in the first aspect.

[0010] In a sixth aspect, embodiments of this application provide a chip including a processor and a communication interface coupled to the processor, the processor being used to run a program to implement the steps of the method as described in the first aspect.

[0011] In a seventh aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method as described in the first aspect.

[0012] In this embodiment, text is identified by a fused independent model. The independent model includes an overall backbone network and lightweight modules connected to the backbone network. The feature vector of the original text is extracted by the backbone network. This vector is used as the input of the named entity recognition module and the classification module, respectively. Therefore, the named entity recognition module and the classification module no longer need to perform feature extraction processing independently. The independent model only needs to perform feature extraction calculation once, which effectively reduces the computational overhead of text feature recognition. Attached Figure Description

[0013] Figure 1 Flowcharts of text recognition methods according to some embodiments of this application are shown; Figure 2 The diagram shows a schematic representation of the structure of the recognition model according to some embodiments of this application; Figure 3 Structural block diagrams of text recognition devices according to some embodiments of this application are shown; Figure 4 A schematic diagram of type embedding vectors for some embodiments of this application is shown; Figure 5 A structural block diagram of an electronic device according to an embodiment of this application is shown; Figure 6 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application. Detailed Implementation

[0014] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0015] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0016] The text recognition method, apparatus, recognition model, electronic device, and storage medium provided in this application will be described in detail below with reference to the accompanying drawings and through specific embodiments and application scenarios.

[0017] A text recognition method is provided. Figure 1 Flowcharts of text recognition methods according to some embodiments of this application are shown, such as... Figure 1 As shown, the method includes: Step 102: Input the text data into the recognition model and extract the feature vector of the text data through the backbone network of the recognition model; Step 104: Input the feature vector into the named entity recognition module of the recognition model, and use the named entity recognition module to identify keywords in the text data based on the feature vector; Step 106: Input the type embedding vector and feature vector corresponding to the keyword into the classification module of the recognition model. The classification module determines the recognition information corresponding to the keyword based on the type embedding vector and feature vector. Step 108: Output the keywords and their corresponding recognition information.

[0018] In the embodiments of this application, Figure 2 The following are schematic diagrams illustrating the structure of the recognition model of some embodiments of this application, such as... Figure 2 As shown, the recognition model 200 includes a backbone network 202, a named entity recognition module 204, and a classification module 206. The backbone network 202 is used to extract feature vectors from the input text data, the named entity recognition module 204 is used to extract keywords from the text data, and the classification module 206 is used to determine the recognition information of the extracted keywords.

[0019] In related technologies, taking public opinion analysis as an example, in public opinion analysis, it is necessary to extract scene words, attribute words, and opinion words from the text, determine whether they match and determine their sentiment, and form a quadruple. In the traditional quadruple extraction scheme, the pipeline mode is generally adopted, namely: (1) using the named entity recognition model to extract all possible scene words, attribute words, and opinion words for the current text; (2) using the relation matching model to select scene words, attribute words, and opinion words that can match each other based on the previously extracted scene words, attribute words, and opinion words, and form relation triples; (3) using the sentiment classification module to classify the sentiment of the triples, add the sentiment polarity to the corresponding triples, and form the final quadruple.

[0020] For example, enter the text: "Since updating to the latest version of the system, the 4G network has become unstable, with the signal strength fluctuating."

[0021] (1) Using a named entity recognition model, extract all possible scene words, attribute words, and opinion words, including: Scene description: After updating the system to the latest version; Attribute words: 4G network, signal; Key words: Unstable, fluctuating strength; (2) Using the relation matching model, triple matching is performed to obtain two reasonable triples, including: a. (After updating to the latest system version; 4G network; unstable). b. (After updating to the latest system version; signal strength is inconsistent.) (3) Using the sentiment classification module, classify the sentiment of the triples and form the final quadruples, including: a. (After updating to the latest system version; 4G network; unstable; negative emotions). b. (After updating to the latest system version; signal; sometimes strong, sometimes weak; negative sentiment).

[0022] In the traditional pipeline extraction scheme described above, each stage uses an independent algorithm model. The output of the previous stage's model, along with the original text, serves as the input for the next stage's model. Therefore, during four-tuple extraction, the same text needs to be input into three different models for inference, resulting in a large overall computational load per sample. In real-world business scenarios, where massive amounts of text need to be processed daily, using a pipeline extraction scheme incurs significant computational resource overhead. Furthermore, because the pipeline scheme uses multiple models, it also requires substantial resources for service deployment.

[0023] On the other hand, in traditional pipeline schemes, prediction errors in the previous stage are propagated to the model in the next stage, causing error accumulation and affecting the final result. Taking the process mentioned above as an example, if the relationship matching model does not match the triple (after updating the latest version of the system; signal; sometimes strong and sometimes weak), then the triple cannot be used for sentiment classification, ultimately leading to a missing quadruple in the identification.

[0024] To address the aforementioned issues, this application's embodiments fuse multiple independent models into a single model. The backbone network of the identification model is used to extract common features required for the three tasks. A single text only requires one forward computation, significantly reducing the computational load in the pipeline scheme (which requires three forward computations). Secondly, it integrates lightweight modules.

[0025] Specifically, taking a public opinion analysis scenario as an example, assuming the keywords include scenario words, attribute words, and opinion words, and the identified information is sentiment information, the explanation is as follows: 1. Suppose the text data is: Since updating to the latest version of the system, the 4G network has become unstable, with the signal strength fluctuating.

[0026] 2. Input the above text data into the recognition model. After forward propagation through the backbone network, the corresponding feature vector h is obtained. b .

[0027] 3. Transfer the feature vector h b As input to the lightweight named entity recognition module, the module's forward propagation yields the extraction results of scene words, attribute words, and opinion words, as well as the location information of the corresponding entities in the text.

[0028] Specifically, it includes: Scene description: After updating the system to the latest version; Attribute words: 4G network, signal; Key words: Unstable, sometimes strong, sometimes weak.

[0029] 4. The extraction results are processed through all possible combinations to obtain the following set of candidate triples: (After updating to the latest system version; 4G network; unstable). (After updating to the latest system version; 4G network; strength fluctuates). (After updating to the latest system version; signal is unstable); (After updating to the latest system version; signal strength is inconsistent).

[0030] 5. For each candidate triplet, determine its corresponding type embedding based on the specific triplet content, and extract the feature vector h from the backbone network. b Based on this, add type embedding vectors corresponding to the scene, attribute, and viewpoint.

[0031] 6. Input the feature vector corresponding to each candidate triplet into the lightweight multi-classification module. After forward propagation of the multi-classification module, the hidden layer state vector is obtained. Extract the state vector corresponding to the [CLS] position, and pass it through the linear classification head of the neural network to finally obtain the classification result, as follows: (After updating to the latest system version; 4G network; unstable) -> Negative; (After updating to the latest system version; 4G network; sometimes strong, sometimes weak) -> Incompatible; (After updating to the latest system version; signal is unstable) -> mismatch; (After updating to the latest system version; signal strength fluctuates) -> negative.

[0032] 7. Based on the allocation results, output the final identified quadruple: (After updating to the latest system version; 4G network; unstable; negative emotions). (After updating to the latest system version; signal; sometimes strong, sometimes weak; negative sentiment).

[0033] The independent recognition model provided in this application includes an overall backbone network and lightweight modules connected to the backbone network. The feature vectors of the original text are extracted through the backbone network. These vectors are used as inputs to the named entity recognition module and the classification module, respectively. Therefore, the named entity recognition module and the classification module no longer need to perform feature extraction processing independently. The independent model only needs to perform feature extraction calculation once, which effectively reduces the computational overhead of text feature recognition.

[0034] In some embodiments of this application, after identifying keywords in text data based on feature vectors using a named entity recognition module, the method further includes: The named entity recognition module identifies the location information of keywords in text data; The type embedding vector and feature vector corresponding to the keyword are input into the classification module of the recognition model, specifically including: Based on the location information, a sub-vector corresponding to the keyword is determined from the feature vector; The type embedding vector and the sub-vector are fused to obtain the fused feature vector; the vector dimension of the type embedding vector is the same as that of the sub-vector, and the type embedding vector is used to indicate the category information of the keyword. The fused feature vector is input into the classification module; the classification module identifies the sub-vectors based on the category information to obtain the identification information.

[0035] In this embodiment, the named entity recognition module can extract the category and location of all entities in the text data based on the feature vector output by the backbone network. For example, assuming the backbone network word embedding is a vector with a dimension of 1024, the entity categories are divided into three types: scene words, attribute words, and opinion words.

[0036] Based on the type of keyword, three type embedding vectors are randomly initialized. The vector dimension of the type embedding vector is the same as the vector dimension of the feature vector output by the backbone network. Continuing with the example above, assuming that the word embedding of the backbone network is a vector with a dimension of 1024, then the vector dimension of the type embedding vector is also a vector with a dimension of 1024.

[0037] Suppose the original input text is "Playing games drains the battery quickly". The named entity recognition module extracts the category and position of all entities in the text. Specifically, the entity category of "playing games" is scene word, with position 0 as the starting position, so the positions of scene words are 1-3; the entity category of "drains the battery" is attribute word, corresponding to positions 7-8; and the entity category of "quickly" is opinion word, corresponding to positions 9-10.

[0038] After obtaining this location information, based on this location information, the sub-vector corresponding to each keyword is found in the feature vector of the entire text. Then, based on the keyword type, the corresponding type embedding vector is fused with the keyword sub-vector. Specifically, following the example above, the two 1024-dimensional feature vectors at positions 7-8 are added to the type embedding vector 'a' corresponding to the attribute word, resulting in two new 1024-dimensional vectors. Similarly, the two 1024-dimensional feature vectors at positions 9-10 are added to the type embedding vector 'o' corresponding to the opinion word, resulting in two new 1024-dimensional vectors.

[0039] These fused feature vectors are input into the classification module. After forward propagation through the multi-classification module and passing through the linear classification head of the neural network, the predicted probabilities corresponding to the classification labels are obtained, and the label with the highest probability is taken as the final prediction result. For example, after label expansion, there are four types of labels: positive, neutral, negative, and mismatch. Assuming the predicted probabilities corresponding to the classification labels are [0.1, 0.05, 0.8, 0.05], the label with the highest probability (0.8) is the third label, i.e., negative. At this point, the final result of the multi-classification module, i.e., the recognition information, is obtained.

[0040] This application embodiment achieves flexible differentiation at the input level by adding type embedding vectors corresponding to different keywords based on their specific positions in the text, on the basis of the feature vectors output by the backbone network.

[0041] In some embodiments of this application, the keywords include N keyword types, and the N keyword types correspond one-to-one with N type embedding vectors, where N is a positive integer; The process of merging type-embedded vectors with sub-vectors includes: The first sub-vector corresponding to the first keyword is determined based on the location information. The keyword type of the first keyword is the first keyword type. Keywords include the first keyword, sub-vectors include the first sub-vector, and N keyword types include the first keyword type. The first type embedding vector corresponding to the first keyword type is fused with the first sub-vector, and the first type embedding vector corresponds to the first keyword type.

[0042] In this application embodiment, the keywords in the text data include multiple keyword types. Taking the public opinion analysis scenario as an example, the possible keyword types in public opinion analysis include scenario word type, attribute word type, and opinion word type.

[0043] For each keyword type, a corresponding type embedding vector is defined. For example, for the three keyword types mentioned above, there are three type embedding vectors s, a, and o, where scenario words correspond to type embedding vector s, attribute words correspond to type embedding vector a, and opinion words correspond to type embedding vector o.

[0044] When fusing sub-vectors and type embedding vectors, each sub-vector is added to its corresponding type embedding vector. Specifically, the feature vector of a scene word is added to its corresponding type embedding vector s, the feature vector of an attribute word is added to its corresponding type embedding vector a, and the feature vector of an opinion word is added to its corresponding type embedding vector o.

[0045] In this model, the vector dimension of each type of embedding vector is the same as the vector dimension of its corresponding sub-vector. Taking a feature vector with 1024 dimensions identified by the backbone network as an example, since the vector dimension of each feature vector output by the backbone network is 1024, the vector dimension of the corresponding type embedding vector is also set to 1024 dimensions. Therefore, after adding the sub-vector and the type embedding vector, the dimension of the new vector remains unchanged. The final recognition result can be obtained through a single forward propagation, reducing the computational load of the recognition model.

[0046] This application embodiment adds type embedding vectors corresponding to different keywords based on the feature vectors output by the backbone network and their specific positions in the text. Therefore, the input features of different types of keywords are different, which realizes flexible differentiation at the input level and improves the classification accuracy.

[0047] In some embodiments of this application, both the named entity recognition module and the classification module are neural network models based on the self-attention mechanism; the classification module includes a neural network classifier.

[0048] In this embodiment, both the named entity recognition module and the classification module are designed as lightweight computing modules. The named entity recognition module can be configured as a single-layer transformer based on a self-attention mechanism. Similarly, the classification module can also be configured as a single-layer transformer based on a self-attention mechanism.

[0049] By setting the named entity recognition module and classification module as lightweight computing modules, i.e., neural network models based on self-attention mechanisms, the amount of additional computation can be reduced, thereby reducing the overall computing power of the recognition model and improving model efficiency.

[0050] The classification module includes a neural network classifier, which can predict whether keywords, such as scene words, attribute words, and opinion words, match and their corresponding sentiment classification.

[0051] Specifically, the following example illustrates the workflow of the classification module: 1. Initialize the corresponding type embedding vectors based on the dimensions of the backbone network word embeddings and the entity categories. Assuming the backbone network word embeddings are vectors with a dimension of 1024, and the entity categories are divided into three types: scene words, attribute words, and opinion words, then three type embedding vectors with a dimension of 1024 can be randomly initialized and used as the type embeddings for scene words, attribute words, and opinion words, respectively. For example, the three type embedding vectors are denoted as s, a, and o, respectively.

[0052] 2. With Figure 2 For example, assuming the original input text is "Playing games drains the battery quickly", two special characters [CLS] and [SEP] are added before and after the original text, resulting in a total of 12 tokens. Each token corresponds to a word embedding vector with a dimension of 1024 in the vocabulary. These 12 1024-dimensional vectors are used as the input to the backbone network.

[0053] 3. After forward propagation through the backbone network, the hidden state of the last layer of the network is obtained, which is also 12 vectors of 1024 dimensions, i.e., feature vectors.

[0054] 4. The hidden layer states (feature vectors) extracted from the backbone network are used as input to the named entity recognition module to extract the category and location of all entities in the text. Figure 2 In the example shown, the entity category of "playing games" is a scene word, starting from position 0, so the positions of scene words are 1-3; the entity category of "consuming power" is an attribute word, corresponding to positions 7-8; and the entity category of "very fast" is an opinion word, corresponding to positions 9-10.

[0055] 5. Add the corresponding type embeddings to the 12 1024-dimensional vectors obtained in step 3, based on the information obtained in step 4. Specifically: Add the three 1024-dimensional feature vectors at positions 1-3 to the type embedding vector *s* corresponding to the scene word, resulting in three new 1024-dimensional vectors. Add the two 1024-dimensional feature vectors at positions 7-8 to the type embedding vector *a* corresponding to the attribute word, resulting in two new 1024-dimensional vectors. Add the two 1024-dimensional feature vectors at positions 9-10 to the type embedding vector *o* corresponding to the opinion word, resulting in two new 1024-dimensional vectors. The feature vectors at other positions do not need to have type embeddings added and remain unchanged. After the above operations, the final result is still 12 1024-dimensional vectors, where the dimensions of those vectors with added type embeddings remain unchanged, but their specific values ​​have changed.

[0056] 6. Input the 12 1024-dimensional vectors obtained in step 5 into the lightweight multi-classification module. After forward propagation of the multi-classification module, the hidden layer state vectors are obtained, which are still 12 1024-dimensional vectors. Extract the first digit (i.e., position 0), i.e., the 1024-dimensional state vector corresponding to the [CLS] position. Through the linear classification head of the neural network, obtain the predicted probability corresponding to the classification label, and take the label with the highest probability as the final prediction result. After label expansion, there are four types of labels: positive, neutral, negative, and mismatch. Assuming that the predicted probabilities corresponding to the classification labels are [0.1, 0.05, 0.8, 0.05], the label with the highest probability (0.8) is the third label, i.e., negative. At this point, the final result of the multi-classification module is obtained.

[0057] The embodiments of this application predict whether keywords match and the corresponding recognition information through a neural network classifier. The computational load is small, which can effectively reduce the overall computational load of the recognition model.

[0058] The text recognition method provided in this application can be executed by a text recognition device. This application uses a text recognition device to perform the text recognition method as an example to illustrate the text recognition device provided in this application.

[0059] In some embodiments of this application, a text recognition device is provided. Figure 3 Structural block diagrams of text recognition devices according to some embodiments of this application are shown, such as... Figure 3 As shown, the text recognition device 300 includes: The input module 302 is used to input text data into the recognition model, extract feature vectors from the text data through the backbone network of the recognition model; input the feature vectors into the named entity recognition module of the recognition model, and identify keywords in the text data based on the feature vectors; input the type embedding vectors and feature vectors corresponding to the keywords into the classification module of the recognition model, and determine the recognition information corresponding to the keywords based on the type embedding vectors and feature vectors. Output module 304 is used to output keywords and their corresponding recognition information.

[0060] The independent recognition model provided in this application includes an overall backbone network and lightweight modules connected to the backbone network. The feature vectors of the original text are extracted through the backbone network. These vectors are used as inputs to the named entity recognition module and the classification module, respectively. Therefore, the named entity recognition module and the classification module no longer need to perform feature extraction processing independently. The independent model only needs to perform feature extraction calculation once, which effectively reduces the computational overhead of text feature recognition.

[0061] In some embodiments of this application, the text recognition device further includes: The recognition module is used to identify the location information of keywords in text data through the named entity recognition module; The determination module is used to determine the sub-vector corresponding to the keyword in the feature vector based on the location information; The fusion module merges the type embedding vector with the sub-vector. The vector dimension of the type embedding vector is the same as that of the sub-vector. The type embedding vector is used to indicate the category information of the keyword, so that the classification module can identify the sub-vector based on the category information to obtain the identification information.

[0062] This application embodiment achieves flexible differentiation at the input level by adding type embedding vectors corresponding to different keywords based on their specific positions in the text, on the basis of the feature vectors output by the backbone network.

[0063] In some embodiments of this application, the keywords include N keyword types, and each of the N keyword types corresponds one-to-one with N type embedding vectors, where N is a positive integer; the text recognition device further includes: The determination module is also used to determine the first sub-vector corresponding to the first keyword based on the location information. The keyword type of the first keyword is the first keyword type, the keyword includes the first keyword, the sub-vector includes the first sub-vector, and the N keyword types include the first keyword type. The fusion module is specifically used to fuse the first type embedding vector corresponding to the first keyword type with the first sub-vector. The first type embedding vector corresponds to the first keyword type.

[0064] This application embodiment adds type embedding vectors corresponding to different keywords based on the feature vectors output by the backbone network and their specific positions in the text. Therefore, the input features of different types of keywords are different, which realizes flexible differentiation at the input level and improves the classification accuracy.

[0065] In some embodiments of this application, both the named entity recognition module and the classification module are neural network models based on a self-attention mechanism; the classification module includes a neural network classifier.

[0066] The embodiments of this application predict whether keywords match and the corresponding recognition information through a neural network classifier. The computational load is small, which can effectively reduce the overall computational load of the recognition model.

[0067] The text recognition method provided in this application can be executed by a text recognition device. This application uses a text recognition device executing the text recognition method as an example to illustrate the text recognition device provided in this application.

[0068] In some embodiments of this application, a recognition model is provided, such as Figure 2 As shown, the recognition model 200 includes: Backbone network 202, wherein the input data of the backbone network is text data, the backbone network is used to extract feature vectors from the text data, and the output data of the backbone network is the feature vectors; Named entity recognition module 204: The input data of the named entity recognition module is a feature vector. The named entity recognition module is used to extract keywords from the text data based on the feature vector. The output data of the named entity recognition module is the keywords. Classification module 206 takes feature vectors and type embedding vectors corresponding to keywords as input data. The classification module determines the recognition information corresponding to the keywords based on the type embedding vectors and feature vectors. The output data of the recognition model consists of keywords and the corresponding recognition information.

[0069] In the embodiments of this application, such as Figure 2 As shown, the recognition model 200 provided in this embodiment includes a backbone network 202, a named entity recognition module 204, and a classification module 206. The fused recognition model 200 extracts text feature vectors through the backbone network 202, and the feature vectors extracted by the backbone network 202 are used as inputs to the named entity recognition module 204 and the classification module 206, respectively. The named entity recognition module 204 is used to extract all possible keywords in the text. Taking a public opinion analysis scenario as an example, keywords include scene words, attribute words, and opinion words.

[0070] While the named entity recognition module 204 extracts keywords, the classification module 206 is used to determine the recognition information corresponding to the keywords. Taking public opinion analysis as an example, the recognition information of the keywords is the sentiment information obtained by matching scene words, attribute words and opinion words.

[0071] The obtained scene words, attribute words, opinion words, and sentiment information are used to form the final required public opinion quadruple.

[0072] Specifically, the pipeline scheme in the relevant technology consists of three models: a named entity recognition model, a relation matching model, and a sentiment classification module. The named entity recognition model comprises a backbone network A and a NER (Named Entity Recognition) classifier head. The relation matching model comprises a backbone network B and a relation classification head. The sentiment classification module comprises a backbone network C and a sentiment classification head.

[0073] The backbone networks A, B, and C are independent models with different parameters. Even if the parameters of the three models are set to the same value at the beginning of training, they will inevitably differ after training for their respective tasks. For the same text input, computation is required through each of the three backbone networks A, B, and C, resulting in high computational costs.

[0074] In this implementation, a single backbone network D performs the functions of the aforementioned backbone networks A, B, and C simultaneously. This is equivalent to backbone networks A, B, and C sharing the same model structure and parameters, effectively merging the three backbone networks into a single new backbone network. The role of backbone network D is to extract the main features of the paper, which serve as input for subsequent modules. A deeper neural network, such as BERT (Eidirectional Encoder Representations from Transformers, a language representation model), can be used. Therefore, this embodiment only requires computation through one backbone network D, reducing the computational load to one-third of the original.

[0075] Following the backbone network, a lightweight named entity recognition module and a lightweight classification module are connected in parallel. The named entity recognition module is used to extract all scene words, attribute words, and opinion words from the input text. The input to the named entity recognition module is the feature vector extracted by the backbone network.

[0076] The classification module is used to predict whether scene words, attribute words, and opinion words match and their corresponding sentiment classifications. The input to the classification module includes feature vectors extracted by the backbone network, as well as three types of embedding vectors corresponding to scene words, attribute words, and opinion words.

[0077] Specifically, Figure 4A schematic diagram of type embedding vectors of some embodiments of this application is shown, such as Figure 4 As shown, based on the feature vectors output by the backbone network, type embedding vectors corresponding to different keywords are added according to their specific positions in the text. Therefore, the final input features corresponding to each different triple are also different, realizing flexible differentiation at the input level.

[0078] In the pipeline approach of related technologies, relationship matching is a binary classification problem with labels of "match" and "mismatch". Sentiment classification, on the other hand, is a multi-class classification problem with labels of "positive", "neutral", and "negative". These two tasks require separate models for processing.

[0079] In this embodiment, the "match" in the relation matching task label is expanded to specific sentiment labels, including: positive, neutral, and negative. All labels are then categorized into four types: positive, neutral, negative, and no match. If the prediction result is one of positive, neutral, or negative, the triple is considered a match by default. This allows the lightweight multi-classification module to simultaneously perform both triple matching and sentiment classification functions.

[0080] During training, the loss function of the above recognition model can be set as a weighted sum of the loss functions of the two lightweight modules, and the structure of the recognition model supports parallel training.

[0081] The independent recognition model provided in this application includes an overall backbone network and lightweight modules connected to the backbone network. The feature vectors of the original text are extracted through the backbone network. These vectors are used as inputs to the named entity recognition module and the classification module, respectively. Therefore, the named entity recognition module and the classification module no longer need to perform feature extraction processing independently. The independent model only needs to perform feature extraction calculation once, which effectively reduces the computational overhead of text feature recognition.

[0082] In some embodiments of this application, the vector dimension of the type embedding vector is the same as the vector dimension of the feature vector.

[0083] In this embodiment, the vector dimension of the type embedding vector is the same as that of the feature vector. Taking the feature vector identified by the backbone network as a vector with 1024 dimensions as an example, if the vector dimension of the type embedding vector is also set to 1024 dimensions, the dimension of the new vector obtained after adding the feature vector and the type embedding vector remains unchanged. Therefore, the final recognition result can be obtained through one forward propagation, reducing the computational load of the recognition model.

[0084] The text recognition device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.

[0085] The text recognition device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.

[0086] The text recognition device provided in this application embodiment can implement all the processes implemented in the above method embodiments, and will not be described again here to avoid repetition.

[0087] Optionally, embodiments of this application also provide an electronic device. Figure 5 A structural block diagram of an electronic device according to an embodiment of this application is shown, such as... Figure 5 As shown, the electronic device 500 includes a processor 502, a memory 504, and a program stored in the memory 504 and executable on the processor 502. When the program is executed by the processor 502, it implements the various processes of the above method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0088] It should be noted that the electronic devices in the embodiments of this application include the aforementioned mobile electronic devices and non-mobile electronic devices.

[0089] Figure 6 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0090] The electronic device 600 includes, but is not limited to, components such as: radio frequency unit 601, network module 602, audio output unit 603, input unit 604, sensor 605, display unit 606, user input unit 607, interface unit 608, memory 609, and processor 610.

[0091] Those skilled in the art will understand that the electronic device 600 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 610 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 6 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0092] The processor 610 inputs text data into the recognition model and extracts feature vectors from the text data through the backbone network of the recognition model; inputs the feature vectors into the named entity recognition module of the recognition model and identifies keywords in the text data based on the feature vectors; inputs the type embedding vectors and feature vectors corresponding to the keywords into the classification module of the recognition model and determines the recognition information corresponding to the keywords based on the type embedding vectors and feature vectors; and outputs the keywords and the recognition information corresponding to the keywords.

[0093] The independent recognition model provided in this application includes an overall backbone network and lightweight modules connected to the backbone network. The feature vectors of the original text are extracted through the backbone network. These vectors are used as inputs to the named entity recognition module and the classification module, respectively. Therefore, the named entity recognition module and the classification module no longer need to perform feature extraction processing independently. The independent model only needs to perform feature extraction calculation once, which effectively reduces the computational overhead of text feature recognition.

[0094] Optionally, the processor 610 is further configured to identify the location information of keywords in text data through a named entity recognition module; input the type embedding vector and feature vector corresponding to the keyword into the classification module of the recognition model, specifically including: determining a sub-vector corresponding to the keyword in the feature vector according to the location information; fusing the type embedding vector and the sub-vector to obtain a fused feature vector; wherein the vector dimension of the type embedding vector is the same as the vector dimension of the sub-vector, and the type embedding vector is used to indicate the category information of the keyword; inputting the fused feature vector into the classification module; wherein the classification module identifies the sub-vector based on the category information to obtain recognition information.

[0095] This application embodiment achieves flexible differentiation at the input level by adding type embedding vectors corresponding to different keywords based on their specific positions in the text, on the basis of the feature vectors output by the backbone network.

[0096] Optionally, the keywords include N keyword types, and each of the N keyword types corresponds one-to-one with N type embedding vectors, where N is a positive integer; The processor 610 is also used to determine the first sub-vector corresponding to the first keyword based on the location information. The keyword type of the first keyword is the first keyword type. The keyword includes the first keyword, the sub-vector includes the first sub-vector, and N keyword types include the first keyword type. The first type embedding vector corresponding to the first keyword type is fused with the first sub-vector. The first type embedding vector corresponds to the first keyword type.

[0097] Optionally, both the named entity recognition module and the classification module are neural network models based on the self-attention mechanism; the classification module includes a neural network classifier.

[0098] The embodiments of this application predict whether keywords match and the corresponding recognition information through a neural network classifier. The computational load is small, which can effectively reduce the overall computational load of the recognition model.

[0099] This application embodiment adds type embedding vectors corresponding to different keywords based on the feature vectors output by the backbone network and their specific positions in the text. Therefore, the input features of different types of keywords are different, which realizes flexible differentiation at the input level and improves the classification accuracy.

[0100] It should be understood that, in this embodiment, the input unit 604 may include a graphics processing unit (GPU) 6041 and a microphone 6042. The GPU 6041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 606 may include a display panel 6061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 607 includes at least one of a touch panel 6071 and other input devices 6072. The touch panel 6071 is also called a touch screen. The touch panel 6071 may include a touch detection device and a touch controller. Other input devices 6072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0101] The memory 609 can be used to store software programs and various data. The memory 609 may primarily include a first storage area for storing programs and a second storage area for storing data. The first storage area may store the operating system, application programs required for at least one function (such as sound playback, image playback, etc.), etc. Furthermore, the memory 609 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (Synchlink DRAM, SLDRAM), and direct memory bus RAM (DRRAM). The memory 609 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0102] Processor 610 may include one or more processing units; optionally, processor 610 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 610.

[0103] This application also provides a readable storage medium storing a program. When the program is executed by a processor, it implements the various processes of the above method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0104] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0105] This application also provides a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled. The processor is used to run a program to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0106] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0107] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above method embodiments and achieve the same technical effects. To avoid repetition, it will not be described again here.

[0108] The methods can be implemented in various ways depending on specific features and / or example applications. For example, these methods can be implemented by a combination of hardware, firmware, and / or software. For example, in a hardware implementation, the processor can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, electronic devices, other device units for performing the functions described above, and / or combinations thereof.

[0109] A computer-readable storage medium can be a tangible device that holds and stores instructions for use by an instruction execution device. A computer-readable storage medium can be an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital universal disk (DVD), memory cards, floppy disks, encoding mechanical devices (e.g., punched cards or grooves with raised structures for recording instructions), and any suitable combination of the foregoing. The computer-readable storage medium used herein should not be construed as the transmission of signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media, or electrical signals transmitted through wires.

[0110] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0111] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.

[0112] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A text recognition method, characterized in that, The method includes: The text data is input into the recognition model, and the feature vector of the text data is extracted through the backbone network of the recognition model; The feature vector is input into the named entity recognition module of the recognition model, and the named entity recognition module identifies keywords in the text data based on the feature vector; The type embedding vector and the feature vector corresponding to the keyword are input into the classification module of the recognition model, and the classification module determines the recognition information corresponding to the keyword based on the type embedding vector and the feature vector. Output the keyword and the corresponding identification information; After the named entity recognition module identifies keywords in the text data based on the feature vector, the method further includes: The named entity recognition module identifies the location information of the keyword in the text data. The step of inputting the type embedding vector corresponding to the keyword and the feature vector into the classification module of the recognition model specifically includes: Based on the location information, a sub-vector corresponding to the keyword is determined from the feature vector; The type embedding vector and the sub-vector are fused to obtain the fused feature vector; wherein the vector dimension of the type embedding vector is the same as the vector dimension of the sub-vector, and the type embedding vector is used to indicate the category information of the keyword; The fused feature vector is input to the classification module; wherein, the classification module identifies the sub-vector based on the category information to obtain the identification information; The keywords include N keyword types, and each of the N keyword types corresponds one-to-one with N type embedding vectors, where N is a positive integer. The process of fusing the type embedding vector with the sub-vector includes: The first sub-vector corresponding to the first keyword is determined based on the location information. The keyword type of the first keyword is the first keyword type. The keyword includes the first keyword. The sub-vector includes the first sub-vector. The N keyword types include the first keyword type. The first type embedding vector corresponding to the first keyword type is fused with the first sub-vector, and the first type embedding vector corresponds to the first keyword type.

2. The text recognition method according to claim 1, characterized in that, Both the named entity recognition module and the classification module are neural network models based on the self-attention mechanism; The classification module includes a neural network classifier.

3. A text recognition device, characterized in that, include: Input module, used for: The text data is input into the recognition model, and the feature vector of the text data is extracted through the backbone network of the recognition model; The feature vector is input into the named entity recognition module of the recognition model, and the named entity recognition module identifies keywords in the text data based on the feature vector; The type embedding vector and the feature vector corresponding to the keyword are input into the classification module of the recognition model, and the classification module determines the recognition information corresponding to the keyword based on the type embedding vector and the feature vector. The output module is used to output the keyword and the identification information corresponding to the keyword; The recognition module is used to identify the location information of the keyword in the text data through the named entity recognition module; The determining module is used to determine a sub-vector corresponding to the keyword in the feature vector based on the location information; The fusion module fuses the type embedding vector with the sub-vector, wherein the vector dimension of the type embedding vector is the same as the vector dimension of the sub-vector, and the type embedding vector is used to indicate the category information of the keyword, so that the classification module can identify the sub-vector based on the category information to obtain the identification information; The keywords include N keyword types, and each of the N keyword types corresponds one-to-one with N type embedding vectors, where N is a positive integer. The determining module is further configured to determine a first sub-vector corresponding to the first keyword based on the location information, wherein the keyword type of the first keyword is a first keyword type, the keyword includes the first keyword, the sub-vector includes the first sub-vector, and the N keyword types include the first keyword type; The fusion module is specifically used to fuse the first type embedding vector corresponding to the first keyword type with the first sub-vector, wherein the first type embedding vector corresponds to the first keyword type.

4. The text recognition device according to claim 3, characterized in that, Both the named entity recognition module and the classification module are neural network models based on the self-attention mechanism; The classification module includes a neural network classifier.

5. A recognition model, characterized in that, include: A backbone network, wherein the input data of the backbone network is text data, the backbone network is used to extract feature vectors from the text data, and the output data of the backbone network is the feature vectors; The named entity recognition module is configured to take the feature vector as its input data, extract keywords from the text data based on the feature vector, output the keywords as its output data, and also identify the position information of the keywords in the text data. A classification module, whose input data includes the feature vector and the type embedding vector corresponding to the keyword, is used to determine the recognition information corresponding to the keyword based on the type embedding vector and the feature vector; and The output data of the recognition model is the keyword and the recognition information corresponding to the keyword; The type embedding vector corresponding to the keyword and the feature vector are input into the classification module of the recognition model, specifically including: Based on the location information, a sub-vector corresponding to the keyword is determined from the feature vector; The type embedding vector and the sub-vector are fused to obtain the fused feature vector; wherein the vector dimension of the type embedding vector is the same as the vector dimension of the sub-vector, and the type embedding vector is used to indicate the category information of the keyword; The fused feature vector is input to the classification module; wherein, the classification module identifies the sub-vector based on the category information to obtain the identification information; The keywords include N keyword types, and each of the N keyword types corresponds one-to-one with N type embedding vectors, where N is a positive integer. The process of fusing the type embedding vector with the sub-vector includes: The first sub-vector corresponding to the first keyword is determined based on the location information. The keyword type of the first keyword is the first keyword type. The keyword includes the first keyword. The sub-vector includes the first sub-vector. The N keyword types include the first keyword type. The first type embedding vector corresponding to the first keyword type is fused with the first sub-vector, and the first type embedding vector corresponds to the first keyword type.

6. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program that can run on the processor, the program being executed by the processor to implement the steps of the method as described in claim 1 or 2.

7. A readable storage medium, characterized in that, A program is stored on the readable storage medium, which, when executed by a processor, implements the steps of the method as described in claim 1 or 2.

Citation Information

Patent Citations

  • Text semantic recognition method and device, electronic equipment and storage medium

    CN114707513A