A text recognition method based on location information and related devices
By using keyword dictionary to determine position weights in short text recognition and combining text feature vectors, the problem of difficult learning context information in short text classification is solved, and the accuracy of text recognition is significantly improved.
Patent Information
- Application Number
- CN202110090663.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-22
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-01-22
AI Technical Summary
The prior art is difficult to effectively learn context information in short text classification, which affects the accuracy of text recognition.
By obtaining the target text, a feature extraction layer in the text recognition model is input to obtain the text feature vector, and the position weight of the text unit is determined based on the keyword dictionary, combining the position weight and the text feature vector input and output layer to obtain the recognition results.
It effectively improves the accuracy of text recognition, and by supplementing the insufficient context information in short text, it makes up for the insufficient learning of context information by the model.
Smart Images

Figure CN113590832B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and in particular, to a text recognition method based on location information and related devices. Background Art
[0002] With the rapid development of Internet technologies, people have higher and higher requirements for pushed information. How to perform fast and accurate information pushing has become a difficult problem. In the process of information pushing, short text classification is an important link in the pushing process, such as the recognition of article titles.
[0003] Generally, the methods for short text classification mainly include term frequency–inverse document frequency (TF-IDF), Latent Dirichlet Allocation (LDA), or using word2vec to sum and average, and then using a deep learning model to classify short texts.
[0004] However, in the above process of short text classification, since the number of words in short text data is small and the contained text information is limited, it is difficult for the above models to directly learn context information from short title data, which affects the accuracy of text recognition. Summary of the Invention
[0005] In view of this, the present application provides a text recognition method based on location information, which can effectively improve the accuracy of text recognition.
[0006] The first aspect of the present application provides a text recognition method based on location information, which can be applied to a system or program with a text recognition function based on location information in a terminal device, and specifically includes:
[0007] Obtain a target text, where the target text includes a plurality of text units;
[0008] Input the target text into a feature extraction layer in a text recognition model to obtain text feature vectors corresponding to each of the text units;
[0009] Determine the position weights corresponding to each of the text units in the target text based on a keyword dictionary, where the keyword dictionary is used to indicate domain words under different text categories, and the position weights are determined based on the position relationship between the text unit and the domain word corresponding to the target text in the target text;
[0010] Input the position weights and the text feature vectors into an output layer in the text recognition model to obtain a recognition result.
[0011] Optionally, in some possible implementation manners of the present application, determining the position weights corresponding to each of the text units in the target text based on the keyword dictionary includes:
[0012] Obtaining corpus data under different preset categories and determining a plurality of candidate words;
[0013] Determining the domain feature values of the candidate words in the preset categories;
[0014] Filtering the candidate words based on the domain feature values to obtain domain dictionaries corresponding to different preset categories;
[0015] Fusing the domain dictionaries corresponding to different preset categories to obtain the keyword dictionary;
[0016] Determining the position weights corresponding to each of the text units in the target text based on the keyword dictionary.
[0017] Optionally, in some possible implementation manners of the present application, obtaining corpus data under different preset categories and determining a plurality of candidate words includes:
[0018] Invoking historical records corresponding to different preset categories;
[0019] Performing text statistics based on the historical records to obtain feature items;
[0020] Editing corresponding preset whitelists according to the feature items;
[0021] Traversing the corpus data based on the preset whitelists to determine a plurality of candidate words.
[0022] Optionally, in some possible implementation manners of the present application, obtaining corpus data under different preset categories and determining a plurality of candidate words includes:
[0023] Obtaining corpus data under different preset categories;
[0024] Determining detection entries in the corpus data;
[0025] Calculating mutual point information corresponding to the detection entries based on the corpus data, where the mutual point information is used to indicate the relevance between the detection entries and the corpus data;
[0026] Determining a plurality of candidate words according to the mutual point information.
[0027] Optionally, in some possible implementation manners of the present application, filtering the candidate words based on the domain feature values to obtain domain dictionaries corresponding to different preset categories includes:
[0028] Screen the candidate words based on the domain feature values to obtain a candidate set;
[0029] Filter the candidate set according to a preset rule to obtain a domain dictionary corresponding to different preset categories, where the preset rule is set based on word parameters corresponding to the preset categories.
[0030] Optionally, in some possible implementation manners of the present application, it is characterized in that the determining the position weight corresponding to each text unit in the target text based on the keyword dictionary includes:
[0031] Obtain the text length of the target text;
[0032] Determine the position distance between each text unit and the domain words in the corresponding text category indicated in the keyword dictionary;
[0033] Determine the position weight according to the relative relationship between the position distance and the text length.
[0034] Optionally, in some possible implementation manners of the present application, the inputting the target text into the feature extraction layer of the text recognition model to obtain a text feature vector corresponding to each text unit includes:
[0035] Input the target text into the feature extraction layer in the text recognition model;
[0036] Extract features from the target text in the feature extraction layer according to the first direction to obtain a first feature vector;
[0037] Extract features from the target text in the feature extraction layer according to the second direction to obtain a second feature vector;
[0038] Concatenate the first feature and the second feature to obtain the text feature vector corresponding to the text unit.
[0039] Optionally, in some possible implementation manners of the present application, the inputting the position weight and the text feature vector into the output layer of the text recognition model to obtain a recognition result includes:
[0040] Perform a dot product on the position weight and the text feature vector to obtain a weighted vector;
[0041] Input the weighted vector into the attention mechanism layer in the text recognition model to obtain an attention vector;
[0042] Input the attention vector into the output layer of the text recognition model to obtain the recognition result.
[0043] Optionally, in some possible implementation manners of the present application, the method further includes:
[0044] Determining a target category corresponding to the target text based on the recognition result;
[0045] Inputting the target text into a push pool corresponding to the target category;
[0046] Recalling the target text from the push pool in response to a push instruction and performing text push.
[0047] Optionally, in some possible implementation manners of the present application, the recalling the target text from the push pool in response to a push instruction and performing text push includes:
[0048] Obtaining a portrait of a target user corresponding to the push instruction;
[0049] Invoking feature tags corresponding to the portrait of the target user;
[0050] Comparing the feature tags with the target category to obtain a comparison result;
[0051] Recalling the target text from the push pool according to the comparison result;
[0052] Determining text data corresponding to the target text and performing text push.
[0053] Optionally, in some possible implementation manners of the present application, the method further includes:
[0054] Extracting a verification text from the text data;
[0055] Performing any one of the above text recognition methods based on the verification text to obtain a verification category;
[0056] Comparing the verification category and the target category to obtain verification information.
[0057] Optionally, in some possible implementation manners of the present application, the target text is a short text, the short text is a title of an article, the feature extraction layer is a bidirectional long short-term memory network, and the output layer is an attention mechanism network.
[0058] A second aspect of the present application provides a text recognition device based on location information, including:
[0059] An obtaining unit, configured to obtain a target text, where the target text includes a plurality of text units;
[0060] An input unit, configured to input the target text into a feature extraction layer in a text recognition model to obtain text feature vectors corresponding to each of the text units;
[0061] A determining unit, configured to determine the position weight corresponding to each of the text units in the target text based on a keyword dictionary, where the keyword dictionary is used to indicate domain words under different text categories, and the position weight is determined based on the positional relationship between the text unit and the domain word corresponding to the target text in the target text;
[0062] An identifying unit, configured to input the position weight and the text feature vector into an output layer in the text recognition model to obtain an identification result.
[0063] Optionally, in some possible implementation manners of this application, the determining unit is specifically configured to obtain corpus data under different preset categories and determine a plurality of candidate words;
[0064] The determining unit is specifically configured to determine the domain feature value of the candidate word in the preset category;
[0065] The determining unit is specifically configured to screen the candidate words based on the domain feature value to obtain domain dictionaries corresponding to different preset categories;
[0066] The determining unit is specifically configured to fuse the domain dictionaries corresponding to different preset categories to obtain the keyword dictionary;
[0067] The determining unit is specifically configured to determine the position weight corresponding to each of the text units in the target text based on the keyword dictionary.
[0068] Optionally, in some possible implementation manners of this application, the determining unit is specifically configured to call historical records corresponding to different preset categories;
[0069] The determining unit is specifically configured to perform text statistics based on the historical records to obtain feature items;
[0070] The determining unit is specifically configured to edit a corresponding preset whitelist according to the feature items;
[0071] The determining unit is specifically configured to traverse the corpus data based on the preset whitelist to determine a plurality of the candidate words.
[0072] Optionally, in some possible implementation manners of this application, the determining unit is specifically configured to obtain corpus data under different preset categories;
[0073] The determining unit is specifically configured to determine detection terms in the corpus data;
[0074] The determining unit is specifically configured to calculate the mutual point information corresponding to the detected term based on the corpus data, where the mutual point information is used to indicate the relevance between the detected term and the corpus data;
[0075] The determining unit is specifically configured to determine multiple candidate words according to the mutual point information.
[0076] Optionally, in some possible implementation manners of the present application, the determining unit is specifically configured to screen the candidate words based on the domain feature values to obtain a candidate set;
[0077] The determining unit is specifically configured to filter the candidate set according to a preset rule to obtain a domain dictionary corresponding to different preset categories, where the preset rule is set based on word parameters corresponding to the preset categories.
[0078] Optionally, in some possible implementation manners of the present application, the input unit is specifically configured to obtain the text length of the target text;
[0079] The determining unit is specifically configured to determine the position distance between each text unit and a domain word in the corresponding text category indicated in the keyword dictionary;
[0080] The determining unit is specifically configured to determine the position weight according to the relative relationship between the position distance and the text length.
[0081] Optionally, in some possible implementation manners of the present application, the input unit is specifically configured to input the target text into the feature extraction layer in the text recognition model;
[0082] The input unit is specifically configured to perform feature extraction on the target text in the feature extraction layer according to a first direction to obtain a first feature vector;
[0083] The input unit is specifically configured to perform feature extraction on the target text in the feature extraction layer according to a second direction to obtain a second feature vector;
[0084] The input unit is specifically configured to splice the first feature and the second feature to obtain the text feature vector corresponding to the text unit.
[0085] Optionally, in some possible implementation manners of the present application, the recognition unit is specifically configured to perform a dot product on the position weight and the text feature vector to obtain a weighted vector;
[0086] The recognition unit is specifically configured to input the weighted vector into the attention mechanism layer in the text recognition model to obtain an attention vector;
[0087] The recognition unit is specifically configured to input the attention vector into an output layer in the text recognition model to obtain the recognition result.
[0088] Optionally, in some possible implementation manners of this application, the recognition unit is specifically configured to determine a target category corresponding to the target text based on the recognition result;
[0089] The recognition unit is specifically configured to input the target text into a push pool corresponding to the target category;
[0090] The recognition unit is specifically configured to recall the target text from the push pool in response to a push instruction and perform text push.
[0091] Optionally, in some possible implementation manners of this application, the recognition unit is specifically configured to obtain a portrait of a target user corresponding to the push instruction;
[0092] The recognition unit is specifically configured to call feature tags corresponding to the portrait of the target user;
[0093] The recognition unit is specifically configured to compare the feature tags with the target category to obtain a comparison result;
[0094] The recognition unit is specifically configured to recall the target text from the push pool according to the comparison result;
[0095] The recognition unit is specifically configured to determine text data corresponding to the target text and perform text push.
[0096] Optionally, in some possible implementation manners of this application, the recognition unit is specifically configured to extract a verification text from the text data;
[0097] The recognition unit is specifically configured to execute the text recognition method described in the first aspect or any item of the first aspect based on the verification text to obtain a verification category;
[0098] The recognition unit is specifically configured to compare the verification category with the target category to obtain verification information.
[0099] A third aspect of this application provides a computer device, including: a memory, a processor, and a bus system; the memory is used to store program code; the processor is used to execute the text recognition method based on location information described in the first aspect or any item of the first aspect according to instructions in the program code.
[0100] A fourth aspect of the present application provides a computer-readable storage medium storing instructions that, when run on a computer, cause the computer to execute the method for text recognition based on location information described in the first aspect or any one of the first aspects above.
[0101] According to one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the method for text recognition based on location information provided in the first aspect or various alternative implementations of the first aspect above.
[0102] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:
[0103] By obtaining a target text, the target text includes a plurality of text units; then inputting the target text into a feature extraction layer in a text recognition model to obtain text feature vectors corresponding to each text unit; further determining the position weights corresponding to each text unit in the target text based on a keyword dictionary, the keyword dictionary is used to indicate domain words under different text categories, and the position weights are determined based on the position relationship between the text unit and the domain word corresponding to the target text in the target text; and then inputting the position weights and the text feature vectors into an output layer in the text recognition model to obtain a recognition result. Thus, a text recognition process that fuses location information and feature information is realized. Since the location information can supplement the deficiency of context information in short texts and make up for the insufficient learning of context information on the model side, the accuracy of text recognition is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0104] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0105] Figure 1 It is a network architecture diagram for the operation of a text recognition system based on location information;
[0106] Figure 2 It is a process architecture diagram for text recognition based on location information provided by an embodiment of the present application;
[0107] Figure 3 It is a flowchart of a method for text recognition based on location information provided by an embodiment of the present application;
[0108] Figure 4 It is a schematic diagram of the scenario of a text recognition method based on location information provided by an embodiment of the present application;
[0109] Figure 5 It is a schematic diagram of the scenario of another text recognition method based on location information provided by an embodiment of the present application;
[0110] Figure 6 It is a schematic diagram of the scenario of another text recognition method based on location information provided by an embodiment of the present application;
[0111] Figure 7 It is a schematic diagram of the scenario of another text recognition method based on location information provided by an embodiment of the present application;
[0112] Figure 8 It is a flowchart of another text recognition method based on location information provided by an embodiment of the present application;
[0113] Figure 9 It is a flowchart of another text recognition method based on location information provided by an embodiment of the present application;
[0114] Figure 10 It is a schematic diagram of the structure of a text recognition device based on location information provided by an embodiment of the present application;
[0115] Figure 11 It is a schematic diagram of the structure of a terminal device provided by an embodiment of the present application;
[0116] Figure 12 It is a schematic diagram of the structure of a server provided by an embodiment of the present application. Specific embodiments
[0117] The embodiments of the present application provide a text recognition method based on location information and related devices, which can be applied to a system or program with a text recognition function based on location information in a terminal device. By obtaining a target text, the target text includes multiple text units; then inputting the target text into the feature extraction layer in the text recognition model to obtain text feature vectors corresponding to each text unit; further determining the position weights corresponding to each text unit in the target text based on a keyword dictionary, the keyword dictionary is used to indicate domain words under different text categories, and the position weights are determined based on the position relationship between the text unit and the domain word corresponding to the target text in the target text; and then inputting the position weights and the text feature vectors into the output layer in the text recognition model to obtain a recognition result. Thus, a text recognition process that integrates location information and feature information is realized. Since the location information can supplement the deficiency of context information in short texts and make up for the insufficient learning of context information on the model side, the accuracy of text recognition is effectively improved.
[0118] In the description and claims of this application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products, or devices.
[0119] First, some terms that may appear in the embodiments of this application are explained.
[0120] Domain term: A representative vocabulary in different categories of corpora, that is, based on this domain term, the corresponding category can be accurately obtained.
[0121] Domain dictionary: An index set composed of different domain terms and their corresponding categories.
[0122] Keyword dictionary: An index set composed of domain dictionaries corresponding to multiple different categories.
[0123] It should be understood that the location information-based text recognition method provided by this application can be applied to a system or program with a location information-based text recognition function in a terminal device. For example, in a news reading software, specifically, the location information-based text recognition system can run in a network architecture as Figure 1 shown, as Figure 1 shown, is a network architecture diagram for the operation of the location information-based text recognition system. As can be seen from the figure, the location information-based text recognition system can provide a location information-based text recognition process for multiple information sources, that is, through the server's parsing and recognition of the information flow, send information that meets the user's needs to the terminal side; it can be understood that Figure 1 shows a variety of terminal devices. The terminal device can be a computer device. In an actual scenario, more or fewer types of terminal devices may participate in the location information-based text recognition process. The specific quantity and types depend on the actual scenario and are not limited here. In addition, Figure 1 shows one server, but in an actual scenario, multiple servers may also participate. The specific number of servers depends on the actual scenario.
[0124] In this embodiment, the server may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through wired or wireless communication means, and the terminal and the server may be connected to form a blockchain network, which is not limited in this application.
[0125] It can be understood that the above text recognition system based on location information can run on a personal mobile terminal, for example, as an application such as a news reading software, or can run on a server, or can also be run on a third-party device to provide text recognition based on location information to obtain the text recognition processing result of the information source based on location information; the specific text recognition system based on location information can run in the above device in the form of a program, or can run as a system component in the above device, or can also be a kind of cloud service program, and the specific operation mode depends on the actual scenario and is not limited here.
[0126] With the rapid development of Internet technology, people's requirements for pushing information are getting higher and higher. How to perform fast and accurate information pushing has become a difficult problem, and in the process of information pushing, natural language processing technology needs to be adopted.
[0127] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can realize effective communication between humans and computers in natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural language, that is, the language that people use in daily life, so it has a close connection with the research of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies. Among them, short text classification is an important link in the pushing process, such as the recognition of article titles.
[0128] Generally, for short text classification methods, there are mainly term frequency–inverse document frequency (TF-IDF), Latent Dirichlet Allocation (LDA), or using word2vec to sum and average, and then using a deep learning model to classify short texts.
[0129] However, in the above process of short text classification, due to the small number of words in short text data and limited text information contained, it is difficult for the above models to directly learn context information from short title data, which affects the accuracy of text recognition.
[0130] To solve the above problems, the present application proposes a text recognition method based on location information, which is applied to Figure 2 the process framework of text recognition based on location information shown in Figure 2 As shown, it is a process architecture diagram of text recognition based on location information provided by an embodiment of the present application. Through the server's parsing of the information flow, short text data (such as titles) therein is obtained, and then a model recognition process that combines keyword dictionaries of short text data with the position weights of domain words is carried out to perform accurate information classification. The above recognition process, according to the common short text classification method, adds domain word information on the basis of the original data and integrates the position information of domain words, thereby effectively improving the short text classification ability. Among them, in addition to titles, the short text data can also be comments, articles with low text content, video subtitles, and so on.
[0131] It can be understood that the biggest problem of the existing methods for short text classification is that due to the small number of words in short text data and limited text information contained, it is very difficult for existing models to directly learn context information from short title data. Most of them learn the information of words or n-grams in sentences and cannot effectively depict the representation of sentences in context semantics; moreover, in title classification, due to the inconsistent information expressed by words corresponding to different domains, it is also very difficult for the model side to learn. The present application can pre-join the mined domain word information data and specifically integrate the position information of domain words, so as to effectively improve the accuracy and recall of short title classification for titles in different domains.
[0132] It can be understood that the method provided in this application can be a program writing to serve as a processing logic in a hardware system, or can be a text recognition device based on location information, and the above-mentioned processing logic is implemented in an integrated or external connection manner. As an implementation manner, the text recognition device based on location information obtains a target text, where the target text includes multiple text units; then inputs the target text into a feature extraction layer in a text recognition model to obtain text feature vectors corresponding to each text unit; further determines the position weights corresponding to each text unit in the target text based on a keyword dictionary, where the keyword dictionary is used to indicate domain words under different text categories, and the position weights are determined based on the position relationship between the text unit and the domain word corresponding to the target text in the target text; and then inputs the position weights and the text feature vectors into an output layer in the text recognition model to obtain a recognition result. Thus, a text recognition process that integrates location information and feature information is realized. Since the location information can supplement the deficiency of context information in short texts and make up for the insufficient learning of context information on the model side, the accuracy of text recognition is effectively improved.
[0133] The solution provided in the embodiments of this application relates to the natural language processing technology of artificial intelligence, and is specifically described through the following embodiments:
[0134] Combined with the above process architecture, the text recognition method based on location information in this application will be introduced below. Please refer to Figure 3 , Figure 3 which is a flowchart of a text recognition method based on location information provided in the embodiments of this application. This management method can be executed by a terminal, or by a server, or jointly by a terminal and a server. The embodiments of this application at least include the following steps:
[0135] 301. Obtain a target text.
[0136] In this embodiment, the target text includes multiple text units; among them, a text unit is a word or phrase obtained by decomposing the target text. For example, if the target text is "Xiaoming won the first prize yesterday", the text units include "Xiaoming", "yesterday", and "the first prize". The specific division method and word granularity depend on the actual scenario and are not limited here.
[0137] It can be understood that the target text in this embodiment is a short text, and the short text is the title of an article, that is, by recognizing the title, the corresponding article content is pushed. Due to the word count limit of short texts, it is difficult for the recognition model to learn the context relationship of short texts. The method of this application can improve the accuracy of short text recognition.
[0138] 302. Input the target text into the feature extraction layer of the text recognition model to obtain the text feature vectors corresponding to each text unit.
[0139] In this embodiment, the process of using the text recognition model can be as Figure 4 shown. Figure 4 FIG. is a schematic diagram of a scenario of a text recognition method based on location information provided by an embodiment of the present application; that is, first input the target text into the text recognition model, and the position weight of the text unit in the target text relative to the domain word will be combined during the recognition process of the text recognition model. This is because the closer a word (text unit) is to the domain word, the greater its impact on whether the article title is of high quality, that is, the more it conforms to the category corresponding to the domain word, thereby improving the accuracy of text classification.
[0140] Specifically, the combination process of the text recognition model and the position weight is based on the feature vector, that is, vector operations are performed on the basis of the text feature vectors corresponding to each text unit obtained in the feature extraction layer; among them, the feature extraction layer can be a model network with any specific text feature extraction ability, such as CNN, RNN, Long Short-Term Memory (LSTM), BERT, etc. The specific feature extraction method depends on the actual scenario and is not limited here.
[0141] In a possible scenario, in order to further improve the extraction of the context relationship of short texts in text recognition, the feature extraction layer can be set to a bidirectional LSTM, specifically as Figure 5 shown. Figure 5 FIG. is a schematic diagram of another scenario of a text recognition method based on location information provided by an embodiment of the present application; the figure shows that the framework of the entire model uses a bidirectional LSTM and an attention mechanism model as the backbone model. In the design of the model, in addition to introducing the domain dictionaries mined under different categories that have been externally mined, the position weight in the title is also fused through the position information memory network. The vector generated by splicing the bidirectional LSTM is then dot-multiplied with the position weight corresponding to each node, and finally the title is vector-represented by introducing the attention mechanism structure.
[0142] For the data conduction process in the bidirectional LSTM, as Figure 6 shown. Figure 6A schematic diagram of another text recognition method based on location information provided by an embodiment of the present application; it can be seen that for the process of extracting text feature vectors by a bidirectional LSTM, the target text (x1, x2, …) can be first input into the feature extraction layer in the text recognition model; then, according to the first direction, the target text is subjected to feature extraction in the feature extraction layer to obtain a first feature vector; and according to the second direction, the target text is subjected to feature extraction in the feature extraction layer to obtain a second feature vector; furthermore, the first feature and the second feature are concatenated to obtain the text feature vector corresponding to the text unit.
[0143] It can be understood that the recurrent neural network in the long short-term memory network has the following characteristics: the recurrent neural network can generate an output at each time node, and the connections between hidden units are recurrent; and the recurrent neural network can generate an output at each time node, and the output at this time node has only a recurrent connection with the hidden unit of the next time node; in addition, the recurrent neural network includes hidden units with recurrent connections and can process sequence data and output a single prediction.
[0144] Specifically, for the internal structure of the LSTM, as Figure 7 shown Figure 7 A schematic diagram of another text recognition method based on location information provided by an embodiment of the present application; by adding an input threshold δ1, a forgetting threshold δ2, and an output threshold δ3 to the LSTM, the weight of the self-loop is changed, so that under the condition of fixed model parameters, the integration scale at different times can be dynamically changed, thereby avoiding the problem of gradient disappearance or gradient explosion. That is, the hidden layer in the LSTM needs to save two values, ht+1 is involved in the forward calculation, ht-1 is involved in the reverse calculation, and the final output value ht depends on ht-1 and ht+1, thereby strengthening the context connection in the text recognition process.
[0145] 303. Determine the position weights of each text unit in the target text based on the keyword dictionary.
[0146] In this embodiment, the keyword dictionary is used to indicate the domain words under different text categories, and the position weights are determined based on the position relationship between the text unit and the domain word corresponding to the target text in the target text.
[0147] Specifically, the determination of the domain word can be carried out through a diff-idf score, that is, the domain characteristics of the candidate word are measured by the diff-idf score, that is, the document frequency of the candidate word appearing in its own domain is much higher than that in other domains. After truncating the sorting by the diff-idf score, post-processing is performed, such as removing noise and entity normalization, etc., and together with some public entries (such as domain words obtained by network retrieval), a keyword dictionary is thus formed.
[0148] For the process of determining candidate words, different preset category corpora can be obtained first, and multiple candidate words can be determined; then, the domain feature values of the candidate words in the preset category can be determined; and the candidate words can be filtered based on the domain feature values to obtain the domain dictionaries corresponding to different preset categories; furthermore, the domain dictionaries corresponding to different preset categories can be integrated to obtain the keyword dictionary; and thus, the position weights corresponding to each text unit in the target text can be determined based on the keyword dictionary.
[0149] Among them, mining candidate words (such as LeBron James, refusal, play) in a specific category corpus (such as sports news) is the mining process of candidate words, which can include whitelist mining and Pointwise Mutual Information (PMI) mining. For whitelist mining, the historical records corresponding to different preset categories can be called first, such as the domain word records in the sports category; then, text statistics can be performed based on the historical records to obtain feature items (such as the item with the most occurrences); and the corresponding preset whitelist can be edited according to the feature items; furthermore, the corpus data can be traversed based on the preset whitelist to determine multiple candidate words, thereby improving the mining efficiency of candidate words.
[0150] For the PMI mining process, different preset category corpora can be obtained first; then, the detection entries in the corpus data can be determined; and the mutual point information (PMI) corresponding to the detection entries can be calculated based on the corpus data, and the mutual point information is used to indicate the relevance between the detection entries and the corpus data; furthermore, multiple candidate words can be determined according to the mutual point information. The calculation process of the mutual point information can be referred to the following formula:
[0151]
[0152] Among them, x and y are the detection entries and the corpus data, respectively, so as to measure the relevance between the two based on the magnitude of the value.
[0153] Then, for each candidate word mined above, the diff_idf value:
[0154]
[0155] Furthermore, the words are sorted and truncated, for example, the 3 words with the largest PMI values are selected.
[0156] Optionally, the screening of candidate words can be based on word parameters, that is, first screening the candidate words based on domain feature values to obtain a candidate set; then filtering the candidate set according to preset rules to obtain domain dictionaries corresponding to different preset categories. The preset rules are set based on the word parameters corresponding to the preset categories, such as part-of-speech filtering (selecting nouns corresponding to the category) and graph alignment, etc., and then fusing the domain dictionaries under multiple different categories to form a keyword dictionary.
[0157] Furthermore, for the calculation process of the position weight, the text length of the target text can be obtained first; then determine the position distance between each text unit and the domain words under the corresponding text category indicated in the keyword dictionary; and then determine the position weight according to the relative relationship between the position distance and the text length. That is, according to the keyword dictionary mined in the previous step, calculate the position weights corresponding to different words in the article title respectively. The position weight calculation formula is as follows:
[0158]
[0159] Among them, θ is the position weight of each word, N is the length of the sentence, l is the position of the domain word in the sentence, and δ is the text unit. This formula is set based on the fact that the closer a word is to the domain word, the greater its impact on whether the article title is of high quality. Other formula deformations based on this process should also be revealed.
[0160] 304. Input the position weight and the text feature vector into the output layer in the text recognition model to obtain the recognition result.
[0161] In this embodiment, combined with the above Figure 5 shown structure, the position weight and the text feature vector can be multiplied to obtain a weighted vector; then input the weighted vector into the attention mechanism layer in the text recognition model to obtain an attention vector; and then input the attention vector into the output layer in the text recognition model to obtain the recognition result.
[0162] Specifically, the attention mechanism layer (Attention Model) is widely used in various different types of deep learning tasks such as natural language processing, image recognition, and speech recognition. The attention model simulates that when a person is looking at something, what can be focused on at the current moment must be a certain place of the thing that can be currently seen. When the gaze moves to another place, the attention also shifts with the movement of the gaze. When people notice a certain target or a certain scene, the attention distribution at each spatial position inside the target and within the scene is different.
[0163] In terms of the calculation process, the attention mechanism layer selectively filters out a small amount of important information from a large amount of information and focuses on this important information, ignoring most of the unimportant information. The focusing process is reflected in the calculation of the weight coefficients. The larger the weight, the more focused on its corresponding Value value. That is, the weight represents the importance of the information, and Value is the corresponding information. By the Attention model, different sizes of attention parameters are assigned to each of the said word vectors. When generating each output intention, the original identical intermediate semantic representation C will be replaced by a C that changes continuously according to the current generated intention. That is, through the Attention mechanism, it is no longer necessary to encode the complete original text sentence into a fixed-length vector. Instead, the decoder can be allowed to participate in different parts of the original text at each step of the output. Finally, the possible discourse intention generated by Attention passes through the softmax classifier to determine the recognition result of the input vector.
[0164] Optionally, after obtaining the recognition result, the target category corresponding to the target text can also be determined based on the recognition result; then the target text is input into the push pool corresponding to the target category; and then in response to the push instruction, the target text is recalled from the push pool and text is pushed, thus realizing an accurate information push process.
[0165] Combined with the above embodiments, by obtaining the target text, the target text includes multiple text units; then the target text is input into the feature extraction layer in the text recognition model to obtain the text feature vectors corresponding to each text unit; further, based on the keyword dictionary, the position weights corresponding to each text unit in the target text are determined. The keyword dictionary is used to indicate the domain words under different text categories, and the position weights are determined based on the position relationship between the text unit and the domain word corresponding to the target text in the target text; then the position weights and the text feature vectors are input into the output layer in the text recognition model to obtain the recognition result. Thus, a text recognition process that fuses position information and feature information is realized. Since the position information can supplement the deficiency of context information in short texts and make up for the insufficient learning of context information on the model side, the accuracy of text recognition is effectively improved.
[0166] The following describes the information push scenario. Please refer to Figure 8 , Figure 8 which is a flowchart of another text recognition method based on position information provided by an embodiment of the present application. The embodiment of the present application at least includes the following steps:
[0167] 801. Obtain the information flow in real time.
[0168] In this embodiment, for short text data with a small number of words or little text information, domain word data mined in advance is added, and a position function is designed according to the positions of the domain words to generate the position weights of each word in the short text, which can be integrated into the model to effectively learn the domain words and improve the accuracy of short text classification.
[0169] It can be understood that the real-time obtained information flow can be a large amount of information flow data every day in the content platform center, and there are many fields, such as category systems in entertainment, sports, technology, etc. In some push scenarios, it is necessary to determine whether the title of an article is of high quality, so as to reduce the amount of data distributed for low-quality titles such as clickbait, and reach high-quality users less. In Lookout, Tencent News Flash, and browser news, it is crucial to mine high-quality title articles and resist low-quality and clickbait titles, which can not only improve the overall article quality of the information flow, but also improve the new user acquisition and retention capabilities of the entire product for some application scenarios.
[0170] In a possible scenario, for the push scenario in the information flow product every day, the operation needs to label a large amount of high-quality title data into the push pool every day, and then the system pushes it to different users personalized according to the recommendation algorithm, mainly playing the role of acquiring new users. The main role of this model is to replace the daily large amount of labeling work of the operation. According to the article titles in the content center every day, it mines the domain words in different categories, and then generates the corresponding position weights for each word according to the position information of the domain words in the title, and integrates it into the short text classification method of deep learning. Therefore, a bidirectional LSTM network and an attention mechanism model are used to classify short texts.
[0171] 802. Identify the short text in the information flow to obtain classification information.
[0172] In this embodiment, the process of identifying the short text refers to Figure 3 the embodiment shown, which will not be elaborated here.
[0173] 803. Based on the classification information, the corresponding push pool of the information flow data.
[0174] In this embodiment, different push pools are set based on different classification information, and the push pool can be set locally or in the cloud to achieve rapid information distribution.
[0175] 804. Push information in response to a push instruction.
[0176] In this embodiment, the push instruction can be initiated when the user logs in to the application, such as when the user logs in to the news application; the push instruction can also be initiated when the user refreshes the interface. The specific method depends on the actual scenario and will not be limited here.
[0177] In the above embodiments, the information volume of short texts is expanded by borrowing external knowledge, words in different category fields under the information flow are mined, and the corresponding position weights are calculated for each word in the title and integrated into the bidirectional LSTM network and the attention mechanism model to classify the text, which can effectively improve the accuracy of short text classification and replace the operation to complete the screening of a large number of high-quality titles; it can also screen out the titles corresponding to more high-quality articles faster, thereby effectively improving the distribution efficiency of the push scenario.
[0178] Optionally, when determining the push information, a data verification process can also be performed. For details, see Figure 9 , Figure 9 which is a flowchart of another text recognition method based on location information provided by the embodiments of the present application. The embodiments of the present application at least include the following steps:
[0179] 901. Push information in response to a push instruction.
[0180] In this embodiment, since the text recognition process is for the article title, the corresponding article can be called through the article title, so as to perform secondary verification on the pushed information based on the article.
[0181] 902. Obtain the portrait of the target user corresponding to the push instruction.
[0182] In this embodiment, the portrait of the target user includes the user's preferences and historical browsing records, etc. Since the user's browsing has a certain correlation, the pushed information can be verified.
[0183] 903. Call the feature tags corresponding to the portrait of the target user.
[0184] In this embodiment, the feature tags are the labeled representations of the portrait of the target user, such as "sports" and "art", etc., so as to facilitate comparison with the classification words in the recognition result.
[0185] 904. Compare the feature tags with the target category to obtain a comparison result.
[0186] In this embodiment, the comparison result can be a comparison result at the text level, such as the feature tags being the same as the target category; the comparison result can also be a comparison result at the semantic level, such as the feature tags being similar in meaning to the target category. There is no limitation here.
[0187] 905. Recall the target text from the push pool according to the comparison result.
[0188] In this embodiment, if the comparison result indicates that the feature tags are the same as or similar to the target category, the target text is recalled from the push pool.
[0189] 906. Determine the text data corresponding to the target text and perform text push.
[0190] In this embodiment, before performing text push, information verification can also be performed based on the text data, that is, first extract the verification text from the text data; then execute Figure 3 the text recognition method shown to obtain the verification category; then compare the verification category with the target category to obtain the verification information.
[0191] Specifically, the verification text can be the subtitle of the article, the beginning of the article, the end of the article, etc., or a text combination of the above representative parts, and the specific form depends on the actual scenario.
[0192] It can be understood that if the verification category is the same as or similar to the target category, it means that the text content also meets the user's push requirements, thus avoiding the push of "clickbait" articles and ensuring the accuracy of information push.
[0193] To better implement the above solutions of the embodiments of the present application, the following also provides related devices for implementing the above solutions. Please refer to Figure 10 , Figure 10 FIG. is a schematic structural diagram of a text recognition device based on location information provided by an embodiment of the present application. The recognition device 1000 includes:
[0194] An acquisition unit 1001, configured to acquire a target text, where the target text includes a plurality of text units;
[0195] An input unit 1002, configured to input the target text into a feature extraction layer in a text recognition model to obtain text feature vectors corresponding to each of the text units;
[0196] A determination unit 1003, configured to determine the position weight corresponding to each of the text units in the target text based on a keyword dictionary, where the keyword dictionary is used to indicate domain words under different text categories, and the position weight is determined based on the position relationship between the text unit and the domain word corresponding to the target text in the target text;
[0197] A recognition unit 1004, configured to input the position weight and the text feature vector into an output layer in the text recognition model to obtain a recognition result.
[0198] Optionally, in some possible implementation manners of the present application, the determination unit 1003 is specifically configured to acquire corpus data under different preset categories and determine a plurality of candidate words;
[0199] The determination unit 1003 is specifically configured to determine the domain feature value of the candidate word in the preset category;
[0200] The determining unit 1003 is specifically configured to screen the candidate words based on the domain feature values to obtain domain dictionaries corresponding to different preset categories;
[0201] The determining unit 1003 is specifically configured to fuse the domain dictionaries corresponding to different preset categories to obtain the keyword dictionary;
[0202] The determining unit 1003 is specifically configured to determine the position weights of each text unit in the target text based on the keyword dictionary.
[0203] Optionally, in some possible implementation manners of the present application, the determining unit 1003 is specifically configured to call historical records corresponding to different preset categories;
[0204] The determining unit 1003 is specifically configured to perform text statistics based on the historical records to obtain feature items;
[0205] The determining unit 1003 is specifically configured to edit corresponding preset whitelists according to the feature items;
[0206] The determining unit 1003 is specifically configured to traverse the corpus data based on the preset whitelist to determine a plurality of candidate words.
[0207] Optionally, in some possible implementation manners of the present application, the determining unit 1003 is specifically configured to obtain corpus data under different preset categories;
[0208] The determining unit 1003 is specifically configured to determine detection entries in the corpus data;
[0209] The determining unit 1003 is specifically configured to calculate mutual point information corresponding to the detection entries based on the corpus data, and the mutual point information is used to indicate the relevance between the detection entries and the corpus data;
[0210] The determining unit 1003 is specifically configured to determine a plurality of candidate words according to the mutual point information.
[0211] Optionally, in some possible implementation manners of the present application, the determining unit 1003 is specifically configured to screen the candidate words based on the domain feature values to obtain a candidate set;
[0212] The determining unit 1003 is specifically configured to filter the candidate set according to a preset rule to obtain domain dictionaries corresponding to different preset categories, and the preset rule is set based on word parameters corresponding to the preset categories.
[0213] Optionally, in some possible implementation manners of the present application, it is characterized in that the determining unit 1003 is specifically configured to obtain the text length of the target text;
[0214] The determining unit 1003 is specifically configured to determine the position distance between each text unit and the domain word in the corresponding text category indicated in the keyword dictionary;
[0215] The determining unit 1003 is specifically configured to determine the position weight according to the relative relationship between the position distance and the text length.
[0216] Optionally, in some possible implementation manners of the present application, the input unit 1002 is specifically configured to input the target text into the feature extraction layer in the text recognition model;
[0217] The input unit 1002 is specifically configured to perform feature extraction on the target text in the feature extraction layer according to a first direction to obtain a first feature vector;
[0218] The input unit 1002 is specifically configured to perform feature extraction on the target text in the feature extraction layer according to a second direction to obtain a second feature vector;
[0219] The input unit 1002 is specifically configured to splice the first feature and the second feature to obtain the text feature vector corresponding to the text unit.
[0220] Optionally, in some possible implementation manners of the present application, the recognition unit 1004 is specifically configured to perform a dot product on the position weight and the text feature vector to obtain a weighted vector;
[0221] The recognition unit 1004 is specifically configured to input the weighted vector into the attention mechanism layer in the text recognition model to obtain an attention vector;
[0222] The recognition unit 1004 is specifically configured to input the attention vector into the output layer in the text recognition model to obtain the recognition result.
[0223] Optionally, in some possible implementation manners of the present application, the recognition unit 1004 is specifically configured to determine the target category corresponding to the target text based on the recognition result;
[0224] The recognition unit 1004 is specifically configured to input the target text into the push pool corresponding to the target category;
[0225] The recognition unit 1004 is specifically configured to recall the target text from the push pool and perform text push in response to a push instruction.
[0226] Optionally, in some possible implementation manners of the present application, the recognition unit 1004 is specifically configured to obtain a portrait of a target user corresponding to the push instruction;
[0227] The recognition unit 1004 is specifically configured to call a feature tag corresponding to the portrait of the target user;
[0228] The recognition unit 1004 is specifically configured to compare the feature tag with the target category to obtain a comparison result;
[0229] The recognition unit 1004 is specifically configured to recall the target text from the push pool according to the comparison result;
[0230] The recognition unit 1004 is specifically configured to determine text data corresponding to the target text and perform text push.
[0231] Optionally, in some possible implementation manners of the present application, the recognition unit 1004 is specifically configured to extract a verification text from the text data;
[0232] The recognition unit 1004 is specifically configured to execute the text recognition method described in the first aspect or any item of the first aspect based on the verification text to obtain a verification category;
[0233] The recognition unit 1004 is specifically configured to compare the verification category and the target category to obtain verification information.
[0234] By obtaining a target text, the target text includes a plurality of text units; then inputting the target text into a feature extraction layer in a text recognition model to obtain text feature vectors corresponding to the respective text units; further determining a position weight corresponding to each text unit in the target text based on a keyword dictionary, the keyword dictionary is used to indicate domain words under different text categories, and the position weight is determined based on the position relationship between the text unit and the domain word corresponding to the target text in the target text; and then inputting the position weight and the text feature vectors into an output layer in the text recognition model to obtain a recognition result. Thereby realizing a text recognition process that fuses position information and feature information. Since the position information can supplement the deficiency of context information in short texts and make up for the insufficient learning of context information on the model side, the accuracy of text recognition is effectively improved.
[0235] The embodiments of the present application further provide a terminal device, such as Figure 11As shown in the figure, it is a schematic structural diagram of another terminal device provided by an embodiment of the present application. For the convenience of description, only the parts related to the embodiment of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The terminal can be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS), an in-vehicle computer, etc. Taking the terminal as a mobile phone as an example:
[0236] Figure 11 Shown is a block diagram of a part of the structure of a mobile phone related to the terminal provided by an embodiment of the present application. Refer to Figure 11 , the mobile phone includes: a radio frequency (RF) circuit 1110, a memory 1120, an input unit 1130, a display unit 1140, a sensor 1150, an audio circuit 1160, a wireless fidelity (WiFi) module 1170, a processor 1180, and a power supply 1190 and other components. Those skilled in the art can understand that Figure 11 the structure of the mobile phone shown in
[0237] does not limit the mobile phone, and may include more or fewer components than shown in the figure, or combine certain components, or arrange different components. Figure 11 The following specifically introduces each component of the mobile phone in combination with
[0238] The RF circuit 1110 can be used for receiving and transmitting information or signals during communication. Specifically, after receiving the downlink information from the base station, it is sent to the processor 1180 for processing. Additionally, the uplink data is sent to the base station. Generally, the RF circuit 1110 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. Moreover, the RF circuit 1110 can also communicate with the network and other devices via wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile Communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0239] The memory 1120 can be used to store software programs and modules. The processor 1180 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 1120. The memory 1120 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.). In addition, the memory 1120 can include a high-speed random access memory and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.
[0240] The input unit 1130 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 1130 can include a touch panel 1131 and other input devices 1132. The touch panel 1131, also known as a touch screen, can collect touch operations of the user thereon or nearby (such as operations of the user using any suitable object or accessory such as a finger, a stylus, etc. on or near the touch panel 1131, and air touch operations within a certain range on the touch panel 1131), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 1131 can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 1180, and can receive and execute the commands sent by the processor 1180. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 1131. In addition to the touch panel 1131, the input unit 1130 can also include other input devices 1132. Specifically, the other input devices 1132 can include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.
[0241] The display unit 1140 can be used to display information input by the user or provided to the user, as well as various menus of the mobile phone. The display unit 1140 can include a display panel 1141. Optionally, the display panel 1141 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 1131 can cover the display panel 1141. When the touch panel 1131 detects a touch operation thereon or nearby, it transmits it to the processor 1180 to determine the type of touch event. Subsequently, the processor 1180 provides a corresponding visual output on the display panel 1141 according to the type of touch event. Although in Figure 11 the touch panel 1131 and the display panel 1141 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 1131 and the display panel 1141 can be integrated to realize the input and output functions of the mobile phone.
[0242] The mobile phone may further include at least one sensor 1150, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 1141 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 1141 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that the mobile phone can also be configured with, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be elaborated here.
[0243] The audio circuit 1160, the speaker 1161, and the microphone 1162 can provide an audio interface between the user and the mobile phone. The audio circuit 1160 can transmit the electrical signal converted from the received audio data to the speaker 1161, and the speaker 1161 converts it into a sound signal for output; on the other hand, the microphone 1162 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1160 and then converted into audio data. After the audio data is output to the processor 1180 for processing, it is sent through the RF circuit 1110 to, for example, another mobile phone, or the audio data is output to the memory 1120 for further processing.
[0244] WiFi belongs to short - range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 1170, which provides users with wireless broadband Internet access. Although Figure 11 the WiFi module 1170 is shown, it can be understood that it does not belong to an essential component of the mobile phone and can be omitted completely within the scope of not changing the essence of the invention according to needs.
[0245] The processor 1180 is the control center of the mobile phone, connecting various parts of the entire mobile phone using various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 1120, and by calling the data stored in the memory 1120, it executes various functions of the mobile phone and processes data. Optionally, the processor 1180 may include one or more processing units; optionally, the processor 1180 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above - mentioned modem processor may not be integrated into the processor 1180 either.
[0246] The mobile phone further includes a power supply 1190 (such as a battery) for powering each component. Optionally, the power supply can be logically connected to the processor 1180 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system.
[0247] Although not shown, the mobile phone may further include a camera, a Bluetooth module, etc., which will not be elaborated herein.
[0248] In the embodiment of the present application, the processor 1180 included in the terminal further has the function of executing each step of the page processing method as described above.
[0249] The embodiment of the present application further provides a server. Please refer to Figure 12 , Figure 12 FIG. is a schematic structural diagram of a server provided by the embodiment of the present application. The server 1200 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1222 (for example, one or more processors) and a memory 1232, and one or more storage media 1230 for storing application programs 1242 or data 1244 (for example, one or more mass storage devices). Among them, the memory 1232 and the storage media 1230 may be transient storage or persistent storage. The program stored in the storage media 1230 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processor 1222 may be configured to communicate with the storage media 1230 and execute a series of instruction operations in the storage media 1230 on the server 1200.
[0250] The server 1200 may further include one or more power supplies 1226, one or more wired or wireless network interfaces 1250, one or more input / output interfaces 1258, and / or one or more operating systems 1241, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0251] The steps performed by the management device in the above embodiments may be based on the Figure 12 server structure shown.
[0252] The embodiment of the present application further provides a computer-readable storage medium, in which text recognition instructions based on location information are stored. When it runs on a computer, it enables the computer to execute the steps performed by the text recognition device based on location information in the method described in the foregoing Figures 3 to 9 embodiment shown.
[0253] In an embodiment of the present application, there is also provided a computer program product including a text recognition instruction based on location information. When it runs on a computer, it causes the computer to execute the steps performed by the text recognition device based on location information in the method described in the foregoing Figures 3 to 9 embodiment shown.
[0254] The embodiment of the present application also provides a text recognition system based on location information. The text recognition system based on location information may include Figure 10 the text recognition device based on location information in the described embodiment, or Figure 11 the terminal device in the described embodiment, or Figure 12 the server described.
[0255] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0256] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0257] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0258] In addition, in each embodiment of the present application, the various functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0259] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a location information-based text recognition device, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0260] As described above, the above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of this application.
Claims
1. A text recognition method based on location information, characterized in that, it includes: Obtain a target text, where the target text includes multiple text units; Input the target text into the feature extraction layer in the text recognition model to obtain text feature vectors corresponding to each of the text units; Determine the position weight corresponding to each of the text units in the target text based on a keyword dictionary; the keyword dictionary is used to indicate domain words under different text categories; the closer the position of the text unit to the domain word corresponding to the target text in the target text, the greater the position weight corresponding to the text unit in the target text; Input the position weight and the text feature vector into the output layer in the text recognition model to obtain a recognition result; The keyword dictionary is obtained in the following manner: Obtain corpus data under different preset categories and determine multiple candidate words; Determine the domain feature value of the candidate word in the preset category; Filter the candidate words based on the domain feature value to obtain domain dictionaries corresponding to different preset categories; Fuse the domain dictionaries corresponding to different preset categories to obtain the keyword dictionary.
2. The method according to claim 1, characterized in that, the obtaining corpus data under different preset categories and determining multiple candidate words includes: Call the historical records corresponding to different preset categories; Perform text statistics based on the historical records to obtain feature items; Edit a corresponding preset whitelist according to the feature items; Traverse the corpus data based on the preset whitelist to determine multiple candidate words.
3. The method according to claim 1, characterized in that, the obtaining corpus data under different preset categories and determining multiple candidate words includes: Obtain corpus data under different preset categories; Determine the detection entries in the corpus data; Calculate the mutual click information corresponding to the detection entries based on the corpus data, where the mutual click information is used to indicate the relevance between the detection entries and the corpus data; Determine multiple candidate words according to the mutual click information.
4. The method according to claim 1, characterized in that, the filtering the candidate words based on the domain feature value to obtain domain dictionaries corresponding to different preset categories includes: Filter the candidate words based on the domain feature value to obtain a candidate set; Filter the candidate set according to a preset rule to obtain domain dictionaries corresponding to different preset categories, where the preset rule is set based on the word parameters corresponding to the preset category.
5. The method according to claim 1, characterized in that, the determining the position weight corresponding to each of the text units in the target text based on the keyword dictionary includes: Obtain the text length of the target text; Determine the position distance between each of the text units and the domain words under the corresponding text category indicated in the keyword dictionary; Determine the position weight according to the proportional relationship between the position distance and the text length.
6. The method according to claim 1, characterized in that, Inputting the target text into a feature extraction layer in a text recognition model to obtain text feature vectors corresponding to each of the text units, includes: Inputting the target text into the feature extraction layer in the text recognition model; Performing feature extraction on the target text in the feature extraction layer according to a first direction to obtain a first feature vector; Performing feature extraction on the target text in the feature extraction layer according to a second direction to obtain a second feature vector; Concatenating the first feature vector and the second feature vector to obtain the text feature vector corresponding to the text unit.
7. The method according to claim 1, wherein, Inputting the position weight and the text feature vector into an output layer in the text recognition model to obtain a recognition result, includes: Performing a dot product on the position weight and the text feature vector to obtain a weighted vector; Inputting the weighted vector into an attention mechanism layer in the text recognition model to obtain an attention vector; Inputting the attention vector into the output layer in the text recognition model to obtain the recognition result.
8. The method according to any one of claims 1-7, wherein, The method further includes: Determining a target category corresponding to the target text based on the recognition result; Inputting the target text into a push pool corresponding to the target category; Recalling the target text from the push pool in response to a push instruction and performing text push.
9. The method according to claim 8, wherein, The recalling the target text from the push pool in response to a push instruction and performing text push includes: Obtaining a portrait of a target user corresponding to the push instruction; Invoking feature tags corresponding to the portrait of the target user; Comparing the feature tags with the target category to obtain a comparison result; Recalling the target text from the push pool according to the comparison result; Determining text data corresponding to the target text and performing text push.
10. The method according to claim 9, wherein, The method further includes: Extracting a verification text from the text data; Performing the text recognition method according to any one of claims 1-7 based on the verification text to obtain a verification category; Comparing the verification category and the target category to obtain verification information.
11. The method according to claim 1, wherein, The target text is a short text, the short text is a title of an article, the feature extraction layer is a bidirectional long short-term memory network, and the output layer is an attention mechanism network.
12. A text recognition device based on location information, wherein, includes: An acquisition unit for acquiring a target text, the target text including a plurality of text units; An input unit for inputting the target text into a feature extraction layer in a text recognition model to obtain text feature vectors corresponding to each of the text units; A determining unit, configured to determine the position weight corresponding to each of the text units in the target text based on a keyword dictionary; the keyword dictionary is used to indicate domain words under different text categories; the closer the position of the text unit and the domain word corresponding to the target text in the target text, the greater the position weight corresponding to the text unit in the target text; An identifying unit, configured to input the position weight and the text feature vector into an output layer in the text recognition model to obtain an identification result; The determining unit is further configured to obtain corpus data under different preset categories and determine a plurality of candidate words; determine the domain feature value of the candidate words in the preset category; Filter the candidate words based on the domain feature value to obtain a domain dictionary corresponding to different preset categories; Fuse the domain dictionaries corresponding to different preset categories to obtain the keyword dictionary.
13. The apparatus according to claim 12, wherein, the determining unit is specifically configured to: invoke historical records corresponding to different preset categories; perform text statistics based on the historical records to obtain feature items; edit a corresponding preset whitelist according to the feature items; traverse the corpus data based on the preset whitelist to determine a plurality of the candidate words.
14. The apparatus according to claim 12, wherein, the determining unit is specifically configured to: obtain corpus data under different preset categories; determine detection entries in the corpus data; calculate mutual point information corresponding to the detection entries based on the corpus data, where the mutual point information is used to indicate the relevance between the detection entries and the corpus data; determine a plurality of the candidate words according to the mutual point information.
15. The apparatus according to claim 12, wherein, the determining unit is specifically configured to: filter the candidate words based on the domain feature value to obtain a candidate set; filter the candidate set according to a preset rule to obtain a domain dictionary corresponding to different preset categories, where the preset rule is set based on word parameters corresponding to the preset category.
16. The apparatus according to claim 12, wherein, the determining unit is specifically configured to: obtain the text length of the target text; determine the position distance between each text unit and the domain word in the corresponding text category indicated in the keyword dictionary; determine the position weight according to the proportional relationship between the position distance and the text length.
17. The apparatus according to claim 12, wherein, the input unit is specifically configured to: input the target text into the feature extraction layer in the text recognition model; perform feature extraction on the target text in the feature extraction layer according to a first direction to obtain a first feature vector; perform feature extraction on the target text in the feature extraction layer according to a second direction to obtain a second feature vector; concatenate the first feature vector and the second feature vector to obtain the text feature vector corresponding to the text unit.
18. The apparatus according to claim 12, wherein, The recognition unit is specifically configured to: Perform a dot product of the position weight and the text feature vector to obtain a weighted vector; Input the weighted vector into the attention mechanism layer in the text recognition model to obtain an attention vector; Input the attention vector into the output layer in the text recognition model to obtain the recognition result.
19. The apparatus according to any one of claims 12 - 18, wherein, The recognition unit is specifically configured to: Determine the target category corresponding to the target text based on the recognition result; Input the target text into the push pool corresponding to the target category; Recall the target text from the push pool in response to a push instruction and perform text push.
20. The apparatus according to claim 19, wherein, The recognition unit is specifically configured to: Obtain the portrait of the target user corresponding to the push instruction; Invoke the feature tags corresponding to the portrait of the target user; Compare the feature tags with the target category to obtain a comparison result; Recall the target text from the push pool according to the comparison result; Determine the text data corresponding to the target text and perform text push.
21. The apparatus according to claim 20, wherein, The recognition unit is specifically configured to: Extract a verification text from the text data; Execute the text recognition method according to any one of claims 1 - 7 based on the verification text to obtain a verification category; Compare the verification category and the target category to obtain verification information.
22. The apparatus according to claim 12, wherein, The target text is a short text, the short text is the title of an article, the feature extraction layer is a bidirectional long short - term memory network, and the output layer is an attention mechanism network.
23. A computer device, wherein, The computer device includes a processor and a memory: The memory is used to store program codes; the processor is used to execute the position - information - based text recognition method according to any one of claims 1 to 11 according to the instructions in the program codes.
24. A computer - readable storage medium stores instructions, which when running on a computer, cause the computer to execute the position - information - based text recognition method according to any one of claims 1 to 11 above.
25. A computer program product, wherein, The computer program product includes computer instructions, and the computer instructions are stored in a computer - readable storage medium; a processor of a computer device reads the computer instructions from the computer - readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the position - information - based text recognition method according to any one of claims 1 to 11 above.
Citation Information
Patent Citations
A text classification method based on multi-angle capsule network
CN109241283A
Semantic recognition method and device, medium and electronic equipment
CN110633464A