Text recognition method, device and program product
By extracting links from text and recognizing fusion matrices, combined with a deep learning model, the problem of inaccurate abnormal text recognition in existing technologies has been solved, achieving efficient recognition and security protection of complex and diverse texts.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- ZTE CORP
- Filing Date
- 2025-09-30
- Publication Date
- 2026-04-23
AI Technical Summary
Existing technologies are not very effective in identifying complex and diverse abnormal texts, making it difficult to accurately identify abnormal information in the text, which leads to increased information security threats.
By extracting links from the text to be processed, determining the link extraction results, and determining the fusion matrix based on the text to be processed, the link extraction results, and the segmentation dimensions, the text type is identified using the fusion matrix. The recognition accuracy is improved by combining a deep learning model with a self-attention mechanism and Stratified 10-Fold cross-validation.
It improves the accuracy of identifying complex and diverse abnormal text, reduces information loss, enhances information security, and can promptly intercept and alert users to abnormal text, thus protecting user information security.
Smart Images

Figure CN2025125642_23042026_PF_FP_ABST
Abstract
Description
Text recognition methods, devices and software products Technical Field
[0001] This application relates to the field of information security technology, and in particular to a text recognition method, device and program product. Background Technology
[0002] With the widespread use of mobile devices and text messaging services, the information security threat posed by anomalous text is increasing. Most anomalous texts contain links that entice users to click, and due to the complex and varied semantics of these texts, current methods for identifying anomalous information within them are not very effective, presenting significant limitations in dealing with complex and diverse anomalous texts. Summary of the Invention
[0003] This application provides a text recognition method, device, and program product, which improves the effectiveness of anomaly recognition for complex text information and reduces the limitations of text anomaly recognition when facing complex and diverse abnormal text.
[0004] This application provides a text recognition method, including:
[0005] Extract links from the text to be processed and determine the link extraction results;
[0006] A fusion matrix is determined based on the text to be processed, the link extraction results, and the segmentation dimension corresponding to the text to be processed, and the text type of the text to be processed is identified based on the fusion matrix.
[0007] This application also provides an electronic device, including: a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for connecting and communicating between the processor and the memory. The program being executed by the processor is a step to implement the text recognition method provided in any of the above embodiments.
[0008] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the text recognition method provided in any of the above embodiments.
[0009] This application also provides a storage medium for computer-readable storage, which stores one or more programs that can be executed by one or more processors to implement the steps of the text recognition method provided in any of the above embodiments.
[0010] Further details regarding the above embodiments and other aspects of this application, as well as their implementations, are provided in the accompanying drawings, detailed description, and claims. Attached Figure Description
[0011] Figure 1 is a flowchart illustrating a text recognition method provided in an embodiment of this application;
[0012] Figure 2 is a flowchart illustrating another text recognition method provided in an embodiment of this application;
[0013] Figure 3 is a flowchart illustrating another text recognition method provided in an embodiment of this application;
[0014] Figure 4 is a flowchart illustrating a method for determining a fusion matrix according to an embodiment of this application;
[0015] Figure 5 is a flowchart illustrating a method for determining the segmentation dimension provided in an embodiment of this application;
[0016] Figure 6 is a flowchart illustrating another text recognition method provided in an embodiment of this application;
[0017] Figure 7 is a flowchart illustrating a text recognition method provided in an embodiment of this application;
[0018] Figure 8 is a structural schematic diagram of a text recognition device provided in an embodiment of this application;
[0019] Figure 9 is a schematic diagram of the structure of a text recognition device provided in an embodiment of this application. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in detail below with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be arbitrarily combined with each other.
[0021] The steps illustrated in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases the steps shown or described may be performed in a different order than that presented here.
[0022] In one exemplary embodiment, Figure 1 is a flowchart illustrating a text recognition method provided in this application. This method is applicable to situations where the text type is identified based on the information contained in the text. This method can be executed by a text recognition device, which can be executed by software and / or hardware and integrated on an electronic device, such as a mobile phone, laptop, desktop computer, or smart tablet. This application does not impose any limitations on this.
[0023] As shown in Figure 1, the text recognition method provided in this application embodiment may include the following steps:
[0024] S101. Extract links from the text to be processed and determine the link extraction results.
[0025] In this embodiment, the text to be processed can be understood as text whose content needs to be semantically extracted and identified in order to classify its type. The text to be processed can be short text with a fixed length range, such as SMS, or long text with a length greater than a certain range, such as news, emails, and articles. This embodiment does not impose any restrictions on this.
[0026] In this embodiment, the link extraction result can be understood as the result obtained after filtering and extracting characters in the text to be processed that appear as links, used to indicate whether there are links in the text to be processed. In some examples, the link extraction result may also include the link legality determination result obtained by verifying the legality of the extracted links.
[0027] In an exemplary implementation, when it is necessary to determine and classify whether a piece of text contains abnormal links and whether it contains abnormal text expressions, the piece of text that needs to be processed can be taken as the text to be processed, and the content related to the link expression in the text to be processed can be extracted, and the extracted result can be determined as the link extraction result.
[0028] S102. Determine the fusion matrix based on the text to be processed, the link extraction results, and the segmentation dimension corresponding to the text to be processed, and identify the text type of the text to be processed based on the fusion matrix.
[0029] In this embodiment, the segmentation dimension corresponding to the text to be processed can be understood as a segmentation point used to maximize the reflection of the heterogeneity between texts of different length intervals in the text to be processed; it can also be understood as a segmentation point used to divide the segmented text to be processed into at least two text intervals containing the same or different number of segmented words based on statistical features. In some examples, the segmentation dimension can be obtained based on statistical feature analysis.
[0030] In this embodiment, the fusion matrix can be understood as an information matrix containing more comprehensive and accurate semantic information, obtained by fusing the semantic information extracted from the text to be processed after being segmented according to the segmentation dimension. The text type of the text to be processed can be specifically understood as reflecting whether there are any anomalies in the text to be processed, and the severity of the anomalies. In some examples, the text type can be legal, illegal, high-risk, or low-risk, etc., and this embodiment does not limit this.
[0031] In a specific example, when it is necessary to identify the text type of the text to be processed, the link extraction results can be fused with the text to be processed to extract and analyze the text semantic features. The identification of the corresponding text type is then completed based on whether there are anomalies in the extracted text semantic features, or the severity of such anomalies. During the text type identification process, to fully consider the impact of links contained in the text on the semantic feature extraction, as well as the impact of heterogeneity between texts of different length ranges on the semantic feature extraction, the link extraction results and segmentation dimensions can be applied to the semantic feature extraction and fusion process. This results in a fusion matrix that contains more comprehensive and accurate semantic information, allowing for a more accurate identification and determination of the text type based on the semantic features contained in the fusion matrix.
[0032] This application provides a text recognition method that involves extracting links from the text to be processed to determine the link extraction results; determining a fusion matrix based on the text to be processed, the link extraction results, and the segmentation dimensions corresponding to the text to be processed; and identifying the text type of the text to be processed based on the fusion matrix. By adopting the above technical solution, the impact of the link extraction results on the semantic information extraction of the text to be processed is fully considered in the extraction process. By using the segmentation dimensions corresponding to the text to be processed, the text to be processed is segmented during the semantic information extraction process for texts where anomalies are difficult to determine through links. This allows for the full utilization of the length features of each segmented part of the text to be processed when extracting information. Consequently, the fusion matrix obtained by processing the link extraction results and the segmentation matrix can retain more semantic information of the text to be processed, reducing information loss during the semantic information extraction process and improving the accuracy of the text type of the text to be processed determined based on the fusion matrix.
[0033] In one exemplary embodiment, Figure 2 is a flowchart illustrating another text recognition method provided by an embodiment of this application. This embodiment further optimizes the above-mentioned optional technical solutions. As shown in Figure 2, the text recognition method provided by this embodiment specifically includes the following steps:
[0034] S201. Extract links from the text to be processed and determine the link extraction results.
[0035] S202. Perform word segmentation and encoding on the text to be processed, and determine the encoding row vector.
[0036] In this embodiment, word segmentation can be understood as a basic step based on natural language processing, which is used to decompose a text string into smaller units, where each unit can be a word, phrase, symbol, or other meaningful character sequence. Encoding processing can be understood as a processing method that replaces each word segment with the corresponding position encoding in a pre-constructed word library table based on the position of different word segments in the table. The encoded row vector can be specifically understood as a vector obtained by sorting each encoding according to the position of the word segments in the text to be processed after all the word segments of the text to be processed are segmented and all the word segments are replaced with the corresponding encodings.
[0037] In a specific example, when it is necessary to identify the text type of the text to be processed, unnecessary characters and some meaningless common words in the text to be processed can be removed first through text preprocessing, and then the preprocessed text to be processed can be split into multiple independent word segments by using text word segmentation technology. For each word segment, determine the position of the word segment in the pre-constructed word library table and the corresponding ID encoding (Identification Code) at that position, and use this ID encoding as the encoding of the word segment. Then, the encoding can be replaced at that position according to the position of each word segment in the text to be processed, and an encoded row vector corresponding to the text to be processed can be obtained.
[0038] In one embodiment, performing word segmentation and encoding processing on the text to be processed and determining the encoded row vector may include:
[0039] Performing word segmentation processing on the text to be processed to determine a word segment set corresponding to the text to be processed;
[0040] Encoding the word segment set according to the word segment encoding mapping relationship between each word segment in the word segment set and the pre-constructed word library table to determine the encoded row vector.
[0041] In some examples, the text word segmentation technologies for performing word segmentation processing on the text to be processed include, but are not limited to, Jieba word segmentation, THULAC word segmentation, HanLP word segmentation, etc. In the embodiments of this application, Jieba word segmentation technology is used as an example to perform word segmentation on the text to be processed, and the accurate mode in Jieba word segmentation can be used for word segmentation.
[0042] In some examples, the accurate mode word segmentation in Jieba word segmentation is more suitable for text analysis compared with the full mode word segmentation used in fields such as search engines, reducing redundancy and solving the ambiguity problem, making the word segmentation result more accurate. For example, when segmenting the sentence "I came to Tsinghua University in Beijing", the word segmentation result in the accurate mode can be expressed as "I came to Beijing Tsinghua University", while the word segmentation result in the full mode can be expressed as "I came to Beijing Tsinghua Tsinghua University Hua Da University". Based on the above comparison, it can be clearly determined that the word segmentation result in the accurate mode is closer to the actual required word segmentation result.
[0043] In some examples, the pre-built lexicon can be constructed by fully segmenting a text dataset obtained from the internet, text libraries, and other data sources capable of acquiring text data. All segments are then sorted in descending order of frequency, and the ID of the most frequent segment is encoded as 1, with subsequent segments following the same pattern. Therefore, when encoding the text to be processed, if the text is divided into M segments, the final text will be encoded as an M-dimensional row vector. The pre-built lexicon in this application is not limited to the above construction method, and the embodiments of this application do not restrict the specific construction method.
[0044] In a specific example, the text to be processed can be segmented to obtain a segmentation set corresponding to the position of each segment in the text. Then, based on the segmentation encoding mapping relationship between each segment in the segmentation set and the pre-built lexicon table, that is, the mapping relationship between each segment in the segmentation set and the segmentation ID encoding in the pre-built lexicon table, the ID encoding corresponding to each segment can be determined. By replacing each segment in the segmentation set with its corresponding ID encoding, the encoding row vector corresponding to the segmentation set can be obtained.
[0045] S203. Based on the segmentation dimension corresponding to the text to be processed, the encoded row vector is segmented to determine at least two vector sub-intervals.
[0046] In this embodiment, a vector sub-interval can be understood as a portion of the encoded row vector after being segmented.
[0047] In a specific example, after determining the encoded line vector of the text to be processed and the corresponding segmentation dimension of the text to be processed, the segmentation point in the encoded line vector can be determined based on the segmentation dimension. Then, the encoded line vector is segmented based on the segmentation point. Each interval containing the encoding of different positions in the encoded line vector is determined as a vector sub-interval, resulting in at least two vector sub-intervals.
[0048] In some examples, assuming the encoded line vector of the text to be processed is M-dimensional and the corresponding segmentation dimension is 50-dimensional, the encoded line vector can be segmented into 50 dimensions based on this segmentation dimension, resulting in a 50-dimensional vector sub-interval and an (M-50)-dimensional vector sub-interval. Alternatively, the segmentation can be performed every 50 dimensions. If M is greater than 50 and less than 100, the result is a 50-dimensional vector sub-interval and an (M-50)-dimensional vector sub-interval; if M is greater than 100 and less than 150, the result is two 50-dimensional vector sub-intervals and an (M-100)-dimensional vector sub-interval; and so on. This clearly shows the resulting vector sub-intervals after segmenting the encoded line vector of the text to be processed for different values of M, meaning that when M is greater than 100, more than two vector sub-intervals can be obtained.
[0049] S204. Based on the link extraction results, supplement the dimensions of each vector sub-interval and determine the supplemented vector sub-intervals.
[0050] In this embodiment, dimension supplementation can be understood as a process that uses the presence or absence of links in the text to be processed as a feature based on the link extraction results, and adds this feature to each vector sub-interval, thereby increasing the dimension of each vector sub-interval. Supplemented vector sub-intervals can be understood as vector sub-intervals after completing the dimension supplementation related to link information.
[0051] In one embodiment, supplementing the dimensions of each vector sub-interval based on the link extraction results may include:
[0052] If the link extraction results indicate that a link exists, the first dimension value is added to each vector sub-interval; and / or if the link extraction results indicate that a link does not exist, the second dimension value is added to each vector sub-interval.
[0053] In this embodiment, the first dimension value can be understood as a feature dimension indicating the presence of links in the text to be processed. The second dimension value can be understood as a feature dimension indicating the absence of links in the text to be processed.
[0054] In a specific example, if the link extraction results indicate that a link exists, the first dimension value, used to indicate the presence of a link in the text to be processed, is added to each corresponding vector sub-interval, resulting in supplementary vector sub-intervals. Conversely, if the link extraction results indicate that a link does not exist, the second dimension value, used to indicate the absence of a link in the text to be processed, is added to each corresponding vector sub-interval, resulting in supplementary vector sub-intervals. This dimension addition method ensures that the resulting supplementary vector sub-intervals contain feature information representing the presence or absence of links, facilitating subsequent extraction and retention of semantic features.
[0055] In some examples, the first dimension value used to characterize the existence of a link can be set to 1, and the second dimension value used to characterize the non-existence of a link can be set to 0. This application does not impose any restrictions on this.
[0056] In some examples, taking the above text to be processed as a 50-dimensional vector sub-interval and a (M-50)-dimensional vector sub-interval as an example, after supplementing the dimensions of each vector sub-interval according to the link extraction results, the dimensions can be supplemented to 51-dimensional and (M-49)-dimensional vector sub-intervals respectively.
[0057] S205. Perform word embedding on each supplementary vector sub-interval to determine the information matrix corresponding to each supplementary vector sub-interval.
[0058] In this embodiment, word embedding can be understood as a technique that maps words to high-dimensional vector spaces. This allows words with similar meanings to be similar in the vector space; in other words, word embedding can map semantic information of words to high-dimensional vector spaces for computer processing. The information matrix can be understood as a matrix obtained after word embedding of each supplementary vector sub-interval, containing the semantic information of each word's text and in a numerical form that can be processed by a computer.
[0059] In a specific example, in order to represent the semantic information of the text to be processed in a computer-processable form, the ID encoding in each supplementary vector sub-interval corresponding to the text to be processed needs to be converted into word vectors. At this time, word embedding technology can be used to process each supplementary vector sub-interval to obtain the semantic information corresponding to each supplementary vector sub-interval, which contains the corresponding dimension of each supplementary vector sub-interval.
[0060] In some examples, taking the supplementary vector sub-intervals of the above-mentioned text to be processed, corresponding to 51-dimensional and (M-49)-dimensional sub-intervals, an embedding layer with parameters n*m can be used to process each supplementary vector sub-interval, obtaining information matrices Q1 and Q2 corresponding to the above two supplementary vector sub-intervals. Q1 is the information matrix of dimension (51*m) corresponding to the 51-dimensional supplementary vector sub-interval, containing the first 50 dimensions of semantic information of the text to be processed; Q2 is the information matrix of dimension [(M-49)*m] corresponding to the (M-49)-dimensional supplementary vector sub-interval, containing the last (M-50) dimensions of semantic information of the text to be processed.
[0061] S206. Determine the dimensional fusion matrix based on each information matrix, and multiply each information matrix by the dimensional fusion matrix to determine the fusion matrix.
[0062] In this embodiment, the dimension fusion matrix can be understood as a matrix used to integrate data from different modalities at the feature level.
[0063] In a specific example, in order to fuse the information in the information matrices corresponding to each vector sub-interval, a dimension fusion matrix can be generated based on the dimension of each information matrix. This matrix can be used to perform dimension fusion processing on information matrices of different dimensions. Then, each information matrix is multiplied by the dimension fusion matrix based on matrix multiplication, so that each information matrix is mapped to a dimension-compatible space to obtain the fusion matrix.
[0064] In some examples, taking the information matrices corresponding to the two supplementary vector sub-intervals mentioned above as Q1 and Q2, a hidden layer can be trained based on the dimensions of Q1 and Q2. After training, this hidden layer generates a dimension fusion matrix A, which can be used to map the input Q1 to a space compatible with the dimension of Q2. The dimension of the dimension fusion matrix A is [m*(M-49)]. Through matrix multiplication Q1×A×Q2, Q1 and Q2 can be fused in the same dimension to obtain an output matrix of dimension 51*m. This output matrix is the fusion matrix. This process ensures the full integration of the information from Q1 and Q2, so that the fusion matrix can contain all the semantic information in the text to be processed, providing a more comprehensive and accurate data basis for subsequently determining the text type of the text to be processed based on the fusion matrix.
[0065] In one example, the output of the hidden layer described above can be a weight matrix. The weights in this matrix are automatically learned from the data during training. By inputting the input matrices Q1 and Q2 into the neural network, the network learns a weight matrix, A, based on the relationship between them. The goal of the training process is to minimize the error between the fused output matrix and the desired target matrix. Therefore, through multiple iterations of training, the neural network can gradually optimize matrix A, enabling it to effectively fuse information across different data samples.
[0066] S207. Identify the text type of the text to be processed based on the fusion matrix.
[0067] In one embodiment, identifying the text type of the text to be processed based on the fusion matrix includes: identifying the text type of the text to be processed based on the text semantic features in the fusion matrix; wherein the text type includes abnormal text and / or non-abnormal text.
[0068] In a specific example, text semantic features can be extracted based on the fusion matrix. Since the correspondence between different text semantic features and text types can be pre-defined, after extracting the text semantic features from the fusion matrix, the text type of the text to be processed can be identified and determined based on this correspondence. In some examples, abnormal text and non-abnormal text can be identified as text types when determining the text type. Non-abnormal text can be understood as text that will not affect user information security, while abnormal text can be understood as text that will affect user information security. Abnormal text can be further subdivided into high-risk, medium-risk, low-risk, and illegal types according to actual needs; this application embodiment does not impose such limitations.
[0069] In some examples, the text type of the text to be processed can be identified through a pre-trained text recognition model. This pre-trained text recognition model can be a deep learning model incorporating a self-attention mechanism, and during training, it can also use Stratified 10-Fold cross-validation to divide the training and validation sets. Optionally, the text recognition model in this embodiment can be a deep learning model with classification capabilities that incorporates a self-attention mechanism, or it can be other types of deep learning models capable of achieving text recognition; this embodiment does not impose any limitations on this.
[0070] The SeqSelfAttention layer can be placed after the convolutional layers of a deep learning model. Since the self-attention mechanism is effective in handling long sequence data, especially suitable for natural language processing tasks, when applied to text recognition models, it allows the model to automatically focus on important information segments in the input text sequence while ignoring noise or irrelevant data. This mechanism dynamically adjusts the importance of these words in the context by assigning weights to each word in the input sequence. Because ordinary convolutional layers have limitations in capturing long-distance dependencies in sequence data, the SeqSelfAttention layer can capture dependencies between distant elements in the input sequence through the self-attention mechanism, thereby improving the model's ability to understand contextual information. Therefore, in this application, by placing the SeqSelfAttention layer after the convolutional layer and performing self-attention processing on the features output by the convolutional layer, the model can better extract and represent the features of the input sequence, improving the performance of downstream tasks. Furthermore, by adding the SeqSelfAttention layer after the convolutional layer, the model can reallocate attention based on local features, reducing information loss and retaining more useful information.
[0071] To ensure the stability and generalization ability of the trained text recognition model, Stratified 10-Fold cross-validation is employed during training to partition the dataset. By dividing the dataset used for training the text recognition model into 10 folds, training and validation are performed on each fold, reducing the risk of overfitting and providing a more reliable evaluation of model performance. Stratified sampling ensures that the data distribution in each fold is consistent with the class distribution in the original dataset. This means that the ratio of anomalous text types to non-anomalous text types remains constant in each fold, ensuring that the model learns fairly and effectively when processing different types of text. The model undergoes a complete training and validation cycle on each fold of the 10-Fold cross-validation, and the final model performance is the average of all fold results. This approach gives the model better robustness and generalization ability when facing different types of text.
[0072] In one exemplary embodiment, Figure 3 is a flowchart illustrating another text recognition method provided by an embodiment of this application. This embodiment further optimizes the above-mentioned optional technical solutions. As shown in Figure 3, the text recognition method provided by this embodiment specifically includes the following steps:
[0073] S301. Extract links from the text to be processed and determine the link extraction results.
[0074] S302. Perform word segmentation and encoding on the text to be processed, and determine the encoding row vector.
[0075] S303. Based on the segmentation dimension corresponding to the text to be processed, the encoded row vector is segmented to determine at least two vector sub-intervals.
[0076] It is understood that the processing methods in S301-S303 are the same as those in S201-S203 above, and will not be described in detail in this application embodiment.
[0077] S304. Determine the information entropy of each vector sub-interval based on the number of word segments within each vector sub-interval and the probability of each word segment appearing in the pre-built lexicon. Determine the weight factor corresponding to each vector sub-interval based on the information entropy of each vector sub-interval.
[0078] In this embodiment, information entropy can be understood as a value used to quantify the uncertainty or randomness of information.
[0079] In a specific example, since different vector sub-intervals correspond to different dimensions, it can be assumed that different vector sub-intervals contain different amounts of information. In order to save more information into the subsequently determined fusion matrix, the information entropy of each word can be determined based on the occurrence probability of the word segments contained in different vector sub-intervals in the pre-built lexicon. Then, the information entropy of each vector sub-interval can be determined based on the information entropy of the word segments contained in each vector sub-interval. The information entropy of each vector sub-interval is weighted and used as the weight factor corresponding to each vector sub-interval, so that each weight factor can be applied to the generation process of the fusion matrix.
[0080] In some examples, it is assumed that the i-th word in a vector subinterval is represented as , where n is the total number of word segments contained in the vector sub-interval. Let be the probability of the i-th word appearing in the pre-built lexicon. Then, the information entropy of this word can be expressed as: Therefore, the sum of the information entropy of all word segments within a vector sub-interval can be determined as the information entropy of that vector sub-interval. The information entropy of a vector sub-interval can be expressed by the following formula:
[0081]
[0082] It is understood that the construction method of the pre-built thesaurus table in this application embodiment is consistent with the construction method of the thesaurus table shown in S202 above, and this application embodiment will not describe it in detail.
[0083] In some examples, taking a 50-dimensional vector sub-interval and an (M-50)-dimensional vector sub-interval as examples, the information entropy of the 50-dimensional vector sub-interval can be expressed as the sum of the information entropies of each word segment, h1, and the information entropy of the (M-50)-dimensional vector sub-interval can be expressed as the sum of the information entropies of each word segment, h2. Therefore, the weight factor of the 50-dimensional vector sub-interval can be expressed as H1 = h1 / (h1 + h2), and the weight factor of the (M-50)-dimensional vector sub-interval can be expressed as H2 = h2 / (h1 + h2). It is understood that the embodiments of this application only illustrate one method for determining the weight factor and do not limit the specific method for determining the weight factor.
[0084] S305. Based on the link extraction results, supplement the dimensions of each vector sub-interval and determine the supplemented vector sub-intervals.
[0085] S306. Perform word embedding on each supplementary vector sub-interval to determine the information matrix corresponding to each supplementary vector sub-interval.
[0086] It is understood that the processing methods in S305-S306 are the same as those in S204-S205 above, and will not be described in detail in this application embodiment.
[0087] S307. Determine the dimensional fusion matrix based on each information matrix, and multiply each information matrix, each weight factor, and the dimensional fusion matrix together to determine the fusion matrix.
[0088] In some examples, taking the information matrices corresponding to the two supplementary vector sub-intervals mentioned above as Q1 and Q2, the weight factor corresponding to Q1 as H1, and the weight factor corresponding to Q2 as H2, a hidden layer can be trained based on the dimensions of Q1 and Q2. After training, the hidden layer generates a dimension fusion matrix A. Then, the information matrices, weight factors, and dimension fusion matrix are multiplied together to obtain the dimension fusion matrix H1×Q1×A×Q2×H2.
[0089] In some examples, Figure 4 is a flowchart of a method for determining a fusion matrix provided by an embodiment of this application. As shown in Figure 4, taking the encoded row vector of the text to be processed as an M-dimensional encoded row vector and the segmentation dimension of the text to be processed as 50-dimensional as an example: First, it is determined whether M is greater than 50. If M is not greater than 50, then the encoded row vector needs to be supplemented with dimension 0 and the dimension is expanded. After the dimension is expanded, the Embedding layer with parameter n*m is used to perform word embedding processing on the dimension-expanded encoded row vector, and the information matrix obtained after processing is determined as the fusion matrix. If M is greater than 50, the encoded row vector can be segmented into 50-dimensional nodes. Then, based on whether links exist in the text to be processed, the dimensions of each segmented vector sub-interval are expanded. Dimension 1 is added for links, and dimension 0 is added for non-linked links. Figure 4 shows an example with only one segmentation point. After dimensional expansion, the M-dimensional encoded row vector can be divided into a 51-dimensional vector space and an (M-49)-dimensional vector space. Then, word embedding is performed on these two vector spaces using an embedding layer with parameters n*m, resulting in information matrices Q1 and Q2 with dimensions (51*m). An information matrix Q2 with dimension [(M-49)*m] is generated based on the dimensions of Q1 and Q2. A dimensional fusion matrix A with dimension [m*(M-49)] is generated by training based on the dimensions of Q1 and Q2. Q1 and Q2 can be fused in the same dimension by matrix multiplication Q1×A×Q2 to obtain an output matrix with dimension 51*m, thus completing the determination of the fusion matrix. Optionally, when fusing Q1 and Q2, weight factors H1 and H2 corresponding to Q1 and Q2 can be introduced respectively. A 51*m dimensional fusion matrix H1×Q1×A×Q2×H2 can be obtained by matrix multiplication. Both of the above methods can realize the determination of the fusion matrix. This application embodiment does not limit this.
[0090] S308. Identify the text type of the text to be processed based on the fusion matrix.
[0091] It is understood that the processing methods in S308 and S207 are the same, and this application embodiment will not describe them in detail.
[0092] In this embodiment, word vector intervals with high entropy values receive greater weight during the multiplication of information matrices Q1 and Q2, while word vector intervals with low entropy values receive less weight. This process allows the model to focus more on the key information-dense parts of the input text, thereby improving the accuracy of semantic representation. The encoded row vectors are segmented according to the segmentation dimension corresponding to the text to be processed, allowing the information entropy of each word segment in the resulting vector sub-intervals to be calculated independently. This effectively assesses which parts of the text to be processed are richer in information. The determination of weight factors based on information entropy, and the application of these weight factors to the fusion matrix, can adapt to the actual characteristics of the data corresponding to the text to be processed. If the information distribution in the text to be processed is uneven, dynamic weight adjustment can be achieved through weight factors, ensuring that the parts with high information content receive more attention. By combining the calculation of information entropy and information matrix, the fusion matrix in this embodiment can better capture and understand the core semantic information in the input text. Information entropy assigns different weight factors to different vector sub-intervals, so that the part containing more semantic information in the final fusion matrix can be more effectively expressed and utilized in the process of determining the type of text to be processed, thereby improving the recognition accuracy of abnormal text types.
[0093] In one embodiment, after determining the link extraction result, the method may further include: if the link extraction result indicates that the link is valid or the link does not exist, determining a fusion matrix based on the text to be processed, the link extraction result, and the segmentation dimension corresponding to the text to be processed, and identifying the text type of the text to be processed based on the fusion matrix.
[0094] In a specific example, if the link extraction result indicates that the link is valid or does not exist, it can be assumed that the semantic features of the entire text to be processed can be directly extracted and analyzed. Then, the text type identification of the text to be processed can be completed based on the fusion matrix obtained from the extraction and analysis. That is, if the link extraction result indicates that the link is valid or does not exist, it can be achieved through steps such as S202-S207 or S302-S308 described above. It is only necessary to correspond the case where the link extraction result indicates that the link is valid with the case where the link exists when performing dimension supplementation as in S204 or S305. If the link extraction result indicates that the link is valid, the first dimension value is used to supplement the dimensions of each vector sub-interval corresponding to the text to be processed. The remaining execution methods are consistent with the above steps, and this embodiment will not be described in detail.
[0095] In one embodiment, after determining the link extraction result, the method may further include: if the link extraction result indicates that the link is invalid, determining the text type of the text to be processed as abnormal text.
[0096] In a specific example, after determining the link extraction results, it can be used to determine whether links exist in the text to be processed. If links exist, their legitimacy can be determined. If the link extraction results indicate that a link is invalid, it can be directly assumed that the link will affect the information security of the text to be processed, and the text type of the text to be processed can be directly determined as abnormal text.
[0097] In one embodiment, if the text type of the text to be processed is determined to be abnormal text, the method may further include: intercepting the text to be processed, and / or prompting the user.
[0098] In a specific example, when the text to be processed is identified as abnormal, to ensure that the user's information security is not affected by the abnormal text, the text to be processed can be intercepted and prompted. The interception process may include directly blocking or masking the text, while the prompting process may include alerting the user to the abnormality of the text through pop-ups, voice prompts, or other perceptible means. These processing methods can be performed individually or simultaneously, and this application embodiment does not impose any limitations on this.
[0099] In one exemplary embodiment, Figure 5 is a flowchart illustrating a method for determining the segmentation dimension provided by an embodiment of this application. This embodiment further optimizes the above-mentioned optional technical solutions. As shown in Figure 5, the method for determining the segmentation dimension provided by this embodiment specifically includes the following steps:
[0100] S401. Determine the number of word segments for each text in the text set to be processed, based on the text set to which the text to be processed belongs.
[0101] In this embodiment, since text recognition is often achieved using batch processing, meaning that multiple texts to be processed are input simultaneously in a single input, the set of multiple texts to be processed in a batch can be defined as the set of texts to be processed. The number of word segments can be understood as the total number of words obtained after a text to be processed has been fully segmented.
[0102] S402. Determine at least one candidate segmentation point based on the number of each word segmentation.
[0103] In this embodiment, candidate segmentation points can be understood as nodes that can be used to divide the text to be processed into two or more segments with different numbers of segments.
[0104] In some examples, the entire set of texts to be processed can be used as a sample, or a portion of the texts to be processed can be selected from the set. Statistical analysis is performed on each sample text, using the number of word segments (x) as the x-axis and the number of sample texts (n) corresponding to the number of word segments (x) as the y-axis. The number of texts within each interval is counted, and a statistical distribution chart is plotted to determine the distribution of the total number of words obtained from segmenting each text in the set. The intervals can be determined based on the maximum and minimum number of word segments among the multiple word segments corresponding to the set, as well as pre-set segmentation requirements based on actual conditions. For example, assuming the maximum number of word segments obtained after segmentation is 200 and the minimum is 50, texts with fewer than 50 word segments but more than 200 word segments are considered non-existent. In the statistical analysis, only the distribution of sample texts with word segments within the interval [50, 200] is considered. Assuming the segmentation requirement is 5 segments, the quantity range can be set to 30. That is, when performing the current statistical analysis, the number of words in the five intervals [50,80), [80,110), [110,140), [140,170), and [170,200] needs to be counted. The node of each segment interval can be determined as a candidate segmentation point, that is, 50, 80, 110, 140, 170, and 200 can be determined as candidate segmentation points.
[0105] S403. Determine the divergence value of the candidate split point based on the first cumulative distribution function value before the candidate split point and the second cumulative distribution function value after the candidate split point.
[0106] In this embodiment, the cumulative distribution function (CDF) can be understood as a function used to describe the probability that a random variable takes a value less than or equal to a certain specific value.
[0107] In a specific example, the ratio of the number of texts in the text set to be processed with a word count less than or equal to that of the candidate segmentation point to the total number of texts in the text set to be processed is determined as the first cumulative distribution function value before the candidate segmentation point; the ratio of the number of texts in the text set to be processed with a word count greater than that of the candidate segmentation point to the total number of texts in the text set to be processed is determined as the second cumulative distribution function value after the candidate segmentation point. Then, the difference distribution between the first and second cumulative distribution function values can be determined using KL divergence, thus obtaining the divergence value of the candidate segmentation point.
[0108] In some examples, it is assumed that the value of the first cumulative distribution function is expressed as The value of the second cumulative distribution function is expressed as The divergence value of the candidate split point can be expressed as:
[0109]
[0110] S404. The candidate split point with the largest divergence value is determined as the split dimension corresponding to each text in the text set to be processed.
[0111] In a specific example, the candidate segmentation point with the largest divergence value represents the point where the distribution of word segmentation in each text to be processed changes most significantly. This also means that the statistical characteristics of word segmentation change significantly near this point. To better extract text information from different regions in the text to be processed, the candidate segmentation point with the largest divergence value can be determined as the segmentation dimension corresponding to each text to be processed in the set of texts to be processed.
[0112] In this embodiment, since the candidate segmentation point with the largest divergence value can reveal the heterogeneity of different word segmentation ranges to the greatest extent, the feature extraction of these ranges is more meaningful. By selecting the candidate segmentation point with the largest divergence value as the segmentation dimension of each text to be processed, it is helpful to understand and utilize the information contained in the text to be processed more accurately. This enables the subsequent determination of the text type of the text to be processed based on the fusion matrix to make full use of the word segmentation feature in the text to be processed, thereby improving the accuracy of text type determination.
[0113] In one exemplary embodiment, Figure 6 is a flowchart illustrating another text recognition method provided by an embodiment of this application. This embodiment further optimizes the above-mentioned optional technical solutions. As shown in Figure 6, the text recognition method provided by this embodiment specifically includes the following steps:
[0114] S501. Extract links from the text to be processed using pre-configured regular expressions and determine the link extraction results.
[0115] In this embodiment, the pre-configured regular expression can be understood as a rule string that is pre-configured according to the actual situation and is used to extract links that may exist in the input text.
[0116] In a specific example, the text to be processed is filtered using a pre-configured regular expression. Based on the filtering results, it is determined whether there are links in the text to be processed, and the link extraction results are obtained.
[0117] In one embodiment, the pre-configured regular expression includes at least: legal character matching information, string matching information, domain name format matching information, and domain name path matching information.
[0118] In this embodiment, the legal character matching information can be understood as the matching information used to match legal characters such as hostnames, paths, and query parameters of Uniform Resource Locators (URLs).
[0119] In some examples, valid character matching information can be represented as:
[0120] (?:[a-zA-Z0-9]|[$#-_@.&+]|[!*\\(\\),]|(?:%[0-9a-fA-F][0-9a-fA-F]))+
[0121] The following sections explain each part of the legal character matching information from the following perspectives:
[0122] 1) (?:...) represents a non-capturing group, used for grouping but not creating sub-matches;
[0123] 2) [a-zA-Z0-9] matches letters and numbers;
[0124] 3) |[$#-_@.&+] matches common characters in URLs, such as the dollar sign $, hash #, underscore _, minus sign -, @, dot ., &, and plus sign +;
[0125] 4) |[!*\\(\\),] matches some special characters, such as the asterisk *, exclamation mark !, backslash \\, parentheses (), and comma,;
[0126] 5) |(?:%[0-9a-fA-F][0-9a-fA-F]) matches URL encoding (% followed by two hexadecimal digits, such as %20 representing a space);
[0127] 6) + indicates that the preceding part can be repeated once or multiple times.
[0128] In this embodiment, string matching information can be understood as matching information used to match the part after "www." that forms the complete URL.
[0129] The representation of string matching information is the same as that of legal character matching information, and will not be described in detail here.
[0130] In this embodiment, the domain name format matching information can be understood as matching information used to match domain name format URLs without protocols.
[0131] In some examples, domain name format matching information can be represented as:
[0132] |(?:[a-zA-Z0-9-]+\.)+[a-zA-Z]{2,}
[0133] The following sections explain each part of the domain name format matching information from the following perspectives:
[0134] 1) `(?:[a-zA-Z0-9-]+\.)+` matches one or more subdomains consisting of letters, numbers, or hyphens (-), followed by a period (.). This part can be repeated multiple times to match domains like sub.domain.com;
[0135] 2) [a-zA-Z]{2,} matches a top-level domain (TLD) consisting of at least two letters, such as ".com" or ".cn".
[0136] In this embodiment, domain name path matching information can be understood as matching information used to match the path portion following the domain name.
[0137] In some examples, domain path matching information can be represented as:
[0138] (?: / [^\s\,\,\.\>\<\}\{\)\(\]\[\]\[\(\)\,\(\u4e00-\u9fa5]*)?
[0139] The following sections explain each part of the domain path matching information from the following perspectives:
[0140] 1) (?:...) indicates a non-capture group;
[0141] 2) / Matches the beginning of a path, i.e., the first forward slash / in a URL;
[0142] 3) [^\s\,\,\。 \>\<\}\{\)\(\]\[\】\
\(\)\、\(\u4e00-\u9fa5]* matches valid characters in the path; where ^\s represents non-whitespace characters (excluding spaces, newlines, etc.); \,\,\。\>\<\}\{\)\(\]\[\
[0143] 4) * indicates that it can match zero or more times;
[0144] 5) The question mark (?) indicates that the path part is optional.
[0145] In one embodiment, the pre-configured regular expression may further include: protocol matching information for matching the URL protocol portion, and string matching information for matching URLs that begin with "www."
[0146] The protocol matching information used to match the URL protocol portion can be represented as: http[s]?: / / . Here, http matches the string "http"; [s]? indicates that "s" is optional, matching either "http" or "https"; and : / / matches the string ": / / ", signifying the end of the protocol portion.
[0147] The string matching information used to match URLs that begin with "www." can be represented as: |www\.. Here, www\. matches the string "www.", where "." is an escaped period, indicating an exact match of a single dot.
[0148] In some examples, the pre-configured regular expression can be fully represented as:
[0149] (http[s]?: / / (?:[a-zA-Z0-9]|[$#-_@.&+]|[!*\\(\\),]|(?:%[0-9a-fA-F][0-9a-fA-F]))+|www\.(?:[a-zA-Z0-9]|[$#-_@.&+]|[!*\\(\\),] |(?:%[0-9a-fA-F][0-9a-fA-F]))+|(?:[a-zA-Z0-9-]+\.)+[a-zA-Z]{2 ,}(?: / [^\s\,\,\.\>\<\}\{\)\(\]\[\]\[\(\)\,\(\u4e00-\u9fa5]*)?)
[0150] S502. Determine the fusion matrix based on the text to be processed, the link extraction results, and the segmentation dimension corresponding to the text to be processed, and identify the text type of the text to be processed based on the fusion matrix.
[0151] In this embodiment, since the pre-configured regular expression takes into account various URL formats, it can effectively match common URL formats, including complete URLs starting with "http" or "https", URLs starting with "www", and domain names without protocols. It can also include special characters, encodings, and path parts that may be contained in the URL, which has strong adaptability and flexibility, making the link extraction results more accurate.
[0152] In some examples, Figure 7 is a flowchart illustrating a text recognition method provided in an embodiment of this application. As shown in Figure 7, the method may specifically include the following steps:
[0153] S601. Obtain the text to be processed.
[0154] S602. Extract links using regular expressions.
[0155] S603. Determine whether there is a link in the text to be processed based on the extraction results. If yes, proceed to S604; otherwise, proceed to S606.
[0156] S604. Call the interface to determine if the connection is valid. If yes, execute S606; otherwise, execute S605.
[0157] S605. Determine the text type of the text to be processed as abnormal text.
[0158] S606, precise pattern segmentation based on Jieba.
[0159] S607. Convert the word segmentation results into an encoding vector.
[0160] S608. Perform vector space partitioning on the encoded vector.
[0161] S609. Perform dimensional fusion on the information matrix converted from the vector sub-intervals obtained after segmentation to obtain the fusion matrix.
[0162] S610. Input the fusion matrix into the text recognition model with a self-attention layer.
[0163] In one exemplary embodiment, FIG8 is a schematic diagram of the structure of a text recognition device provided in an embodiment of the present application. As shown in FIG8, the text recognition device may include a link extraction module 710 and a text type recognition module 720.
[0164] The link extraction module 710 is used to extract links from the text to be processed and determine the link extraction results; the text type recognition module 720 is used to determine the fusion matrix based on the text to be processed, the link extraction results, and the segmentation dimension corresponding to the text to be processed, and to recognize the text type of the text to be processed based on the fusion matrix.
[0165] The text recognition device of this application fully considers the impact of link extraction results on the extraction of semantic information of the text to be processed during the extraction process. By segmenting the text to be processed according to the segmentation dimension corresponding to the text to be processed, the text to be processed is segmented during the semantic information extraction process for text to be processed that is difficult to determine as abnormal through links. This allows the length features of each segmented part of the text to be processed to be fully utilized when extracting information from each part of the text to be processed. As a result, the fusion matrix obtained by processing the link extraction results and the segmentation matrix can retain more text semantic information of the text to be processed, reduce information loss during the text semantic information extraction process, and improve the accuracy of the text type of the text to be processed determined by the fusion matrix.
[0166] In one embodiment, the text type recognition module 720 can be used for:
[0167] The text to be processed is segmented and encoded to determine the encoded row vector;
[0168] Based on the segmentation dimension corresponding to the text to be processed, the encoded row vector is segmented to determine at least two vector sub-intervals;
[0169] Based on the link extraction results, the dimensions of each vector sub-interval are supplemented to determine the supplemented vector sub-intervals;
[0170] Word embedding is performed on each supplementary vector sub-interval to determine the information matrix corresponding to each supplementary vector sub-interval;
[0171] The dimensional fusion matrix is determined based on each information matrix, and the fusion matrix is obtained by multiplying each information matrix with the dimensional fusion matrix.
[0172] The text type of the text to be processed is identified based on the fusion matrix.
[0173] In one embodiment, the text type recognition module 720 can be used for:
[0174] The text to be processed is segmented and encoded to determine the encoded row vector;
[0175] Based on the segmentation dimension corresponding to the text to be processed, the encoded row vector is segmented to determine at least two vector sub-intervals;
[0176] The information entropy of each vector sub-interval is determined based on the number of words in each vector sub-interval and the probability of each word appearing in the pre-built lexicon. The weight factor corresponding to each vector sub-interval is then determined based on the information entropy of each vector sub-interval.
[0177] Based on the link extraction results, the dimensions of each vector sub-interval are supplemented to determine the supplemented vector sub-intervals;
[0178] Word embedding is performed on each supplementary vector sub-interval to determine the information matrix corresponding to each supplementary vector sub-interval;
[0179] The dimension fusion matrix is determined by multiplying each information matrix, each weight factor and the dimension fusion matrix together to determine the fusion matrix.
[0180] The text type of the text to be processed is identified based on the fusion matrix.
[0181] In one embodiment, the method for determining the segmentation dimension includes:
[0182] Based on the text set to which the text to be processed belongs, determine the number of words segmented for each text in the text set to be processed.
[0183] Determine at least one candidate segmentation point based on the number of each word segmentation;
[0184] The divergence value of the candidate split point is determined based on the first cumulative distribution function value before the candidate split point and the second cumulative distribution function value after the candidate split point;
[0185] The candidate split point with the largest divergence value is determined as the split dimension corresponding to each text in the text set to be processed.
[0186] In one embodiment, the text to be processed is segmented and encoded to determine the encoded line vector, including:
[0187] The text to be processed is segmented into words to determine the set of words corresponding to the text to be processed.
[0188] Based on the word segmentation encoding mapping relationship between each word in the word segmentation set and the pre-built lexicon table, the word segmentation set is encoded to determine the encoding row vector.
[0189] In one embodiment, the dimension of each vector sub-interval is supplemented based on the link extraction results, including:
[0190] If the link extraction results indicate that a link exists, the first dimension value is added to each vector sub-interval respectively;
[0191] and / or
[0192] If the link extraction result indicates that the link does not exist, a second dimension value is added to each vector sub-interval.
[0193] In one embodiment, link extraction is performed on the text to be processed, and the link extraction result is determined, including:
[0194] Link extraction is performed on the text to be processed using pre-configured regular expressions to determine the link extraction results.
[0195] In one embodiment, the link extraction module 710 can be used to:
[0196] Link extraction is performed on the text to be processed using pre-configured regular expressions to determine the link extraction results.
[0197] In one embodiment, the pre-configured regular expression includes at least: legal character matching information, string matching information, domain name format matching information, and domain name path matching information.
[0198] In one embodiment, identifying the text type of the text to be processed based on the fusion matrix includes:
[0199] The text type of the text to be processed is identified based on the text semantic features in the fusion matrix;
[0200] The text types include abnormal text and / or text without abnormalities.
[0201] In one embodiment, after determining the link extraction results, the method may further include:
[0202] If the link extraction results indicate that the link is invalid, the text type of the text to be processed will be determined as abnormal text.
[0203] In one embodiment, if the text type of the text to be processed is determined to be abnormal text, the method may further include:
[0204] Intercept the text to be processed and / or prompt the user.
[0205] The text recognition device provided in this application can execute the text recognition method provided in any embodiment of this application, and has the corresponding module and beneficial effects of the method execution.
[0206] This application embodiment also provides a text recognition device. Figure 9 is a structural schematic diagram of a text recognition device provided in this application embodiment. As shown in Figure 9, the text recognition device provided in this application embodiment includes a memory 820, a processor 810, and a computer program stored in the memory and executable on the processor. When the processor 810 executes the program, it implements the above-mentioned text recognition method.
[0207] The text recognition device may also include a memory 820; the processor 810 in the text recognition device may be one or more, with one processor 810 as an example in FIG9; the memory 820 is used to store one or more programs; the one or more programs are executed by the one or more processors 810, so that the one or more processors 810 implement the text recognition method as described in the embodiments of this application.
[0208] The text recognition device may also include: a communication device 830, an input device 840, and an output device 850.
[0209] The processor 810, memory 820, communication device 830, input device 840 and output device 850 in the text recognition device can be connected by a bus or other means. Figure 9 shows an example of connection by bus.
[0210] Input device 840 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the text recognition device. Output device 850 may include display devices such as a display screen.
[0211] The communication device 830 may include a receiver and a transmitter. The communication device 830 is configured to perform information transmission and reception communication under the control of the processor 810.
[0212] The memory 820, as a computer-readable storage medium, can be configured to store software programs, computer-executable programs, and modules, such as program instructions / modules corresponding to the text recognition method described in the embodiments of this application (e.g., link extraction module 710 and text type recognition module 720). The memory 820 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on the use of the text recognition device, etc. Furthermore, the memory 820 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some instances, the memory 820 may further include memory remotely located relative to the processor 810, and these remote memories can be connected to the text recognition device via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0213] This application also provides a storage medium storing a computer program, which, when executed by a processor, implements any of the text recognition methods described in this application.
[0214] Optionally, the text recognition method includes: extracting links from the text to be processed and determining the link extraction results; if the links are determined to be valid or non-existent based on the link extraction results, determining a fusion matrix based on the text to be processed, the link extraction results, and the segmentation dimension corresponding to the text to be processed, and determining the text type of the text to be processed based on the fusion matrix.
[0215] The computer storage medium in this application embodiment can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. The computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0216] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit programs for use by or in connection with an instruction execution system, apparatus, or device.
[0217] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, radio frequency (RF), etc., or any suitable combination thereof.
[0218] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and also conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0219] Optionally, embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the text recognition method as provided in any embodiment of this application.
[0220] The above description is merely an exemplary embodiment of this application and is not intended to limit the scope of protection of this application.
[0221] Those skilled in the art will understand that the term user terminal encompasses any suitable type of wireless user equipment, such as mobile phones, portable data processing devices, portable web browsers, or vehicle-mounted mobile stations.
[0222] Generally, the various embodiments of this application can be implemented in hardware or dedicated circuitry, software, logic, or any combination thereof. For example, some aspects can be implemented in hardware, while others can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device, although this application is not limited thereto.
[0223] Embodiments of this application can be implemented by executing computer program instructions through the data processor of a mobile device, for example, in a processor entity, or through hardware, or through a combination of software and hardware. The computer program instructions can be assembly instructions, Instruction Set Architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages.
[0224] Any block diagram of logical flow in the accompanying drawings of this application may represent program steps, or may represent interconnected logic circuits, modules, and functions, or may represent a combination of program steps and logic circuits, modules, and functions. The computer program may be stored on memory. The memory may be of any type suitable to the local technical environment and may be implemented using any suitable data storage technology, such as, but not limited to, read-only memory (ROM), random access memory (RAM), optical storage devices and systems (Digital Video Disc (DVD) or Compact Disk (CD), etc.). Computer-readable media may include non-transitory storage media. The data processor may be of any type suitable to the local technical environment, such as, but not limited to, general-purpose computers, special-purpose computers, microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and processors based on multi-core processor architectures.
[0225] A detailed description of exemplary embodiments of this application has been provided above through exemplary and non-limiting examples. However, various modifications and adjustments to the above embodiments will be apparent to those skilled in the art when considered in conjunction with the accompanying drawings and claims, without departing from the scope of this application. Therefore, the proper scope of this application will be determined by the claims.
Claims
1. A text recognition method, comprising: Extract links from the text to be processed and determine the link extraction results; A fusion matrix is determined based on the text to be processed, the link extraction results, and the segmentation dimension corresponding to the text to be processed, and the text type of the text to be processed is identified based on the fusion matrix.
2. The text recognition method according to claim 1, wherein, The step of determining the fusion matrix based on the text to be processed, the link extraction results, and the segmentation dimension corresponding to the text to be processed includes: The text to be processed is segmented and encoded to determine the encoded row vector; The encoded row vector is segmented according to the segmentation dimension corresponding to the text to be processed, and at least two vector sub-intervals are determined. Based on the link extraction results, the dimensions of each vector sub-interval are supplemented to determine the supplemented vector sub-intervals; Word embedding is performed on each of the supplementary vector sub-intervals to determine the information matrix corresponding to each of the supplementary vector sub-intervals; The dimensional fusion matrix is determined based on each of the information matrices, and the fusion matrix is obtained by multiplying each of the information matrices with the dimensional fusion matrix.
3. The text recognition method according to claim 1, wherein, The step of determining the fusion matrix based on the text to be processed, the link extraction results, and the segmentation dimension corresponding to the text to be processed includes: The text to be processed is segmented and encoded to determine the encoded row vector; The encoded row vector is segmented according to the segmentation dimension corresponding to the text to be processed, and at least two vector sub-intervals are determined. The information entropy of each vector sub-interval is determined based on the number of words in each vector sub-interval and the probability of each word appearing in the pre-built lexicon. The weight factor corresponding to each vector sub-interval is determined based on the information entropy of each vector sub-interval. Based on the link extraction results, the dimensions of each vector sub-interval are supplemented to determine the supplemented vector sub-intervals; Word embedding is performed on each of the supplementary vector sub-intervals to determine the information matrix corresponding to each of the supplementary vector sub-intervals; The dimensional fusion matrix is determined based on each of the information matrices, and the fusion matrix is obtained by multiplying each of the information matrices, each of the weight factors, and the dimensional fusion matrix.
4. The text recognition method according to claim 1, wherein, The method for determining the segmentation dimension includes: Based on the set of texts to be processed to which the text to be processed belongs, determine the number of word segments for each text in the set of texts to be processed; At least one candidate segmentation point is determined based on the number of each segmentation term; The divergence value of the candidate split point is determined based on the first cumulative distribution function value before the candidate split point and the second cumulative distribution function value after the candidate split point; The candidate segmentation point with the largest divergence value is determined as the segmentation dimension corresponding to each text in the set of texts to be processed.
5. The text recognition method according to claim 2 or 3, wherein, The step of performing word segmentation and encoding on the text to be processed, and determining the encoded line vector, includes: The text to be processed is segmented into words to determine the segmentation set corresponding to the text to be processed; Based on the word segmentation encoding mapping relationship between each word in the word segmentation set and the pre-built lexicon table, the word segmentation set is encoded to determine the encoding row vector.
6. The text recognition method according to claim 2 or 3, wherein, The step of supplementing the dimensions of each of the vector sub-intervals based on the link extraction results includes: If the link extraction result indicates that a link exists, a first dimension value is added to each of the vector sub-intervals; and / or If the link extraction result indicates that the link does not exist, a second dimension value is added to each of the vector sub-intervals.
7. The text recognition method according to claim 1, wherein, The link extraction process for the text to be processed, and the determination of the link extraction results, includes: Link extraction is performed on the text to be processed using pre-configured regular expressions to determine the link extraction results.
8. The text recognition method according to claim 7, wherein, The pre-configured regular expressions include: legal character matching information, string matching information, domain name format matching information, and domain name path matching information.
9. The text recognition method according to any one of claims 1-3, wherein, The step of identifying the text type of the text to be processed based on the fusion matrix includes: The text type of the text to be processed is identified based on the text semantic features in the fusion matrix; The text types include abnormal text and / or text without abnormalities.
10. The text recognition method according to claim 9, further comprising, after determining the link extraction result: If the link extraction result indicates that the link is invalid, the text type of the text to be processed is determined to be abnormal text.
11. The text recognition method according to claim 10, further comprising at least one of the following when the text type of the text to be processed is determined to be abnormal text: Intercept the text to be processed, and / or prompt the user.
12. An electronic device, comprising: The program includes a memory, a processor, a program stored in the memory and executable on the processor, and a data bus for establishing communication between the processor and the memory, wherein the program, when executed by the processor, implements the text recognition method as described in any one of claims 1-11.
13. A computer program product comprising a computer program that, when executed by a processor, implements the text recognition method according to any one of claims 1-11.
Citation Information
Patent Citations
Long text classification method and device
CN112100389A
Risk prediction method, device and equipment and readable storage medium
CN114330966A
Feature extraction-based phishing mail detection method and system
CN114465780A
Text processing method and related equipment
CN115270788A
Detection method of phishing mail and training method of phishing mail detection model
CN115987658A