Text recognition method and device, nonvolatile storage medium and electronic equipment
By calculating the weight coefficient using the position of word participle results and the frequency of feature words in the text recognition method, and combining the logistic regression model for classification processing, the problem of inability to effectively utilize keyword position information in the prior art is solved, and the effect of improving the recognition accuracy is achieved.
Patent Information
- Application Number
- CN202510330485.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-01
AI Technical Summary
Existing text recognition methods cannot effectively utilize keyword position information, resulting in limited recognition accuracy.
By obtaining the training sample set and the text to be processed, word segmentation is performed, and the weight coefficient is determined based on the position of the word segmentation result and the frequency of occurrence of preset feature words, and classification is performed in combination with the logistic regression model to improve the recognition accuracy.
Effectively utilizing keyword position information improves the accuracy of text recognition and solves the problem of limited recognition accuracy.
Smart Images

Figure CN120234418A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of machine learning. Specifically, it relates to a text recognition method, an apparatus, a non-volatile storage medium, and an electronic device. Background Art
[0003] Related comment anomaly recognition methods often rely on the TF-IDF algorithm for text vectorization, and then use machine learning models such as Naive Bayes, linear regression, and KNN for classification. However, these methods have certain limitations. Although the TF-IDF algorithm is simple and easy to implement, when dealing with natural language texts, it ignores the position information of keywords in the text, which is particularly important in comment analysis because the first and last sentences of a comment usually can more accurately reflect the core view of the comment and should be given higher weights.
[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] The present application provides a text recognition method, an apparatus, a non-volatile storage medium, and an electronic device to at least solve the technical problem of limited recognition accuracy caused by the inability to effectively utilize keyword position information in related text recognition.
[0006] According to one aspect of the present application, a text recognition method is provided, including: obtaining a training sample set, where the training sample set includes: historical comment texts carrying preset labels, and the historical comment texts are from a target data source; obtaining a text to be processed in the target data source, performing word segmentation processing on the text to be processed to obtain a plurality of word segmentation results; determining a weight coefficient for each word segmentation result according to the text position of each word segmentation result in the text to be processed and the occurrence frequency of a preset feature word in each word segmentation result in a first text, where the first text is all the text included in the target label corresponding to the preset feature word; for each training sample, determining a relevance index between each word segmentation result and the training sample, and performing weighted summation on the relevance index according to the weight coefficient to obtain a correlation score between the text to be processed and each training sample; using a logistic regression model to perform classification processing on the correlation score to obtain probability values of the text to be processed belonging to different preset labels, and determining the label corresponding to the maximum probability value as the label of the text to be processed.
[0007] Optionally, the weight coefficient is determined by the following formula: where posw is the weight coefficient; W is the first weight coefficient determined according to the text position of each word segmentation result in the text to be processed; P(m q ) is the occurrence frequency of the preset feature word in each word segmentation result in the first text; P(m q)′ is the occurrence frequency of the preset feature word in each word segmentation result in the second text, where the second text is all the text included in the other preset labels except the target label corresponding to the preset feature word among the multiple preset labels in the training sample set.
[0008] Optionally, the first weight coefficient is determined by the following method: when the word segmentation result is in the first sentence of the first paragraph or the last sentence of the last paragraph in the text to be processed, assign the first value to the first weight coefficient; when the word segmentation result is in the first sentence of a non-first paragraph in the text to be processed, assign the second value to the first weight coefficient, where the second value is less than the first value; when the word segmentation result is not in the first sentence of the first paragraph or the last sentence of the last paragraph or the first sentence of a non-first paragraph in the text to be processed, assign the third value to the first weight coefficient, where the third value is less than the second value.
[0009] Optionally, determine the relevance index between each word segmentation result and each training sample, including: determining the relevance index between the target word segmentation result and the target training sample according to the first probability of the target word segmentation result appearing in the target training sample and the second probability of the target word segmentation result appearing in the word sequence composed of multiple word segmentation results, where the target training sample is any one training sample in the training sample set, and the target word segmentation result is any one word segmentation result among the multiple word segmentation results.
[0010] Optionally, determine the relevance index between each word segmentation result and each training sample, including: determining the first parameter according to the length of the target training sample and the average length of all the text in the training sample set, where the target training sample is any one training sample in the training sample set; determining the relevance index between the target word segmentation result and the target training sample according to the first probability of the target word segmentation result appearing in the target training sample, the second probability of the target word segmentation result appearing in the word sequence composed of multiple word segmentation results, and the first parameter, where the target word segmentation result is any one word segmentation result among the multiple word segmentation results.
[0011] Optionally, the first parameter is determined by the following formula: where, K is the first parameter; k i is an adjustable constant; b is a natural number from 0.5 to 1; dl is the length of the target training sample; avgdl is the average length of all the text in the training sample set.
[0012] Optionally, the relevance index between the word segmentation result and the training sample is determined by the following formula: where, q i is the i-th word segmentation result, i is a positive integer; R(q i, d) is the relevance index between the i-th word segmentation result and the training sample; f i is the first probability; qf i is the second probability; k1 and k2 are adjustable constants.
[0013] According to another aspect of the present application, there is also provided a text recognition device, including: a first acquisition module, configured to acquire a training sample set, where the training sample set includes: historical review texts carrying preset labels, and the historical review texts are from a target data source; a second acquisition module, configured to acquire a text to be processed in the target data source, perform word segmentation processing on the text to be processed, and obtain a plurality of word segmentation results; a first determination module, configured to determine a weight coefficient for each word segmentation result according to the text position of each word segmentation result in the text to be processed and the occurrence frequency of a preset feature word in the first text in each word segmentation result, where the first text is all the text included in the target label corresponding to the preset feature word; a second determination module, configured to, for each training sample, determine the relevance index between each word segmentation result and the training sample, and perform weighted summation on the relevance index according to the weight coefficient to obtain the correlation score between the text to be processed and each training sample; a processing module, configured to perform classification processing on the correlation score by using a logistic regression model to obtain the probability values of the text to be processed belonging to different preset labels, and determine the label corresponding to the maximum probability value as the label of the text to be processed.
[0014] According to another aspect of the present application, there is also provided a non-volatile storage medium, where the storage medium includes a stored program, and when the program runs, it controls the device where the storage medium is located to execute the above text recognition method.
[0015] According to another aspect of the present application, there is also provided an electronic device, including: a memory and a processor, where the processor is configured to run the program stored in the memory, and when the program runs, it executes the above text recognition method.
[0016] According to another aspect of the present application, there is also provided a computer program, where when the computer program is executed by a processor, it implements the above text recognition method.
[0017] According to another aspect of the present application, there is also provided a computer program product, where the computer program product includes a non-volatile computer-readable storage medium, and the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above text recognition method.
[0018] In this application, a training sample set is obtained, where the training sample set includes: historical review texts carrying preset labels, and the historical review texts are from a target data source; the text to be processed in the target data source is obtained, and the text to be processed is segmented to obtain a plurality of segmentation results; according to the text position of each segmentation result in the text to be processed and the occurrence frequency of a preset feature word in each segmentation result in the first text, the weight coefficient of each segmentation result is determined, where the first text is all the text included in the target label corresponding to the preset feature word; for each training sample, the relevance index between each segmentation result and the training sample is determined, and according to the weight coefficient, the relevance indexes are weighted and summed to obtain the correlation score between the text to be processed and each training sample; the logical regression model is used to classify the correlation scores to obtain the probability values of the text to be processed belonging to different preset labels, and the label corresponding to the maximum probability value is determined as the label of the text to be processed. In this way, the purpose of effectively using the keyword position information is achieved, thereby realizing the technical effect of improving the recognition accuracy, and further solving the technical problem of limited recognition accuracy caused by the inability to effectively use the keyword position information in the related text recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0020] Figure 1 is a flowchart of a text recognition method according to an embodiment of the present application;
[0021] Figure 2 is a structural diagram of a text recognition device according to an embodiment of the present application;
[0022] Figure 3 is a hardware structure block diagram of a computer terminal of a text recognition method according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0024] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0025] According to an embodiment of the present application, a method embodiment of a text recognition method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0026] Figure 1 is a flowchart of a text recognition method according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:
[0027] Step S101, obtain a training sample set, where the training sample set includes: historical review texts carrying preset labels, and the historical review texts are from a target data source.
[0028] Automatically or manually crawl historical review texts in the target data source (such as a specific forum, social media platform, etc.) through web crawler technology, and manually classify and label these review texts with preset labels. The preset labels can include but are not limited to categories such as normal and involving multiple different abnormal fields. These labeled review texts constitute the training sample set, providing basic data for subsequent model training.
[0029] Step S102, obtain the text to be processed in the target data source, perform word segmentation processing on the text to be processed, and obtain a plurality of word segmentation results.
[0030] Real-time or regularly crawl new review texts from the target data source as the text to be processed. First, perform Chinese word segmentation processing on each text to be processed, and use tools such as jieba word segmentation to cut the text into multiple word segmentation results to prepare for subsequent feature extraction.
[0031] Step S103: Determine the weight coefficient of each word segmentation result based on the text position of each word segmentation result in the text to be processed and the occurrence frequency of the preset feature word in each word segmentation result in the first text, where the first text is all the text included in the target label corresponding to the preset feature word.
[0032] Based on the word segmentation results obtained in S102 and combined with their position information in the text, calculate the weight coefficient for each word segmentation result. For example, if the word segmentation result is in the first or last sentence of the text, its weight coefficient is 7; if it is in the first sentence of a certain paragraph, the weight coefficient is 4; for word segmentation results in other positions, the weight coefficient is 2.
[0033] In addition, for the preset feature word in each word segmentation result, calculate its occurrence frequency in all the text (the first text) under the corresponding target label, as well as its frequency in the text under other labels, to more accurately reflect the importance of the word.
[0034] Step S104: For each training sample, determine the correlation index between each word segmentation result and the training sample, and perform a weighted sum on the correlation index according to the weight coefficient to obtain the correlation score between the text to be processed and each training sample.
[0035] For each training sample, calculate the correlation index between each word segmentation result in the text to be processed and the sample. This index reflects the relative importance and occurrence frequency of the word in the text. Then, perform a weighted sum on the correlation index according to the weight coefficient determined in S103 to obtain the correlation score between the text to be processed and each training sample.
[0036] The correlation score between the text to be processed and each training sample can be obtained through the following formula:
[0037]
[0038] where q i is the i-th word segmentation result, i is a positive integer, and d is a certain training sample.
[0039] Step S105: Use the logistic regression model to perform classification processing on the correlation score to obtain the probability values of the text to be processed belonging to different preset labels, and determine the label corresponding to the maximum probability value as the label of the text to be processed.
[0040] Use the pre-trained logistic regression model to perform classification prediction on the correlation score calculated in S104 to obtain the probability values of the text to be processed belonging to different preset labels (such as normal, politics-related, gambling-related, etc.). Finally, determine the label with the maximum probability value as the final classification label of the text to be processed, thus realizing the anomaly recognition of the new review text.
[0041] The logistic regression model is pre-trained through the following method: performing gradient descent operations based on the objective loss function and the objective derivative corresponding to the objective loss function until the objective loss function meets the preset convergence condition, obtaining the trained logistic regression model, where the objective loss function is: m is the number of the training samples; σ is the activation function; is the i-th training sample after vectorization, i is a positive integer; θ is the model parameter; y (i) is the true label corresponding to the i-th test sample; the objective derivative is: X b is a matrix containing all the feature vectors of the training samples. Each row represents a training sample, and each column corresponds to a feature. For each sample, its feature vector contains all the variable values that may affect the classification result, such as the correlation score, word position information, and other features. To enable the logistic regression model to handle the intercept term, a column of all 1s is usually added at the leftmost end of the feature matrix, and this column is called the bias term or intercept term. Therefore, the feature matrix is denoted as X b . is the transpose of X b .
[0042] The above logistic regression algorithm successfully solves the binary classification problem. Next, the OvR and OvO concepts are introduced to achieve the multi-classification problem. For example, if there are five classes to be recognized, then five classifications are performed.
[0043] Furthermore, the above text recognition method can be integrated into a third-party application through the http interface for the third-party application to call, to monitor the target website, and after abnormal remarks appear, push them to the responsible person through methods such as text messages and WeChat robots.
[0044] According to the above steps, a training sample set is obtained, where the training sample set includes: historical review texts carrying preset labels, and the historical review texts are from a target data source; the text to be processed in the target data source is obtained, and the text to be processed is segmented to obtain a plurality of segmentation results; according to the text position of each segmentation result in the text to be processed and the occurrence frequency of the preset feature word in each segmentation result in the first text, the weight coefficient of each segmentation result is determined, where the first text is all the texts included in the target label corresponding to the preset feature word; for each training sample, the relevance index between each segmentation result and the training sample is determined, and according to the weight coefficient, the relevance indexes are weighted and summed to obtain the correlation score between the text to be processed and each training sample; the logical regression model is used to classify the correlation scores to obtain the probability values of the text to be processed belonging to different preset labels, and the label corresponding to the maximum probability value is determined as the label of the text to be processed, achieving the purpose of effectively using the keyword position information, thereby realizing the technical effect of improving the recognition accuracy.
[0045] The following gives an exemplary illustration and explanation of Figure 1 the steps shown.
[0046] According to some optional embodiments of the present application, the weight coefficient is determined by the following formula: where posw is the weight coefficient; W is the first weight coefficient determined according to the text position of each segmentation result in the text to be processed; P(m q ) is the occurrence frequency of the preset feature word in each segmentation result in the first text; P(m q )′ is the occurrence frequency of the preset feature word in each segmentation result in the second text, where the second text is all the texts included in the other preset labels except the target label corresponding to the preset feature word among the multiple preset labels in the training sample set.
[0047] Specifically, P(m q ) represents the occurrence frequency of the preset feature word in the first text (i.e., all the texts included in the target label). The calculation of this frequency helps to evaluate the importance of the word in a specific category.
[0048] P(m q )′ represents the occurrence frequency of the preset feature word in the second text (i.e., all the texts included in the other preset labels except the target label). By comparing P(m q ) and P(m q )′, it is possible to more accurately judge whether the word is a key feature of a certain specific category.
[0049] Through the above formula, the weight coefficient of the preset feature words in each word segmentation result can be calculated, and this coefficient will be used to adjust the importance of the feature words in the subsequent steps. During the calculation process, the +0.5 part in the formula plays a role in smoothing processing, avoiding the situation where the calculation of the log function becomes infinite or undefined when P(m q ) or P(m q )′ is 0, ensuring the stability of the weight coefficient calculation.
[0050] Furthermore, the first weight coefficient can be determined by the following method: when the word segmentation result is in the first sentence of the first paragraph or the last sentence of the last paragraph in the text to be processed, assign the first numerical value to the first weight coefficient; when the word segmentation result is in the first sentence of a non-first paragraph in the text to be processed, assign the second numerical value to the first weight coefficient, where the second numerical value is less than the first numerical value; when the word segmentation result is not in the first sentence of the first paragraph or the last sentence of the last paragraph or the first sentence of a non-first paragraph in the text to be processed, assign the third numerical value to the first weight coefficient, where the third numerical value is less than the second numerical value.
[0051] Specifically, when the word segmentation result appears in the first sentence of the first paragraph or the last sentence of the last paragraph of the text to be processed, it is considered that these words are crucial for expressing the core idea of the text, so a higher weight should be given. In this case, assign the first numerical value, such as 7, to the first weight coefficient to highlight the importance and influence of these words.
[0052] For the word segmentation results that appear in the first sentence of a non-first paragraph in the text to be processed, although these words are not as crucial as those in the first and last sentences, they may still provide important clues to the overall meaning of the paragraph or text. Therefore, assign the second numerical value, such as 4, to the first weight coefficient of the words in these positions. This numerical value is less than the first numerical value, reflecting their relatively lower but still important position weight.
[0053] For the word segmentation results in the text to be processed that are neither in the first sentence of the first paragraph or the last sentence of the last paragraph nor in the first sentence of a non-first paragraph, their position influence is relatively small. In this case, assign the third numerical value, such as 2, to the first weight coefficient. This numerical value is less than the second numerical value, indicating their lower priority in position weight.
[0054] In some alternative embodiments of the present application, the relevance index between each word segmentation result and each training sample is determined, including: determining the relevance index between the target word segmentation result and the target training sample according to the first probability that the target word segmentation result appears in the target training sample and the second probability that the target word segmentation result appears in the word sequence composed of multiple word segmentation results, where the target training sample is any one training sample in the training sample set, and the target word segmentation result is any one word segmentation result among the multiple word segmentation results.
[0055] As some alternative embodiments of the present application, determining the relevance index between each word segmentation result and each training sample includes: determining a first parameter according to the length of the target training sample and the average length of all texts in the training sample set, where the target training sample is any one training sample in the training sample set; determining the relevance index between the target word segmentation result and the target training sample according to the first probability that the target word segmentation result appears in the target training sample, the second probability that the target word segmentation result appears in the word sequence composed of multiple word segmentation results, and the first parameter, where the target word segmentation result is any one word segmentation result among the multiple word segmentation results.
[0056] Specifically, the first parameter can be determined by the following formula: where K is the first parameter; k i is an adjustable constant; b is a natural number from 0.5 to 1; dl is the length of the target training sample; avgdl is the average length of all texts in the training sample set.
[0057] Furthermore, the relevance index between the word segmentation result and the training sample can be determined by the following formula: where q i is the i-th word segmentation result, and i is a positive integer; R(q i , d) is the relevance index between the i-th word segmentation result and the training sample; f i is the first probability; qf i is the second probability; k1, k2 are adjustable constants.
[0058] Modify the parameters k1, k2, b in the above formula, and the following is the accuracy display under different constant value cases:
[0059]
[0060] According to the above experimental results, k1 = 2, k2 = 1, b = 0.5 can be selected as the value of the BM25-POS formula. The accuracy value of 97.1% can effectively meet the needs of abnormal comment recognition.
[0061] As some alternative implementation manners of the present application, the text recognition method further includes the following steps: optimizing the parameters k1, k2, b in the BM25-POS algorithm so that it can be automatically adjusted according to different data sets and application scenarios. Hyperparameter optimization methods such as grid search, random search, or Bayesian optimization can be used, combined with cross-validation to evaluate the model performance, and the best parameter combination can be found. In addition, an adaptive adjustment mechanism can be designed to dynamically modify the parameters according to the performance of the model in actual applications to adapt to the changing network comment environment. Specifically, it includes the following steps:
[0062] 1. Grid search. First, define the search range of k1, k2, and b parameters. For example, k1 can be between [1.0, 3.0] with a step size of 0.1; k2 can be between [0, 2] with a step size of 0.1; b can be between [0.5, 0.9] with a step size of 0.05. For each set of parameter combinations, use cross-validation (such as k-fold cross-validation) to evaluate model performance. Divide the dataset into k subsets, take one of the subsets as the test set in turn, and the rest as the training set, and calculate the average accuracy. Compare the cross-validation accuracy under all parameter combinations, and select the parameter combination with the highest accuracy as the best parameter.
[0063] 2. Random search. Similar to grid search, but parameter values are randomly drawn from a defined distribution. This method is more effective when the parameter space is large. Cross-validation is also used to evaluate the model performance under randomly drawn parameter combinations. After multiple iterations, the best performing parameter combination is selected.
[0064] 3. Bayesian optimization. Use Gaussian process or tree Pareto estimation to build a surrogate model to predict the model performance under different parameter combinations. Under the guidance of the surrogate model, search for areas in the parameter space that may perform better, continue cross-validation evaluation, and iteratively optimize parameter combinations. When the performance improvement space predicted by the surrogate model is small, stop searching and select the current optimal parameter combination.
[0065] In addition, the text recognition method also includes the following steps: Based on the original logistic regression model, a deep learning model such as a convolutional neural network or a long short-term memory network is introduced to capture more complex text structures and semantic information. CNN can effectively extract local features in text, while LSTM is good at processing sequence data and can better understand the sequential relationship between words. By combining the deep learning model with the logistic regression model to form a hybrid model, the performance of abnormal comment recognition can be improved. The specific steps include the following:
[0066] 1. Data preprocessing. First, the original review data is preprocessed, including word segmentation, stop word removal, word embedding, etc. Word embedding can use pre-trained word vectors, such as Word2Vec, GloVe, or BERT, to convert words into fixed-length vector representations, which can capture the semantic information of the words.
[0067] 2. Build a deep learning model. Input layer: Take the review text after word embedding processing as input and convert it into a word vector matrix. Convolutional layer: Use convolutional kernels of different sizes (such as 1, 2, 3, etc.) to capture local patterns in the text. Each convolutional kernel corresponds to a feature map, and multiple convolutional kernels can be used to extract n-gram features of different lengths. Pooling layer: Apply max pooling or average pooling operations to reduce the dimension of the feature vector and extract the most important eigenvalue at the same time. Fully connected layer: Flatten the pooled feature vector and pass it through one or more fully connected layers to integrate all local features and output a feature vector of a fixed length.
[0068] 3. Model fusion. Combine the output of the deep learning model with the input of the logistic regression model, and there are two main ways: Hybrid model: Combine the features extracted by the deep learning model with the features calculated by the BM25-POS algorithm in parallel as the input of the logistic regression. This is equivalent to increasing the feature dimension of the model, enabling the logistic regression model to better utilize the local and global structural information of the text. Cascade model: Let the output of the logistic regression model be further processed by the deep learning model, or vice versa. First, extract features through the deep learning model and then input them into the logistic regression model. In this way, the deep learning model can be used as a feature extractor, and the logistic regression model classifies based on these features.
[0069] 4. Training and optimization. Use deep learning frameworks such as TensorFlow or PyTorch to train the CNN model. You can use pre-trained word vectors for fine-tuning or train word vectors from scratch. Train the logistic regression model based on the features extracted by the deep learning model, and you can use the LogisticRegression class in sklearn. If a cascade model is adopted, the deep learning model and the logistic regression model can be embedded in a larger neural network for joint training. This can ensure smoother feature transformation between models and improve the overall performance.
[0070] Figure 2 is a structural diagram of a text recognition device according to an embodiment of the present application, as Figure 2 shown, the device includes:
[0071] The first acquisition module 21 is used to acquire a training sample set, where the training sample set includes: historical review texts carrying preset labels, and the historical review texts come from a target data source.
[0072] The second acquisition module 22 is used to acquire the text to be processed in the target data source, perform word segmentation on the text to be processed, and obtain multiple word segmentation results.
[0073] The first determination module 23 is configured to determine the weight coefficient of each word segmentation result according to the text position of each word segmentation result in the text to be processed and the occurrence frequency of the preset feature word in the first text, where the first text is all the text included in the target label corresponding to the preset feature word.
[0074] The second determination module 24 is configured to, for each training sample, determine the relevance index between each word segmentation result and the training sample, and perform weighted summation on the relevance index according to the weight coefficient to obtain the correlation score between the text to be processed and each training sample.
[0075] The processing module 25 is configured to perform classification processing on the correlation score by using a logistic regression model to obtain the probability values of the text to be processed belonging to different preset labels, and determine the label corresponding to the maximum probability value as the label of the text to be processed.
[0076] Optionally, the weight coefficient is determined by the following formula: where posw is the weight coefficient; W is the first weight coefficient determined according to the text position of each word segmentation result in the text to be processed; P(m q ) is the occurrence frequency of the preset feature word in the first text in each word segmentation result; P(m q )′ is the occurrence frequency of the preset feature word in the second text in each word segmentation result, where the second text is all the text included in the other preset labels except the target label corresponding to the preset feature word among the multiple preset labels in the training sample set.
[0077] Optionally, the first weight coefficient is determined by the following method: when the word segmentation result is in the first sentence of the first paragraph or the last sentence of the last paragraph in the text to be processed, assign a first value to the first weight coefficient; when the word segmentation result is in the first sentence of a non-first paragraph in the text to be processed, assign a second value to the first weight coefficient, where the second value is less than the first value; when the word segmentation result is not in the first sentence of the first paragraph or the last sentence of the last paragraph or the first sentence of a non-first paragraph in the text to be processed, assign a third value to the first weight coefficient, where the third value is less than the second value.
[0078] Optionally, determining the relevance index between each word segmentation result and each training sample includes: determining the relevance index between the target word segmentation result and the target training sample according to the first probability of the target word segmentation result appearing in the target training sample and the second probability of the target word segmentation result appearing in the word sequence composed of multiple word segmentation results, where the target training sample is any one training sample in the training sample set, and the target word segmentation result is any one word segmentation result among the multiple word segmentation results.
[0079] Optionally, determine the relevance index between each word segmentation result and each training sample, including: determining a first parameter according to the length of the target training sample and the average length of all texts in the training sample set, where the target training sample is any one training sample in the training sample set; determining the relevance index between the target word segmentation result and the target training sample according to the first probability that the target word segmentation result appears in the target training sample, the second probability that the target word segmentation result appears in the word sequence composed of multiple word segmentation results, and the first parameter, where the target word segmentation result is any one word segmentation result among the multiple word segmentation results.
[0080] Optionally, determine the first parameter through the following formula: where K is the first parameter; k i is an adjustable constant; b is a natural number from 0.5 to 1; dl is the length of the target training sample; avgdl is the average length of all texts in the training sample set.
[0081] Optionally, determine the relevance index between the word segmentation result and the training sample through the following formula: where q i is the i-th word segmentation result, and i is a positive integer; R(q i , d) is the relevance index between the i-th word segmentation result and the training sample; f i is the first probability; qf i is the second probability; k1 and k2 are adjustable constants.
[0082] It should be noted that the above Figure 2 each module can be a program module (for example, a set of program instructions that implement a specific function), or a hardware module. For the latter, it can be presented in the following forms, but not limited to: the manifestation form of each of the above modules is a processor, or the functions of each of the above modules are implemented by a processor.
[0083] It should be noted that Figure 2 the preferred implementation manner of the illustrated embodiment can refer to the relevant description of the Figure 1 illustrated embodiment, which will not be elaborated here.
[0084] Figure 3 shows a hardware structure block diagram of a computer terminal for implementing a text recognition method. As Figure 3As shown, the computer terminal 30 may include one or more processors 302 (shown as 302a, 302b, ……, 302n in the figure) (the processor 302 may include, but is not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA), a memory 304 for storing data, and a transmission module 306 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 3 the structure shown is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 30 may further include more or fewer components than those Figure 3 shown in, or have a different configuration from that Figure 3 shown.
[0085] It should be noted that the above one or more processors 302 and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 30. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistor terminal path connected to an interface).
[0086] The memory 304 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the text recognition method in the embodiments of the present application. The processor 302 executes various functional applications and data processing by running the software programs and modules stored in the memory 304, that is, implements the above-mentioned text recognition method. The memory 304 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 304 may further include a memory remotely set relative to the processor 302, and these remote memories can be connected to the computer terminal 30 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0087] The transmission module 306 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the communication provider of the computer terminal 30. In one example, the transmission module 306 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission module 306 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0088] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 30.
[0089] It should be noted here that in some alternative embodiments, the above Figure 3 shown computer terminal may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware elements and software elements. It should be pointed out that Figure 3 is only an example of a specific specific instance and is intended to show the types of components that may exist in the above computer terminal.
[0090] It should be noted that Figure 3 the shown computer terminal is used to execute Figure 1 the shown text recognition method. Therefore, the relevant explanations in the execution method of the above commands also apply to this electronic device, which will not be elaborated here.
[0091] The embodiment of the present application also provides a non-volatile storage medium. The non-volatile storage medium includes a stored program, wherein when the program runs, it controls the device where the storage medium is located to execute the above text recognition method.
[0092] A program for a non-volatile storage medium to perform the following functions: obtaining a set of training samples, where the set of training samples includes: historical review texts carrying preset labels, and the historical review texts are from a target data source; obtaining a text to be processed in the target data source, performing word segmentation on the text to be processed to obtain a plurality of word segmentation results; determining a weight coefficient for each word segmentation result according to the text position of each word segmentation result in the text to be processed and the occurrence frequency of a preset feature word in each word segmentation result in a first text, where the first text is all the text included in the target label corresponding to the preset feature word; for each training sample, determining a relevance index between each word segmentation result and the training sample, and performing weighted summation on the relevance index according to the weight coefficient to obtain a correlation score between the text to be processed and each training sample; using a logistic regression model to perform classification processing on the correlation score to obtain probability values of the text to be processed belonging to different preset labels, and determining the label corresponding to the maximum probability value as the label of the text to be processed.
[0093] An embodiment of the present application further provides an electronic device, including: a memory and a processor, where the processor is used to run a program stored in the memory, and when the program runs, it executes the above text recognition method.
[0094] The processor is used to run a program that performs the following functions: obtaining a set of training samples, where the set of training samples includes: historical review texts carrying preset labels, and the historical review texts are from a target data source; obtaining a text to be processed in the target data source, performing word segmentation on the text to be processed to obtain a plurality of word segmentation results; determining a weight coefficient for each word segmentation result according to the text position of each word segmentation result in the text to be processed and the occurrence frequency of a preset feature word in each word segmentation result in a first text, where the first text is all the text included in the target label corresponding to the preset feature word; for each training sample, determining a relevance index between each word segmentation result and the training sample, and performing weighted summation on the relevance index according to the weight coefficient to obtain a correlation score between the text to be processed and each training sample; using a logistic regression model to perform classification processing on the correlation score to obtain probability values of the text to be processed belonging to different preset labels, and determining the label corresponding to the maximum probability value as the label of the text to be processed.
[0095] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0096] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0097] In the above embodiments of the present application, the collected information is information and data authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with relevant laws, regulations, and standards, necessary protection measures are taken, it does not violate public order and good customs, and a corresponding operation entry is provided for the user to choose to authorize or reject.
[0098] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0099] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0100] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0101] If the above-mentioned integrated units are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the relevant technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs, etc., which can store program codes.
[0102] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A text recognition method, characterized in that: include: Acquire a training sample set, wherein the training sample set includes: historical review texts with preset labels, and the historical review texts come from a target data source; Acquire the text to be processed in the target data source, perform word segmentation processing on the text to be processed, and obtain multiple word segmentation results; Determine a weight coefficient of each of the word segmentation results according to the text position of each of the word segmentation results in the text to be processed and the frequency of occurrence of a preset feature word in each of the word segmentation results in a first text, wherein the first text is all texts included in the target tag corresponding to the preset feature word; For each training sample, determining a correlation index between each of the word segmentation results and the training sample, and performing weighted summation on the correlation indexes according to the weight coefficients to obtain a correlation score between the text to be processed and each training sample; The relevance scores are classified using a logistic regression model to obtain probability values of the text to be processed belonging to different preset labels, and the label corresponding to the maximum probability value is determined as the label of the text to be processed.
2. The method according to claim 1, characterized in that The weight coefficient is determined by the following formula: Wherein, posw is the weight coefficient; W is the first weight coefficient determined according to the text position of each word segmentation result in the text to be processed; P(m q ) is the frequency of occurrence of each of the preset feature words in the word segmentation results in the first text; P(m q )′ is the frequency of occurrence of each preset feature word in the word segmentation result in the second text, wherein the second text is all texts included in other preset tags among multiple preset tags in the training sample set except the target tag corresponding to the preset feature word.
3. The method according to claim 2, characterized in that The first weight coefficient is determined by the following method: When the word segmentation result is located in the first sentence of the first paragraph or the last sentence of the last paragraph in the text to be processed, assigning a first numerical value to the first weight coefficient; When the word segmentation result is located in the first sentence of the non-first paragraph in the text to be processed, assigning a second value to the first weight coefficient, wherein the second value is smaller than the first value; When the word segmentation result is not located in the first sentence of the first paragraph, the last sentence of the last paragraph, or the first sentence of the non-first paragraph in the text to be processed, a third numerical value is assigned to the first weight coefficient, wherein the third numerical value is smaller than the second numerical value.
4. The method according to claim 1, characterized in that Determining the correlation index between each of the word segmentation results and each of the training samples includes: According to a first probability of a target word segmentation result appearing in a target training sample and a second probability of the target word segmentation result appearing in a word sequence formed by a plurality of the word segmentation results, a correlation index between the target word segmentation result and the target training sample is determined, wherein the target training sample is any one of the training samples in the training sample set, and the target word segmentation result is any one of the plurality of word segmentation results.
5. The method according to claim 1, characterized in that Determining the correlation index between each of the word segmentation results and each of the training samples includes: Determine a first parameter according to the length of a target training sample and the average length of all texts in the training sample set, wherein the target training sample is any one training sample in the training sample set; According to a first probability of a target word segmentation result appearing in the target training sample, a second probability of the target word segmentation result appearing in a word sequence formed by multiple word segmentation results, and the first parameter, a correlation index between the target word segmentation result and the target training sample is determined, wherein the target word segmentation result is any one of the multiple word segmentation results.
6. The method according to claim 5, characterized in that The first parameter is determined by the following formula: Wherein, K is the first parameter; k i is an adjustable constant; b is a natural number ranging from 0.5 to 1; dl is the length of the target training sample; avgdl is the average length of all texts in the training sample set.
7. The method according to claim 5, characterized in that The correlation index between the word segmentation result and the training sample is determined by the following formula: Among them, q i is the i-th word segmentation result, i is a positive integer; R(q i , d) is the correlation index between the i-th word segmentation result and the training sample; f i is the first probability; qf i is the second probability; k1 and k2 are adjustable constants.
8. A text recognition device, characterized in that: include: A first acquisition module is used to acquire a training sample set, wherein the training sample set includes: historical comment texts carrying preset tags, and the historical comment texts come from a target data source; A second acquisition module is used to acquire the text to be processed in the target data source, perform word segmentation processing on the text to be processed, and obtain multiple word segmentation results; A first determination module is used to determine a weight coefficient of each of the word segmentation results according to a text position of each of the word segmentation results in the text to be processed and an occurrence frequency of a preset feature word in each of the word segmentation results in a first text, wherein the first text is all texts included in a target tag corresponding to the preset feature word; A second determination module is used to determine, for each training sample, a correlation index between each of the word segmentation results and the training sample, and perform weighted summation on the correlation index according to the weight coefficient to obtain a correlation score between the text to be processed and each training sample; The processing module is used to classify the relevance scores using a logistic regression model to obtain probability values of the text to be processed belonging to different preset labels, and determine the label corresponding to the maximum probability value as the label of the text to be processed.
9. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored program, wherein when the program is executed, the device where the non-volatile storage medium is located is controlled to execute the text recognition method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program executes the text recognition method according to any one of claims 1 to 7 when running.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the text recognition method according to any one of claims 1 to 7 is implemented.