Method, apparatus, electronic device, storage medium and program product for text correction
Correcting the OCR recognition results through vector transformation model and unsupervised clustering algorithm, the problem of inaccurate identification of specific entities in the text is solved, and the recognition accuracy and consistency is improved. It is suitable for scenarios where data is scarce or resources are limited.
Patent Information
- Application Number
- CN202411621518.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-13
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2044-11-13
AI Technical Summary
When the prior art passes optical character recognition (OCR), it is difficult to accurately identify specific entities in the text such as human names, place names and proper nouns, especially when the distribution is sparse, resulting in inaccurate recognition.
The vector transformation model and unsupervised clustering algorithm are used to correct the text, and the pre-trained model is fine-tuned through mask processing, text fragment interception, vector transformation and clustering processing, combined with transfer learning, and optimize the text recognition results.
It improves the recognition accuracy and consistency of specific objects in the text, reduces manual intervention, and reduces error rates, and is especially suitable for scenarios where data is scarce or resource limited.
Smart Images

Figure CN119646214B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of text correction, and in particular, to a method, apparatus, electronic device, storage medium, and program product for text correction. Background Art
[0002] In some data sorting scenarios, certain contents in text are usually prone to errors. For example, when performing text recognition on a text image through image recognition technologies such as Optical Character Recognition (OCR), specific entities in the image recognition text usually have inaccurate recognition problems due to reasons such as their sparse distribution. Among them, entities mainly include personal names, place names, organization names, and proper nouns, etc. Therefore, how to accurately correct text is a problem that needs to be solved. Summary of the Invention
[0003] To solve the above problems, embodiments of the present disclosure provide a method, apparatus, electronic device, storage medium, and program product for text correction.
[0004] On the one hand, an embodiment of the present disclosure provides a method for text correction, including:
[0005] Performing mask processing on multiple target objects of specified types included in the text to be processed respectively to obtain target text;
[0006] According to the mask positions in the target text, segmenting the target file to obtain multiple text segments containing masks;
[0007] According to the vector conversion model, performing vector conversion on each text segment respectively to obtain text vectors corresponding to each text segment respectively;
[0008] Performing unsupervised clustering processing on each text vector to obtain at least one text vector cluster and its corresponding reference text vector respectively;
[0009] Respectively correcting the target objects corresponding to the text vectors included in the corresponding text vector cluster according to the reference objects corresponding to each reference text vector.
[0010] In one implementation, before performing mask processing on multiple target objects of specified types included in the text to be processed respectively, the method further includes:
[0011] Performing text recognition on the text image to be processed to obtain the text to be processed;
[0012] Determining the target objects of specified types in the target text.
[0013] In one implementation, before performing vector conversion on each text segment according to the vector conversion model to obtain text vectors corresponding to each text segment respectively, the method further includes:
[0014] Performing target object recognition on multiple initial samples in the target task training data respectively to obtain target objects of a specified type included in each initial sample;
[0015] Performing masking processing on each target object in each initial sample respectively to obtain multiple target samples;
[0016] Constructing multiple text construction data corresponding to each target sample respectively according to the masking positions in each target sample;
[0017] Fine-tuning the pre-trained model according to the text construction data corresponding to each target sample respectively to obtain a trained vector conversion model; the pre-trained model is a model for vector conversion trained well based on general data.
[0018] In one implementation, constructing multiple text construction data corresponding to each target sample respectively according to the masking positions in each target sample includes:
[0019] Performing segmentation on each target sample respectively according to the masking positions in each target sample to obtain multiple text segments corresponding to each target sample respectively;
[0020] For each text segment in each target sample respectively, generating corresponding text construction data according to the text segment, the target object masked in the text segment, and the negative sample label; the negative sample label is different from the target object masked in the text segment.
[0021] In one implementation, correcting the target objects corresponding to the text vectors included in the corresponding text vector cluster respectively according to the reference object corresponding to each reference text vector includes:
[0022] For each target object corresponding to each text vector cluster respectively, performing the following steps:
[0023] Determining the similarity between the reference object corresponding to the text vector cluster and the target object;
[0024] If it is determined that the similarity meets the text correction condition, correcting the target object to the reference object corresponding to the text vector cluster.
[0025] In one implementation, determining the similarity between the reference object corresponding to the text vector cluster and the target object includes:
[0026] Determining the number of characters with differences between the reference object corresponding to the text vector cluster and the target object; the number is negatively correlated with the similarity;
[0027] Determine that the similarity meets the text correction conditions, including:
[0028] If the quantity is not higher than the quantity threshold, it is determined that the text correction condition is met.
[0029] On the one hand, an embodiment of the present disclosure provides a text correction device, including:
[0030] A mask unit for respectively performing mask processing on multiple target objects of a specified type included in the text to be processed to obtain a target text;
[0031] A truncation unit for truncating the target file into multiple text segments containing masks according to the mask positions in the target text;
[0032] A conversion unit for respectively performing vector conversion on each text segment according to a vector conversion model to obtain text vectors respectively corresponding to each text segment; the vector conversion model is obtained by fine-tuning a pre-trained model;
[0033] A clustering unit for performing unsupervised clustering processing on each text vector to obtain at least one text vector cluster and its corresponding reference text vector respectively;
[0034] A correction unit for respectively correcting the target objects corresponding to the text vectors included in the corresponding text vector cluster according to the reference objects corresponding to each reference text vector.
[0035] In one implementation, the mask unit is further configured to:
[0036] Perform text recognition on the image to be processed to obtain the text to be processed;
[0037] Determine the target objects of the specified type in the target text.
[0038] In one implementation, the conversion unit is further configured to:
[0039] Perform target object recognition on multiple initial samples in the target task training data respectively to obtain the target objects of the specified type included in each initial sample;
[0040] Perform mask processing on each target object in each initial sample respectively to obtain multiple target samples;
[0041] Construct multiple text construction data respectively corresponding to each target sample according to the mask positions in each target sample;
[0042] Fine-tune the pre-trained model according to the text construction data corresponding to each target sample to obtain a trained vector conversion model; the pre-trained model is a model for vector conversion trained based on general data.
[0043] In one implementation, the conversion unit is further configured to:
[0044] Segment each target sample according to the masked positions in each target sample to obtain multiple text segments corresponding to each target sample;
[0045] For each text segment in each target sample, generate corresponding text construction data according to the text segment, the target object masked in the text segment, and the negative sample label; the negative sample label is different from the target object masked in the text segment.
[0046] In one implementation, the correction unit is configured to:
[0047] For each target object corresponding to each text vector cluster, perform the following steps:
[0048] Determine the similarity between the reference object corresponding to the text vector cluster and the target object;
[0049] If it is determined that the similarity meets the text correction condition, correct the target object to the reference object corresponding to the text vector cluster.
[0050] In one implementation, the correction unit is configured to:
[0051] Determine the number of characters in which the reference object corresponding to the text vector cluster differs from the target object; the number is negatively correlated with the similarity;
[0052] The correction unit is configured to:
[0053] If the number is not higher than the number threshold, it is determined that the text correction condition is met.
[0054] On the one hand, an electronic device is provided in an embodiment of the present disclosure, including:
[0055] A processor; and
[0056] A memory storing computer instructions for causing the processor to execute the steps of the method provided in any of the various optional implementations of text correction as described above.
[0057] On the one hand, a computer-readable storage medium is provided in an embodiment of the present disclosure, storing computer instructions for causing a computer to execute the steps of the method provided in any of the various optional implementations of text correction as described above.
[0058] On the one hand, an embodiment of the present disclosure provides a computer program product, including computer-readable code or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the steps of the method provided in any of the above-mentioned various alternative implementations of text correction.
[0059] The method for text correction in the embodiments of the present disclosure includes respectively performing masking processing on multiple target objects of specified types included in the text to be processed to obtain target text; intercepting fragments of the target file according to the masking positions in the target text to obtain multiple text fragments containing masks; respectively performing vector conversion on each text fragment according to a vector conversion model to obtain text vectors corresponding to each text fragment; performing unsupervised clustering processing on each text vector to obtain at least one text vector cluster and its corresponding reference text vector respectively; and respectively correcting the target objects corresponding to the text vectors included in the corresponding text vector cluster according to the reference objects corresponding to each reference text vector. In this way, the vector conversion model and the unsupervised clustering algorithm can be combined to correct the target objects in the text, solving the problem that it is difficult to correct the text. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] Figure 1 is a flowchart of a method for training a vector conversion model in the embodiments of the present disclosure.
[0061] Figure 2 is a flowchart of a method for text correction in the embodiments of the present disclosure.
[0062] Figure 3 is a flowchart of a method for correcting target objects in the embodiments of the present disclosure.
[0063] Figure 4 is a structural block diagram of a device for text correction in the embodiments of the present disclosure.
[0064] Figure 5 is a schematic structural diagram of an electronic device in the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] The technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present disclosure belong to the scope of protection of the present disclosure. In addition, the technical features involved in different embodiments of the present disclosure described below can be combined with each other as long as they do not conflict with each other.
[0066] In practical applications, text recognition is usually performed on text images through OCR for data collation. However, during the recognition process of technologies such as OCR, the error correction of recognition results is usually limited to single-line text content and it is difficult to consider the context as a whole, so it is impossible to take overall consideration of the full text. Therefore, some specific content often has inaccurate recognition problems due to reasons such as sparse distribution.
[0067] Based on the deficiencies of the above related technologies, in the embodiments of the present disclosure, a method, device, electronic device, storage medium and program product for text correction are provided, aiming to solve the problem of difficult accurate text error correction.
[0068] In the embodiments of the present disclosure, a method for text correction is provided. This method can be applied to an electronic device. The present disclosure does not limit the type of the electronic device, and it can be any device type suitable for implementation, such as a terminal device and a server, etc. The present disclosure will not elaborate on this.
[0069] In the embodiments of the present application, model training is first performed, and then text correction is performed through the trained vector conversion model.
[0070] Refer to Figure 1 As shown, it is a flowchart of a method for training a vector conversion model in the embodiments of the present disclosure. The following will describe this method in combination with Figure 1 This method will be described below. The specific implementation process of this method is as follows:
[0071] Step 101: Perform target object recognition on multiple initial samples in the target task training data respectively, and obtain the target objects of the specified type included in each initial sample.
[0072] In one implementation, obtain the target task training data containing multiple initial samples, and perform target object recognition on each initial sample respectively through a target object recognition model or other means, and obtain the target objects of the specified type included in each initial sample.
[0073] For example, the initial samples can be open-source e-books, etc. The specified type of target object can be set according to the actual application scenario. The target object can be an entity in the text. For example, the target object can be a person's name and a work name, etc. The target object recognition model can be a Named Entity Recognition (NER) model. In practical applications, the target object recognition model can be selected according to the actual application scenario, and no limitation is made here.
[0074] In practical applications, both the specified type and the target task training data can be selected specifically according to the actual target task scenario, and no limitation is made here.
[0075] Step 102: Perform masking processing on each target object in each initial sample to obtain multiple target samples.
[0076] For example, the recognized person names and work names can be masked (MASK) processed, that is, marked with [MASK] at their positions in the text.
[0077] Step 103: According to the masking positions in each target sample, construct multiple text construction data corresponding to each target sample respectively.
[0078] In one implementation, when performing Step 103, the following steps can be adopted:
[0079] S1031: According to the masking positions in each target sample, split each target sample respectively to obtain multiple text segments corresponding to each target sample respectively.
[0080] In the embodiments of the present disclosure, only one masking position in one target sample is taken as an example to illustrate the interception of text segments. A text segment with a specified length is intercepted from the target sample, and a masking position in the target sample is taken as the center.
[0081] For example, from the target sample, a text segment with a length of 512 words (i.e., the specified length) and centered on a certain [MASK] is intercepted.
[0082] Wherein, each text segment contains at least one masking position corresponding to a target object. The lengths of different text segments can be the same or different.
[0083] In practical applications, the specified length can be set according to the actual application scenario and is not limited here. The specified length can be determined according to the input length of the vector conversion model.
[0084] S1032: For each text segment in each target sample respectively, generate corresponding text construction data according to the text segment, the target object masked in the text segment, and the negative sample label; the negative sample label is different from the target object masked in the text segment.
[0085] Optionally, the data format of the text construction data can be expressed as: {"query": str, "pos": List[str], "neg": List[str]}.
[0086] Wherein, "query": str represents the text segment parameter: the text segment, "pos": List[str] represents the positive sample label variable: the positive sample label, and "neg": List[str] represents the negative sample label variable: the negative sample label.
[0087] Among them, the positive sample labels and negative sample labels can both be one or more, and there is no limit here. The positive sample label is the target object masked in the text segment.
[0088] Step 104: Construct data according to the text corresponding to each target sample, and fine-tune the pre-trained model to obtain a trained vector conversion model; the pre-trained model is a model for vector conversion trained based on general data.
[0089] In one implementation, the contrast loss can be determined by comparing the positive sample labels, negative sample labels, and their corresponding model output results in the text construction data, and the model parameters can be adjusted according to the contrast loss until a trained vector conversion model is obtained.
[0090] Among them, the pre-trained model can be the trained bge-large-zh 1.5 model. The bge-large-zh 1.5 model is a Chinese text embedding model designed to convert Chinese text into high-dimensional vector representations. The text vectors obtained by text conversion can capture the semantic information in the text, making texts with similar semantics closer in the vector space, thus facilitating tasks such as similarity search and clustering. In addition to using the bge-large-zh 1.5 model, other pre-trained models can also be considered, such as Bidirectional Encoder Representations from Transformers (BERT) and Generative Pretrained Transformer (GPT), etc. These models have wide applications in the field of natural language processing, can capture the deep semantic information of texts, and are suitable for the recognition of names and work names in OCR. In practical applications, the pre-trained model can be selected according to the actual application scenario, and there is no limit here.
[0091] Among them, fine-tuning refers to the process of continuing to train the model for a specific task on the basis of the pre-trained model in combination with transfer learning. The purpose of doing this is to utilize the knowledge learned by the pre-trained model on large-scale data (i.e., general data), and at the same time adjust the model to adapt to the new task or dataset. Fine-tuning is usually carried out on small-scale and task-specific data, so that the model can learn the characteristics and rules of the specific task.
[0092] Among them, one of the main advantages of transfer learning is that it can reduce the training data requirements and improve the training efficiency. In traditional machine learning, a large amount of labeled data is required to train a model. However, in transfer learning, the pre-trained model has already learned the basic features, so less data is needed to train for a specific task. This is very beneficial for situations where a model needs to be quickly deployed or where training resources are limited.
[0093] In the embodiments of the present disclosure, the model training process includes selecting a pre-trained model, preprocessing the training data, and fine-tuning the model, and then an optimized vector conversion model can be obtained. Taking the bge-large-zh 1.5 model, etc. as the pre-trained model, this pre-trained model can convert Chinese text into a high-dimensional vector representation, capture the semantic information in the text, so that it can combine the context information to achieve the proximity of semantically similar texts in the vector space. By fine-tuning the pre-trained model, a trained vector conversion model is obtained, so that on the basis of maintaining the original pre-trained knowledge, the vector conversion model further optimizes the representation ability of the context information of the entire text recognition target object, improves the accuracy of semantic embedding, enables the model to adapt to new tasks or data sets, and improves the performance of the model on specific tasks.
[0094] In this way, after the model training is completed, the vector conversion of the text can be performed through the trained vector conversion model.
[0095] Refer to Figure 2 As shown, it is a flowchart of a text correction method in the embodiments of the present disclosure. The specific implementation process of this method is as follows:
[0096] Step 201: Mask each of the multiple target objects of the specified type included in the text to be processed to obtain the target text.
[0097] Optionally, the text to be processed can be the image recognition text obtained by recognizing an image through OCR technology, or other types of texts such as books.
[0098] In one implementation manner, the implementation process of step 201 further includes:
[0099] Perform text recognition on the image to be processed to obtain the text to be processed; determine the target objects of the specified type in the target text through the target object recognition model.
[0100] Among them, for text recognition and target object recognition, reference can be made to step 101 above.
[0101] Step 202: According to the mask positions in the target text, intercept the target file into multiple text segments containing masks.
[0102] Among them, when performing step 202, it can be based on a principle similar to the text segment interception in step 103, which will not be elaborated here.
[0103] Step 203: According to the vector conversion model, perform vector conversion on each text segment respectively to obtain text vectors corresponding to each text segment respectively.
[0104] In one implementation, each text segment is input into the vector conversion model to obtain text vectors (embeddings) corresponding to each text segment respectively.
[0105] Among them, the vector conversion model can start from the entire text to be processed (such as, the whole book), make full use of the context information of the target object (such as, an entity) in the entire text to be processed, and represent the masked target object, so as to ensure that the cosine similarity of vectors of different text segments containing the same target object is higher.
[0106] Step 204: Perform unsupervised clustering processing on each text vector to obtain at least one text vector cluster and its corresponding reference text vector respectively.
[0107] In one implementation, through an unsupervised clustering algorithm, perform unsupervised clustering interception on each text vector to obtain one or more text vector clusters, and the representative reference text vectors in each text vector cluster can be identified.
[0108] Among them, a text vector cluster includes multiple text vectors. Optionally, the unsupervised clustering algorithm can adopt the Affinity Propagation algorithm. The Affinity Propagation algorithm is a clustering algorithm based on graph theory, aiming to identify clusters and exemplars in the data, that is, the text vector clusters and their corresponding reference text vectors in the embodiments of the present disclosure. Different from traditional clustering algorithms, the Affinity Propagation does not require specifying the number of clusters in advance, nor does it require randomly initializing the cluster centers, but obtains the final clustering result by calculating the similarity between data points (that is, the text vectors in the embodiments of the present disclosure).
[0109] Optionally, the unsupervised clustering algorithm can also adopt the density-based clustering algorithm (Density-Based Spatial Clustering of Applications with Noise, DBSCAN), or the K-means algorithm. Different unsupervised clustering algorithms have their own advantages when dealing with large-scale data sets. For example, K-means has better results when dealing with spherical data sets, while DBSCAN is suitable for dealing with data sets with different densities. In practical applications, the unsupervised clustering algorithm can be set according to the actual application scenario, which is not limited here.
[0110] In this way, using an unsupervised clustering algorithm such as the Affinity Propagation algorithm for unsupervised clustering can automatically select the number of clusters from the data, without the need to specify the number of clusters in advance, nor the need to randomly initialize the cluster centers. Instead, the final clustering result is obtained by calculating the similarity between data points, improving the accuracy and efficiency of clustering.
[0111] Step 205: Respectively, according to the reference object corresponding to each reference text vector, correct the target objects corresponding to the text vectors included in the corresponding text vector cluster.
[0112] In one implementation manner, when executing step 205, for each target object corresponding to each text vector cluster, the following steps are executed:
[0113] S2051: Determine the similarity between the reference object corresponding to the text vector cluster and the target object.
[0114] S2052: If it is determined that the similarity meets the text correction condition, then correct the target object to the reference object corresponding to the text vector cluster.
[0115] The following combines Figure 3 , and gives an example explanation of the method for correcting the target object. Refer to Figure 3 shown, which is a flowchart of a method for correcting the target object. The specific implementation process of this method includes:
[0116] Step 301: Determine the number of characters with differences between the reference object corresponding to the text vector cluster and the target object; the number is negatively correlated with the similarity.
[0117] Step 302: If the number is not higher than the number threshold, it is determined that the text correction condition is met.
[0118] For example, set the quantity threshold to 1, the reference object is the work name "Remembrance of Things Past", and the work name corresponding to a text vector in the text vector cluster is "Remembrance of Flowing Water Years". Then the number of different characters between the two is 1, which is not higher than the quantity threshold 1, so it is determined that the text correction condition is met.
[0119] Optionally, the character can also be a Chinese character. That is, it is possible to determine whether the text correction condition is met according to the number of different Chinese characters.
[0120] After clustering, use the reference object (such as a person's name or a work name) corresponding to the representative point (i.e., the reference object) in each category (i.e., the text vector cluster) as the standard answer. If the number of different characters between other target objects in this category and the standard answer does not exceed a quantity threshold gamma, then correct them to the standard answer. By introducing an unsupervised clustering algorithm, it is possible to automatically select the number of clusters from the data, thereby more accurately correcting the errors of target objects (such as misspelled person's names and work names) in the text to be processed (such as OCR recognition text).
[0121] For example, "Spring River Flower Moon Night" appears 7 times in a book, 5 times are correct, and 2 times are incorrect, which are: "Spring River Flower Moon Night", "River Water Reflects Moonlight", and "Flower Fragrance Drifts Far Away with the Wind". In the beautiful scenery of "Spring River Flower Moon Night", we stroll by the river, feeling the tranquility of nature. "Spring River Flower Moon Night" is an eternal theme in the poet's works and an immortal scene on the painter's canvas. That night, in the "Spring River Flower Moon Night", the moonlight was like water and the flowers were in full bloom, making people linger. "Spring River Flower Moon Night" is a poem passed down through the ages and a heart-stirring painting. In the romantic atmosphere of "Spring River Flower Moon Night", lovers made eternal vows. "Spring River Flower Moon Night" is not only a display of the beauty of nature but also people's longing for a better life in their hearts. Then it is possible to perform target object recognition, masking, text segment extraction, text vector conversion, clustering, and text correction based on the clustering results on this book. It is determined that "Spring River Flower Moon Night" and "Spring River Flower Moon Night" differ from "Spring River Flower Moon Night" by only one character, so both "Spring River Flower Moon Night" and "Spring River Flower Moon Night" can be corrected to "Spring River Flower Moon Night".
[0122] An application scenario of an embodiment of the present disclosure is a scenario of text correction for text recognized based on OCR technology, which can be applied to fields such as e-book digitization and document management. It can optimize the OCR recognition result, ensure the consistency of target object recognition, and maintain the integrity and accuracy of data. It is particularly suitable for fields such as law, finance, and medical care where extremely high requirements are placed on data accuracy. In practical applications, it can also be applied to other text correction scenarios, which will not be elaborated here.
[0123] In the embodiments of the present disclosure, the target object in the text to be processed is masked, and the masked target text is vector-converted through a vector conversion model, so that the semantic information of each text segment in the target text can be accurately captured through contrastive learning of context information. Furthermore, text segments with similar semantics will be closer in the vector space, facilitating tasks such as similarity search and clustering. In addition, an unsupervised clustering algorithm is combined to cluster and intercept the text vectors of multiple text segments of the target text output by the vector conversion model, improving the accuracy and efficiency of clustering. Moreover, the target object is corrected through the representative points of the clustering, which can reduce problems such as inaccurate and inconsistent target object recognition, and improve the efficiency of data collation.
[0124] Furthermore, transfer learning and context contrast training techniques are combined, that is, a pre-trained model is fine-tuned using specific target task data. Thus, based on the knowledge of the pre-trained model, the model can be optimized for specific target tasks. The trained vector conversion model can quickly adapt to new target tasks, improving the accuracy of semantic embedding, reducing the consumption of training data and other resources. It is particularly suitable for situations where the data annotation cost is high or the data is scarce, because it can further optimize the representation ability of the context information of the target object in the whole text while maintaining the original pre-trained knowledge, reducing the training data requirements and improving the training efficiency. It is particularly suitable for scenarios with limited resources and scarce data, having broad application prospects and significant practical value.
[0125] In addition, in the OCR text recognition and error correction scenario, the target object is recognized through a target object recognition model, followed by masking, unsupervised clustering, and error correction in sequence, which can optimize the recognition result of the OCR recognized text, improve the automation degree of OCR text recognition and error correction, increase the efficiency, reduce manual intervention, and lower the error rate.
[0126] Based on the same inventive concept, an apparatus for text correction is also provided in the embodiments of the present disclosure. Since the principle of the above apparatus and device for solving problems is similar to that of a text correction method, the implementation of the above apparatus can refer to the implementation of the method, and the repeated parts will not be elaborated. The apparatus can be applied to an electronic device. The present disclosure does not limit the type of the electronic device, which can be any suitable device type, such as a terminal device and a server, etc., and the present disclosure will not elaborate on this. The apparatus embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful apparatus, it is formed by the processor of the electronic device where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory for operation.
[0127] Refer to Figure 4As shown, it is a structural block diagram of a text correction device in an embodiment of the present disclosure. In some embodiments, the text correction device of the present disclosure example includes:
[0128] A masking unit 401, configured to perform masking processing on multiple target objects of specified types included in the text to be processed respectively, to obtain a target text;
[0129] A truncation unit 402, configured to truncate the target file into multiple text segments including masks according to the mask positions in the target text;
[0130] A conversion unit 403, configured to perform vector conversion on each text segment respectively according to a vector conversion model to obtain text vectors corresponding to each text segment respectively; the vector conversion model is obtained by fine-tuning a pre-trained model;
[0131] A clustering unit 404, configured to perform unsupervised clustering processing on each text vector to obtain at least one text vector cluster and its corresponding reference text vector respectively;
[0132] A correction unit 405, configured to correct the target objects corresponding to the text vectors included in the corresponding text vector cluster respectively according to the reference objects corresponding to each reference text vector.
[0133] In one embodiment, the masking unit 401 is further configured to:
[0134] Perform text recognition on the image to be processed to obtain the text to be processed;
[0135] Determine the target objects of the specified type in the target text.
[0136] In one embodiment, the conversion unit 403 is further configured to:
[0137] Perform target object recognition on multiple initial samples in the target task training data respectively to obtain the target objects of the specified type included in each initial sample respectively;
[0138] Perform masking processing on each target object in each initial sample respectively to obtain multiple target samples;
[0139] Construct multiple text construction data corresponding to each target sample respectively according to the mask positions in each target sample;
[0140] Fine-tune the pre-trained model according to the text construction data corresponding to each target sample respectively to obtain a trained vector conversion model; the pre-trained model is a model for vector conversion trained based on general data.
[0141] In one embodiment, the conversion unit 403 is further configured to:
[0142] Segment each target sample according to the masked positions in the target samples, to obtain multiple text segments respectively corresponding to each target sample;
[0143] For each text segment in each target sample, generate corresponding text construction data according to the text segment, the target object masked in the text segment, and the negative sample label; the negative sample label is different from the target object masked in the text segment.
[0144] In one implementation, the correction unit 405 is configured to:
[0145] For each target object corresponding to each text vector cluster, perform the following steps:
[0146] Determine the similarity between the reference object corresponding to the text vector cluster and the target object;
[0147] If it is determined that the similarity meets the text correction condition, correct the target object to the reference object corresponding to the text vector cluster.
[0148] In one implementation, the correction unit 405 is configured to:
[0149] Determine the number of characters where there are differences between the reference object corresponding to the text vector cluster and the target object; the number is negatively correlated with the similarity;
[0150] The correction unit 405 is configured to:
[0151] If the number is not higher than the number threshold, it is determined that the text correction condition is met.
[0152] The method for text correction in the embodiments of the present disclosure includes performing masking processing on multiple target objects of specified types included in the text to be processed, to obtain the target text; segmenting the target file according to the masked positions in the target text, to obtain multiple text segments including masks; performing vector conversion on each text segment according to the vector conversion model, to obtain text vectors respectively corresponding to each text segment; performing unsupervised clustering processing on the text vectors, to obtain at least one text vector cluster and its corresponding reference text vectors respectively; and correcting the target objects corresponding to the text vectors included in the corresponding text vector cluster respectively according to the reference object corresponding to each reference text vector. In this way, the vector conversion model and the unsupervised clustering algorithm can be combined to correct the target objects in the text, and the problem of difficult text correction is solved.
[0153] The method for text correction in the embodiments of the present disclosure includes
[0154] In the embodiments of the present disclosure, there is also provided an electronic device, including:
[0155] A processor; and
[0156] A memory storing computer instructions for causing the processor to execute the method according to any of the above embodiments.
[0157] In an embodiment of the present disclosure, a computer-readable storage medium is provided, storing computer instructions for causing a computer to execute the method according to any of the above embodiments.
[0158] An embodiment of the present disclosure further provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the method according to any of the above embodiments.
[0159] Figure 5 A schematic structural diagram of an electronic device 5000 is shown. Refer to Figure 5 As shown, the electronic device 5000 includes: a processor 5010 and a memory 5020. Optionally, it may further include a power supply 5030, a display unit 5040, and an input unit 5050.
[0160] The processor 5010 is the control center of the electronic device 5000, connecting each component through various interfaces and lines, and executing various functions of the electronic device 5000 by running or executing software programs and / or data stored in the memory 5020, so as to perform overall monitoring of the electronic device 5000.
[0161] In an embodiment of the present disclosure, when the processor 5010 calls the computer program stored in the memory 5020, it executes each step in the above embodiment.
[0162] Optionally, the processor 5010 may include one or more processing units; preferably, the processor 5010 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, applications, etc., and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 5010. In some embodiments, the processor and the memory may be implemented on a single chip, and in some embodiments, they may also be separately implemented on independent chips.
[0163] The memory 5020 may mainly include a program storage area and a data storage area. Among them, the program storage area may store the operating system, various applications, etc.; the data storage area may store data created according to the use of the electronic device 5000, etc. In addition, the memory 5020 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices, etc.
[0164] The electronic device 5000 further includes a power supply 5030 (such as a battery) for powering each component. The power supply can be logically connected to the processor 5010 through a power management system, so as to manage functions such as charging, discharging, and power consumption through the power management system.
[0165] The display unit 5040 can be used to display information input by the user or information provided to the user, as well as various menus of the electronic device 5000. In the embodiments of the present disclosure, it is mainly used to display the display interfaces of various applications in the electronic device 5000 and objects such as text and pictures displayed in the display interfaces. The display unit 5040 may include a display panel 5041. The display panel 5041 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.
[0166] The input unit 5050 can be used to receive information such as numbers or characters input by the user. The input unit 5050 may include a touch panel 5051 and other input devices 5052. Among them, the touch panel 5051, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 5051).
[0167] Specifically, the touch panel 5051 can detect the touch operation of the user, detect the signals brought by the touch operation, convert these signals into contact coordinates, send them to the processor 5010, and receive and execute the commands sent by the processor 5010. In addition, the touch panel 5051 can be implemented in multiple types such as resistive, capacitive, infrared, and surface acoustic wave. The other input devices 5052 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.
[0168] Of course, the touch panel 5051 can cover the display panel 5041. After the touch panel 5051 detects a touch operation on or near it, it transmits it to the processor 5010 to determine the type of touch event. Subsequently, the processor 5010 provides a corresponding visual output on the display panel 5041 according to the type of touch event. Although in Figure 5 the touch panel 5051 and the display panel 5041 are implemented as two independent components to realize the input and output functions of the electronic device 5000, in some embodiments, the touch panel 5051 and the display panel 5041 can be integrated to realize the input and output functions of the electronic device 5000.
[0169] The electronic device 5000 may further include one or more sensors, such as a pressure sensor, a gravitational acceleration sensor, a proximity light sensor, etc. Of course, according to the needs in specific applications, the above-mentioned electronic device 5000 may further include other components such as a camera. Since these components are not the key components used in the embodiments of the present disclosure, therefore, in Figure 5 it is not shown and will not be elaborated further.
[0170] Those skilled in the art can understand that Figure 5 this is merely an example of an electronic device and does not constitute a limitation on the electronic device. It may include more or fewer components than shown in the figure, or combine certain components, or different components.
[0171] For the convenience of description, the above parts are divided into respective modules (or units) according to functions and described separately. Of course, when implementing the present disclosure, the functions of the respective modules (or units) can be implemented in the same or multiple software or hardware.
Claims
1. A method for text correction, characterized in that, The method includes: Performing mask processing on multiple target objects of specified types included in the text to be processed respectively to obtain a target text; According to the mask positions in the target text, segmenting the target file to obtain multiple text segments containing masks; According to a vector conversion model, performing vector conversion on each text segment respectively to obtain text vectors corresponding to each text segment respectively; Performing unsupervised clustering processing on each text vector to obtain at least one text vector cluster and its corresponding reference text vectors respectively; Respectively according to the reference objects corresponding to each reference text vector, correcting the target objects corresponding to the text vectors included in the corresponding text vector cluster.
2. The method according to claim 1, wherein Before performing mask processing on multiple target objects of specified types included in the text to be processed respectively, the method further includes: Performing text recognition on the image to be processed to obtain the text to be processed; Determining the target objects of the specified types in the target text.
3. The method according to claim 1 or 2, characterized in that, Before performing vector conversion on each text segment respectively according to the vector conversion model to obtain text vectors corresponding to each text segment respectively, the method further includes: Performing target object recognition on multiple initial samples in the target task training data respectively to obtain the target objects of the specified types included in each initial sample respectively; Performing mask processing on each target object in each initial sample respectively to obtain multiple target samples; According to the mask positions in each target sample, constructing multiple text construction data corresponding to each target sample respectively; According to the text construction data corresponding to each target sample respectively, fine-tuning a pre-trained model to obtain the trained vector conversion model; the pre-trained model is a model for vector conversion trained based on general data.
4. The method according to claim 3, wherein The constructing multiple text construction data corresponding to each target sample respectively according to the mask positions in each target sample includes: According to the mask positions in each target sample, segmenting each target sample respectively to obtain multiple text segments corresponding to each target sample respectively; For each text segment in each target sample respectively, generating corresponding text construction data according to the text segment, the target object masked in the text segment, and a negative sample label; the negative sample label is different from the target object masked in the text segment.
5. The method according to claim 1 or 2, characterized in that, The correcting the target objects corresponding to the text vectors included in the corresponding text vector cluster respectively according to the reference objects corresponding to each reference text vector includes: For each target object corresponding to each text vector cluster respectively, performing the following steps: Determining the similarity between the reference object corresponding to the text vector cluster and the target object; If it is determined that the similarity meets the text correction condition, correcting the target object to the reference object corresponding to the text vector cluster.
6. The method according to claim 5, characterized in that, The determining the similarity between the reference object corresponding to the text vector cluster and the target object includes: Determining the number of characters with differences between the reference object corresponding to the text vector cluster and the target object; the number is negatively correlated with the similarity; The determining that the similarity meets the text correction condition includes: If the quantity is not higher than the quantity threshold, it is determined that the text correction condition is met.
7. An apparatus for text correction, characterized in that, The device includes: A masking unit, configured to perform masking processing on multiple target objects of specified types included in the text to be processed respectively, to obtain a target text; A truncating unit, configured to truncate the target file into multiple text segments containing masks according to the mask positions in the target text; A conversion unit, configured to perform vector conversion on each text segment respectively according to a vector conversion model, to obtain text vectors respectively corresponding to the text segments; the vector conversion model is obtained by fine-tuning a pre-trained model; A clustering unit, configured to perform unsupervised clustering processing on the text vectors, to obtain at least one text vector cluster and a reference text vector respectively corresponding thereto; A correction unit, configured to correct the target objects corresponding to the text vectors included in the corresponding text vector cluster respectively according to the reference objects corresponding to each reference text vector.
8. An electronic device, characterized in that, Includes: A processor; And A memory, storing computer instructions, the computer instructions being used to cause the processor to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, Storing computer instructions, the computer instructions being used to cause a computer to execute the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, Includes computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, when the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Text error correction method, text error correction device, storage medium and electronic equipment
CN111046652A
Text error correction method and device, storage medium and electronic device
CN112861518A