Text chunking method, device, electronic device and storage medium
By detecting and tiling text images in photo translation technology, combining position, area images and semantic features, the problem of semantic breakage in single-line text translation is solved, and more accurate and coherent translation results are achieved.
Patent Information
- Application Number
- CN202111570376.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-21
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-21
AI Technical Summary
In the existing photo translation technology, directly translating a single line of text cannot obtain semantic context information with the current text line, resulting in semantic breakage and inconsistency in the translation results, affecting accuracy and actual effects.
By performing text detection on the translated text image, the positions of each text line are determined, and then text blocking is performed based on the position, area image and recognition text to obtain text blocking results. This method combines position features, area image features and semantic features to evaluate the statement integrity and semantic coherence of text lines, thereby performing text chunking.
By connecting text lines with semantic relationships together, the execution effect of translation tasks and the accuracy of translation results are improved, and the actual experience of users and the evaluation indicators of photo translation technology are improved.
Smart Images

Figure CN114255466B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic information technology, and in particular, to a text chunking method, apparatus, electronic device, and storage medium. Background Art
[0002] The photo translation technology occupies a very important position in people's daily lives, bringing great convenience to people's lives and studies, enabling people to easily obtain relatively familiar text when facing unfamiliar pictures and texts.
[0003] In the current photo translation process, most directly translate the single-line text recognized by OCR and present the translation result to the user. However, directly translating single-line text cannot obtain the context information semantically related to the text content in the current text line, which will cause the translation result to be unable to connect with the context, and there will be semantic breaks and unsmooth sentences in the translated text, affecting the accuracy and actual effect of the translation. Summary of the Invention
[0004] The present invention provides a text chunking method, apparatus, electronic device, and storage medium to solve the defects in the prior art that directly translating a single text line results in low accuracy and poor actual effect of the translation result.
[0005] The present invention provides a text chunking method, including:
[0006] Performing text detection on the text image to be chunked to obtain the positions of each text line in the text image;
[0007] Based on the positions of each text line, the regional image of each text line in the text image, and the recognized text of each text line, performing text chunking on the text image to obtain a text chunking result.
[0008] According to the text chunking method provided by the present invention, the performing text chunking on the text image based on the positions of each text line, the regional image of each text line in the text image, and the recognized text of each text line to obtain a text chunking result includes:
[0009] Based on the position of any text line, the regional image of the any text line in the text image, and the recognized text of the any text line, determining the text line feature of the any text line;
[0010] Based on the text line features of each text line, performing text chunking on the text image to obtain a text chunking result.
[0011] A text chunking method provided by the present invention, which performs text chunking on the text image based on the text line features of each text line to obtain a text chunking result, includes:
[0012] Based on the positions of each text line, splice the text line features of each text line to obtain a text line feature sequence;
[0013] Based on a decoding dictionary, perform text chunking decoding on the text line feature sequence to obtain the text chunking result. The decoding dictionary includes the text line features of each text line, as well as the encoding features of a start symbol, an end symbol, and a chunking symbol.
[0014] A text chunking method provided by the present invention, which performs text chunking decoding on the text line feature sequence based on a decoding dictionary to obtain the text chunking result, includes:
[0015] Based on the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment, determine the decoding result and the hidden state at the current moment of the text line feature sequence, remove the decoding result from the decoding dictionary, and update the text chunking result based on the decoding result;
[0016] Update the next moment of the current moment to the current moment until the decoding result at the current moment is the encoding feature of the end symbol or the text chunking result reaches a preset length.
[0017] A text chunking method provided by the present invention, which determines the decoding result and the hidden state at the current moment of the text line feature sequence based on the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment, includes:
[0018] Based on the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment, determine the output probability of the text line features of each text line in the decoding dictionary at the current moment;
[0019] Based on the output probability of the text line features of each text line in the decoding dictionary at the current moment, determine the decoding result and the hidden state at the current moment of the text line feature sequence.
[0020] A text chunking method provided by the present invention, which determines the text line feature of any text line based on the position of any text line, the regional image of any text line in the text image, and the recognized text of any text line, includes:
[0021] Encode the position of any text line to obtain the position feature of any text line;
[0022] Based on the position of any one of the text lines, determine the region image features corresponding to the region image of the any one text line in the text image from the image features of the text image;
[0023] Extract semantic features from the recognized text of the any one text line to obtain the semantic features of the any one text line;
[0024] Based on the position features, region image features, and semantic features of the any one text line, determine the text line features of the any one text line.
[0025] According to a text chunking method provided by the present invention, the region images of the respective text lines in the text image and the recognized texts of the respective text lines are determined based on the following steps:
[0026] Based on the positions of the respective text lines, perform image segmentation on the text image to obtain the region images of the respective text lines in the text image;
[0027] Perform text recognition on the region images of the respective text lines in the text image to obtain the recognized texts of the respective text lines.
[0028] The present invention also provides a text chunking device, including:
[0029] A text detection unit for performing text detection on a text image to be chunked to obtain the positions of the respective text lines in the text image;
[0030] A text chunking unit for performing text chunking on the text image based on the positions of the respective text lines, the region images of the respective text lines in the text image, and the recognized texts of the respective text lines to obtain a text chunking result.
[0031] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps of any one of the above-mentioned text chunking methods are implemented.
[0032] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above-mentioned text chunking methods are implemented.
[0033] The text chunking method, device, electronic device, and storage medium provided by the present invention perform text chunking on a text image according to the positions of each text line, the regional images of each text line in the text image, and the recognized text of each text line, obtaining a text chunking result. By evaluating the statement integrity of each text line and the semantic coherence between each text line in the text image from three different perspectives, and performing text chunking based on the evaluation results, it can overcome the defects in the traditional solution of directly translating a single text line, resulting in low accuracy of the translation result and poor actual effect. In the embodiments of the present invention, text lines with semantic relationships can be concatenated together, thus providing strong assistance for the execution of subsequent translation tasks and the improvement of the accuracy rate and actual effect of the translation result, and further promoting the actual experience of users and the improvement process of relevant evaluation indicators of the photo translation technology. Description of the Drawings
[0034] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0035] Figure 1 It is a flowchart of the text chunking method provided by the present invention;
[0036] Figure 2 It is a flowchart of step 120 in the text chunking method provided by the present invention;
[0037] Figure 3 It is a flowchart of step 122 in the text chunking method provided by the present invention;
[0038] Figure 4 It is a flowchart of step 1222 in the text chunking method provided by the present invention;
[0039] Figure 5 It is a flowchart of step 1222-1 in the text chunking method provided by the present invention;
[0040] Figure 6 It is a flowchart of step 121 in the text chunking method provided by the present invention;
[0041] Figure 7 It is a structural diagram of the BiLSTM provided by the present invention;
[0042] Figure 8 It is a schematic diagram of the determination process of the regional image and the recognized text provided by the present invention;
[0043] Figure 9It is the network structure diagram of Attention ED provided by the present invention;
[0044] Figure 10 It is the overall framework diagram of the text chunking method provided by the present invention;
[0045] Figure 11 It is the structural schematic diagram of the text chunking device provided by the present invention;
[0046] Figure 12 It is the structural schematic diagram of the electronic device provided by the present invention. Detailed implementation manners
[0047] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without making creative efforts fall within the scope of protection of the present invention.
[0048] The photo translation technology is a technical means for translating a captured image and converting it into corresponding text, which brings great convenience to people's daily life, work and study. At present, the OCR recognition technology that is in a rapid development stage can directly recognize the text lines in an image, obtain the recognition result, and translate the recognition result, so that people can easily obtain relatively familiar text when facing unfamiliar pictures and texts.
[0049] In the current photo translation technology, first, a text detection model is used to detect the text lines in the text image to obtain the text line detection result; then, based on the text line detection result, the regional images of each text line are cropped from the text image; thereafter, the OCR recognition technology is used to perform text recognition on the regional images of each text line cropped in the text image; finally, the translation model is used to translate each single-line text obtained by text recognition one by one, and the translation result is presented to the user.
[0050] However, in the above solution, when translating the single-line text obtained by text recognition, since the semantic relationship of the text content in a single text line is broken, there will also be semantic breaks and unsmooth sentences in the translation result. At the same time, since the context information having a semantic relationship with the text content in the current text line cannot be obtained, the actual effect and accuracy of the translation are greatly reduced, and as a result, there are likely to be a large number of unsmooth, non-flowing and even incomprehensible sentences in the finally obtained translation result, seriously affecting the actual experience of the user.
[0051] For example, when translating the second line in a certain natural paragraph, since this text line is very likely to be only the second half of a sentence, and the first half of the sentence is the first line in the natural paragraph. At this time, if the second line is translated alone, due to the absence of the first half of the sentence, the resulting translation is very difficult to meet the user's expectations and cannot satisfy the user.
[0052] In view of the above situation, the present invention provides a text chunking method, aiming to concatenate text lines with semantic relationships and then perform translation to improve the accuracy and actual effect of translation. Figure 1 It is a schematic flowchart of the text chunking method provided by the present invention, as Figure 1 shown, the method includes:
[0053] Step 110, perform text detection on the text image to be chunked to obtain the positions of each text line in the text image.
[0054] Specifically, before text chunking, it is first necessary to determine the text image to be chunked, that is, the text image to be processed; subsequently, text detection can be performed on the text image to be chunked to detect the text regions in the text image to be chunked, so as to obtain the positions of each text line in the text image to be chunked. It should be noted that the process of performing text detection on the text image to be chunked here can be implemented through a text detection model. The specific process can be to input the text image to be chunked into the text detection model, and the text detection model extracts features from the input text image and performs text detection based on the image features of the text image obtained by feature extraction, and finally obtains the positions of each text line in the text image output by the text detection model.
[0055] Before inputting the text image to be chunked into the text detection model, a text detection model can also be pre-trained according to the sample text image and the position labels of each text line in the sample text image. The training process of the text detection model includes the following steps: First, collect a large number of sample text images and label the positions of each sample text line in the sample text image to obtain the position labels of each text line; then, based on the sample text image and the position labels of each sample text line in the sample text image, train the initial text detection model to obtain the trained text detection model.
[0056] It should be noted that the initial text detection model here is a single-line text detection model, which can be constructed based on PSENet (Progressive Scale Expansion Network), where PSENet is composed of a residual structure and a feature pyramid structure.
[0057] Step 120: Based on the positions of each text line, the regional images of each text line in the text image, and the recognized text of each text line, perform text chunking on the text image to obtain a text chunking result.
[0058] Considering that the semantic relationship of the text content in a single text line is fragmented, if text chunking is directly performed based on the position of each text line, it is very likely to split text lines with semantic relationships. To avoid the problem of poor translation effects in subsequent translation tasks caused by the fragmentation of semantic relationships, incomplete and incoherent sentence information, in the embodiments of the present invention, after obtaining the positions of each text line, it is also necessary to determine the regional images of each text line in the text image and the recognized text of each text line. Performing text chunking based on these three can completely overcome the above defects. Moreover, connecting text lines with semantic relationships together can provide strong assistance for the execution of subsequent translation tasks and the improvement of the accuracy of translation results.
[0059] Specifically, after obtaining the positions of each text line in the text image through Step 110, if text chunking is to be performed on the text image to be chunked, it is also necessary to determine the regional images of each text line in the text image and the recognized text of each text line.
[0060] Among them, the regional images of each text line in the text image can be cropped from the text image based on the positions of each text line. The specific process can be to perform image segmentation on the text image to be chunked according to the positions of each text line to obtain the regional images of each text line in the text image.
[0061] And the recognized text of each text line can be determined based on the regional images of each text line in the text image. Specifically, text recognition can be performed on the regional images of each text line in the text image to obtain the recognized text of each text line. The text recognition process here can be implemented through a text recognition model or other text recognition methods. The embodiments of the present invention do not make specific limitations on this.
[0062] After determining the regional images of each text line in the text image and the recognized text of each text line, the text image can be text-blocked by combining the positions of each text line and the above two. Specifically, the process can be as follows: First, according to the positions of each text line, the regional images of each text line in the text image, and the recognized text of each text line, the corresponding position features, regional image features, and semantic features are determined respectively; then, based on the position features of each text line, the regional image features corresponding to the regional images of each text line in the text image, and the semantic features of the recognized text of each text line, the text image is text-blocked to obtain the text-blocking result. By combining the features at three different levels to text-block the text image, the semantic relationships between each text line in the text image, the statement integrity of each text line, and the semantic coherence can be fully considered, thereby avoiding the situation of splitting text lines with semantic relationships, and further ensuring a better accuracy of the translation result of the subsequent translation task.
[0063] The text-blocking method provided by the present invention text-blocks the text image according to the positions of each text line, the regional images of each text line in the text image, and the recognized text of each text line, obtains the text-blocking result, evaluates the statement integrity of each text line in the text image and the semantic coherence between each text line from three different perspectives, and text-blocks according to the evaluation result, which can overcome the defects of directly translating a single text line in the traditional solution, resulting in low accuracy of the translation result and poor actual effect. In the embodiments of the present invention, text lines with semantic relationships can be concatenated together, thereby providing strong assistance for the execution of the subsequent translation task and the improvement of the accuracy and actual effect of the translation result, and further promoting the improvement process of the user's actual experience and the relevant evaluation indicators of the photo translation technology.
[0064] Based on the above embodiments, Figure 2 is a schematic flowchart of step 120 in the text-blocking method provided by the present invention, as Figure 2 shown, step 120 includes:
[0065] Step 121, based on the position of any text line, the regional image of this text line in the text image, and the recognized text of this text line, determine the text line feature of this text line;
[0066] Step 122, based on the text line features of each text line, text-block the text image to obtain the text-blocking result.
[0067] Specifically, in step 120, the process of text-blocking the text image according to the positions of each text line, the regional images of each text line in the text image, and the recognized text of each text line specifically includes the following steps:
[0068] First, perform step 121. According to the position of any text line, the regional image of the text line in the text image, and the recognized text of the text line, determine the text line feature of the text line. The specific process can be to determine the corresponding position feature, regional image feature, and semantic feature respectively according to the position of the text line, the regional image of the text line in the text image, and the recognized text of the text line. Then, fuse the position feature, regional image feature, and semantic feature of the text line to obtain the text line feature of the text line. The text line feature obtained in this way can better represent the statement integrity of the text line and the semantic coherence between the text line and the context information.
[0069] Furthermore, for the case of fusing the above three features, the regional image feature of any text line can complement the semantic feature of the text line, making up for the missing features of each other. On this basis, combine the position feature of the text line to determine the text line feature, which ensures the representation ability of the text line feature obtained by this process for the statement integrity of the text line and the semantic coherence with the context information, that is, it can better ensure the statement integrity of the text and the semantic coherence with the context. The fusion method of these three can be splicing, addition, weighted fusion, etc., and the embodiments of the present invention do not make specific limitations on this. Preferably, in the embodiments of the present invention, the fusion method is selected as splicing, that is, splice the position feature, regional image feature, and semantic feature of any text line to obtain the text line feature of the text.
[0070] Immediately afterwards, step 122 can be executed. According to the text line features of each text line, perform text chunking on the text image to obtain the text chunking result. This process can be implemented through a text chunking decoding network. Specifically, input the text line features of each text line into the text chunking decoding network, and the text chunking decoding network performs text chunking decoding according to the input text line features of each text line to obtain the text chunking result. It should be noted that the text chunking decoding network here can be an LSTM (Long Short-Term Memory) decoding network.
[0071] Based on the above embodiments, Figure 3 is a schematic flowchart of step 122 in the text chunking method provided by the present invention. As Figure 3 shown, step 122 includes:
[0072] Step 1221, based on the positions of each text line, splice the text line features of each text line to obtain a text line feature sequence.
[0073] Step 1222: Based on the decoding dictionary, perform text chunk decoding on the text line feature sequence to obtain the text chunk result. The decoding dictionary includes the text line features of each text line, as well as the encoding features of the start symbol, end symbol, and chunk symbol.
[0074] Specifically, in step 122, the process of text chunking the text image according to the text line features of each text line may specifically include the following steps:
[0075] First, execute step 1221 to splice the text line features of each text line to obtain the text line feature sequence. It should be noted that the process of feature splicing here needs to be guided by the positions of each text line, that is, according to the sequence and / or vertical order of the arrangement positions of each text line in the text image, splice the text line features of each text line to obtain the text line feature sequence;
[0076] Immediately afterwards, execute step 1222. According to the text line features of each text line and the encoding features of the start symbol, end symbol, and chunk symbol, construct a decoding dictionary; according to the constructed decoding dictionary, perform text chunk decoding on the text line feature sequence to obtain the text chunk result. The process of decoding the text line feature sequence here is actually a process of predicting the output probability of each element in the text line feature sequence at the current moment. The greater the output probability at the current moment, the greater the possibility that the element belongs to the text line feature that should be output at the current moment; conversely, the smaller the output probability at the current moment, the smaller the possibility that the element belongs to the text line feature that should be output at the current moment; finally, the text chunk result can be determined according to the output probabilities predicted at each moment.
[0077] Based on the above embodiments, the constructed decoding dictionary can be expressed as:
[0078] [h 1 , h 2 , h 3 , …, h i , h s , h e , h j 1≤i≤n
[0079] where h i represents the text line feature of the i-th text line, n represents the number of text lines in the text image, h s represents the encoding feature of the start symbol, h e represents the encoding feature of the end symbol, h j represents the encoding feature of the chunk symbol. Here, h s , h e , h jThey are all parameters of the text block decoding network and are learned during the training process of the text block decoding network.
[0080] Based on the above embodiments, Figure 4 It is a schematic flowchart of step 1222 in the text block method provided by the present invention. As Figure 4 shown, step 1222 includes:
[0081] Step 1222-1: Based on the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment, determine the decoding result and hidden state of the text line feature sequence at the current moment, remove the decoding result from the decoding dictionary, and update the text block result based on the decoding result;
[0082] Step 1222-2: Update the next moment of the current moment to the current moment until the decoding result at the current moment is the encoding feature of the end symbol or the text block result reaches the preset length.
[0083] Specifically, in step 1222, when performing text block decoding on the text line feature sequence according to the decoding dictionary, the correlation between the text line features of each text line and the hidden state at the previous moment can be further considered. This correlation can better represent the semantic association degree between each text line and the text line corresponding to the text line feature output at the previous moment, that is, it can more accurately evaluate the semantic relationship between each text line and the context information. The process of text block based on this correlation specifically includes the following steps:
[0084] Step 1222-1: First, determine the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment, that is, determine the attention weight of each text line feature in the decoding dictionary at the current moment, and this attention weight is determined by the text line feature of the corresponding text line, the dimension of the text line feature, and the hidden state at the previous moment;
[0085] Immediately afterwards, based on the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment, determine the decoding result and hidden state of the text line feature sequence at the current moment;
[0086] After that, the text block result can be updated according to the decoding result at the current moment, that is, add the text line corresponding to the text line feature output at the current moment to the text block result at the previous moment to obtain the updated text block result. At the same time, the decoding dictionary also needs to be updated according to the decoding result at the current moment, that is, remove the decoding result at the current moment from the decoding dictionary to obtain the updated decoding dictionary;
[0087] Execute step 1222-2 to update the next moment of the current moment to the current moment. At this time, the hidden state of the current moment obtained in step 1222-1 has been updated to the hidden state of the previous moment, and the decoding result obtained in step 1222-1 has also been removed from the decoding dictionary. At this time, step 1222-1 and step 1222-2 can be executed again until the decoding result of the current moment is the encoding feature of the end symbol, indicating that the decoding process of the text line feature sequence has come to an end, that is, the decoding is completed, or until the text chunking result reaches the preset length, indicating that the text chunking process of the text image is over, that is, each text line in the text image has been chunked. The preset length here is the length of the decoding dictionary and is determined according to the number of text lines in the text image.
[0088] It should be noted that the process of text chunking and decoding the text line feature sequence here can be implemented by a text chunking decoding network. The text chunking decoding network can be an LSTM decoding network. The LSTM decoding network decodes the serial number of the text line feature that should be output at the current moment according to the serial number of the text line feature output at the previous moment and the positions of each text line in the text image. At this time, the text chunking result of the previous moment can be updated according to the serial number of the text line feature output at the current moment to obtain the text chunking result of the current moment. Then, the text line feature corresponding to the serial number output at the current moment is removed from the decoding dictionary, and the next moment of the current moment is updated to the current moment until the LSTM decoding network outputs the serial number corresponding to the encoding feature of the end symbol or the text chunking result reaches the preset length.
[0089] When performing text chunking and decoding on the text line feature sequence, the LSTM decoding network outputs the serial number corresponding to the encoding feature of the start symbol at the initial moment; when the LSTM decoding network outputs the serial number corresponding to the encoding feature of the chunking symbol, it indicates that one process of the text chunking process of the text image has been completed, that is, the segmentation of one text chunk has been completed. Subsequently, the next text chunking process can be started until the LSTM decoding network outputs the serial number corresponding to the encoding feature of the end symbol or the text chunking result reaches the preset length. Thus, the text chunking and decoding process of the text line feature sequence is completed.
[0090] Based on the above embodiments, Figure 5 is a schematic flowchart of step 1222-1 in the text chunking method provided by the present invention. As Figure 5 shown, in step 1222-1, based on the correlation between the text line features of each text line in the decoding dictionary and the hidden state of the previous moment, determine the decoding result and hidden state of the text line feature sequence at the current moment, including:
[0091] Step 1222-11: Determine the output probability of the text line features of each text line in the decoding dictionary at the current moment based on the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment.
[0092] Step 1222-12: Determine the decoding result and the hidden state of the text line feature sequence at the current moment based on the output probability of the text line features of each text line in the decoding dictionary at the current moment.
[0093] Specifically, in Step 1222-1, the process of determining the decoding result and the hidden state of the text line feature sequence at the current moment according to the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment may specifically include the following steps:
[0094] Step 1222-11: First, determine the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment; then, based on the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment, determine the output probability of the text line features of each text line in the decoding dictionary at the current moment. The output probability at the current moment can represent the likelihood of the corresponding text line feature being output at the current moment. The greater the output probability at the current moment, the stronger the correlation between the text line feature and the text line feature output at the previous moment, that is, the greater the likelihood that the text line feature belongs to the text line feature that should be output at the current moment; conversely, the smaller the output probability at the current moment, the weaker the correlation between the text line feature and the text line feature output at the previous moment, that is, the smaller the likelihood that the text line feature belongs to the text line feature that should be output at the current moment.
[0095] After determining the output probability of each text line feature at the current moment, Step 1222-12 can be executed. According to the output probability of the text line features of each text line in the decoding dictionary at the current moment, determine the decoding result and the hidden state of the text line feature sequence at the current moment, that is, determine the maximum output probability from the output probabilities of the text line features of each text line at the current moment, and use the text line feature corresponding to the maximum output probability as the decoding result at the current moment; then, according to the decoding results of each moment, the text segmentation result can be determined.
[0096] Based on the above embodiments, the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment, and the output probability of the text line features of each text line in the decoding dictionary at the current moment can be expressed by the following formula:
[0097] Among them, the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment can be calculated by the following formula:
[0098]
[0099] Among them, a t,i represents the attention weight of the text line feature of the i-th text line in the decoding dictionary at time t, that is, the correlation between the text line feature of the i-th text line and the hidden state at the previous moment of time t, q t represents the hidden state at the previous moment of time t, and d represents the dimension of the text line feature of the i-th text line.
[0100] The output probability of the text line feature of each text line in the decoding dictionary at the current moment can be calculated from the correlation between the text line feature of each text line in the decoding dictionary and the hidden state at the previous moment, and its calculation formula can be expressed as:
[0101]
[0102] Among them, p i represents the output probability of the text line feature of the i-th text line in the decoding dictionary at time t, exp represents the exponential function with the natural constant e as the base, exp(a t,i ) is l is the length of the decoding dictionary, that is, the preset length.
[0103] Based on the above embodiments, Figure 6 is the schematic flowchart of step 121 in the text chunking method provided by the present invention. As Figure 6 shown, step 121 includes:
[0104] Step 1211, encoding the position of the text line to obtain the position feature of the text line;
[0105] Step 1212, based on the position of the text line, determining the region image feature corresponding to the region image of the text line in the text image from the image features of the text image;
[0106] Step 1213, extracting semantic features from the recognized text of the text line to obtain the semantic features of the text line;
[0107] Step 1214, determining the text line feature of the text line based on the position feature, region image feature, and semantic feature of the text line.
[0108] Specifically, in step 121, the process of determining the text line feature of a text line according to the position of any text line, the region image of the text line in the text image, and the recognized text of the text line specifically includes the following steps:
[0109] First, execute step 1211, step 1212, and step 1213;
[0110] Step 1211: Encode the position of the text line to obtain the position feature of the text line.
[0111] Step 1212: According to the position of the text line, determine the region image feature corresponding to the region image of the text line in the text image from the image features of the text image. It should be noted that the region image features extracted from the image features of the text image contain the global information of the text image, that is, they can represent some attributes of the text image to a certain extent.
[0112] Step 1213: Extract semantic features from the recognized text of the text line to obtain the semantic features of the text line. It should be noted that the recognized text of the text line can be determined based on the region image of the text line in the text image, that is, perform text recognition on the region image of the text line in the text image to obtain the recognized text of the text line.
[0113] Here, the process of extracting semantic features from the recognized text can be implemented by a semantic extraction model. The specific process can be to input the recognized text of each text line into the semantic extraction model, and the semantic extraction model extracts semantic features from the input recognized text of each text line to obtain the semantic features of each text line output by the semantic extraction model.
[0114] Before inputting the recognized text of each text line into the semantic extraction model, a semantic extraction model can also be pre-trained according to the sample text and the semantic features of the sample text. The training process of the semantic extraction model includes the following steps: First, collect a large number of sample texts and determine the semantic features of the sample texts; then, based on the sample texts and the semantic features of the sample texts, train the initial semantic extraction model to obtain a trained semantic extraction model. It should be noted that the initial semantic extraction model here can be constructed based on BiLSTM (Bi-directional Long Short-Term Memory). Figure 7 is the structural schematic diagram of the BiLSTM provided by the present invention. As Figure 7 shown, the input of BiLSTM is the sentence corpus, that is, the text line of the sample text. After passing through the double-layer LSMT, the hidden state at the last moment is used as Sentence Embedding, that is, the sample semantic feature of the text line in the sample text.
[0115] Subsequently, step 1214 can be executed to fuse the position feature, region image feature, and semantic feature of the text line to obtain the text line feature of the text line. The text line feature obtained in this way can better represent the sentence integrity of the text line and the semantic coherence with the context information.
[0116] It should be noted that the fusion methods of these three can be splicing, addition, weighted fusion, etc., and the embodiments of the present invention do not make specific limitations in this regard. Preferably, in the embodiments of the present invention, the fusion method is selected as splicing, that is, the position feature, region image feature, and semantic feature of any text line are spliced to obtain the text line feature of the text.
[0117] The method provided by the present invention can more completely represent the statement integrity of the corresponding text line and the semantic coherence with the context information by combining the position feature, region image feature, and semantic feature to determine the text line feature. Text chunking based on this feature can avoid splitting text lines with semantic relationships, improve the integrity of the statement and the coherence of the semantics, and contribute to the improvement of the translation result accuracy.
[0118] Based on the above embodiments, Figure 8 is a schematic diagram of the determination process of the region image and the recognized text provided by the present invention, as Figure 8 shown, the region image of each text line in the text image and the recognized text of each text line are determined based on the following steps:
[0119] Step 810, based on the positions of each text line, perform image segmentation on the text image to obtain the region image of each text line in the text image;
[0120] Step 820, perform text recognition on the region image of each text line in the text image to obtain the recognized text of each text line.
[0121] Specifically, after obtaining the positions of each text line in the text image through step 110, before performing text chunking on the text image to be chunked, it is also necessary to determine the region image of each text line in the text image and the recognized text of each text line. Among them, the region image of each text line in the text image can be cropped from the text image based on the positions of each text line, and the recognized text of each text line can be determined based on the region image of each text line in the text image.
[0122] Among them, the determination process of the region image of each text line in the text image can specifically be step 810, perform image segmentation on the text image to be chunked according to the positions of each text line, that is, extract the region image of each text line in the text image from the text image according to the positions of each text line, so as to obtain the region image of each text line in the text image.
[0123] Further, after determining the regional images of each text line in the text image, step 820 can be executed to perform text recognition on the regional images of each text line in the text image, recognize the text content of the regional image of each text line in the text image, so as to obtain the recognition text of each text line.
[0124] It should be noted that the process of performing text recognition on the regional images of each text line in the text image here can be implemented through a text recognition model. The specific process can be to input the regional images of each text line in the text image into the text recognition model, and the text recognition model performs text recognition on the input regional images of each text line in the text image, and finally obtains the recognition text of each text line output by the text recognition model.
[0125] Before inputting the regional images of each text line in the text image into the text recognition model, a text recognition model can also be pre-trained according to the sample text line images and the transcription texts of the sample text line images. The training process of the text recognition model includes the following steps: First, collect a large number of sample text line images and determine the transcription texts of the sample text line images; then, based on the sample text line images and the transcription texts of the sample text line images, train the initial text recognition model to obtain a trained text recognition model. It should be noted that the initial text recognition model here can be constructed based on Attention ED (Attention Encoder Decoder).
[0126] Figure 9 is the network structure diagram of Attention ED provided by the present invention. As Figure 9 shown, this network includes an Encoder (encoder) and a Decoder (decoder). The Encoder includes a CNN (Convolutional Neural Networks, convolutional neural network) and a BiLSTM, and the Decoder includes an RNN (Recurrent Neural Network, recurrent neural network); the input of this network is the sample text line image ("CAT"), and the feature encoding is performed through the CNN and BiLSTM in the Encoder, and then the features obtained by encoding are decoded in the Decoder by combining Attention (attention mechanism) and RNN, and the sample recognition text ("SOS"CAT"EOS") can be obtained. Among them, "SOS" represents the start symbol, and "EOS" represents the end symbol.
[0127] Figure 10 is the overall framework diagram of the text chunking method provided by the present invention. As Figure 10As shown in the figure, the process of text chunking for a text image can be specifically as follows: First, text detection is performed on the text image to obtain the positions of each text line in the text image;
[0128] Subsequently, according to the positions of each text line, image segmentation is performed on the text image to determine the regional images of each text line in the text image, and text recognition is performed on the regional images of each text line in the text image to obtain the recognized text of each text line;
[0129] Then, the positions of each text line are encoded to obtain the position features of each text line; according to the positions of each text line, the regional image features corresponding to the regional images of each text line in the text image are determined from the image features of the text image, that is, the regional image features of each text line; semantic feature extraction is performed on the recognized text of each text line to obtain the semantic features of each text line;
[0130] After that, the position features, regional image features, and semantic features of each text line are fused to obtain the text line features of each text line, and according to the positions of each text line, the text line features of each text line are concatenated to obtain a text line feature sequence;
[0131] Finally, according to the decoding dictionary, text chunking decoding is performed on each text line feature sequence to obtain a text chunking result, where the decoding dictionary includes the text line features of each text line, as well as the encoding features of the start symbol, end symbol, and chunking symbol.
[0132] The method provided by the embodiments of the present invention extracts the position features, regional image features, and semantic features of each text line, and fuses them to obtain the text line features of each text line. The text line features obtained thereby can better represent the statement integrity of the text line and the semantic coherence with the context information, so that the text chunking process based on the text line features can concatenate text lines with semantic relationships together, avoiding the situation of splitting text lines with semantic relationships, and providing strong assistance for the execution of subsequent translation tasks and the improvement of the accuracy and actual effect of translation results; in addition, integrating multiple technologies can further improve the accuracy and actual effect of translation results.
[0133] Next, the text chunking device provided by the present invention will be described. The text chunking device described below can be correspondingly referred to the text chunking method described above.
[0134] Figure 11 is a schematic structural diagram of the text chunking device provided by the present invention. As Figure 11 shown, the device includes:
[0135] A text detection unit 1110, configured to perform text detection on the text image to be chunked, and obtain the positions of each text line in the text image;
[0136] A text chunking unit 1120, configured to perform text chunking on the text image based on the positions of each text line, the regional images of each text line in the text image, and the recognized text of each text line, so as to obtain a text chunking result.
[0137] The text chunking device provided by the present invention performs text chunking on a text image according to the positions of each text line, the regional images of each text line in the text image, and the recognized text of each text line, and obtains a text chunking result. It evaluates the statement integrity of each text line in the text image and the semantic coherence between each text line from three different perspectives, and performs text chunking according to the evaluation result. It can overcome the defects in the traditional solution that directly translates a single text line, resulting in low accuracy of the translation result and poor actual effect. In the embodiments of the present invention, text lines with semantic relationships can be concatenated together, thereby providing strong assistance for the execution of subsequent translation tasks and the improvement of the accuracy rate and actual effect of the translation result, and further promoting the actual experience of users and the improvement process of relevant evaluation indicators of photo translation technology.
[0138] Based on the above embodiments, the text chunking unit 1120 is configured to:
[0139] Based on the position of any text line, the regional image of this text line in the text image, and the recognized text of this text line, determine the text line feature of this text line;
[0140] Based on the text line features of each text line, perform text chunking on the text image to obtain a text chunking result.
[0141] Based on the above embodiments, the text chunking unit 1120 is configured to:
[0142] Based on the positions of each text line, splice the text line features of each text line to obtain a text line feature sequence;
[0143] Based on a decoding dictionary, perform text chunking decoding on the text line feature sequence to obtain the text chunking result. The decoding dictionary includes the text line features of each text line, as well as the encoding features of a start symbol, an end symbol, and a chunking symbol.
[0144] Based on the above embodiments, the text chunking unit 1120 is configured to:
[0145] Determine the decoding result and hidden state of the text line feature sequence at the current moment based on the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment, remove the decoding result from the decoding dictionary, and update the text chunking result based on the decoding result;
[0146] Update the next moment of the current moment to the current moment until the decoding result at the current moment is the encoding feature of the end symbol or the text chunking result reaches the preset length.
[0147] Based on the above embodiments, the text chunking unit 1120 is configured to:
[0148] Determine the output probability of the text line features of each text line in the decoding dictionary at the current moment based on the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment;
[0149] Determine the decoding result and hidden state of the text line feature sequence at the current moment based on the output probability of the text line features of each text line in the decoding dictionary at the current moment.
[0150] Based on the above embodiments, the apparatus further includes a feature determination unit, configured to:
[0151] Encode the position of the text line to obtain the position feature of the text line;
[0152] Based on the position of the text line, determine the region image feature corresponding to the region image of the text line in the text image from the image features of the text image;
[0153] Extract semantic features from the recognized text of the text line to obtain the semantic features of the text line;
[0154] Based on the position feature, region image feature, and semantic feature of the text line, determine the text line feature of the text line.
[0155] Based on the above embodiments, the apparatus further includes an image segmentation unit and a text recognition unit. The image segmentation unit is configured to:
[0156] Perform image segmentation on the text image based on the positions of the respective text lines to obtain the region images of the respective text lines in the text image;
[0157] The text recognition unit is configured to:
[0158] Perform text recognition on the region images of the respective text lines in the text image to obtain the recognized texts of the respective text lines.
[0159] Figure 12Illustrates a schematic diagram of the physical structure of an electronic device, such as Figure 12 shown. The electronic device may include: a processor 1210, a communications interface 1220, a memory 1230, and a communication bus 1240. Among them, the processor 1210, the communications interface 1220, and the memory 1230 complete mutual communication through the communication bus 1240. The processor 1210 can call the logical instructions in the memory 1230 to execute a text chunking method, which includes: performing text detection on the text image to be chunked to obtain the positions of each text line in the text image; based on the positions of each text line, the regional images of each text line in the text image, and the recognized text of each text line, performing text chunking on the text image to obtain a text chunking result.
[0160] In addition, when the logical instructions in the above-mentioned memory 1230 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0161] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the text chunking method provided by the above-mentioned various methods. The method includes: performing text detection on the text image to be chunked to obtain the positions of each text line in the text image; based on the positions of each text line, the regional images of each text line in the text image, and the recognized text of each text line, performing text chunking on the text image to obtain a text chunking result.
[0162] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the text chunking method provided by the above-mentioned various methods. The method includes: performing text detection on the text image to be chunked to obtain the positions of each text line in the text image; based on the positions of each text line, the regional image of each text line in the text image, and the recognized text of each text line, performing text chunking on the text image to obtain a text chunking result.
[0163] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0164] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A text chunking method, characterized in that, it includes: performing text detection on the text image to be chunked to obtain the positions of each text line in the text image; performing text chunking on the text image based on the positions of each text line, the regional images of each text line in the text image, and the recognized texts of each text line to obtain a text chunking result; The performing text chunking on the text image based on the positions of each text line, the regional images of each text line in the text image, and the recognized texts of each text line to obtain a text chunking result includes: determining the position features corresponding to the positions of each text line, the regional image features corresponding to the regional images of each text line in the text image, and the semantic features corresponding to the recognized texts of each text line; performing text chunking on the text image based on the position features, regional image features, and semantic features of each text line to obtain a text chunking result; The performing text chunking on the text image based on the position features, regional image features, and semantic features of each text line to obtain a text chunking result includes: determining the text line features of each text line based on the position features, regional image features, and semantic features of each text line; performing text chunking on the text image based on the text line features of each text line to obtain a text chunking result; The performing text chunking on the text image based on the text line features of each text line to obtain a text chunking result includes: concatenating the text line features of each text line based on the positions of each text line to obtain a text line feature sequence; performing text chunking decoding on the text line feature sequence based on a decoding dictionary to obtain the text chunking result.
2. The text chunking method according to claim 1, characterized in that, the decoding dictionary includes the text line features of each text line, as well as the encoding features of a start symbol, an end symbol, and a chunking symbol.
3. The text chunking method according to claim 2, characterized in that, the performing text chunking decoding on the text line feature sequence based on a decoding dictionary to obtain the text chunking result includes: determining the decoding result and the hidden state at the current moment of the text line feature sequence based on the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment, removing the decoding result from the decoding dictionary, and updating the text chunking result based on the decoding result; updating the next moment of the current moment to the current moment until the decoding result at the current moment is the encoding feature of the end symbol or the text chunking result reaches a preset length.
4. The text chunking method according to claim 3, characterized in that, the determining the decoding result and the hidden state at the current moment of the text line feature sequence based on the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment includes: Determine the output probability of the text line features of each text line in the decoding dictionary at the current moment based on the correlation between the text line features of each text line in the decoding dictionary and the hidden state at the previous moment; Determine the decoding result and hidden state of the text line feature sequence at the current moment based on the output probability of the text line features of each text line in the decoding dictionary at the current moment.
5. The text chunking method according to any one of claims 1 to 4, characterized in that the determination of the position features corresponding to the positions of the respective text lines, the region image features corresponding to the region images of the respective text lines in the text image, and the semantic features corresponding to the recognized texts of the respective text lines includes: Encode the positions of the respective text lines to obtain the position features corresponding to the positions of the respective text lines; Based on the positions of the respective text lines, determine, from the image features of the text image, the region image features corresponding to the region images of the respective text lines in the text image; Extract semantic features from the recognized texts of the respective text lines to obtain the semantic features corresponding to the recognized texts of the respective text lines.
6. The text chunking method according to any one of claims 1 to 4, characterized in that the region images of the respective text lines in the text image and the recognized texts of the respective text lines are determined based on the following steps: Based on the positions of the respective text lines, perform image segmentation on the text image to obtain the region images of the respective text lines in the text image; Perform text recognition on the region images of the respective text lines in the text image to obtain the recognized texts of the respective text lines.
7. A text chunking device, characterized in that comprising: a text detection unit configured to perform text detection on a text image to be chunked to obtain the positions of the respective text lines in the text image; a text chunking unit configured to perform text chunking on the text image based on the positions of the respective text lines, the region images of the respective text lines in the text image, and the recognized texts of the respective text lines to obtain a text chunking result; The performing text chunking on the text image based on the positions of the respective text lines, the region images of the respective text lines in the text image, and the recognized texts of the respective text lines to obtain a text chunking result includes: Determine the position features corresponding to the positions of the respective text lines, the region image features corresponding to the region images of the respective text lines in the text image, and the semantic features corresponding to the recognized texts of the respective text lines; Perform text chunking on the text image based on the position features, region image features, and semantic features of the respective text lines to obtain a text chunking result; The performing text chunking on the text image based on the position features, region image features, and semantic features of the respective text lines to obtain a text chunking result includes: Determine the text line features of the respective text lines based on the position features, region image features, and semantic features of the respective text lines; Based on the text line features of each of the text lines, perform text chunking on the text image to obtain a text chunking result; The performing text chunking on the text image based on the text line features of each of the text lines to obtain a text chunking result includes: Based on the positions of each of the text lines, splice the text line features of each of the text lines to obtain a text line feature sequence; Based on a decoding dictionary, perform text chunking decoding on the text line feature sequence to obtain the text chunking result.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, the steps of the text chunking method according to any one of claims 1 to 6 are implemented.
9. A non-transitory computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, the steps of the text chunking method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method for detecting and identifying continuous segmented texts in image
CN110399845A
Image character recognition method, device and equipment and storage medium
CN110569846A