Text processing method, device, electronic device and storage medium
Through parallel calculation and optimization of data processing methods, the problem of slow text detection processing speed in the prior art is solved, and a more efficient text detection post-processing algorithm is realized.
Patent Information
- Application Number
- CN202080000735.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-15
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2040-08-27
AI Technical Summary
In the text detection processing, the number of register access times increases due to the logic of sequential determination, which reduces the processing speed and computing efficiency of the algorithm.
Through parallel calculation, the coordinate value of the central pixel, the spatial offset of adjacent pixels and the width of the pixel processing area are obtained, and the preset vector calculation formula and mask mapping relationship are used to optimize the calculation of data judgment conditions and improve the processing speed.
It effectively improves the processing speed of text detection post-processing algorithm, reduces the number of register memory accesses, and improves computing efficiency.
Smart Images

Figure CN114026613B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a text processing method, device, electronic device and storage medium. Background Art
[0002] Text translation is divided into two steps: text detection and text recognition. Among them, text detection mainly classifies the pixels in the image, distinguishes the text part from the background part, and frames the selected difference to determine the boundary of the partition; text recognition mainly recognizes the framed text image and translates it through the trained Chinese and English content.
[0003] Application Contents
[0004] The present application partly provides a text recognition method, device, electronic device and storage medium to solve the problem in the related art that when performing text detection processing, the logic of sequential judgment is adopted when judging pixel edges, which leads to an increase in the number of register accesses, reduces the processing speed of the entire algorithm, and greatly reduces the computing efficiency.
[0005] The first aspect of the present application provides a text processing method, comprising the following steps: determining multiple central pixels and adjacent pixels of each central pixel of the text image according to the coordinate value of each pixel in the text image; obtaining the coordinate value of each central pixel, the spatial offset of each adjacent pixel of the central pixel, and the width of a pixel processing area; obtaining the position information of each adjacent pixel according to the coordinate value of each central pixel, the spatial offset of each adjacent pixel of the central pixel, the width of the pixel processing area, and a preset vector calculation formula; and obtaining the text of the text image according to the position information and a preset mask mapping relationship.
[0006] The second aspect of the present application provides a text processing device, including: a determination module, used to determine multiple central pixels of the text image and adjacent pixels of each central pixel according to the coordinate value of each pixel in the text image; a first acquisition module, used to obtain the coordinate value of each central pixel, the spatial offset of each adjacent pixel of the central pixel, and the width of the pixel processing area; a calculation module, used to obtain the position information of each adjacent pixel according to the coordinate value of each central pixel, the spatial offset of each adjacent pixel of the central pixel, the width of the pixel processing area and a preset vector calculation formula; a second acquisition module, used to obtain the text of the text image according to the position information and a preset mask mapping relationship.
[0007] The third aspect of the present application provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the text processing method as described in the above embodiment.
[0008] The fourth aspect of the present application provides a computer-readable storage medium on which a computer program is stored, characterized in that the program is executed by a processor to implement the text processing method as described above.
[0009] In this method, after obtaining the coordinate value of the central pixel, the spatial offset of the adjacent pixels of the central pixel and the width of the pixel processing area, the position information of each adjacent pixel can be determined by parallel calculation, so as to obtain the text of the text image according to the position information and the preset mask mapping relationship. Therefore, by rearranging the data during text detection processing, optimizing the calculation of data judgment conditions, and combining the underlying operation with parallel processing, the processing speed of the text detection post-processing algorithm is effectively improved, and the problem that in the related technology, when performing text detection processing, the logic of sequential judgment is adopted when judging the edge of the pixel, which increases the number of register accesses, reduces the processing speed of the entire algorithm, and greatly reduces the calculation efficiency is solved.
[0010] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0012] Figure 1 The figure is a schematic diagram of the text detection workflow based on Pixel ink;
[0013] Figure 2 It is a schematic diagram of positive pixel text link based on 8 neighborhoods;
[0014] Figure 3 The figure is a schematic diagram of the link position process for post-processing of text detection;
[0015] Figure 4 is a flowchart of a text processing method according to an embodiment of the present application;
[0016] Figure 5 A schematic diagram of edge determination based on multiple data streams according to an embodiment of the present application;
[0017] Figure 6A schematic diagram of coordinate mapping according to an embodiment of the present application;
[0018] Figure 7 Schematic diagram of a text processing method according to an embodiment of the present application. DETAILED DESCRIPTION
[0019] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0020] The text processing method, device, electronic device and storage medium of the embodiments of the present application are described below with reference to the accompanying drawings.
[0021] Before introducing the text processing method of the embodiment of the present application, a text processing method in the related art is briefly introduced.
[0022] The translation pen is a text learning assistance device. Users input the corresponding image of the text, and the processor loaded on the pen performs offline OCR processing. In this process, the deep learning model is mainly used to perform network calculations on the terminal. For the deep learning model, two classification predictions are made for the input image pixels: text and non-text prediction. Text and non-text prediction divides the pixels in the image into "positive pixels" (representing text) and "negative pixels" (representing non-text). Then, through positive links, the text belonging to the positive pixels is linked together to achieve instance segmentation of text and non-text, and then the bounding box of the text is extracted from the segmented text results.
[0023] Specifically, Figure 1 As indicated, Figure 1 The figure is a schematic diagram of the text detection workflow based on Pixel Link in the related art. The text and non-text prediction is completed by the CNN network (Convolutional Neural Networks), and the backbone network can select VGG16 (Oxford Visual Geometry Group, VGG network). After the predicted result is detected, the classification of the full image is obtained through thresholding. Each pixel is divided into "text-positive pixel" and "non-text-negative pixel", and marked as 1 and 0 respectively. After processing, the 0 and 1 masks of the whole image are obtained to distinguish the text and non-text parts.
[0024] Among them, Figure 2As shown, in the link postive process, the CNN network can be used to search for the 8 adjacent pixels of the central pixel and determine whether there are positive pixels. Each linking operation will detect the linking status in 8 directions, so that the text and non-text parts in the image are separated in turn, and the circumscribed rectangular frame of the resulting text area is extracted. It should be noted that the pixels after the neural network calculation can be divided into "text" or "background", where "text" is marked as positive and "background" is marked as negative; link postive is to store the pixels belonging to "text" in a fixed data structure according to a certain sequence relationship for subsequent processing and calculation. In addition, since there will be a certain amount of noise in the processing process, further filtering is required. In this solution, geometric shape screening is used to perform a filter to obtain the text box area in the image.
[0025] In the above calculations, except for the neural network calculation part of VGG16, the work of text box extraction needs to be performed pixel by pixel. Specifically, Figure 3 As shown in the figure, in the post-processing part of text detection, that is, in the entire detection process, starting from the image input, after the calculation of the neural network, the initial classification result is obtained, and the pixel is determined to belong to "text" or "background"; after obtaining the information, the "text" pixels are further classified, and the "text" belonging to a certain word is stored in the same storage space. This process is called the post-processing part of the entire process. The first is to perform edge determination. The image window size processed is an area of 81x81. In this space, the processed pixel must have a complete 8-neighborhood. Therefore, when processing a pixel, first determine whether its coordinates are in the range of (1, 1) to (80, 80); At the same time, it is also necessary to determine that the processed pixel does not belong to the edge of the entire image. If the pixel meets the requirements, the 8-neighborhood space near the central pixel is determined in turn, and 1 multiplication and 1 addition are performed each time the coordinates are moved to determine the target to be processed. After processing, the coordinates are sent to the lookup table. Since the image analyzed by the neural network has marked each point as "positive pixel": 1 and "negative pixel": 0, the lookup will determine whether the pixel being checked belongs to "text" or "non-text" and update it to the connectivity list. The whole process is decode, which is equivalent to segmenting the text box using the connectivity relationship of the 0 / 1 image encoded by the neural network.
[0026] However, in this process, the pixel judgment in the 8-neighborhood space is performed sequentially, starting from the first pixel in the upper left corner, first row and then column, and continuing to calculate until the last pixel in the lower right corner. This sequential execution method is inefficient. First, for a single pixel, it is necessary to search the 8 neighboring pixels around it one by one; second, for each pixel, the above operations will be performed one by one. The entire sequential execution of judgment slows down the processing speed of the entire algorithm. In actual operation, if the calculation is performed on a relatively weak CPU (about 900Mhz), this part of the work will occupy nearly 50% of the running time.
[0027] The present application provides a text recognition method, in which after obtaining the coordinate value of the central pixel, the spatial offset of the adjacent pixels of the central pixel and the width of the pixel processing area, the position information of each adjacent pixel can be determined by parallel calculation, so as to obtain the text of the text image according to the position information and the preset mask mapping relationship. Therefore, by rearranging the data during text detection processing, optimizing the calculation of data judgment conditions, and combining the underlying operation with parallel processing, the processing speed of the text detection post-processing algorithm is effectively improved, and the problem that in the related technology, when performing text detection processing, the logic of sequential judgment is adopted when performing pixel edge judgment, which leads to an increase in the number of register accesses, reduces the processing speed of the entire algorithm, and greatly reduces the calculation efficiency is solved.
[0028] Figure 4 A flowchart of a text processing method provided in an embodiment of the present application.
[0029] like Figure 4 As shown, the text processing method includes the following steps:
[0030] S1, determining a plurality of central pixels of the text image and adjacent pixels of each central pixel according to the coordinate value of each pixel in the text image.
[0031] In one embodiment of the present application, multiple central pixels of a text image and adjacent pixels of each central pixel are determined based on the coordinate value of each pixel in the text image, including: obtaining the coordinate value of each pixel in the text image and the width and height of a pixel processing area; determining whether the coordinate value of each pixel in the text image meets preset conditions; and taking the pixel that meets the preset conditions as the central pixel.
[0032] It can be understood that, assuming that the pixel processing area size is 81×81, within this area, the processed pixel must have a complete 8-neighborhood, that is, the pixel is taken as the center pixel. Therefore, when processing a pixel, it is necessary to first determine whether its coordinates are in the range of (1, 1) to (80, 80); at the same time, it is also necessary to determine that the processed pixel does not belong to the edge part of the entire image.
[0033] Therefore, the embodiments of the present application can make a judgment according to the coordinate values of each pixel in the text image and the width and height of the pixel processing area. Among them, in an embodiment of the present application, the preset conditions can be: the abscissa of the pixel is greater than or equal to 0; the ordinate of the pixel is greater than or equal to 0; the abscissa and ordinate of the pixel are respectively less than the width and height of the pixel processing area. Among them, it can be determined whether the preset conditions are met through neon registers.
[0034] Optionally, in an embodiment of the present application, the neon register is any one of a 128-bit register or a 64-bit register.
[0035] For example, as Figure 5 shown, Figure 5 is a schematic diagram of edge determination based on multiple data streams. Assume that the central pixel coordinates are (x, y), then the neighboring pixel 1 is (x - 1, y - 1), the neighboring pixel 2 is (x, y - 1), the neighboring pixel 3 is (x + 1, y - 1), the neighboring pixel 4 is (x - 1, y), the neighboring pixel 5 is (x + 1, y), the neighboring pixel 6 is (x - 1, y + 1), the neighboring pixel 7 is (x, y + 1), and the neighboring pixel 8 is (x + 1, y + 1).
[0036] Therefore, when performing edge determination, the conditions that neighboring pixel 1 needs to meet can be: x - 1 >= 0, x - 1 <= w, y - 1 >= 0, y - 1 < h; the conditions that neighboring pixel 2 needs to meet can be: x >= 0, x <= w, y - 1 >= 0, y - 1 < h; the conditions that neighboring pixel 3 needs to meet can be: x + 1 >= 0, x + 1 <= w, y - 1 >= 0, y - 1 < h; the conditions that neighboring pixel 4 needs to meet can be: x - 1 >= 0, x - 1 <= w, y >= 0, y < h; the conditions that neighboring pixel 5 needs to meet can be: x + 1 >= 0, x + 1 <= w, y >= 0, y < h; the conditions that neighboring pixel 6 needs to meet can be: x - 1 >= 0, x - 1 <= w, y + 1 >= 0, y + 1 < h; the conditions that neighboring pixel 7 needs to meet can be: x >= 0, x <= w, y + 1 >= 0, y + 1 < h; the conditions that neighboring pixel 8 needs to meet can be: x + 1 >= 0, x + 1 <= w, y + 1 >= 0, y + 1 < h.
[0037] It should be noted that since the registers used in the relevant technology are generally 32-bit ARM processors, which can store up to 32 bits of data at a time, and the ARM processor adopts the von Neumann structure, each instruction is read once and processed once, and it does not have the ability to process multiple data in parallel. Each time the CPU can only take one data, calculate, and return a calculation result, so that when performing the above calculation in the relevant technology, a total of 8 calculations are required, which takes a long time overall and greatly reduces the processing speed.
[0038] Therefore, the embodiment of the present application takes into account the characteristics of the neon register. The hardware is implemented by armv7 and armv8, and can divide 64-bit and 128-bit data, and perform 4-division and 8-division calculations respectively. For example, when designing a program, 64 bits can be divided into 4 16-bit data or 128 bits can be divided into 4 32-bit or 16 8-bit data in an artificial way. When the embodiment of the present application adopts the neon register, multiple data can be obtained at the same time, and the processing unit responds at the same time, and the data taken out can be processed at the same time and written back. In the above example, the edge judgment of the 8 neighborhoods is directly stored by an 8-bit array, and the vector calculation characteristics of the neon register are used to calculate the judgment results of the columns in parallel, and then the judgment results of the rows are calculated in parallel, that is, the embodiment of the present application can calculate the 4 conditions of 8 adjacent pixels at the same time, and the judgment can also be calculated in groups at the same time, thereby greatly reducing the calculation time. Compared with the successive calculation in the related art, the judgment reduces the number of register accesses and improves the running speed of the module.
[0039] S2, obtaining the coordinate value of each central pixel, the spatial offset of the adjacent pixels of each central pixel, and the width of the pixel processing area.
[0040] S3, obtaining the position information of each adjacent pixel according to the coordinate value of each central pixel, the spatial offset of the adjacent pixels of each central pixel, the width of the pixel processing area and a preset vector calculation formula.
[0041] In one embodiment of the present application, the position information of the adjacent pixels of each central pixel is obtained according to the coordinate value of each central pixel, the spatial offset of the adjacent pixels of each central pixel, the width of the pixel processing area and a preset vector calculation formula, including: obtaining the horizontal coordinate and the vertical coordinate of each central pixel; obtaining a first vector product according to the width of the pixel processing area and the spatial offset of the adjacent pixels; obtaining a second vector product according to the horizontal coordinate of each central pixel and the spatial offset of the adjacent pixels; obtaining a third vector product according to the first vector product and the vertical coordinate of each central pixel; and obtaining the position information of each adjacent pixel according to the second vector product and the third vector product.
[0042] It can be understood that although the parallel computing method used in the embodiment of the present application is more complicated than the single computing method in the related art, because the vector calculations of rows and columns can be completed simultaneously in a single register, the embodiment of the present application uses neon registers, so that the algorithm can run multi-vector calculations.
[0043] Specifically, in the embodiment of the present application, when executing the calculation of multiple vectors, the calculation can be expanded into multiplication vectors and addition vectors, wherein the calculation formula can be:
[0044] pos it ion=y*w*neighbours+x*neighbours;
[0045] Among them, x is the horizontal coordinate of the center pixel, y is the vertical coordinate of the center pixel, w is the width of the pixel processing area, and neighbors is the offset of the 8-neighborhood space, which is 1-8 respectively. After calculation through the above formula, the corresponding coordinate values of all pixels can be obtained, that is, the position information of each adjacent pixel can be obtained.
[0046] Among them, in one embodiment of the present application, the adjacent pixels of each central pixel include 8 pixels immediately adjacent to the central pixel, 2 pixels located in the row where the central pixel is located, 3 pixels located in the row above the row where the central pixel is located, and 3 pixels located in the row below the row where the central pixel is located; the spatial offsets of the 8 pixels are 1, 2, 3, 4, 5, 6, 7, and 8 in the order from left to right in each row and from top to bottom in adjacent row pixels.
[0047] That is to say, Figure 5 As shown, the central pixel is located at (x, y), and the 8 adjacent pixels to the central pixel are: 2 pixels located in the same row as the central pixel, such as (x-1, y) and (x+1, y), and their corresponding spatial offsets can be 4 and 5; 3 pixels located in the upper row of the row where the central pixel is located, such as (x-1, y-1), (x, y-1) and (x+1, y-1), and their corresponding spatial offsets can be 1, 2 and 3; 3 pixels located in the lower row of the row where the central pixel is located, such as (x-1, y+1), (x, y+1) and (x+1, y+1), and their corresponding spatial offsets can be 6, 7 and 8.
[0048] Furthermore, combined with Figure 5 and Figure 6 , Figure 6 The boxes are all vector representations. Assuming the width of the pixel processing area is 80, the results of the parallel operation can be as follows:
[0049] pos it ion(x-1,y-1)=y*80*1+x*1;
[0050] pos it ion(x,y-1)=y*80*2+x*2;
[0051] pos it ion(x+1,y-1)=y*80*3+x*3;
[0052] pos it ion(x-1,y)=y*80*4+x*4;
[0053] pos it ion(x+1,y)=y*80*5+x*5;
[0054] pos it ion(x-1,y+1)=y*80*6+x*6;
[0055] pos it ion(x,y+1)=y*80*7+x*7;
[0056] pos it ion(x+1,y+1)=y*80*8+x*8.
[0057] It should be noted that the multiplication vector calculation and the addition vector calculation in the above calculation. It can be seen that in the related art, when performing 8-neighborhood calculation, 8 loop operations need to be performed, that is, 8*3 multiplications and 8*1 additions need to be calculated. However, after the current data rearrangement and vectorization, the present application performs a total of 3 multiplications and 1 addition to obtain the position information of each adjacent pixel. Compared with the calculation of the related art, the method of the embodiment of the present application greatly reduces the number of multiplications and additions, and at the same time reduces the number of memory access results. For example, on a 900Mhz arm cpu, the post-processing of a 640*480 image can reduce the processing time by nearly 30%, effectively accelerating text detection.
[0058] S4, obtaining the text of the text image according to the position information and the preset mask mapping relationship.
[0059] It is understandable that, since the pixel data is two-dimensional data, and when it is stored in the storage space, it is generally stored as a one-dimensional continuous address, it is necessary to obtain the address of each pixel in order to obtain the correct pixel value. The embodiment of the present application may preset a mask mapping relationship between the position information and the coordinates of the adjacent pixels in each direction on the matrix. Among them, as shown in Table 1, Table 1 is the coordinate value of the adjacent pixels in each direction on the matrix in the embodiment of the present application, and the coordinate value can be a fixed value. The mask mapping relationship between the position information and the coordinates of the adjacent pixels in each direction on the matrix can be as follows:
[0060] For example, coordinates (1, 1) correspond to position information position (x-1, y-1); coordinates (1, 2) correspond to position information position (x, y-1); coordinates (1, 3) correspond to position information position (x+1, y-1); coordinates (2, 1) correspond to position information position (x-1, y); coordinates (2, 3) correspond to position information position (x+1, y); coordinates (3, 1) correspond to position information position (x-1, y+1); coordinates (3, 2) correspond to position information position (x, y+1); and coordinates (3, 3) correspond to position information position (x+1, y+1).
[0061] Table 1
[0062] (1,1) (1,2) (1,3) (2,1) (2,2) (2,3) (3,1) (3,2) (3,3)
[0063] Therefore, after the position information of each adjacent pixel is obtained, the text of the text image can be obtained according to the position information and the preset mask mapping relationship.
[0064] According to the text processing method proposed in the embodiment of the present application, after obtaining the coordinate value of the central pixel, the spatial offset of the adjacent pixels of the central pixel and the width of the pixel processing area, the position information of each adjacent pixel can be determined by parallel calculation, so as to obtain the text of the text image according to the position information and the preset mask mapping relationship. Therefore, by rearranging the data during text detection processing, optimizing the calculation of data judgment conditions, and combining the underlying operation with parallel processing, the processing speed of the text detection post-processing algorithm is effectively improved, and the problem that in the related technology, when performing text detection processing, the logic of sequential judgment is adopted when performing pixel edge judgment, which leads to an increase in the number of register accesses, reduces the processing speed of the entire algorithm, and greatly reduces the calculation efficiency.
[0065] Next, the text processing device proposed according to the embodiment of the present application is described with reference to the accompanying drawings.
[0066] Figure 7 It is a block diagram of a text processing device according to an embodiment of the present application.
[0067] like Figure 7 As shown, the text processing device 10 includes: a determination module 100 , a first acquisition module 200 , a calculation module 300 and a second acquisition module 400 .
[0068] Among them, the determination module 100 is used to determine multiple central pixels of the text image and the adjacent pixels of each central pixel according to the coordinate value of each pixel in the text image. The first acquisition module 200 is used to obtain the coordinate value of each central pixel, the spatial offset of the adjacent pixels of each central pixel, and the width of the pixel processing area. The calculation module 300 is used to obtain the position information of each adjacent pixel according to the coordinate value of each central pixel, the spatial offset of the adjacent pixels of each central pixel, the width of the pixel processing area and a preset vector calculation formula. The second acquisition module 400 is used to obtain the text of the text image according to the position information and the preset mask mapping relationship.
[0069] It should be noted that the above explanations of the text processing method embodiment are also applicable to the text processing device of this embodiment and will not be repeated here.
[0070] According to the text processing device proposed in the embodiment of the present application, after obtaining the coordinate value of the central pixel, the spatial offset of the adjacent pixels of the central pixel and the width of the pixel processing area, the position information of each adjacent pixel can be determined by parallel calculation, so as to obtain the text of the text image according to the position information and the preset mask mapping relationship. Therefore, by rearranging the data during text detection processing, optimizing the calculation of data judgment conditions, and combining the underlying operation with parallel processing, the processing speed of the text detection post-processing algorithm is effectively improved, and the problem that in the related technology, when performing text detection processing, the logic of sequential judgment is adopted when performing pixel edge judgment, which increases the number of register accesses, reduces the processing speed of the entire algorithm, and greatly reduces the calculation efficiency is solved.
[0071] In order to implement the above embodiment, the present application further proposes an electronic device, including: at least one processor and a memory. The memory is communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the text processing method of the above embodiment, such as:
[0072] Determine a plurality of central pixels of the text image and adjacent pixels of each central pixel according to the coordinate value of each pixel in the text image;
[0073] Obtaining the coordinate value of each central pixel, the spatial offset of the adjacent pixels of each central pixel, and the width of the pixel processing area;
[0074] Obtaining position information of each adjacent pixel according to the coordinate value of each central pixel, the spatial offset of the adjacent pixels of each central pixel, the width of the pixel processing area and a preset vector calculation formula; and
[0075] The text of the text image is obtained according to the position information and the preset mask mapping relationship.
[0076] In order to implement the above-mentioned embodiment, the present application also proposes a computer-readable storage medium on which a computer program is stored, characterized in that the program is executed by a processor to implement the above-mentioned text processing method.
[0077] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0078] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0079] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0080] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or N wirings (electronic devices), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways if necessary, and then stored in a computer memory.
[0081] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0082] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0083] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0084] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A text processing method, It is characterized in that include: Determine a plurality of central pixels of the text image and adjacent pixels of each central pixel according to the coordinate value of each pixel in the text image; Acquire the coordinate value of each central pixel, the spatial offset of each adjacent pixel of the central pixel, and the width of the pixel processing area; According to the coordinate value of each of the central pixels, the spatial offset of the adjacent pixels of each of the central pixels, the width of the pixel processing area and the preset vector calculation formula, the position information of each of the adjacent pixels is obtained, wherein the position information of the adjacent pixels is calculated by the following formula: position=y*w*neighbors+x*neighbors; Where x is the horizontal coordinate of the center pixel, y is the vertical coordinate of the center pixel, w is the width of the pixel processing area, and neighbors is the offset of the 8-neighborhood space, which are 1-8 respectively; and The text of the text image is obtained according to the position information and a preset mask mapping relationship.
2. The method according to claim 1, It is characterized in that The step of obtaining the position information of the adjacent pixels of each central pixel according to the coordinate value of each central pixel, the spatial offset of the adjacent pixels of each central pixel, the width of the pixel processing area and a preset vector calculation formula includes: Obtaining the horizontal coordinate and the vertical coordinate of each of the central pixels; Obtaining a first vector product according to the width of the pixel processing area and the spatial offset of the adjacent pixels; Obtaining a second vector product according to the abscissa of each of the central pixels and the spatial offset of the adjacent pixels; Obtaining a third vector product according to the first vector product and the ordinate of each of the central pixels; The position information of each of the adjacent pixels is obtained according to the second vector product and the third vector product.
3. The method according to claim 1 or 2, It is characterized in that The adjacent pixels of each central pixel include 8 pixels immediately adjacent to the central pixel, 2 pixels located in the row where the central pixel is located, 3 pixels located in the row above the row where the central pixel is located, and 3 pixels located in the row below the row where the central pixel is located; the spatial offsets of the 8 pixels are 1, 2, 3, 4, 5, 6, 7, and 8 respectively in the order from left to right of each row and from top to bottom of adjacent row pixels.
4. The method according to claim 1, It is characterized in that Determining a plurality of central pixels of the text image and adjacent pixels of each central pixel according to the coordinate value of each pixel in the text image comprises: Obtaining the coordinate value of each pixel in the text image and the width and height of the pixel processing area; Determining whether the coordinate value of each pixel in the text image meets a preset condition; The pixel that meets the preset condition is used as the central pixel.
5. The method according to claim 4, It is characterized in that The preset conditions are: The horizontal coordinate of the pixel is greater than or equal to 0; The vertical coordinate of the pixel is greater than or equal to 0; The horizontal coordinate and the vertical coordinate of the pixel are respectively smaller than the width and the height of the pixel processing area.
6. The method according to claim 5, It is characterized in that Whether the preset condition is met is determined by the neon register.
7. The method according to claim 6, It is characterized in that The neon register is a 128-bit register or any one of a 64-bit register.
8. A text processing device, It is characterized in that include: A determination module, used for determining a plurality of central pixels of the text image and adjacent pixels of each central pixel according to the coordinate value of each pixel in the text image; A first acquisition module, used for acquiring the coordinate value of each central pixel, the spatial offset of each adjacent pixel of the central pixel and the width of a pixel processing area; A calculation module is used to obtain the position information of each of the adjacent pixels according to the coordinate value of each of the central pixels, the spatial offset of the adjacent pixels of each of the central pixels, the width of the pixel processing area and a preset vector calculation formula, wherein the position information of the adjacent pixels is calculated by the following formula: position=y*w*neighbors+x*neighbors; Where x is the horizontal coordinate of the center pixel, y is the vertical coordinate of the center pixel, w is the width of the pixel processing area, and neighbors is the offset of the 8-neighborhood space, which are 1-8 respectively; and The second acquisition module is used to obtain the text of the text image according to the position information and a preset mask mapping relationship.
9. An electronic device, It is characterized in that include: at least one processor; And, a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the text processing method as described in any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, It is characterized in that The program is executed by a processor to implement the text processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Character detection method based on deformable convolutional neural network
CN110399882A
Text detection method and device, electronic equipment and storage medium
CN110717486A