Text data processing method and device
By detecting and deleting abnormalities in the target text in the page image of the reading page, the semantic confusion caused by incomplete text in the page image is solved, and the accuracy and semantic integrity of the reading are improved.
Patent Information
- Application Number
- PCT/CN2023/138449
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-13
- Publication Date
- 2025-06-19
AI Technical Summary
When reading a page aloud by terminal devices such as picture book robots, if the text in the page image is incomplete, it may lead to the playback of semantic misunderstanding audio.
By obtaining the target text and its text features in the page image to be read aloud, angle abnormality detection, position abnormality detection and occlusion abnormality detection are performed, and abnormal text is deleted to improve the accuracy of the text.
Improves the accuracy of text to be converted into audio, and avoids playing semantic misunderstanding audio.
Smart Images

Figure CN2023138449_19062025_PF_FP_ABST
Abstract
Description
Text data processing method and device Technical Field
[0001] The present invention relates to the technical field of text data processing, and in particular to a text data processing method and device. Background Art
[0002] With the development of artificial intelligence technology, intelligent reading has become a trend. At present, terminal devices such as reading and picture book robots have been widely loved by people because they can read the contents of picture books and other pages aloud.
[0003] When reading aloud page content, devices like picture book robots typically require users to place the page to be read within camera range. After capturing the page image, they perform text recognition on the page image, convert the recognized text into audio, and then play it. During this process, if the text in the page image is incomplete, the audio may be misleading.
[0004] Summary of the Invention
[0005] The embodiments of the present application provide a text data processing method and device, which can, at least to a certain extent, improve the accuracy of the text to be converted into audio during the page reading process, and avoid playing audio with semantic confusion.
[0006] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.
[0007] According to a first aspect of an embodiment of the present application, a text data processing method is provided, comprising:
[0008] Obtain target text in a page image of a page to be read aloud, and text features of the target text in the page image; based on at least one detection dimension, perform anomaly detection on the target text according to the text features and a detection threshold corresponding to the detection dimension to determine abnormal text in the target text, wherein the detection dimension includes angle anomaly detection, position anomaly detection, and occlusion anomaly detection; and delete the abnormal text from the target text.
[0009] In some embodiments, the detection dimension includes angle anomaly detection, the detection threshold corresponding to the detection dimension includes an image angle threshold, the text feature includes the number of the first line of text of the target text, and the typeset direction and the number of line text characters of each line of text in the target text, and the abnormality detection of the target text based on the text feature and the detection threshold corresponding to the detection dimension includes: determining the line text angle of each line of text according to the typeset direction of each line of text; determining the paragraph text angle of each paragraph text in the page image according to the number of line text characters and the line text angle of each line of text; determining the image angle of the page image according to the first line of text number and the line text angle of each line of text; and performing abnormality detection on the target text based on the paragraph text angle, the image angle and the image angle threshold.
[0010] In some embodiments, determining the paragraph text angle of each paragraph text in the page image based on the number of line text characters and the line text angle of each line text includes: determining a target line text in each paragraph text in the target text whose number of line text characters is greater than a first preset value; and determining the average value of the line text angles of the target line text in each paragraph text as the paragraph text angle of each paragraph text.
[0011] In some embodiments, the image angle of the page image is determined based on the number of the first lines of text and the line text angles of each line of text, including: if the number of the first lines of text is less than a second preset value, then the average value of the line text angles of all lines of text in the target text is determined as the image angle; if the number of the first lines of text is greater than or equal to the second preset value, then the median of the line text angles of all lines of text in the target text is determined as the image angle.
[0012] In some embodiments, the target text is identified from the page image by a text recognition model, and the target text is detected for abnormality based on the paragraph text angle, the image angle and the image angle threshold, including: determining a difference index based on the number of the first lines of text and the line text angles of each line of text, wherein the difference index is used to characterize the degree of difference between the angles of each line of text; if the difference index is greater than the difference index threshold, the target text is determined to be the abnormal text; if the difference index is less than or equal to the difference index threshold, and the image angle is greater than or equal to the image angle threshold, and all paragraph text angles in the target text are greater than the image angle threshold, the target text is determined to be the abnormal text, wherein the image angle threshold is the recognition angle corresponding to the text recognition model.
[0013] In some embodiments, the text features also include the number of paragraph text characters and the number of second lines of text in each paragraph text in the target text, the text features are identified by a text feature recognition model from the page image, and the target text is detected for anomalies based on the paragraph text angle, the image angle and the image angle threshold, and further includes: if the difference index is less than or equal to the difference index threshold, and the image angle is greater than or equal to the image angle threshold, and there is at least one paragraph text angle in the target text that is less than or equal to the image angle threshold, then the target paragraph text is determined from the various paragraph texts, wherein the target paragraph text The ratio between the number of paragraph text characters in the text and the number of corresponding second lines of text is less than or equal to a third preset value; if the maximum paragraph text angle among the paragraph text angles of each paragraph text is greater than or equal to a fourth preset value, and there is at least one target paragraph text whose paragraph text angle is greater than the fourth preset value, then the target text is determined to be the abnormal text, wherein the fourth preset value is the recognition angle corresponding to the text feature recognition model; if the maximum paragraph text angle is less than the fourth preset value, and the paragraph text angles of all target paragraph texts in the target text are greater than the image angle threshold, then the target text is determined to be the abnormal text.
[0014] In some embodiments, the text features also include the paragraph area of each paragraph text in the target text, and the paragraph area is the area occupied by the paragraph text in the page image. The abnormality detection of the target text based on the text features also includes: determining the two paragraph texts with the largest paragraph area from the various paragraph texts; if the difference between the paragraph text angles of the two paragraph texts is greater than a fifth preset value, the target text is determined as the abnormal text.
[0015] In some embodiments, the two paragraph texts include a first paragraph text and a second paragraph text, the paragraph area of the second paragraph text is smaller than the paragraph area of the first paragraph text, and the proportion of the paragraph area of the second paragraph text to the page image area is greater than or equal to a preset proportion.
[0016] In some embodiments, the detection dimension includes position anomaly detection, the detection threshold corresponding to the detection dimension includes a boundary distance threshold, the text feature includes the distance between each line of text in the target text and each image boundary in the page image, and the anomaly detection of the target text based on the text feature and the detection threshold corresponding to the detection dimension includes: performing anomaly detection on the target text based on the distance between each line of text and each image boundary in the page image and the boundary distance threshold.
[0017] In some embodiments, the text features also include language information of each line of text in the target text, and the target text is detected for abnormality based on the distance between each line of text and each image boundary in the page image and the boundary distance threshold, including: determining candidate text in the target text and determining the abnormality type of the candidate text based on the distance between each line of text and each image boundary in the page image and the boundary distance threshold; based on the language information, if the language of the candidate text is a preset language, determining abnormal text from the candidate text based on the abnormality type and preset screening rules; if the language of the candidate text is not the preset language, determining the candidate text as the abnormal text.
[0018] In some embodiments, the text features further include the height of each line of text in the target text, the image boundary includes a lower image boundary, and determining candidate text in the target text based on the distance between each line of text and each image boundary in the page image and the boundary distance threshold includes: determining a target distance between each line of text and the lower image boundary; if the target distance is less than the boundary distance threshold, determining the corresponding line of text as the candidate text;
[0019] If the target distance is greater than or equal to the boundary distance threshold, determine whether the corresponding line of text is the candidate text based on the line text height of the corresponding line of text, the target distance and the height of the target area in the page image, wherein the maximum distance between the target area and the lower image boundary is less than a sixth preset value.
[0020] In some embodiments, determining whether the corresponding line of text is the candidate text based on the line text height of the corresponding line of text, the target distance and the height of the target area in the page image includes: determining multiple distance ranges based on the height of the target area; determining the line text height threshold corresponding to each distance range, wherein the upper limit value of the distance range is negatively correlated with the line text height threshold corresponding to the distance range; using the line text height threshold corresponding to the distance range to which the target distance belongs as the target height threshold, and if the line text height of the corresponding line of text is less than or equal to the target height threshold, determining that the corresponding line of text is the candidate text.
[0021] In some embodiments, the text data processing method further includes: determining the line text angle of each line of text in the target text; determining a target ranging point on the detection box of each line of text according to the line text angle of each line of text; and determining the distance between the target ranging point and the lower image boundary as the distance between each line of text and the lower image boundary.
[0022] In some embodiments, the image boundary includes a left image boundary, a right image boundary, an upper image boundary and a lower image boundary, and determining the abnormal text from the candidate text according to the abnormality type and preset screening rules includes: if the abnormality type is exceeding the left image boundary, the candidate text containing the preset characters representing the beginning of the sentence is determined as the text to be retained; if the abnormality type is exceeding the right image boundary, the text to be retained is determined based on whether the candidate text contains the preset characters representing the end of the sentence; if the abnormality type is exceeding the upper image boundary, the text to be retained is determined from the candidate text whose line text angle is greater than the seventh preset value; if the abnormality type is exceeding the lower image boundary, the text to be retained is determined from the candidate text whose line text angle is greater than the eighth preset value; determine the abnormal text from the text to be retained, and determine the line text of the candidate text other than the text to be retained as abnormal text.
[0023] In some embodiments, determining abnormal text from the text to be retained includes: determining whether the text to be retained is abnormal text based on position information of abnormal text in the same paragraph as the text to be retained.
[0024] In some embodiments, the text data processing method also includes: responding to a shadow obstruction judgment instruction sent by a terminal device, judging whether each line of text is obscured by a shadow area according to the shadow area parameters of the terminal device; and determining the line of text obscured by the shadow area as the abnormal text.
[0025] In some embodiments, the detection dimension includes occlusion anomaly detection, the detection threshold corresponding to the detection dimension includes an occlusion ratio threshold, the text feature includes the average character area of each line of text in the target text, and the abnormality detection of the target text based on the text feature and the detection threshold corresponding to the detection dimension includes: determining the occlusion area from the page image; determining the occluded paragraph text corresponding to the occlusion area from the target text; and performing abnormality detection on each line of text based on the average character area of each line of text in the occluded paragraph text and the occlusion ratio threshold.
[0026] In some embodiments, determining the occluded paragraph text corresponding to the occluded area from the target text includes: expanding the detection frame of the paragraph text in the target text according to a preset size to obtain an extended detection frame; determining a first intersection-and-union ratio (IOR) between the extended detection frame and the occluded area; and determining the paragraph text corresponding to the extended detection frame whose first IOR is greater than a ninth preset value as the occluded paragraph text corresponding to the occluded area.
[0027] In some embodiments, the text features also include the character width of each line of text in the target text, and the abnormality detection of each line of text based on the average character area of each line of text in the occluded paragraph text and the occlusion ratio threshold includes: determining the second intersection-and-union ratio of the detection box of each line of text and the occlusion area based on the average character area of each line of text in the occluded paragraph text; if the second intersection-and-union ratio is greater than or equal to the occlusion ratio threshold, determining that each line of text is the abnormal text; if the second intersection-and-union ratio is less than the occlusion ratio threshold, determining whether each line of text is the abnormal text based on the minimum distance from the boundary of the occlusion area to the detection box of each line of text and the character width of each line of text.
[0028] In some embodiments, the text features also include the number of characters in each line of text in the target text, and determining whether each line of text is the abnormal text based on the minimum distance from the boundary of the occluded area to the detection box of each line of text and the character width of each line of text includes: for each line of text, determining the occlusion distance and non-occlusion distance based on the character width, and the occlusion distance is less than the non-occlusion distance; if the minimum distance is less than the occlusion distance, determining that the line of text is the abnormal text; if the minimum distance is greater than or equal to the occlusion distance, determining whether the line of text is the abnormal text based on the non-occlusion distance and / or the number of characters in the line of text; if the minimum distance is less than the non-occlusion distance, or the minimum distance and the number of characters in the line of text both meet preset conditions, determining whether the line of text is the abnormal text based on whether the adjacent lines of text in the same paragraph as the line of text are occluded.
[0029] In some embodiments, deleting the abnormal text from the target text includes: determining a first ratio of the number of characters in a line of text of the abnormal text to the number of characters in a line of text of the target text;
[0030] Determine a second ratio of the number of lines of text in the abnormal text to the number of first lines of text in the target text; if the larger of the first ratio and the second ratio is less than or equal to a tenth preset value, delete the abnormal text; if the larger of the first ratio and the second ratio is greater than the tenth preset value, delete the target text.
[0031] In some embodiments, before obtaining the target text in the page image of the page to be read aloud and the text features of the target text in the page image, the method further includes: based on a pre-built text recognition model and / or text feature recognition model, respectively identifying the target text and / or text features in the page image of the page to be read aloud; if the target text and / or the text features are not identified, outputting an abnormal prompt; if the target text and the text features are identified and the trapezoidal correction parameters sent by the terminal device are received, performing inverse deformation processing on the page image according to the trapezoidal correction parameters.
[0032] According to a second aspect of an embodiment of the present application, there is provided a text data processing apparatus, comprising:
[0033] A data acquisition unit is used to acquire the target text in the page image of the page to be read aloud, as well as the text features of the target text in the page image; an abnormal text detection unit is used to perform abnormality detection on the target text based on at least one detection dimension and according to the text features and the detection threshold corresponding to the detection dimension to determine the abnormal text in the target text, wherein the detection dimension includes angle abnormality detection, position abnormality detection and occlusion abnormality detection; an abnormal text deletion unit is used to delete the abnormal text from the target text.
[0034] According to a third aspect of an embodiment of the present application, a text data processing device is provided, comprising a processor and a memory, wherein the memory stores computer program instructions that can be executed by the processor, and when the processor executes the computer program instructions, the steps of the method described in any one of the first aspects above are implemented.
[0035] According to a fourth aspect of an embodiment of the present application, a computer-readable storage medium is provided, in which computer program instructions are stored. When the computer program instructions are executed by a processor, the processor is prompted to implement the steps of the method described in any one of the first aspects above.
[0036] In the present application, the target text in the page image of the page to be read aloud and the text features of the target text in the page image are obtained; based on at least one detection dimension, the target text is subjected to anomaly detection according to the text features and the detection threshold corresponding to the detection dimension to determine the abnormal text in the target text, wherein the detection dimension includes at least one of angle anomaly detection, position anomaly detection, and occlusion anomaly detection; and the abnormal text is deleted from the target text. The above scheme can delete the abnormal text in the target text, so that the abnormal text will not be converted into audio, thereby improving the accuracy of the text to be converted into audio and avoiding the playback of audio with disordered semantics.
[0037] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The accompanying drawings are incorporated into and constitute a part of the specification, illustrating embodiments consistent with the present application and, together with the specification, explaining the principles of the present application. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without inventive effort. In the drawings:
[0039] FIG1 shows a flowchart of an application scenario of a text data processing method according to an embodiment;
[0040] FIG2 shows a schematic flow chart of a text data processing method according to an embodiment;
[0041] FIG3 shows a detailed schematic diagram of step 202 in one embodiment;
[0042] FIG4 shows a schematic diagram of target paragraph text and non-target paragraph text in one embodiment;
[0043] FIG5 shows a detailed schematic diagram of step 202 in another embodiment;
[0044] FIG6 is a schematic diagram showing a page image containing a line of text in an abnormal position in one embodiment;
[0045] FIG7 is a schematic diagram showing a page image including a shadow area in one embodiment;
[0046] FIG8 shows a detailed schematic diagram of step 202 in yet another embodiment;
[0047] FIG9 is a schematic diagram showing an image of a page containing text obscured by a hand in one embodiment;
[0048] FIG10 shows a block diagram of a text data processing apparatus according to an embodiment;
[0049] FIG11 shows a schematic structural diagram of a text data processing device in one embodiment. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0051] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner.In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps etc. can be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0052] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0053] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0054] It should also be noted that the terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that shown or described.
[0055] In order to enable those skilled in the art to better understand the present application, the application scenario involved in the present application is first briefly described with reference to FIG1 .
[0056] Figure 1 shows a flowchart of an application scenario for a text data processing method according to one embodiment. As shown in Figure 1, after a terminal device captures an image of the page to be read aloud, the cloud performs preprocessing, text recognition, layout analysis, and text-to-audio conversion on the image before sending the audio to the terminal device. The text data processing method in this embodiment of the application can be performed after the layout analysis step and before the text-to-audio step.
[0057] Understandably, due to the limited field of view of the terminal device's camera, if the page to be read is too large, text truncation may occur. If the page to be read is placed in an offset position, text truncation may also occur. If the user places their hand on the page to be read and does not lift it, part of the text may be obscured. In these situations, directly converting text to audio may result in the playback of audio with garbled meanings.
[0058] It should be noted that the text data processing method provided in the embodiment of the present application can be executed by the cloud. However, in other embodiments of the present application, the terminal device can also have similar functions to the cloud to execute the text data processing solution provided in the embodiment of the present application.
[0059] In an embodiment of the present application, the target text in the page image of the page to be read aloud and the text features of the target text in the page image are obtained; based on at least one detection dimension, anomaly detection is performed on the target text according to the text features to determine abnormal text in the target text; and the abnormal text is deleted from the target text. The above scheme not only deletes the abnormal text in the target text, preventing the abnormal text from being converted into audio, improving the accuracy of the text to be converted into audio, and avoiding the playback of semantically disordered audio, but also outputs a corresponding abnormal prompt based on the determined abnormal text to remind the user of the reason for the abnormal reading of the page.
[0060] FIG2 is a flow chart of a text data processing method according to an embodiment of the present invention. As shown in FIG2 , the text data processing method may include the following steps 201 to 203 .
[0061] Step 201 : Acquire target text in a page image of a page to be read aloud, and text features of the target text in the page image.
[0062] The page to be read aloud refers to a page placed within the field of view of the camera of the terminal device, which can be a picture book, chapter book or other reading material, and the embodiment of the present application is not limited to this.
[0063] In some embodiments, the target text and / or text features in the page image of the page to be read aloud can be identified based on a pre-built text recognition model and / or text feature recognition model; if the target text and / or text features are not identified, an abnormal prompt is output.
[0064] During implementation, the text recognition model may be an Adaptive Bezier-Curve Network (ABCNet), a Pyramid Grafting Network (PGNet), or other models, and the present application is not limited thereto. If the target text is not recognized based on the text recognition model, an exception prompt is output.
[0065] The text feature recognition model may include models such as a layout analysis model. The text features may be features such as the number of the first line of text in the target text, the typeset direction of each line of text in the target text, the number of characters in a line of text, and the height of a line of text. This is not limited in the embodiments of the present application. The typeset direction of each line of text can be determined by the layout analysis model; and features such as the number of the first line of text, the number of characters in a line of text, and the height of a line of text can be determined by analyzing the output results of the text recognition model. If the text feature is not recognized based on the text feature recognition model, an exception prompt is output.
[0066] In some embodiments, the abnormal prompt may refer to an abnormal code, and the terminal device may store audio corresponding to different abnormal codes. When the terminal device receives the abnormal code, it may play the audio corresponding to the abnormal code to remind the user that an abnormality occurs during the reading process.
[0067] It is understandable that if the target text is recognized based on the text recognition model and the text features are recognized based on the text feature model, the target text and the text features are obtained, and anomaly detection is performed based on the target text and the text features.
[0068] Step 202: Based on at least one detection dimension, perform anomaly detection on the target text according to the text features and the detection threshold corresponding to the detection dimension to determine abnormal text in the target text, wherein the detection dimension includes angle anomaly detection, position anomaly detection, and occlusion anomaly detection.
[0069] During the implementation process, the detection dimension is at least one of angle anomaly detection, position anomaly detection, and occlusion anomaly detection. The embodiment of the present application can use a single detection dimension for anomaly detection, or can use multiple detection dimensions for anomaly detection. Through anomaly detection, abnormal text can be determined from the target text, thereby improving the accuracy of the text to be converted into audio.
[0070] In some embodiments, anomaly detection can be performed on the target text based on the text features in three dimensions: angle anomaly detection, position anomaly detection, and occlusion anomaly detection.
[0071] It should be noted that different detection dimensions correspond to different detection thresholds. For example, when the detection dimension is angle anomaly detection, the corresponding detection threshold can be the image angle threshold of the page image; when the detection dimension is position anomaly detection, the corresponding detection threshold can be the boundary distance threshold between the line text and the boundaries of each image in the page image; when the detection dimension is occlusion anomaly detection, the corresponding detection threshold can be the occlusion ratio threshold between the line text and the occluded area in the page image.
[0072] The specific method of performing anomaly detection based on the above-mentioned detection dimensions will be described below in the form of an example, and the embodiments of the present application are not limited thereto.
[0073] Step 203: Delete abnormal text from the target text.
[0074] After the abnormal text is determined, it can be determined whether to delete only the abnormal text or delete all the target text according to whether the proportion of the abnormal text reaches a preset value.
[0075] In some embodiments, a first ratio of the number of characters in a line of text of the abnormal text to the number of characters in a line of text of the target text can be determined; a second ratio of the number of characters in a line of text of the abnormal text to the number of characters in the first line of text of the target text can be determined; if the larger of the first ratio and the second ratio is less than or equal to a tenth preset value, the abnormal text is deleted; if the larger of the first ratio and the second ratio is greater than the tenth preset value, the target text is deleted.
[0076] Considering that the number of characters in a line of text is greater when the target text is English text, in this case, the first ratio may be the ratio of the number of words in the abnormal text to the number of words in the target text.
[0077] If the larger of the first ratio and the second ratio is less than or equal to the tenth preset value, only the abnormal text needs to be deleted, and the target text after deleting the abnormal text is converted into audio, which is used to be played on the terminal device.
[0078] If the larger of the first ratio and the second ratio is greater than the tenth preset value, for user experience considerations, the target text can be deleted completely, that is, the target text is not converted into audio, and a specific exception code is output to the terminal device, so that the terminal device plays the audio corresponding to the exception code.
[0079] The embodiment of the present application obtains the target text in the page image of the page to be read aloud, as well as the text features of the target text in the page image; performs anomaly detection on the target text based on at least one detection dimension according to the text features and the detection threshold corresponding to the detection dimension to determine abnormal text in the target text, wherein the detection dimension includes angle anomaly detection, position anomaly detection, and occlusion anomaly detection; and deletes the abnormal text from the target text. The above scheme can delete the abnormal text in the target text, so that the abnormal text will not be converted into audio, thereby improving the accuracy of the text to be converted into audio and avoiding the playback of audio with disordered semantics.
[0080] FIG3 illustrates a detailed schematic diagram of step 202 in one embodiment. In this embodiment, the detection dimension includes angle anomaly detection, the detection threshold corresponding to the detection dimension includes an image angle threshold, and the text features include the number of the first line of text in the target text, the layout direction of each line of text in the target text, and the number of characters in each line of text. As shown in FIG3 , performing anomaly detection on the target text based on the text features and the detection threshold corresponding to the detection dimension may include the following steps:
[0081] Step 301, determining the text angle of each line of text according to the typesetting direction of each line of text;
[0082] Step 302, determining the paragraph text angle of each paragraph text in the page image according to the number of text line characters and the text line angle of each text line;
[0083] Step 303: determining an image angle of the page image based on the number of first lines of text and the text angles of each line of text;
[0084] Step 304 : performing anomaly detection on the target text based on the paragraph text angle, the image angle, and the image angle threshold.
[0085] It should be noted that due to the randomness of the angle at which the page to be read is positioned, the page to be read may be significantly tilted from the perspective of the terminal device's camera, resulting in problems such as text out of bounds and exceeding the optimal support angle of the text recognition model. Furthermore, text sorting during layout analysis also has an optimal support angle range. Exceeding this range can easily lead to errors in text sorting. Therefore, to avoid text recognition errors or text sorting errors, it is necessary to use the line text angle to assess whether the image angle is too large.
[0086] In step 301, since the page to be read aloud may contain text arranged vertically or horizontally, it is possible to first determine whether the detection frame is vertical or horizontal relative to the X-axis of the page image coordinate system based on the aspect ratio of the detection frame of the line text, so as to determine the layout direction of the line text, and then calculate the line text angle based on the layout direction.
[0087] If the typesetting direction is horizontal relative to the X-axis of the page image coordinate system, the line text angle is the acute angle formed by the detection box of the line text and the X-axis of the page image coordinate system; if the typesetting direction is vertical relative to the X-axis of the page image coordinate system, the line text angle is the acute angle formed by the detection box of the line text and the Y-axis of the page image coordinate system.
[0088] In step 302, paragraph text refers to text present in the form of paragraphs within a page image, and is composed of lines of text within the paragraphs. Within each paragraph of the target text, target lines of text having a line text character count greater than a first preset value can be determined; and the average of the line text angles of the target lines of text within each paragraph of the target text can be determined as the paragraph text angle of each paragraph.
[0089] Among them, the first preset value can be set according to actual conditions. In some embodiments, the first preset value can be 1. At this time, the paragraph text angle is the average value of the line text angles of all non-single-character line texts in the paragraph.
[0090] It's important to note that if a line of text has fewer characters, the detection box length will be smaller, resulting in inaccurate angle estimation. By excluding the angles of lines of text with fewer characters when calculating paragraph angles, the accuracy of paragraph angle estimation is improved.
[0091] In step 303, if the number of the first line of text is less than the second preset value, the average value of the line text angles of all the lines of text in the target text is determined as the image angle; if the number of the first line of text is greater than or equal to the second preset value, the median of the line text angles of all the lines of text in the target text is determined as the image angle.
[0092] It can be understood that if the first line of text is small, the difference between the line text angles is large, and the image angle calculated using the average value of the line text angles is more accurate; if the first line of text is large, the image angle calculated using the median of the line text angles is more accurate.
[0093] In step 304, due to the distortion of the page image captured by the camera, text that is parallel on the page to be read aloud may not be parallel in the page image. Therefore, the image angle estimation has certain inaccuracies. If the image angle is directly used to detect anomalies of the target text, the detection results will be inaccurate. Based on this, the embodiment of the present application combines the paragraph text angle and the image angle to jointly detect anomalies of the target text, so as to provide tolerance for the estimation error of the image angle.
[0094] In some embodiments, a difference index is determined based on the number of first lines of text and the line text angles of each line of text, wherein the difference index is used to characterize the degree of difference between the angles of each line of text; if the difference index is greater than a difference index threshold, the target text is determined to be abnormal text; if the difference index is less than or equal to the difference index threshold, and the image angle is greater than or equal to the image angle threshold, and all paragraph text angles in the target text are greater than the image angle threshold, then the target text is determined to be abnormal text, wherein the image angle threshold is the recognition angle corresponding to the text recognition model.
[0095] The difference index may be an index such as text line angle variance, text line angle standard deviation, or the like. For example, if the difference index is text line angle variance, if the number of first text lines is less than a second preset value, the text line angle variance may be determined as the text line angle variance. If the number of first text lines is greater than or equal to the second preset value, the text line angle variance may be determined as the text line angle variance after removing a maximum value and a minimum value from the text line angles of all text lines.
[0096] It can be understood that if the difference index is greater than the difference index threshold, it is considered that the difference between the line text angles is large, and the target text can be directly determined to be abnormal text, and there is no need to make subsequent image angle judgments. Otherwise, the image angle needs to be compared with the image angle threshold.
[0097] If the image angle is less than the image angle threshold, the image angle is normal and the target text contains no abnormal text. If the image angle is greater than or equal to the image angle threshold, the angles of all paragraphs are compared with the image angle threshold. If the angles of all paragraphs are greater than the image angle threshold, the image angle is abnormal and the target text is identified as abnormal. If the angle of at least one paragraph is less than or equal to the image angle threshold, further analysis is required.
[0098] In some embodiments, the text features also include the number of paragraph text characters and the number of second lines of text in each paragraph text in the target text. The text features are identified by a text feature recognition model from the page image. If the difference index is less than or equal to the difference index threshold, and the image angle is greater than or equal to the image angle threshold, and there is at least one paragraph text angle in the target text that is less than or equal to the image angle threshold, then the target paragraph text is determined from each paragraph text, wherein the ratio between the number of paragraph text characters in the target paragraph text and the number of second lines of text corresponding to it is less than or equal to a third preset value; if the maximum paragraph text angle among the paragraph text angles of each paragraph text is greater than or equal to a fourth preset value, and there is at least one target paragraph text whose paragraph text angle is greater than the fourth preset value, then the target text is determined as abnormal text, wherein the fourth preset value is the recognition angle corresponding to the text feature recognition model; if the maximum paragraph text angle is less than the fourth preset value, and the paragraph text angles of all target paragraph texts in the target text are greater than the image angle threshold, then the target text is determined as abnormal text.
[0099] During the implementation process, the recognition angle corresponding to the text recognition model can be the angle that the text recognition model supports better, and the recognition angle corresponding to the text feature recognition model can be the maximum angle supported by the text feature recognition model. These two recognition angles can be used to design multi-level angle thresholds to perform angle anomaly detection on the target text.
[0100] Figure 4 shows a schematic diagram of target paragraph text and non-target paragraph text in one embodiment. As shown in Figure 4, non-target paragraph text is text independent of the target paragraph text, such as text within a bubble box. In picture books, non-target paragraph text is usually presented as dialogue or modal particles in illustrations. Its semantics are independent of the semantics of the target paragraph text, and it is necessary to avoid being read aloud in the target paragraph text.
[0101] It can be understood that if the ratio between the number of paragraph text characters in the paragraph text and the number of its corresponding second line text is less than or equal to the third preset value, the paragraph text is determined to be the target paragraph text; otherwise, the paragraph text is determined to be a non-target paragraph text.
[0102] During implementation, using Chinese picture books as an example, if the ratio between the number of characters in a paragraph and the number of text in the second line is less than or equal to 1.2, the paragraph is determined to be the target paragraph. Using English picture books as an example, if the ratio between the number of words in a paragraph and the number of text in the second line is less than or equal to 1.2, the paragraph is determined to be the target paragraph.
[0103] If the maximum paragraph text angle among the paragraph text angles of each paragraph text is greater than or equal to the fourth preset value, and there is at least one target paragraph text whose paragraph text angle is greater than the fourth preset value, it means that the paragraph text angle of the target paragraph text does not meet the angle corresponding to the text feature recognition model, and the target text needs to be determined as an abnormal text; if the paragraph text angles of all target paragraph texts are greater than the fourth preset value, it is determined that there is no abnormal text in the target text.
[0104] If the maximum paragraph text angle is less than the fourth preset value, and the paragraph text angles of all target paragraph texts are greater than the image angle threshold, it means that the paragraph text angles of all target paragraph texts do not meet the angle corresponding to the text recognition model, and the target text needs to be determined as abnormal text; if there is at least one target paragraph text whose paragraph text angle is less than or equal to the image angle threshold, it is determined that there is no abnormal text in the target text.
[0105] In some embodiments, the text features also include the paragraph area of each paragraph text in the target text. The paragraph area is the area occupied by the paragraph text in the page image. The two paragraph texts with the largest paragraph area can be determined from the various paragraph texts; if the difference between the paragraph text angles of the two paragraph texts is greater than a fifth preset value, the target text is determined as abnormal text.
[0106] It is understandable that the page to be read aloud may not be fully expanded because it is lifted up by the user. If the number of paragraph texts in the target text is greater than or equal to 2, and the paragraph text angles of the two paragraph texts with the largest paragraph area are too different, it means that the page to be read aloud is not fully expanded, and the target text needs to be determined as abnormal text.
[0107] In some embodiments, the two paragraph texts include a first paragraph text and a second paragraph text, the paragraph area of the second paragraph text is smaller than the paragraph area of the first paragraph text, and the proportion of the paragraph area of the second paragraph text to the page image area is greater than or equal to a preset proportion.
[0108] Among them, the preset ratio can be set according to the specific situation, which can be 1 / 10, 1 / 8 or other ratios, and the embodiment of the present application is not limited to this.
[0109] By limiting the paragraph area of the second paragraph text, the text in the bubble box and the text in the page picture can be excluded. This type of text is generally small in area and has random angles. Using the paragraph area of the paragraph text after excluding non-this type of text to perform abnormal judgment can improve the accuracy of the judgment.
[0110] The embodiment of the present application realizes angle anomaly detection of the target text through the above-mentioned scheme. During the angle anomaly detection process, the line text angle is used to estimate the image angle, and the paragraph text angle and the image angle are used to jointly determine the abnormal text. The target text with an image angle that is too large but a paragraph text angle that meets the conditions is retained, thereby achieving fault tolerance for image angle estimation errors and improving the accuracy of angle anomaly detection.
[0111] FIG5 illustrates a detailed schematic diagram of step 202 in another embodiment. In this embodiment, the detection dimension includes position anomaly detection, the detection threshold corresponding to the detection dimension includes a boundary distance threshold, and the text feature includes the distance between each line of text in the target text and each image boundary in the page image. As shown in FIG5 , performing anomaly detection on the target text based on the text feature and the detection threshold corresponding to the detection dimension may include the following steps:
[0112] Step 501 : performing abnormality detection on the target text according to the distance between each line of text and each image boundary in the page image and a boundary distance threshold.
[0113] Figure 6 shows a schematic diagram of a page image containing abnormally positioned lines of text in one embodiment. As shown in Figure 6, due to the limited camera range, when the page to be read aloud is large or positioned offset, some text may exceed the camera range. In this case, such truncated text needs to be deleted.
[0114] It can be understood that the image boundaries in the page image include the upper image boundary, the lower image boundary, the left image boundary and the right image boundary, and the boundary distance thresholds between the line text and each boundary can be the same or different, and the embodiments of the present application do not limit this.
[0115] In some embodiments, if the distance between a line of text and each image boundary in the page image is greater than a boundary distance threshold, the line of text may be determined to be abnormal text.
[0116] In some other embodiments, the text features further include language information of each line of text in the target text, and step 501 may include the following steps:
[0117] Step 5011, determining candidate texts in the target text based on the distances between each line of text and each image boundary in the page image and a boundary distance threshold, and determining the abnormality type of the candidate texts;
[0118] Step 5012: If the language of the candidate text is a preset language based on the language information, then determine the abnormal text from the candidate text based on the abnormality type and the preset screening rules;
[0119] Step 5013: If the language of the candidate text is not the preset language, the candidate text is determined to be an abnormal text.
[0120] In step 5011, if the distance between the line text and the left image boundary is less than the boundary distance threshold, the target text is determined to be a candidate text, and the exception type is beyond the left image boundary; if the distance between the line text and the right image boundary is less than the boundary distance threshold, the target text is determined to be a candidate text, and the exception type is beyond the right image boundary; if the distance between the line text and the upper image boundary is less than the boundary distance threshold, the target text is determined to be a candidate text, and the exception type is beyond the upper image boundary.
[0121] Referring back to Figure 6, since the camera has a certain pitch angle, the text closer to the bottom of the shooting range appears smaller in the page image and has a lower resolution. This type of text is the line text in the error-prone area at the bottom of the page image and is easily misidentified. It can be deleted as abnormal text.
[0122] In some embodiments, the text features also include the line text height of each line text in the target text, which can determine the target distance between each line text and the lower image boundary; if the target distance is less than the boundary distance threshold, the corresponding line text is determined as candidate text; if the target distance is greater than or equal to the boundary distance threshold, then based on the line text height of the corresponding line text, the target distance and the height of the target area in the page image, it is determined whether the corresponding line text is a candidate text, and the maximum distance between the target area and the lower image boundary is less than a sixth preset value.
[0123] It can be understood that the height of the target area in the page image can be understood as the height of the error-prone area. This error-prone area is the area at the bottom of the page image that the camera tends to misidentify. This height can be calculated by multiplying the height of the page image by the bottom position limit ratio. The boundary distance threshold and the height of the target area in the page image are used to decouple the bottom line text truncation judgment and the line text judgment in the error-prone area. The bottom line text truncation is still determined using the boundary distance threshold, while the line text judgment in the error-prone area is determined using the height of the target area in the page image. This avoids indiscriminate deletion of line text.
[0124] In some embodiments, the line text angle of each line of text in the target text can be determined; the target ranging point is determined on the detection frame of each line of text based on the line text angle of each line of text; and the distance between the target ranging point and the lower image boundary is determined as the distance between each line of text and the lower image boundary.
[0125] It should be noted that when calculating the distance between the line text and the lower image boundary, considering that the detection box of some line texts may be enlarged or reduced, in order to avoid errors in the line text detection box, the target distance measurement point can be determined based on the line text angle. If the line text angle is greater than the threshold, the midpoint of the two vertices in the line text detection box closest to the lower image boundary is used as the target distance measurement point; if the line text angle is less than or equal to the threshold, the four vertices of the detection box are used as the target distance measurement points, and the minimum value of the distance from the four vertices to the lower image boundary is used as the distance between the line text and the lower image boundary.
[0126] It is understandable that the distance between the line text and other image boundaries can also be calculated through similar steps as above, and the embodiments of the present application will not be repeated here.
[0127] In some embodiments, candidate texts can be determined by the following steps: determining multiple distance ranges based on the height of the target area; determining the line text height threshold corresponding to each distance range, wherein the upper limit value of the distance range is negatively correlated with the line text height threshold corresponding to the distance range; using the line text height threshold corresponding to the distance range to which the target distance belongs as the target height threshold, if the line text height of the corresponding line text is less than or equal to the target height threshold, then determining that the corresponding line text is a candidate text.
[0128] Taking the example that the height of the target area is bottom_border_thre and there are two distance ranges, it can be determined that the first distance range is bottom_border_thre / 2 to bottom_border_thre, and the second distance range is less than bottom_border_thre / 2. The first line text height threshold corresponding to the first distance range is bottom_box_h_thre1, which corresponds to the height limit of the inline text in the area defined by the first distance range. The second line text height threshold corresponding to the second distance range is bottom_box_h_thre2, which corresponds to the height limit of the inline text in the area defined by the second distance range, that is, the height limit of the inline text in the area at the bottom of the page image, wherein the second line text height threshold needs to be greater than the first line text height threshold in order to retain text with a higher inline text height in the area.
[0129] During the implementation process, if the target distance belongs to the first distance range and the line text height of the corresponding line text is less than or equal to the first line text height threshold, the corresponding line text is determined to be candidate text; if the target distance belongs to the second distance range and the line text height of the corresponding line text is less than or equal to the second line text height threshold, the corresponding line text is determined to be candidate text.
[0130] In step 5012, the preset language refers to a Latin language, such as English, French, etc., in which the first letter of a sentence is usually capitalized. If the language of the candidate text is the preset language, the candidate text needs to be screened again.
[0131] In some embodiments, if the exception type is exceeding the left image boundary, the candidate text containing preset characters representing the beginning of a sentence is determined as the text to be retained; if the exception type is exceeding the right image boundary, the text to be retained is determined based on whether the candidate text contains preset characters representing the end of a sentence; if the exception type is exceeding the upper image boundary, the text to be retained is determined from the candidate texts whose line text angle is greater than a seventh preset value; if the exception type is exceeding the lower image boundary, the text to be retained is determined from the candidate texts whose line text angle is greater than an eighth preset value; abnormal text is determined from the text to be retained, and the line text in the candidate text other than the text to be retained is determined as abnormal text.
[0132] For candidate text that extends beyond the left image boundary, if it contains a predefined character that indicates the beginning of a sentence, such as a capital letter or a punctuation mark that indicates the beginning of a sentence, then it is determined to be the text to be retained. Of course, if the candidate text contains a capital letter that indicates the beginning of a sentence, but that capital letter is a bullet point, then the candidate text cannot be determined to be the text to be retained.
[0133] For candidate text that exceeds the right image boundary, if it contains preset characters representing the end of a sentence, such as a period, and the next line of text of the candidate text in the same paragraph starts with an uppercase letter, it is determined to be text to be retained.
[0134] For candidate text that exceeds both the left and right image boundaries, two judgments are required to determine whether it is the text to be retained.
[0135] For candidate text that exceeds the upper image boundary, if the line text angle is greater than the seventh preset value, and if the reading order of the page to be read is from left to right, it can be determined whether the candidate text is truncated at the beginning or end of the sentence based on whether the line text angle is greater than 90 degrees. If it is truncated at the beginning of the sentence, further judgment can be made based on whether it exceeds the left image boundary; if it is truncated at the end of the sentence, further judgment can be made based on whether it exceeds the right image boundary to determine the text to be retained; if the candidate text only has one word, the preset word library can be used to search for the word based on the remaining characters in the word. If the word is found, the candidate text is determined to be the text to be retained.
[0136] For candidate text that exceeds the lower image boundary, similar steps as those for exceeding the upper image boundary can be used to determine the text to be retained. During implementation, if the candidate text is located in an area prone to recognition errors, it can be directly determined as abnormal text. For candidate text with a line text angle less than or equal to the eighth preset value and a line text height greater than a threshold, a word search can be performed in a preset vocabulary. If all words in the candidate text can be found, the candidate text is determined to be the text to be retained.
[0137] In some embodiments, whether the text to be retained is abnormal text may be determined based on location information of abnormal text in the same paragraph as the text to be retained.
[0138] Specifically, for the text to be retained that exceeds the left image boundary or the right image boundary, if the previous line of text and the next line of text of the text to be retained in the same paragraph are both abnormal text, the text to be retained is determined to be abnormal text.
[0139] If there is abnormal text that exceeds the upper image boundary within the same paragraph, the lines of text preceding the abnormal text in the candidate text will remain abnormal. Therefore, if the text to be retained is arranged before the abnormal text that exceeds the upper image boundary within the same paragraph, the text to be retained will be determined as abnormal text.
[0140] If there is abnormal text that exceeds the lower image boundary within the same paragraph, all lines of text following the abnormal text in the candidate text will remain abnormal. Therefore, if the text to be retained is arranged after the abnormal text that exceeds the lower image boundary within the same paragraph, the text to be retained will be determined as abnormal text.
[0141] In some embodiments, in response to a shadow occlusion judgment instruction sent by the terminal device, it is also possible to judge whether each line of text is obscured by the shadow area according to the shadow area parameters of the terminal device; and determine the line of text obscured by the shadow area as abnormal text.
[0142] Figure 7 shows a schematic diagram of a page image containing a shaded area in one embodiment. As shown in Figure 7, some terminal devices may block the camera to a certain extent due to reasons such as camera model or appearance design, and thus there will be a shaded area in the page image, which may obstruct the text line.
[0143] The terminal device can send shadow area parameters to the cloud at the same time as sending the shadow occlusion judgment instruction. When the cloud performs anomaly detection on the line text corresponding to the shadow area, it will determine the line text occluded by the shadow area as abnormal text.
[0144] The shadow area parameters may be the coordinates of edge positioning points of the shadow area in different directions, and the directions may be upper left, upper right, lower left, lower right, etc., which are specifically determined according to the design of the camera.
[0145] During implementation, whether each line of text is blocked by the shadow area can be determined based on whether the minimum distance between the four vertices of the detection box of each line of text and the shadow area exceeds a threshold.
[0146] The embodiment of the present application realizes position anomaly detection of the target text through the above-mentioned scheme. During the position anomaly detection process, after determining the candidate text in the target text by using the distance between the line text and the boundaries of each image in the page image, a deletion protection mechanism is established, that is, the candidate text is screened twice according to the language information, anomaly type and preset screening rules, thereby reducing the amount of line text deleted and avoiding deleting too many line texts.
[0147] FIG8 shows a detailed schematic diagram of step 202 in another embodiment. In this embodiment, the detection dimension includes occlusion anomaly detection, the detection threshold corresponding to the detection dimension includes an occlusion ratio threshold, and the text feature includes the average character area of each line of text in the target text. As shown in FIG8 , performing anomaly detection on the target text based on the text feature and the detection threshold corresponding to the detection dimension may include the following steps:
[0148] Step 801, determining a blocked area from a page image;
[0149] Step 802, determining the blocked paragraph text corresponding to the blocked area from the target text;
[0150] Step 803 : performing anomaly detection on each line of text in the obscured paragraph text according to the average character area of each line of text and the obscured ratio threshold.
[0151] Figure 9 shows a schematic diagram of a page image containing text obscured by a hand in one embodiment. As shown in Figure 9, if the page to be read aloud is obscured by a hand, part of the line text will be obscured and truncated. If the line text is directly converted into audio, the problem of playing audio with disordered semantics will arise. Therefore, it is necessary to delete the line text obscured or truncated by the hand.
[0152] In step 801, the occluded area can be a hand area, or an area within the page image that is obstructed by other obstructions. For example, if the occluded area is a hand area, the page image can be fed into a hand segmentation model for prediction to obtain a hand mask. If the proportion of the hand mask to the page image meets a threshold, the hand mask is dilated to obtain the hand area. Otherwise, occlusion anomaly detection can be terminated.
[0153] In step 802 , if the occlusion area is located within the detection frame of a certain paragraph text, the paragraph text is the occlusion paragraph text; otherwise, the occlusion anomaly detection may be terminated.
[0154] In some embodiments, the detection frame of the paragraph text in the target text can be expanded according to a preset size to obtain an extended detection frame; the first intersection-and-union ratio (IoU) of the extended detection frame and the occluded area is determined; and the paragraph text corresponding to the extended detection frame whose first IoU is greater than a ninth preset value is determined as the occluded paragraph text corresponding to the occluded area.
[0155] The preset size may be determined based on an average value of the heights of the text lines in the paragraph text. For example, the average value of the heights of the text lines may be multiplied by a coefficient to obtain the preset size.
[0156] The first intersection-and-union ratio refers to the ratio of the intersection and union of the extended detection frame and the occluded area. If the first intersection-and-union ratio is greater than the ninth preset value, it means that the overlapping area between the extended detection frame and the occluded area is large, and the paragraph text corresponding to the extended detection frame can be determined as the occluded paragraph text.
[0157] In step 803, the text features also include the character width of each line of text in the target text. The second intersection-and-union (IoU) of the detection frame of each line of text and the occluded area can be determined based on the average character area of each line of text in the occluded paragraph text; if the second IoU is greater than or equal to the occlusion ratio threshold, each line of text is determined to be abnormal text; if the second IoU is less than the occlusion ratio threshold, whether each line of text is abnormal text is determined based on the minimum distance from the boundary of the occluded area to the detection frame of each line of text and the character width of each line of text.
[0158] It is understandable that camera distortion can cause horizontal lines of text to bend. To improve accuracy when determining occlusion anomalies, a polygonal detection frame can be used for the line text detection. This polygonal detection frame can be obtained in a variety of ways. For example, if the text recognition model is a segmentation model, the line text mask output by the segmentation model can be fitted with a polygon to obtain a polygonal detection frame. For English word detection, a polygonal detection frame can be obtained by concatenating word frames. For Chinese character detection, a polygonal detection frame can be obtained by concatenating single-character frames.
[0159] When a line of text is blocked by an occluder, the occluder may be inside or outside the detection frame of the line of text (for example, when the line of text is truncated, there may be multiple detection frames). Therefore, it is necessary to calculate the second intersection-over-union ratio and combine it with distance-assisted judgment to determine whether the line of text is abnormal text.
[0160] It should be noted that, considering that the lines of text may be arranged closely, expanding the detection frame of the line of text will lead to misjudgment of the line of text located above the obstruction. When calculating the second intersection-over-union ratio in step 803, the detection frame of the line of text does not need to be expanded. When performing abnormal detection of the line of text, two judgments are made in combination with the second intersection-over-union ratio and distance.
[0161] Since the scale of text lines on different pages to be read aloud varies greatly and the number of characters in a line of text is not fixed, the second intersection-in-parallel ratio can be calculated using the average character area. The second intersection-in-parallel ratio can be calculated using the following formula: Where IoU is the second intersection over union ratio, A1∩B1 is the intersection of the line text detection box and the occlusion area, n is a positive number, and char_area is the average character area. char_area can be obtained by dividing the area of the polygonal detection box by the number of line text characters in the detection box.
[0162] In some embodiments, the text features also include the number of characters in each line of text in the target text. For each line of text, the occlusion distance and the non-occlusion distance can be determined based on the character width, and the occlusion distance is less than the non-occlusion distance; if the minimum distance is less than the occlusion distance, the line of text is determined to be abnormal text; if the minimum distance is greater than or equal to the occlusion distance, then based on the non-occlusion distance and / or the number of characters in the line of text, it is determined whether the line of text is abnormal text.
[0163] Among them, the occlusion distance dist_thre1 can take the larger value between the preset value and 0.5 character width, for example, dist_thre1=max(6,charw0.5), and the non-occlusion distance dist_thre2 can take 2 character widths, that is, dist_thre2=charw2.
[0164] It should be noted that, considering that the scales of line texts in different pages to be read aloud vary greatly and the number of characters in a line text is not fixed, the embodiment of the present application uses character width to determine the occlusion distance and non-occlusion distance instead of using a fixed distance threshold, thereby increasing the accuracy of the judgment.
[0165] Occlusion can cause the text recognition model to identify outlier text boxes whose minimum distance is greater than or equal to the occlusion distance. For example, when a line of text is occluded by multiple fingers, the text recognition model will output a detection box corresponding to the text between the fingers. This outlier text box can be determined based on the unoccluded distance and / or the number of characters in the line of text.
[0166] In some embodiments, if the minimum distance is less than the non-obstructed distance, or the minimum distance and the number of characters in the line of text both meet preset conditions, whether the line of text is abnormal text is determined based on whether the adjacent line of text in the same paragraph as the line of text is obstructed.
[0167] During implementation, a line of text whose minimum distance is less than the non-occlusion distance, or a line of text whose minimum distance is greater than 1.5 times the non-occlusion distance and whose number of characters is less than 0.7 times the maximum number of characters in a line of text within a paragraph, can be determined as the line of text corresponding to a candidate outlier text box. For a line of text corresponding to a candidate outlier text box, whether the line of text is occluded in the adjacent lines of text in the same paragraph as the line of text is determined to determine whether it is the line of text corresponding to the outlier text box. If so, it is determined to be abnormal text.
[0168] If, within the same paragraph, both the previous and next lines of text corresponding to the line of text of the candidate outlier text box are blocked, the line of text is determined to be abnormal text; if neither is blocked, the line of text is determined not to be abnormal text; if only one line of text is blocked, based on whether the angle between the line between the center point of the detection box of the previous line of text and the center point of the detection box of the next line of text and the X-axis of the page image coordinate system is less than a threshold, it is determined whether the previous line of text and the next line of text are in the same line; if they are in the same line and other lines of text in the same line are also blocked, the line of text corresponding to the candidate outlier text box is determined to be abnormal text.
[0169] The embodiment of the present application realizes occlusion anomaly detection of the target text through the above scheme. During the occlusion anomaly detection process, the scale difference of the line text in the page to be read aloud is taken into account, the second intersection-over-union ratio of the detection box of the line text and the occlusion area is calculated using the average character area, and an adaptive distance threshold is designed based on the character width to determine the line text corresponding to the outlier text box, avoiding the problem of inaccurate judgment caused by using a fixed distance threshold for judgment. During the occlusion anomaly detection process, for the line text corresponding to the candidate outlier text box, the occlusion judgment of the adjacent line text in the same paragraph is added, and then the line text corresponding to the outlier text box generated by the occlusion and relatively far away from the occluding object can be determined as abnormal text, and then deleted, avoiding the conversion of the line text into speech and causing confusion in the semantics of the played audio.
[0170] In some embodiments, if the target text is recognized based on the text recognition model, the text features are recognized based on the text feature model, and the trapezoidal correction parameters sent by the terminal device are received, the page image is inversely deformed according to the trapezoidal correction parameters.
[0171] It should be noted that if the page image is a trapezoidal-corrected image based on the device parameters of the terminal device (such as the pitch angle, height, resolution, etc. of the camera), there will be a higher bottom text detection rate when text recognition is performed on the image.
[0172] After obtaining the target text and text features, the page image can be inversely deformed to obtain the original image, and then anomaly detection can be performed based on the original image, so that the boundary distance threshold can be accurately determined according to the original image during position anomaly detection, and the occlusion area can be segmented according to the original image during occlusion anomaly detection, thereby improving the accuracy of detection.
[0173] In some embodiments, the text features include the language information and / or typesetting direction of each line of text in the target text. The text data processing method may further include: in response to the target language information sent by the terminal device, determining the line of text in the target text whose language information is different from the target language information as abnormal text; and / or, in response to the target typesetting direction sent by the terminal device, determining the line of text in the target text whose typesetting direction is different from the target typesetting direction as abnormal text.
[0174] During implementation, for example, if the target language information corresponds to English, if there is at least one character in the line of text that is not English, and that character is not a number, punctuation mark, or special symbol, then the language information of the line of text is determined to be not English, and it can be identified as abnormal text. Of course, if there is no English line of text in the target text, an abnormality prompt can be output.
[0175] By determining the language information of each line of text in the target text, it is possible to restrict the language of the content read aloud, thereby improving the user experience.
[0176] For example, if the target text orientation is horizontal relative to the X-axis of the page image coordinate system, and the acute angle between the detection frame of a non-single-character line of text and the X-axis of the page image coordinate system is greater than 60 degrees, then the text orientation of the line of text is determined to be vertical, which is different from the target text orientation and can be identified as abnormal text. Of course, if there is no text with a horizontal orientation among the target text lines, an abnormality prompt can be output.
[0177] By determining the typesetting direction of each line of text in the target text, the typesetting direction of the content read aloud can be restricted, thereby improving the user experience.
[0178] The following describes an embodiment of the device of the present application, which can be used to execute the text data processing method in the above embodiment of the present application. For details not disclosed in the embodiment of the device of the present application, please refer to the embodiment of the text data processing method in the above embodiment of the present application.
[0179] Referring to Figure 10, a block diagram of a text data processing device in an embodiment of the present application is shown. As shown in Figure 10, the text data processing device in an embodiment of the present application may include: a data acquisition unit 1001, an anomaly detection unit 1002, and an abnormal text deletion unit 1003, wherein the data acquisition unit 1001 is used to acquire the target text in the page image of the page to be read aloud, as well as the text features of the target text in the page image; the anomaly detection unit 1002 is used to perform anomaly detection on the target text based on at least one detection dimension and according to the text features and the detection threshold corresponding to the detection dimension to determine abnormal text in the target text, wherein the detection dimension includes angle anomaly detection, position anomaly detection, and occlusion anomaly detection; and the abnormal text deletion unit 1003 is used to delete abnormal text from the target text.
[0180] In some embodiments, the detection dimension includes angle anomaly detection, the detection threshold corresponding to the detection dimension includes an image angle threshold, the text feature includes the number of the first line of text of the target text, and the typesetting direction and the number of characters in each line of text in the target text, and the anomaly detection unit 1002 is further configured to determine the line text angle of each line of text according to the typesetting direction of each line of text;
[0181] The paragraph text angle of each paragraph text in the page image is determined based on the number of line text characters and the line text angle of each line text; the image angle of the page image is determined based on the number of first line texts and the line text angle of each line text; and anomaly detection is performed on the target text based on the paragraph text angle, image angle and image angle threshold.
[0182] In some embodiments, the anomaly detection unit 1002 is also used to determine a target line text in each paragraph text in the target text whose line text character number is greater than a first preset value; and determine the average value of the line text angle of the target line text in each paragraph text as the paragraph text angle of each paragraph text.
[0183] In some embodiments, the anomaly detection unit 1002 is also used to determine the average value of the line text angles of all lines of text in the target text as the image angle if the number of the first line of text is less than a second preset value; if the number of the first line of text is greater than or equal to the second preset value, then determine the median of the line text angles of all lines of text in the target text as the image angle.
[0184] In some embodiments, the target text is identified from the page image by a text recognition model, and the anomaly detection unit 1002 is further used to determine a difference index based on the number of first lines of text and the line text angles of each line of text, wherein the difference index is used to characterize the degree of difference between the angles of each line of text; if the difference index is greater than a difference index threshold, the target text is determined to be an abnormal text; if the difference index is less than or equal to the difference index threshold, and the image angle is greater than or equal to the image angle threshold, and all paragraph text angles in the target text are greater than the image angle threshold, then the target text is determined to be an abnormal text, wherein the image angle threshold is the recognition angle corresponding to the text recognition model.
[0185] In some embodiments, the text features also include the number of paragraph text characters and the number of second lines of text in each paragraph text in the target text. The abnormality detection unit 1002 is also used to determine the target paragraph text from each paragraph text if the difference index is less than or equal to the difference index threshold, and the image angle is greater than or equal to the image angle threshold, and there is at least one paragraph text angle in the target text that is less than or equal to the image angle threshold, wherein the ratio between the number of paragraph text characters in the target paragraph text and the number of its corresponding second lines of text is less than or equal to a third preset value; if the maximum paragraph text angle among the paragraph text angles of each paragraph text is greater than or equal to a fourth preset value, and there is at least one target paragraph text whose paragraph text angle is greater than the fourth preset value, then the target text is determined to be abnormal text, wherein the fourth preset value is the recognition angle corresponding to the text feature recognition model; if the maximum paragraph text angle is less than the fourth preset value, and the paragraph text angles of all target paragraph texts in the target text are greater than the image angle threshold, then the target text is determined to be abnormal text.
[0186] In some embodiments, the text features also include the paragraph area of each paragraph text in the target text, where the paragraph area is the area occupied by the paragraph text in the page image. The abnormality detection unit 1002 is also used to determine the two paragraph texts with the largest paragraph area from the various paragraph texts; if the difference between the paragraph text angles of the two paragraph texts is greater than a fifth preset value, the target text is determined to be abnormal text.
[0187] In some embodiments, the two paragraph texts include a first paragraph text and a second paragraph text, the paragraph area of the second paragraph text is smaller than the paragraph area of the first paragraph text, and the proportion of the paragraph area of the second paragraph text to the page image area is greater than or equal to a preset proportion.
[0188] In some embodiments, the detection dimension includes position anomaly detection, the detection threshold corresponding to the detection dimension includes a boundary distance threshold, the text feature includes the distance between each line of text in the target text and each image boundary in the page image, and the anomaly detection unit 1002 is also used to perform anomaly detection on the target text based on the distance between each line of text and each image boundary in the page image and the boundary distance threshold.
[0189] In some embodiments, the text features also include language information of each line of text in the target text. The anomaly detection unit 1002 is further used to determine candidate text in the target text and determine the anomaly type of the candidate text based on the distance between each line of text and each image boundary in the page image and the boundary distance threshold; based on the language information, if the language of the candidate text is a preset language, the abnormal text is determined from the candidate text based on the anomaly type and preset screening rules; if the language of the candidate text is not the preset language, the candidate text is determined as abnormal text.
[0190] In some embodiments, the text features also include the line text height of each line text in the target text, the image boundary includes the lower image boundary, and the anomaly detection unit 1002 is also used to determine the target distance between each line text and the lower image boundary; if the target distance is less than the boundary distance threshold, the corresponding line text is determined as a candidate text; if the target distance is greater than or equal to the boundary distance threshold, it is determined whether the corresponding line text is a candidate text based on the line text height of the corresponding line text, the target distance and the height of the target area in the page image, wherein the maximum distance between the target area and the lower image boundary is less than a sixth preset value.
[0191] In some embodiments, the anomaly detection unit 1002 is also used to determine multiple distance ranges based on the height of the target area; determine the line text height threshold corresponding to each distance range, wherein the upper limit value of the distance range is negatively correlated with the line text height threshold corresponding to the distance range; use the line text height threshold corresponding to the distance range to which the target distance belongs as the target height threshold, and if the line text height of the corresponding line text is less than or equal to the target height threshold, determine that the corresponding line text is a candidate text.
[0192] In some embodiments, the data acquisition unit 1001 is also used to determine the line text angle of each line of text in the target text; determine the target ranging point on the detection frame of each line of text according to the line text angle of each line of text; and determine the distance between the target ranging point and the lower image boundary as the distance between each line of text and the lower image boundary.
[0193] In some embodiments, the image boundary includes a left image boundary, a right image boundary, an upper image boundary, and a lower image boundary. The anomaly detection unit 1002 is further configured to, if the anomaly type is exceeding the left image boundary, determine the candidate text containing the preset character representing the beginning of a sentence as the text to be retained;
[0194] If the exception type is exceeding the right image boundary, the text to be retained is determined based on whether the candidate text contains the preset character representing the end of the sentence; if the exception type is exceeding the upper image boundary, the text to be retained is determined from the candidate texts whose line text angle is greater than the seventh preset value; if the exception type is exceeding the lower image boundary, the text to be retained is determined from the candidate texts whose line text angle is greater than the eighth preset value; the abnormal text is determined from the text to be retained, and the line text in the candidate text other than the text to be retained is determined as the abnormal text.
[0195] In some embodiments, the anomaly detection unit 1002 is further configured to determine whether the text to be retained is an abnormal text based on position information of abnormal text in the same paragraph as the text to be retained.
[0196] In some embodiments, the anomaly detection unit 1002 is also used to respond to the shadow occlusion judgment instruction sent by the terminal device, and judge whether each line of text is obscured by the shadow area according to the shadow area parameters of the terminal device; and determine the line of text obscured by the shadow area as abnormal text.
[0197] In some embodiments, the detection dimension includes occlusion anomaly detection, the detection threshold corresponding to the detection dimension includes an occlusion ratio threshold, the text feature includes the average character area of each line of text in the target text, and the anomaly detection unit 1002 is also used to determine the occlusion area from the page image; determine the occluded paragraph text corresponding to the occlusion area from the target text; and perform anomaly detection on each line of text based on the average character area of each line of text in the occluded paragraph text and the occlusion ratio threshold.
[0198] In some embodiments, the anomaly detection unit 1002 is also used to expand the detection frame of the paragraph text in the target text according to a preset size to obtain an extended detection frame; determine a first intersection-and-union ratio (IoU) between the extended detection frame and the occluded area; and determine the paragraph text corresponding to the extended detection frame whose first IoU is greater than a ninth preset value as the occluded paragraph text corresponding to the occluded area.
[0199] In some embodiments, the text features also include the character width of each line of text in the target text, and the anomaly detection unit 1002 is also used to determine the second intersection-and-union (IoU) of the detection frame of each line of text and the occluded area based on the average character area of each line of text in the occluded paragraph text; if the second IoU is greater than or equal to the occlusion ratio threshold, each line of text is determined to be an abnormal text; if the second IoU is less than the occlusion ratio threshold, then based on the minimum distance from the boundary of the occlusion area to the detection frame of each line of text and the character width of each line of text, it is determined whether each line of text is an abnormal text.
[0200] In some embodiments, the text features also include the number of characters in each line of text in the target text. The anomaly detection unit 1002 is also used to determine the occlusion distance and non-occlusion distance for each line of text based on the character width, and the occlusion distance is less than the non-occlusion distance; if the minimum distance is less than the occlusion distance, the line of text is determined to be abnormal text; if the minimum distance is greater than or equal to the occlusion distance, whether the line of text is abnormal text is determined based on the non-occlusion distance and / or the number of characters in the line of text; if the minimum distance is less than the non-occlusion distance, or the minimum distance and the number of characters in the line of text both meet the preset conditions, whether the line of text is abnormal text is determined based on whether the adjacent lines of text in the same paragraph as the line of text are occluded.
[0201] In some embodiments, the abnormal text deletion unit 1003 is also used to determine a first ratio of the number of characters in a line of abnormal text to the number of characters in a line of target text; determine a second ratio of the number of lines of abnormal text to the number of the first line of text of the target text; if the larger of the first ratio and the second ratio is less than or equal to the tenth preset value, delete the abnormal text; if the larger of the first ratio and the second ratio is greater than the tenth preset value, delete the target text.
[0202] In some embodiments, the text data processing device also includes an image preprocessing unit (not shown), which is used to identify the target text and / or text features in the page image of the page to be read aloud based on a pre-built text recognition model and / or text feature recognition model; if the target text and / or text features are not identified, an abnormal prompt is output; if the target text and text features are identified, and the trapezoidal correction parameters sent by the terminal device are received, the page image is inversely deformed according to the trapezoidal correction parameters.
[0203] Based on the same inventive concept, an embodiment of the present application further provides a text data processing device. Referring to Figure 11, a structural schematic diagram of the text data processing device in an embodiment of the present application is shown. The text data processing device includes one or more memories 1104, one or more processors 1102, and at least one computer program (computer program instruction) stored in the memory 1104 and executable on the processor 1102. When the processor 1102 executes the computer program, the method described above is implemented.
[0204] In FIG11 , a bus architecture (represented by bus 1100) is shown. Bus 1100 may include any number of interconnected buses and bridges. Bus 1100 links various circuits, including one or more processors represented by processor 1102 and memory represented by memory 1104. Bus 1100 may also link various other circuits, such as peripherals, voltage regulators, and power management circuits, all of which are well known in the art and, therefore, will not be described further herein. Bus interface 1105 provides an interface between bus 1100 and receiver 1101 and transmitter 1103. Receiver 1101 and transmitter 1103 may be the same component, namely a transceiver, providing a unit for communicating with various other devices over a transmission medium. Processor 1102 is responsible for managing bus 1100 and general processing, while memory 1104 may be used to store data used by processor 1102 when performing operations.
[0205] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, in which computer program instructions are stored. When the computer program instructions are executed by a processor, the processor is prompted to implement the steps of the method as described above.
[0206] The functions described herein may be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored as one or more instructions or codes on or transmitted via a computer-readable medium. Other examples and implementations are within the scope and spirit of this application and the appended claims. For example, due to the nature of software, the functions described above may be implemented using software executed by a processor, hardware, firmware, hardwiring, or a combination of any of these. Furthermore, the functional units may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit.
[0207] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0208] The units described as separate components may or may not be physically separate, and the components of the control device may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0209] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store computer program instructions.
[0210] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for processing text data, characterized in that, Including: Obtain the target text in the page image of the page to be read aloud, and the text features of the target text in the page image; Based on at least one detection dimension, perform anomaly detection on the target text according to the text features and the detection threshold corresponding to the detection dimension to determine the abnormal text in the target text, where the detection dimension includes angle anomaly detection, position anomaly detection, and occlusion anomaly detection; Delete the abnormal text from the target text.
2. The method according to claim 1, characterized in that, The detection dimension includes angle anomaly detection, the detection threshold corresponding to the detection dimension includes an image angle threshold, the text features include the number of first-line texts of the target text, and the typesetting direction and the number of characters in each line text of the target text. Performing anomaly detection on the target text according to the text features and the detection threshold corresponding to the detection dimension includes: Correspondingly determine the line text angle of each line text according to the typesetting direction of each line text; Determine the paragraph text angle of each paragraph text in the page image according to the number of characters in each line text and the line text angle; Determine the image angle of the page image according to the number of first-line texts and the line text angle of each line text; Perform anomaly detection on the target text according to the paragraph text angle, the image angle, and the image angle threshold.
3. The method according to claim 2, characterized in that, The determining the paragraph text angle of each paragraph text in the page image according to the number of characters in each line text and the line text angle includes: Determine the target line text in each paragraph text of the target text where the number of characters in the line text is greater than a first preset value; Determine the average value of the line text angles of the target line text in each paragraph text as the paragraph text angle of each paragraph text.
4. The method according to claim 2, characterized in that, The determining the image angle of the page image according to the number of first-line texts and the line text angle of each line text includes: If the number of first-line texts is less than a second preset value, then determine the average value of the line text angles of all line texts in the target text as the image angle; If the number of first-line texts is greater than or equal to the second preset value, then determine the median of the line text angles of all line texts in the target text as the image angle.
5. The method according to claim 2, characterized in that, The target text is recognized from the page image by a text recognition model. Performing anomaly detection on the target text according to the paragraph text angle, the image angle, and the image angle threshold includes: Determine a difference index according to the number of first-line texts and the line text angle of each line text, where the difference index is used to characterize the degree of difference between the line text angles; If the difference index is greater than the difference index threshold, then determine the target text as the abnormal text; If the difference index is less than or equal to the difference index threshold, and the image angle is greater than or equal to the image angle threshold, and the angles of all paragraph texts in the target text are greater than the image angle threshold, then the target text is determined as the abnormal text, where the image angle threshold is the recognition angle corresponding to the text recognition model.
6. The method according to claim 5, characterized in that, The text features further include the number of paragraph text characters and the number of second-line texts in each paragraph text of the target text. The text features are recognized from the page image by a text feature recognition model. The abnormal detection of the target text based on the paragraph text angle, the image angle, and the image angle threshold further includes: If the difference index is less than or equal to the difference index threshold, and the image angle is greater than or equal to the image angle threshold, and there is at least one paragraph text angle in the target text that is less than or equal to the image angle threshold, then the target paragraph text is determined from each paragraph text, where the ratio between the number of paragraph text characters and the number of corresponding second-line texts in the target paragraph text is less than or equal to a third preset value; If the maximum paragraph text angle among the paragraph text angles of each paragraph text is greater than or equal to a fourth preset value, and there is at least one paragraph text angle of the target paragraph text that is greater than the fourth preset value, then the target text is determined as the abnormal text, where the fourth preset value is the recognition angle corresponding to the text feature recognition model; If the maximum paragraph text angle is less than the fourth preset value, and the angles of all target paragraph texts in the target text are greater than the image angle threshold, then the target text is determined as the abnormal text.
7. The method according to any one of claims 2 to 6, characterized in that, The text features further include the paragraph area of each paragraph text in the target text, and the paragraph area is the area occupied by the paragraph text in the page image. The abnormal detection of the target text based on the text features further includes: Determine the two paragraph texts with the largest paragraph areas from each paragraph text; If the difference between the paragraph text angles of the two paragraph texts is greater than a fifth preset value, then the target text is determined as the abnormal text.
8. The method according to claim 7, characterized in that, The two paragraph texts include a first paragraph text and a second paragraph text. The paragraph area of the second paragraph text is less than that of the first paragraph text, and the ratio of the paragraph area of the second paragraph text to the area of the page image is greater than or equal to a preset ratio.
9. The method according to claim 1, characterized in that, The detection dimension includes position abnormal detection. The detection threshold corresponding to the detection dimension includes a boundary distance threshold. The text features include the distances between each line text in the target text and each image boundary in the page image. The abnormal detection of the target text based on the text features and the detection threshold corresponding to the detection dimension includes: Perform abnormal detection on the target text according to the distances between each line text and each image boundary in the page image and the boundary distance threshold.
10. The method according to claim 9, characterized in that, The text feature further includes the language information of each line of text in the target text. The abnormal detection of the target text according to the distances between each line of text and each image boundary in the page image and the boundary distance threshold includes: Determining candidate texts in the target text according to the distances between each line of text and each image boundary in the page image and the boundary distance threshold, and determining the abnormal types of the candidate texts; According to the language information, if the language of the candidate text is a preset language, determining abnormal texts from the candidate texts according to the abnormal type and a preset screening rule; If the language of the candidate text is not the preset language, determining the candidate text as the abnormal text.
11. The method according to claim 10, wherein, The text feature further includes the line text height of each line of text in the target text. The image boundary includes the lower image boundary. Determining candidate texts in the target text according to the distances between each line of text and each image boundary in the page image and the boundary distance threshold includes: Determining the target distance between each line of text and the lower image boundary; If the target distance is less than the boundary distance threshold, determining the corresponding line of text as the candidate text; If the target distance is greater than or equal to the boundary distance threshold, determining whether the corresponding line of text is the candidate text according to the line text height of the corresponding line of text, the target distance, and the height of the target area in the page image, where the maximum distance between the target area and the lower image boundary is less than a sixth preset value.
12. The method according to claim 11, wherein, Determining whether the corresponding line of text is the candidate text according to the line text height of the corresponding line of text, the target distance, and the height of the target area in the page image includes: Determining a plurality of distance ranges according to the height of the target area; Determining the line text height threshold corresponding to each distance range, where the upper limit value of the distance range is negatively correlated with the line text height threshold corresponding to the distance range; Taking the line text height threshold corresponding to the distance range to which the target distance belongs as the target height threshold. If the line text height of the corresponding line of text is less than or equal to the target height threshold, determining the corresponding line of text as the candidate text.
13. The method according to claim 11, wherein, Further includes: Determining the line text angle of each line of text in the target text; Determining a target distance measurement point on the detection frame of each line of text according to the line text angle of each line of text; Taking the distance between the target distance measurement point and the lower image boundary as the distance between each line of text and the lower image boundary.
14. The method according to claim 10, wherein, The image boundary includes the left image boundary, the right image boundary, the upper image boundary, and the lower image boundary. Determining abnormal texts from the candidate texts according to the abnormal type and a preset screening rule includes: If the abnormal type is exceeding the left image boundary, determining the candidate text containing a preset character representing the start of a sentence as the text to be retained; If the abnormal type is exceeding the right image boundary, determining the text to be retained according to whether the candidate text contains a preset character representing the end of a sentence; If the abnormal type is exceeding the upper image boundary, determine the text to be retained from the candidate texts whose line text angle is greater than a seventh preset value; If the abnormal type is exceeding the lower image boundary, determine the text to be retained from the candidate texts whose line text angle is greater than an eighth preset value; Determine the abnormal text from the text to be retained, and determine the line text other than the text to be retained in the candidate texts as the abnormal text.
15. The method according to claim 14, wherein, The determining the abnormal text from the text to be retained includes: Determine whether the text to be retained is abnormal text according to the position information of the abnormal text in the same paragraph as the text to be retained.
16. The method according to any one of claims 10 to 15, wherein, It further includes: In response to a shadow occlusion judgment instruction sent by the terminal device, judge whether each line text is occluded by the shadow area according to the shadow area parameters of the terminal device; Determine the line text occluded by the shadow area as the abnormal text.
17. The method according to claim 1, wherein, The detection dimension includes occlusion anomaly detection, the detection threshold corresponding to the detection dimension includes an occlusion ratio threshold, the text feature includes the average character area of each line text in the target text, and the abnormal detection of the target text according to the text feature and the detection threshold corresponding to the detection dimension includes: Determine the occlusion area from the page image; Determine the occluded paragraph text corresponding to the occlusion area from the target text; Perform abnormal detection on each line text according to the average character area of each line text in the occluded paragraph text and the occlusion ratio threshold.
18. The method according to claim 17, wherein, The determining the occluded paragraph text corresponding to the occlusion area from the target text includes: Expand the detection frame of the paragraph text in the target text according to a preset size to obtain an expanded detection frame; Determine the first intersection over union of the expanded detection frame and the occlusion area; Determine the paragraph text corresponding to the expanded detection frame with the first intersection over union greater than a ninth preset value as the occluded paragraph text corresponding to the occlusion area.
19. The method according to claim 17, wherein, The text feature further includes the character width of each line text in the target text, and the abnormal detection of each line text according to the average character area of each line text in the occluded paragraph text and the occlusion ratio threshold includes: Determine the second intersection over union of the detection frame of each line text and the occlusion area according to the average character area of each line text in the occluded paragraph text; If the second intersection over union is greater than or equal to the occlusion ratio threshold, determine that each line text is the abnormal text; If the second intersection over union is less than the occlusion ratio threshold, determine whether each line text is the abnormal text according to the minimum distance from the occlusion area boundary to the detection frame of each line text and the character width of each line text.
20. The method according to claim 19, wherein, The text feature further includes the number of line text characters of each line text in the target text, and the determining whether each line text is the abnormal text according to the minimum distance from the occlusion area boundary to the detection frame of each line text and the character width of each line text includes: For each line text, determine an occlusion distance and a non-occlusion distance according to the character width, where the occlusion distance is less than the non-occlusion distance; If the minimum distance is less than the occlusion distance, determine that the line text is the abnormal text; If the minimum distance is greater than or equal to the occlusion distance, determine whether the line text is the abnormal text according to the non-occlusion distance and / or the number of characters in the line text; If the minimum distance is less than the non-occlusion distance, or both the minimum distance and the number of characters in the line text meet the preset conditions, determine whether the line text is the abnormal text according to whether the adjacent line text in the same paragraph as the line text is occluded; 21. The method according to claim 1, wherein, Deleting the abnormal text from the target text includes: Determine a first ratio of the number of characters in the line text of the abnormal text to the number of characters in the line text of the target text; Determine a second ratio of the number of lines of the abnormal text to the number of lines of the first line text of the target text; If the larger of the first ratio and the second ratio is less than or equal to a tenth preset value, delete the abnormal text; If the larger of the first ratio and the second ratio is greater than the tenth preset value, delete the target text; 22. The method according to claim 1, wherein, Before obtaining the target text in the page image of the page to be read aloud and the text features of the target text in the page image, the method further includes: Based on a pre-constructed text recognition model and / or text feature recognition model, respectively recognize the target text and / or text features in the page image of the page to be read aloud; If the target text and / or the text features are not recognized, output an abnormal prompt; If the target text and the text features are recognized and trapezoidal correction parameters sent by the terminal device are received, perform an inverse deformation process on the page image according to the trapezoidal correction parameters; 23. A text data processing device, wherein, including: A data acquisition unit for acquiring the target text in the page image of the page to be read aloud and the text features of the target text in the page image; An abnormal text detection unit for performing abnormal detection on the target text based on at least one detection dimension according to the text features and the detection threshold corresponding to the detection dimension to determine the abnormal text in the target text, where the detection dimension includes angle abnormal detection, position abnormal detection, and occlusion abnormal detection; An abnormal text deletion unit for deleting the abnormal text from the target text; 24. A text data processing device, comprising a processor and a memory, wherein, The memory stores computer program instructions that can be executed by the processor. When the processor executes the computer program instructions, the steps of the method according to any one of claims 1 to 22 are implemented; 25. A computer-readable storage medium, wherein, The computer-readable storage medium stores computer program instructions. When the computer program instructions are executed by the processor, the processor is prompted to implement the steps of the method according to any one of claims 1 to 22.
Citation Information
Patent Citations
Electronic picture book generation method and device, computer equipment and storage medium
CN114882505A
Document anomaly detection network model construction method and device, electronic equipment and medium
CN115035539A
Text detection method and device, equipment and storage medium
CN115273098A
OCR character recognition method, electronic equipment and storage medium
CN115457565A
Selective reading-out method having automatic header extracting function, and recording medium recording program therefor
JP2000352988A
Cited By
Method and device for testing page text display of in-vehicle infotainment system, and vehicle
CN120371711A