Page text checking method and device, electronic equipment, medium and product
By determining the text type based on page element tags, and using image processing and recognition technology to combine preset standard databases and translation databases for verification, the problem of manual verification dependence and low accuracy in the existing technology is solved, and efficient and accurate text verification is achieved.
Patent Information
- Application Number
- CN202510073154.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
In the prior art, page text verification relies on manual labor, resulting in accuracy dependent on the professional quality of the checker, which is inefficient and difficult to ensure the accuracy of the checking results.
By determining the text type based on page element labels, the first text is determined using image processing algorithms and image recognition technology, and the preset standard library and translation library are used for verification, improving the accuracy of text extraction and verification.
It improves the accuracy and efficiency of text verification, expands the scope of application of text verification methods, and can effectively process image text and truncated text.
Smart Images

Figure CN120012755A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of cloud computing, and in particular to a page text verification method, device, electronic device, medium and product. Background Art
[0002] With the development of cloud computing, the demand for multi-language internationalized cloud platforms is increasing. Therefore, in order to ensure the correctness of page texts in different languages, page texts need to be checked. In related technologies, page translation texts are usually checked manually after the platform is developed. This means that the accuracy of the check results depends only on the professional quality of the checkers, and the accuracy of the check is difficult to guarantee. At the same time, the check efficiency is low. Summary of the invention
[0003] The present disclosure provides a page text verification method, device, electronic device, medium and product to solve the problems in the related technology and improve the text verification efficiency and text verification accuracy.
[0004] The first aspect of the present disclosure proposes a page text verification method, which includes: determining the text type of the page based on the element tag of the page; determining the first text of the page based on the text type; determining the language environment of the page based on the parameter information of the page; and verifying the first text of the page using a preset standard library based on the first text and the language environment.
[0005] In some embodiments of the present disclosure, determining the text type based on the element tags of the page includes: when the element tag includes an image tag, determining the text type as an image text type; when the element tag includes a truncation tag, determining the text type as a truncation text type; when the element tag does not include an image tag and a truncation tag, determining the text type as a regular text type.
[0006] In some embodiments of the present disclosure, determining the first text of a page based on the text type includes: when the text type is an image text type, obtaining a regional screenshot of a preset area of the page to determine the first text using an image processing algorithm; when the text type is a truncated text type, extracting target text contained in the page text to determine the first text; when the text type is a regular text type, splitting the page text based on punctuation marks contained in the page text to determine the first text.
[0007] In some embodiments of the present disclosure, obtaining a regional screenshot of a preset area of a page to determine the first text using an image processing algorithm includes: based on the regional screenshot, using a preset image processing model to determine the predicted text boundary of the regional screenshot and the text probability corresponding to the predicted text area; based on the predicted text boundary and text probability, using a preset text probability threshold and a non-maximum suppression algorithm to determine the target text boundary; based on the target text boundary, using an image recognition algorithm to determine the first text.
[0008] In some embodiments of the present disclosure, obtaining a regional screenshot of a preset area of a page to determine the first text using an image processing algorithm includes: based on the regional screenshot, using a preset image processing model to determine the predicted text boundary of the regional screenshot and the text probability corresponding to the predicted text area; based on the predicted text boundary and text probability, using a preset text probability threshold and a non-maximum suppression algorithm to determine the target text boundary; based on the target text boundary, using an image recognition algorithm to determine the first text.
[0009] In some embodiments of the present disclosure, determining the language environment of the page based on parameter information of the page includes: determining the language environment of the page based on interface request header parameters and / or page routing path parameters of the page.
[0010] In some embodiments of the present disclosure, based on the first text and the language environment, the first text of the page is checked using a preset standard library, including: based on the language environment and the first data type field in the preset standard library, determining the first data corresponding to the language environment in the preset standard library; traversing the first data to determine whether there is first data matching the first text; when there is first data matching the first text, determining that the translation of the first text is correct; when there is no first data matching the first text, determining that the translation of the first text is abnormal, and checking the first data using the preset translation library, the data amount of the second data in the preset translation library is greater than or equal to the data amount of the first data in the preset standard library.
[0011] In some embodiments of the present disclosure, checking the first data using a preset translation library includes: determining the second data corresponding to the language environment in the preset translation library based on the language environment and the second data type field in the preset translation library; traversing the second data to determine whether there is second data matching the first text; when there is second data matching the first text, querying the first data based on the data identification field of the matching second data to determine whether the first text is translated correctly; when there is no second data matching the first text, determining that the first text translation is incorrect.
[0012] In some embodiments of the present disclosure, querying the first data based on the data identification field of the matched second data to determine whether the first text is translated correctly includes: traversing the first data to determine whether the first data contains the identification field; when the first data contains the identification field, determining that there is an error in the translation of the second data, and modifying the second data; when the first data does not contain the identification field, updating the preset standard library and / or the preset translation library.
[0013] The second aspect of the present disclosure provides a page text checking device, the device comprising: a first determining unit, used to determine the text type of the page based on the element tag of the page; a second determining unit, used to determine the first text of the page based on the text type; a third determining unit, used to determine the language environment of the page based on the parameter information of the page;
[0014] The checking unit is used to check the first text of the page based on the first text and the language environment by using a preset standard library.
[0015] The third aspect embodiment of the present disclosure proposes an electronic device, comprising: a processor and a memory for storing a computer program that can be run on the processor, wherein the processor, when used to run the computer program, executes the method described in the first aspect embodiment of the present disclosure.
[0016] The fourth aspect embodiment of the present disclosure proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the method described in the first aspect embodiment of the present disclosure.
[0017] The fifth aspect embodiment of the present disclosure provides a computer program product, including a computer program, which implements the method described in the first aspect embodiment of the present disclosure when executed by a processor.
[0018] In summary, the page text verification method proposed in the present disclosure includes: determining the text type of the page based on the element tag of the page; determining the first text of the page based on the text type; determining the language environment of the page based on the parameter information of the page; and verifying the first text of the page using a preset standard library based on the first text and the language environment. The method of the present disclosure improves the accuracy of text extraction by determining the text type and determining the first text in different ways according to different text types, and improves the accuracy of text verification by verifying the first text using a preset standard library. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.
[0020] Figure 1 A flowchart of a page text verification method provided by an embodiment of the present disclosure;
[0021] Figure 2 A schematic diagram of a flow chart of a page text verification method provided by an embodiment of the present disclosure;
[0022] Figure 3 A flowchart of another method for checking page text provided by an embodiment of the present disclosure;
[0023] Figure 4 A flowchart of another method for checking page text provided by an embodiment of the present disclosure;
[0024] Figure 5 A schematic diagram of a flow chart of a page text verification method provided by an embodiment of the present disclosure;
[0025] Figure 6 A flowchart of another method for checking page text provided by an embodiment of the present disclosure;
[0026] Figure 7 A flowchart of another method for checking page text provided by an embodiment of the present disclosure;
[0027] Figure 8 A schematic diagram of the structure of a page text checking device provided by an embodiment of the present disclosure;
[0028] Fig. 9 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0029] The embodiments of the present disclosure are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions.
[0030] With the development of cloud computing, the demand for multi-language internationalized cloud platforms is increasing. Therefore, in order to ensure the correctness of page texts in different languages, it is necessary to check the page texts. In the related technology, the page translation texts are usually checked manually after the platform development is completed.
[0031] First, the verification method in the relevant technology is introduced:
[0032] 1. Get the text elements at various locations in the page through the xpath language, including the primary language (Chinese) and the secondary language (English).
[0033] 2. Determine whether the text types of the primary language text and the secondary language text at the same position conform to the preset text type.
[0034] 3. If the text types of the primary language text and the secondary language text at the same position meet the preset text type, then based on the primary language text, determine whether the character length of the corresponding secondary language text meets the preset length requirement.
[0035] 4. If the length of the secondary language text meets the preset length requirement, all text elements with the same primary language text in the page are extracted, and whether the translation is consistent is determined based on the secondary language text of these text elements.
[0036] However, since the same Chinese may have different translations in different contexts, it is not possible to accurately identify whether the translation of the same Chinese in different positions is accurate based only on the length requirement, and the above method is not suitable for checking text in image form.
[0037] Therefore, the present disclosure provides a page text verification method, device, electronic device, medium and product to solve the problems in the related technology, improve the text verification efficiency, text verification accuracy and the scope of application of the text verification method.
[0038] The page text verification method provided by the present disclosure is described in detail below with reference to the accompanying drawings.
[0039] Figure 1 A flow chart of a page text verification method provided by an embodiment of the present disclosure. Figure 2 As shown, the page text verification method includes steps 101-104.
[0040] Step 101, determining the text type of the page based on the element tag of the page.
[0041] In some embodiments, element tags are used to indicate the structure and content of a page, wherein element tags may include: image tags, truncation tags, content block tags, etc.
[0042] For example, an image tag might be , the truncated label is for example 、 wait.
[0043] In some embodiments, different element tags may indicate different text types, wherein, when the element tag includes an image tag, the text type is determined to be an image text type; when the element tag includes a truncation tag, the text type is determined to be a truncated text type; when the element tag does not include an image tag and a truncation tag, the text type is determined to be a regular text type.
[0044] For example, when the element tag of the page contains When the text type of the page is determined to be an image text type, that is, the page contains a picture, and there may be text in the picture.
[0045] For example, when the element tag of the page contains When a text segment of the page is viewed, it can be determined that the text segment of the page may be divided into two or more segments (it should be understood that truncation refers to dividing the same text segment into different locations on the page, or reflecting the different contents of the same text segment in different application styles, rather than dividing the same text segment with punctuation marks), so the text type of the page can be determined as a truncated text type.
[0046] For example, when the element tag of the page does not contain , nor does it include When there is a truncation tag, it means that the page does not contain any image and the text in the page is not truncated. At this time, the text of the page can be determined as a regular text type.
[0047] Step 102: Determine the first text of the page based on the text type.
[0048] In some embodiments, the first text refers to the text used to match with the preset standard library. In other words, the first text is the text used to determine whether the translation of the text in the page is correct.
[0049] In some embodiments, the first text may be all text contained in the page, or may be part of the text contained in the page, which is not limited in the present disclosure.
[0050] In some optional embodiments, when the text type is an image text type, a regional screenshot of a preset region of the page is acquired to determine the first text using an image processing algorithm.
[0051] In some optional embodiments, when the text type is a truncated text type, the target text contained in the page text is extracted to determine the first text.
[0052] In some optional embodiments, when the text type is a regular text type, the page text is split based on punctuation marks contained in the page text to determine the first text.
[0053] Step 103: Determine the language environment of the page based on the parameter information of the page.
[0054] In some embodiments, the language environment of the page can be determined through the interface request header parameters and / or the page routing path parameters of the page.
[0055] In other words, the parameter information may include: interface request header parameters and / or page routing path parameters.
[0056] The language environment refers to the language of the first text, such as English environment, Japanese environment, Korean environment, etc.
[0057] Step 104 , based on the first text and the language environment, the first text of the page is checked using a preset standard library.
[0058] In some embodiments, based on the language environment, the first data of the same language environment in the preset standard library (i.e., the preset standard translation text) is determined, and then the first text and the first data are matched. When the first text can be matched with the first data successfully, it is determined that the text translation of the page is correct; when the first text fails to match the first data, it is determined that there is an abnormality in the text translation of the page.
[0059] For example, the language environment is English, and the first data of English type in the preset standard library includes: data 1, data 2 and data 3. When the first text can match any one of data 1, data 2 and data 3, it means that the translation of the first text is correct.
[0060] In summary, the page text verification method proposed in the present disclosure includes: determining the text type of the page based on the element tag of the page; determining the first text of the page based on the text type; determining the language environment of the page based on the parameter information of the page; and verifying the first text of the page using a preset standard library based on the first text and the language environment. The method of the present disclosure improves the accuracy of text extraction by determining the text type and determining the first text in different ways according to different text types, and improves the accuracy of text verification by verifying the first text using a preset standard library.
[0061] Figure 2 A flowchart of a page text verification method proposed in the present disclosure is further shown. Figure 1 2 is further explained based on the embodiment shown. Figure 2 As shown, the method comprises the following steps:
[0062] Step 201: When the element tag includes an image tag, determine that the text type is an image text type.
[0063] In some embodiments, when the element tag of the page contains When the text type of the page is determined to be an image text type, that is, the page contains a picture, and there may be text in the picture.
[0064] Step 202, obtaining a screenshot of a preset area of the page.
[0065] In some embodiments, the acquisition of the regional screenshot can be achieved through an application programming interface (Application Programming Interface, API), but is not limited to this. The present disclosure does not limit the method for acquiring the regional screenshot.
[0066] For example, the Canvas API of HTML5 is used to capture a screenshot of a specified area of a page (ie, a screenshot of a preset area of the page).
[0067] Step 203: Based on the region screenshot, a preset image processing model is used to determine the predicted text boundary of the region screenshot and the text probability corresponding to the predicted text region.
[0068] In some embodiments, the preset image processing model may be a model including the EAST (Efficient and Accuracy Scene Text) algorithm, but is not limited thereto. Any model that can perform text boundary detection may be used as the preset image processing model.
[0069] In some embodiments, the size and / or color of the area screenshot does not meet the input requirements of the model, so the area screenshot needs to be preprocessed, for example, the area screenshot is converted into a grayscale or color image, and appropriately scaled and cropped to meet the model input requirements.
[0070] In some embodiments, the area screenshot is input into a preset image processing model to obtain the text boundaries detected by the model in the area screenshot (i.e., predicted text boundaries) and the probability of text existing in the area formed by the predicted text boundaries (i.e., predicted text area).
[0071] Step 204, based on the predicted text boundary and text probability, a preset text probability threshold and a non-maximum suppression algorithm are used to determine the target text boundary.
[0072] In some embodiments, predicted text regions with text probabilities less than a preset text probability threshold may be eliminated to improve the accuracy of text position determination.
[0073] In some embodiments, based on the predicted text boundaries, a non-maximum suppression algorithm may be used to screen the predicted text boundaries, eliminate redundant predicted text boundaries, and improve the accuracy of text boundary recognition.
[0074] In some embodiments, the target text boundary is determined by combining the predicted text region eliminated by using a preset text probability threshold with the predicted text boundary eliminated by using a non-maximum suppression algorithm. For example, the target text boundary is determined by determining the region overlapping the predicted text region after elimination and the predicted text boundary after elimination.
[0075] Step 205 , based on the target text boundary, using an image recognition algorithm, determine the first text.
[0076] In some embodiments, the text within the target text boundary is recognized by an image recognition algorithm, such as an optical character recognition (OCR) algorithm, and the recognized text is determined as the first text.
[0077] Step 206: Determine the language environment of the page based on the interface request header parameters and / or the page routing path parameters of the page.
[0078] Step 207: Determine first data corresponding to the language environment in the preset standard library based on the language environment and the first data type field in the preset standard library.
[0079] In some embodiments, the data in the preset standard library may be data that has been reviewed by experts and practitioners in various fields, and has a certain degree of authority and accuracy.
[0080] In some embodiments, the first data type field is used to indicate the language of the data in the preset standard library. For example, English indicates that the data is in English, and Japanese indicates that the data is in Japanese.
[0081] In some embodiments, according to the language environment, all data in the preset standard library are traversed, and data of the same type as the language environment is determined as the first data.
[0082] For example, when the language environment is an English environment, the data containing the English field in the preset standard library is determined as the first data.
[0083] Step 208, traverse the first data to determine whether there is first data matching the first text.
[0084] In some embodiments, the first data is traversed to see whether there is first data that is consistent with the first text content.
[0085] Step 209: When there is first data matching the first text, it is determined that the first text is translated correctly.
[0086] Step 210, when there is no first data matching the first text, it is determined that the first text translation is abnormal, and the first data is checked using a preset translation library, and the data volume of the second data in the preset translation library is greater than or equal to the data volume of the first data in the preset standard library.
[0087] In some embodiments, when there is no first data matching the first text in the preset standard library, it indicates that there is an abnormality in the translation of the first text and further verification of the first text is required.
[0088] In some embodiments, the data in the preset translation library may be data obtained from data sources such as web forums, open source databases, etc., but is not limited thereto, wherein the data in the preset database includes data in the preset standard library, and the data volume of the second data in the preset translation library is greater than or equal to the data volume of the first data in the preset standard library.
[0089] It should be understood that the data in the preset translation database is also audited and has a certain degree of authority, indicating that the source of the data in the preset translation database is wider than the source of the data in the preset standard database.
[0090] In summary, the page text verification method disclosed in the present invention includes: when the element tag includes an image tag, determining that the text type is an image text type; obtaining a regional screenshot of a preset area of the page; based on the regional screenshot, using a preset image processing model, determining the predicted text boundary of the regional screenshot and the text probability corresponding to the predicted text area; based on the predicted text boundary and the text probability, using a preset text probability threshold and a non-maximum suppression algorithm, determining the target text boundary; based on the target text boundary, using an image recognition algorithm, determining the first text; based on the interface request header parameter of the page and / or the page routing path parameter, determining the language environment of the page; based on the language environment and the first data type field in the preset standard library, determining the first data corresponding to the language environment in the preset standard library; traversing the first data to determine whether there is first data matching the first text; when there is first data matching the first text, determining that the first text is translated correctly; when there is no first data matching the first text, determining that the first text translation is abnormal, and using the preset translation library to verify the first data, the data amount of the second data in the preset translation library is greater than or equal to the data amount of the first data in the preset standard library. The method disclosed herein determines the image boundary in a page and performs text recognition on the area within the boundary, thereby obtaining the first text of the page, thereby expanding the scope of application of the text verification method, and uses a preset standard library to verify the first text, thereby improving the accuracy of text verification.
[0091] Figure 3 A flowchart of a page text verification method proposed in the present disclosure is further shown. Figure 2 Step 210 is further explained based on the embodiment shown. Figure 3 As shown, the method comprises the following steps:
[0092] Step 301: based on the language environment and the second data type field in the preset translation library, determine the second data corresponding to the language environment in the preset translation library.
[0093] In some embodiments, the second data type field is used to indicate the language of the data in the preset translation library. For example, English indicates that the data is English, and Japanese indicates that the data is Japanese.
[0094] In some embodiments, according to the language environment, all data in the preset translation library are traversed, and data of the same type as the language environment is determined as the second data.
[0095] For example, when the language environment is an English environment, the data containing the English field in the preset translation library is determined as the second data.
[0096] Step 302, traverse the second data to determine whether there is second data matching the first text.
[0097] In some embodiments, the second data is traversed to see whether there is second data that is consistent with the first text content.
[0098] Step 303: When there is second data matching the first text, query the first data based on the data identification field of the matching second data to determine whether the first text is translated correctly.
[0099] In some embodiments, when there is second data matching the first text, it means that the preset standard library does not record the relevant data, or there is a translation error in the second data in the preset database. Therefore, it is necessary to further determine whether the first text is translated correctly.
[0100] Furthermore, the first data can be traversed through the data identification field of the second data matched with the first text to determine whether the first data contains the identification field; when the first data contains the identification field, it is determined that there is an error in the translation of the second data and the second data is modified; when the first data does not contain the identification field, the preset standard library and / or the preset translation library is updated.
[0101] It should be understood that both the data in the preset standard library and the data in the preset translation library contain fields for indicating data identity, such as a Label field, a Short field, etc., and data with the same content have the same data identification field in the preset standard library and the preset translation library.
[0102] For example, when the second data in the preset translation library matches the first text, based on the Label of the second data, it is queried whether there is data containing the same Label field in the preset standard library. If the preset standard library does not contain it, it means that the data in the preset standard library needs to be updated. At this time, it is determined that the translation of the first data text is correct; if it is contained in the preset standard library, it means that there is a translation error in the second data, the second data is modified, and the translation of the first text is determined to be incorrect.
[0103] Step 304: When there is no second data matching the first text, determine that the first text is incorrectly translated.
[0104] In summary, the page text verification method disclosed in the present invention includes: based on the language environment and the second data type field in the preset translation library, determining the second data corresponding to the language environment in the preset translation library; traversing the second data to determine whether there is second data matching the first text; when there is second data matching the first text, querying the first data based on the data identification field of the matched second data to determine whether the first text is translated correctly; when there is no second data matching the first text, determining that the first text is translated incorrectly. The method disclosed in the present invention further verifies the first text using the preset translation library, thereby improving the accuracy of the text verification result.
[0105] Figure 4 A flowchart of a page text verification method proposed in the present disclosure is further shown. Figure 1 4 is further explained based on the embodiment shown. Figure 4 As shown, the method comprises the following steps:
[0106] Step 401: When the element tag includes a truncation tag, determine that the text type is a truncation text type.
[0107] In some embodiments, when the element tag of the page contains When a tag indicating truncated text is displayed, the text type of the page can be determined as a truncated text type, that is, the same text in the page is divided into different positions on the page, or is embodied in different application styles.
[0108] Step 402: Determine the first text of the page based on the text type.
[0109] In some embodiments, when the text type is a truncated type, the text in the page may be directly determined as the first text.
[0110] It should be understood that each portion of text generated due to truncation in the page is a first text.
[0111] Step 403: determine the language environment of the page based on the interface request header parameters and / or the page routing path parameters of the page.
[0112] Step 404: Determine first data corresponding to the language environment in the preset standard library based on the language environment and the first data type field in the preset standard library.
[0113] In some embodiments, the principle of step 404 is the same as that of step 207, and reference may be made to the relevant contents and embodiments of step 207, which will not be repeated here.
[0114] Step 405, traverse the first data to determine whether there is first data matching the first text.
[0115] In some embodiments, since the first text may be only part of the entire text in the page, direct matching with the first data may result in no match. Therefore, a fuzzy matching algorithm may be used to easily determine whether there is first data that matches the first text.
[0116] Step 406: When there is first data matching the first text, determine that the first text is translated correctly.
[0117] In particular, in an optional embodiment, the first data fuzzily matched to each first text may be merged to see whether they can be merged into a complete data. If they can be merged, it is determined that the first text translation is correct; if they cannot be merged, the first text translation is abnormal.
[0118] Step 407: When there is no first data matching the first text, it is determined that the translation of the first text is abnormal, and the first data is checked using a preset translation library.
[0119] In some embodiments, the principle of step 407 is the same as that of step 210 and steps 301 to 304, and reference may be made to the relevant contents and embodiments, which will not be repeated here.
[0120] In summary, the page text verification method disclosed in the present invention includes: when the element tag includes a truncation tag, determining that the text type is a truncated text type; based on the text type, determining the first text of the page; based on the interface request header parameters of the page and / or the page routing path parameters, determining the language environment of the page; based on the language environment and the first data type field in the preset standard library, determining the first data corresponding to the language environment in the preset standard library; traversing the first data to determine whether there is a first data matching the first text; when there is a first data matching the first text, determining that the first text translation is correct; when there is no first data matching the first text, determining that the first text translation is abnormal, and using the preset translation library to verify the first data. The method disclosed in the present invention can determine whether the first text matches the first data by fuzzy matching, so that the text verification method can be applied to truncated type texts, expanding the scope of application of the method; and, using the preset translation library to further verify the first text, improves the accuracy of the text verification result.
[0121] Figure 5 A flowchart of a page text verification method proposed in the present disclosure is further shown. Figure 1 5 is further explained based on the embodiment shown. Figure 5 As shown, the method comprises the following steps:
[0122] Step 501: When the element tag does not include an image tag and a truncation tag, the text type is determined to be a regular text type.
[0123] In some embodiments, when the element tag of the page does not contain Tags, nor When a tag indicating truncated text is displayed, the text type of the page can be determined as a normal text type.
[0124] Step 402: Determine the first text of the page based on the text type.
[0125] In some embodiments, when the text type is a regular text type, the text in the page may be directly determined as the first text.
[0126] In some embodiments, each portion of text generated by segmenting a page by punctuation marks (eg, ", " and ".", etc.) may be regarded as a first text.
[0127] In some embodiments, the entire paragraph of text may also be determined as a first text.
[0128] Step 403: determine the language environment of the page based on the interface request header parameters and / or the page routing path parameters of the page.
[0129] Step 404: Determine first data corresponding to the language environment in the preset standard library based on the language environment and the first data type field in the preset standard library.
[0130] In some embodiments, the principle of step 404 is the same as that of step 207, and reference may be made to the relevant contents and embodiments of step 207, which will not be repeated here.
[0131] Step 405, traverse the first data to determine whether there is first data matching the first text.
[0132] Step 406: When there is first data matching the first text, determine that the first text is translated correctly.
[0133] In some embodiments, the principle of step 406 and step 209 can be referred to the relevant content and embodiments, and will not be repeated here.
[0134] Step 407: When there is no first data matching the first text, it is determined that the translation of the first text is abnormal, and the first data is checked using a preset translation library.
[0135] In some embodiments, the principle of step 407 is the same as that of step 210 and steps 301 to 304, and reference may be made to the relevant contents and embodiments, which will not be repeated here.
[0136] In summary, the page text verification method disclosed in the present invention includes: when the element tag does not include an image tag and a truncation tag, determining that the text type is a regular text type; based on the text type, determining the first text of the page; based on the interface request header parameter of the page and / or the page routing path parameter, determining the language environment of the page; based on the language environment and the first data type field in the preset standard library, determining the first data corresponding to the language environment in the preset standard library; traversing the first data to determine whether there is a first data matching the first text; when there is a first data matching the first text, determining that the first text translation is correct; when there is no first data matching the first text, determining that the first text translation is abnormal, and using the preset translation library to verify the first data. The method disclosed in the present invention can split the page text by punctuation marks to obtain the first text, thereby determining whether the first text matches the first data, so that the text verification method can be applied to the truncation type text, expanding the scope of application of the method; and further verifying the first text using the preset translation library improves the accuracy of the text verification result.
[0137] The following is an exemplary description of the disclosed method.
[0138] The specific implementation process is as follows:
[0139] 1. Get text: In the browser plug-in content script, get the text of the page element node (that is, the above-mentioned element tag) through the js script, and pre-process the line break characters, etc.
[0140] 2. Determine the language environment: Determine whether it is an English environment or another foreign language environment based on the interface request header parameter or the page routing path parameter; based on the language environment, compare the obtained page text with the corresponding translation text in the standard library. For example, if the language environment is English, traverse the standard library and compare the page text with the English in the standard library. If the consistent text can be matched, it means that the translation is correct, otherwise record the page text.
[0141] 3. Standard library comparison: In the browser plug-in background script, read the data of the standard library (a piece of data in the standard library contains fields such as Label, Short, Chinese, English, Korean, and Japanese. Label and Short represent the identifier of this piece of data, which is generally consistent with Chinese. Chinese is the Chinese text that needs to be translated, and English, Korean, Japanese, etc. are the corresponding foreign language translation texts.)
[0142] 4. According to the language environment, compare the obtained page text with the corresponding translation text in the standard library. For example, if the language environment is English, traverse the standard library and compare the page text with the English in the standard library. If the consistent text can be matched, it means that the translation is correct, otherwise record the page text.
[0143] 5. Translation library comparison: compare the page text that is not matched in the standard library with the project translation library (the data structure is consistent with the standard library).
[0144] 6. If there is no match in the translation database, it means that the translation is missing or the code is incorrect. Record this text and check the corresponding code later. If there is a match in the translation database, it means that the translation content in the translation database is incorrect. Record the corresponding Label of the incorrect translation.
[0145] 7. Go to the standard library for comparison and matching based on the corresponding Label.
[0146] 8. If no match is found, it means that the standard library needs to be updated; if a match is found in the standard library, it means that the translation library is not translated correctly and needs to be updated.
[0147] 9. Update the standard library or update the translation library.
[0148] Furthermore, different processing methods are adopted for different situations of different pages.
[0149] For pages with pictures, the picture content may contain text, and the picture should also be switched to the corresponding environment according to the language environment. When the text in the picture is recognized, image recognition technology is used. Figure 7 As shown, the following steps are included:
[0150] 1) Encapsulated JavaScript library: Introducing deep learning-based text recognition technology (EAST model [2]), encapsulating the EAST model and its reasoning code into an easy-to-use JavaScript library to run in a browser environment. Technologies such as TensorFlow.js can be used to convert the trained model into a format that can be loaded and executed on the front end.
[0151] 2) Image capture and preprocessing: Use HTML5’s Canvas API to capture screenshots of the specified area of the page, convert them into grayscale or color images, and perform appropriate scaling and cropping to suit the model input requirements.
[0152] 3) Model reasoning: The preprocessed image is input into the EAST model for reasoning to obtain the rotated rectangular bounding box of the text area and the corresponding text probability map.
[0153] 4) Post-processing: Apply threshold screening and non-maximum suppression algorithm to generate the final text box list based on the prediction results. Each box contains its position, size and rotation angle in the original image.
[0154] 5) Text extraction: Based on the coordinates of the detected text box in the original image, use OCR technology to identify and extract the text content.
[0155] 6) Record the text and add the img tag.
[0156] 2. A piece of text is etc. label truncation situation.
[0157] In this case, when the page gets the text, one paragraph of text will be truncated into two paragraphs. At this time, it cannot be matched in the translation library and the standard library, so fuzzy matching is performed. If a fuzzy match can be found, the result of the fuzzy matching is judged. If the same data in the corresponding text is matched and is exactly equal to this data after merging, it means that the translation is correct.
[0158] 3. A piece of text is composed of multiple short texts.
[0159] The splicing here is connected by symbols such as "-", " / ", etc., and the text is split according to the symbols, and then compared and matched, that is, the process in (1) is carried out.
[0160] In summary, the present disclosure has the following beneficial effects:
[0161] 1. The method disclosed herein improves the accuracy of text extraction by determining the text type and adopting different methods to determine the first text according to different text types, and improves the accuracy of text verification by checking the first text using a preset standard library.
[0162] 2. The method disclosed in the present invention determines the image boundary in the page and performs text recognition on the area within the boundary to obtain the first text of the page, thereby expanding the scope of application of the text verification method, and uses a preset standard library to verify the first text, thereby improving the accuracy of text verification.
[0163] 3. The method disclosed in the present invention further verifies the first text using a preset translation library, thereby improving the accuracy of the text verification result.
[0164] Corresponding to the methods provided in the above-mentioned embodiments, the present disclosure also provides a page text verification device. Since the device provided in the embodiments of the present disclosure corresponds to the methods provided in the above-mentioned embodiments, the implementation method of the method is also applicable to the device provided in this embodiment and will not be described in detail in this embodiment.
[0165] Figure 8 FIG. 8 is a schematic diagram of a page text checking device 800 provided in an embodiment of the present disclosure. Figure 8 As shown, the page text checking device includes:
[0166] A first determining unit 810 is used to determine the text type of the page based on the element tag of the page;
[0167] A second determining unit 820, configured to determine a first text of a page based on the text type;
[0168] The third determining unit 830 is used to determine the language environment of the page based on the parameter information of the page;
[0169] The checking unit 840 is used to check the first text of the page based on the first text and the language environment using a preset standard library.
[0170] In some embodiments, the first determination unit 810 is also used to determine that the text type is an image text type when the element tag includes an image tag; determine that the text type is a truncated text type when the element tag includes a truncation tag; and determine that the text type is a regular text type when the element tag does not include an image tag and a truncation tag.
[0171] In some embodiments, the second determination unit 820 is also used to obtain a regional screenshot of a preset area of the page when the text type is an image text type, so as to determine the first text by using an image processing algorithm; when the text type is a truncated text type, extract the target text contained in the page text to determine the first text; when the text type is a regular text type, split the page text based on punctuation marks contained in the page text to determine the first text.
[0172] In some embodiments, the second determination unit 820 is also used to obtain a regional screenshot of a preset area of the page, so as to use an image processing algorithm to determine the first text, including: based on the page screenshot, using a preset image processing model to determine the predicted text boundary of the page screenshot and the text probability corresponding to the predicted text area; based on the predicted text boundary and text probability, using a preset text probability threshold and non-maximum suppression algorithm to determine the target text boundary; based on the target text boundary, using an image recognition algorithm to determine the first text.
[0173] In some embodiments, the third determination unit 830 is further used to determine the language environment of the page based on the interface request header parameters and / or the page routing path parameters of the page.
[0174] In some embodiments, the checking unit 840 is further used to determine the first data corresponding to the language environment in the preset standard library based on the language environment and the first data type field in the preset standard library; traverse the first data to determine whether there is first data matching the first text; when there is first data matching the first text, determine that the translation of the first text is correct; when there is no first data matching the first text, determine that the translation of the first text is abnormal, and use the preset translation library to check the first data, and the data amount of the second data in the preset translation library is greater than or equal to the data amount of the first data in the preset standard library.
[0175] In some embodiments, the checking unit 840 is further used to determine the second data corresponding to the language environment in the preset translation library based on the language environment and the second data type field in the preset translation library; traverse the second data to determine whether there is second data matching the first text; when there is second data matching the first text, query the first data based on the data identification field of the matching second data to determine whether the first text is translated correctly; when there is no second data matching the first text, determine that the first text translation is incorrect.
[0176] In some embodiments, based on the language environment and the second data type field in the preset translation library, second data corresponding to the language environment in the preset translation library is determined; the second data is traversed to determine whether there is second data matching the first text; when there is second data matching the first text, based on the data identification field of the matching second data, the first data is queried to determine whether the first text is translated correctly; when there is no second data matching the first text, it is determined that the first text translation is incorrect.
[0177] In summary, the page text verification device includes: a first determination unit for determining the text type of the page based on the element tag of the page; a second determination unit for determining the first text of the page based on the text type; a third determination unit for determining the language environment of the page based on the parameter information of the page; and a verification unit for verifying the first text of the page based on the first text and the language environment using a preset standard library. The device disclosed herein improves the accuracy of text extraction by determining the text type and determining the first text in different ways according to different text types, and improves the accuracy of text verification by verifying the first text using a preset standard library.
[0178] In the above embodiments provided by the present disclosure, the methods and devices provided by the embodiments of the present disclosure are introduced. In order to implement the functions in the methods provided by the above embodiments of the present disclosure, the electronic device may include a hardware structure and a software module, and implement the above functions in the form of a hardware structure, a software module, or a hardware structure plus a software module. One of the above functions may be executed in the form of a hardware structure, a software module, or a hardware structure plus a software module.
[0179] Fig. 9 1 is a block diagram of an electronic device 900 for implementing the above-mentioned page text verification method according to an exemplary embodiment. For example, the electronic device 900 may be a mobile phone, a computer, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0180] Reference Fig. 9 The electronic device 900 may include a communication interface 901, which can interact with other devices; a processor 902, which is connected to the communication interface 901 to interact with other devices and is used to execute the method provided by one or more of the above technical solutions when running a computer program; and a memory 903, on which the computer program is stored. Specifically, the specific processing process of the processor 902 can refer to the page text verification method described in the above embodiment of the present disclosure.
[0181] Of course, in actual application, the various components in the electronic device 900 are coupled together through the bus system 905. It can be understood that the bus system 905 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 905 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Fig. 9 Various buses are labeled as bus system 905.
[0182] The memory 903 in the embodiment of the present disclosure is used to store various types of data to support the operation of the electronic device 900. Examples of such data include: any computer program used to operate on the electronic device 900.
[0183] The method disclosed in the above embodiment of the present disclosure can be applied to the processor 902, or implemented by the processor 902. The processor 902 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 902 or the instruction in the form of software. The above processor 902 may be a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. The processor 902 can implement or execute the disclosed methods, steps and logic block diagrams in the embodiment of the present disclosure. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the embodiment of the present disclosure can be directly embodied as a hardware decoding processor to execute, or the hardware and software modules in the decoding processor are combined and executed. The software module can be located in a storage medium, which is located in the memory 903. The processor 902 reads the information in the memory 903 and completes the steps of the above method in combination with its hardware.
[0184] In an exemplary embodiment, the electronic device 900 can be implemented by one or more application specific integrated circuits (ASIC), DSP, programmable logic device (PLD), complex programmable logic device (CPLD), field programmable gate array (FPGA), general processor, controller, microcontroller (MCU), microprocessor, or other electronic components to execute the aforementioned method.
[0185] The embodiments of the present disclosure further provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the page text verification method described in the above embodiments of the present disclosure.
[0186] The embodiments of the present disclosure further provide a computer program product, including a computer program, which executes the page text verification method described in the above embodiments of the present disclosure when a processor executes the computer program.
[0187] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the attached claims.
[0188] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples" or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0189] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention belong.
[0190] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processing module, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (control method), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing in a suitable manner if necessary, and then stored in a computer memory.
[0191] It should be understood that the various parts of the embodiments of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0192] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0193] In addition, each functional unit in each embodiment of the present invention may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.
[0194] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A page text verification method, characterized in that: The method comprises: Determining the text type of the page based on the element tag of the page; Based on the text type, determining a first text of the page; Determining the language environment of the page based on the parameter information of the page; Based on the first text and the language environment, the first text of the page is checked using a preset standard library.
2. The method according to claim 1, characterized in that Determining the text type of the page based on the element tag of the page includes: When the element tag includes an image tag, determining that the text type is an image text type; When the element tag includes a truncation tag, determining that the text type is a truncation text type; When the element tag does not include the image tag and the truncation tag, the text type is determined to be a regular text type.
3. The method according to claim 1 or 2, characterized in that: The determining, based on the text type, the first text of the page comprises: When the text type is an image text type, obtaining a regional screenshot of a preset region of the page to determine the first text by using an image processing algorithm; When the text type is a truncated text type, extracting target text contained in the page text to determine the first text; When the text type is a regular text type, the page text is split based on punctuation marks contained in the page text to determine the first text.
4. The method according to claim 3, characterized in that The obtaining of a screenshot of a preset area of the page to determine the first text by using an image processing algorithm includes: Based on the area screenshot, using a preset image processing model, determining a predicted text boundary of the area screenshot and a text probability corresponding to the predicted text area; Based on the predicted text boundary and the text probability, a preset text probability threshold and a non-maximum suppression algorithm are used to determine the target text boundary; Based on the target text boundary, the first text is determined using an image recognition algorithm.
5. The method according to claim 1, characterized in that The determining the language environment of the page based on the parameter information of the page includes: The language environment of the page is determined based on the interface request header parameters and / or the page routing path parameters of the page.
6. The method according to claim 1, characterized in that The checking the first text of the page based on the first text and the language environment using a preset standard library includes: Based on the language environment and a first data type field in the preset standard library, determining first data in the preset standard library corresponding to the language environment; Traversing the first data to determine whether there is first data matching the first text; When there is first data matching the first text, determining that the first text is translated correctly; When there is no first data matching the first text, it is determined that the translation of the first text is abnormal, and the first data is checked using a preset translation library, wherein the amount of second data in the preset translation library is greater than or equal to the amount of first data in the preset standard library.
7. The method according to claim 6, characterized in that The checking the first data by using a preset translation library comprises: Based on the language environment and a second data type field in the preset translation library, determining second data in the preset translation library corresponding to the language environment; Traversing the second data to determine whether there is second data matching the first text; When there is second data matching the first text, querying the first data based on the data identification field of the matching second data to determine whether the first text is translated correctly; When there is no second data matching the first text, it is determined that the first text is incorrectly translated.
8. The method according to claim 7, characterized in that The querying the first data based on the data identification field of the matched second data to determine whether the first text is translated correctly comprises: Traversing the first data to determine whether the first data contains the identification field; When the first data includes the identification field, determining that there is an error in the translation of the second data, and modifying the second data; When the first data does not include the identification field, the preset standard library and / or the preset translation library is updated.
9. A page text checking device, characterized in that: The device comprises: A first determining unit, configured to determine a text type of the page based on an element tag of the page; A second determining unit, configured to determine a first text of the page based on the text type; A third determining unit, configured to determine a language environment of the page based on parameter information of the page; The checking unit is used to check the first text of the page based on the first text and the language environment using a preset standard library.
10. An electronic device, characterized in that: include: a processor and a memory for storing a computer program capable of being executed on the processor, Wherein, when the processor is used to run the computer program, it executes the method according to any one of claims 1 to 8.
11. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.
12. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 8.