Confusion text recognition method and device, electronic equipment and storage medium

By extracting the target font file URL from the webpage content and parsing the mapping table to decode the obfuscated text, the problem of recognition failure caused by dynamic changes in font files and URLs on the web platform is solved, and real-time recognition and efficient decoding of obfuscated text is achieved.

CN120873308APending Publication Date: 2025-10-31BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510815384.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies cannot handle the failure to recognize obfuscated text caused by dynamic changes in font files and URLs on web platforms.

Method used

By extracting the URL address of the target font file from the content of the webpage to be identified, obtaining the target font file and parsing its corresponding mapping table, and using the mapping table to decode the obfuscated text, real-time recognition of obfuscated text on different platforms can be achieved.

Benefits of technology

Dynamically tracking and acquiring font files from different platforms improves the efficiency of obfuscated text recognition and reduces hardware costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873308A_ABST
    Figure CN120873308A_ABST
Patent Text Reader

Abstract

The invention relates to a confused text recognition method and device, electronic equipment and a storage medium, and the method comprises the steps: extracting a target font file URL address from to-be-recognized webpage content, obtaining a target font file pointed by the target font file URL address, obtaining a target mapping table corresponding to the target font file, and carrying out the recognition of a confused text in the target font file, and decoding the confused text according to the target mapping table to obtain a target text corresponding to the confused text. According to the mode, the font files of different platforms can be dynamically tracked and obtained, the confused texts in the font files are recognized in real time, meanwhile, the recognition efficiency of the confused texts is improved, and the hardware cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet text recognition technology, and in particular to a method, apparatus, electronic device and storage medium for recognizing obfuscated text. Background Technology

[0002] With the development of the internet, businesses are increasingly relying on data collection platforms for market analysis, competitor research, and pricing strategies. However, to protect critical data assets, mainstream online platforms use custom font files to display key data text as obfuscated text, preventing businesses from obtaining the true critical data.

[0003] While existing data collection methods can identify obfuscated text displayed on web platforms, they cannot address the problem of obfuscated text recognition failure caused by dynamic changes in font files and URLs (Uniform Resource Locators) of web platforms. Summary of the Invention

[0004] This application provides a method, apparatus, electronic device, and storage medium for obfuscated text recognition, in order to solve the problem that existing technologies cannot cope with the failure of obfuscated text recognition caused by dynamic changes in font files and URLs on network platforms.

[0005] Firstly, this application provides a method for identifying obfuscated text, the method comprising:

[0006] Extract the URL address of the target font file from the content of the webpage to be identified;

[0007] Obtain the target font file pointed to by the URL address of the target font file, and obtain the target mapping table corresponding to the target font file. The target mapping table is used to map the Unicode code points of the obfuscated text to the corresponding characters in order to decode the obfuscated text.

[0008] For the obfuscated text in the target font file, the obfuscated text is decoded according to the target mapping table to obtain the target text corresponding to the obfuscated text.

[0009] In one possible implementation, extracting the target font file URL address from the webpage content to be identified includes:

[0010] Based on predefined regular expression matching rules, extract the target font file URL address from the content of the web page to be identified.

[0011] In one possible implementation, obtaining the target font file pointed to by the target font file URL address includes:

[0012] The cache area is located based on the target font file URL address, and the cache area includes the correspondence between font file URL addresses and font files;

[0013] If the target font file URL address is found in the cache area, the font file corresponding to the found target font file URL address is determined as the target font file;

[0014] If the target font file URL address is not found in the cache, the target font file is downloaded using the target font file URL address.

[0015] In one possible implementation, after downloading the target font file using the target font file URL address, the method further includes:

[0016] The target font file is stored in the cache area, and a correspondence between the target font file URL address and the target font file is established in the cache area.

[0017] In one possible implementation, the step of decoding the obfuscated text in the target font file according to the target mapping table to obtain the target text corresponding to the obfuscated text includes:

[0018] For the obfuscated text in the target font file, obtain the Unicode code point of each obfuscated character in the obfuscated text;

[0019] Based on the Unicode code point of each obfuscated character, the corresponding character is found in the target mapping table, and the character is determined as the target text corresponding to the obfuscated text.

[0020] In one possible implementation, the target mapping table includes a character mapping table and a glyph mapping table. The character mapping table includes a correspondence between Unicode code points and glyph names, and the glyph mapping table includes a correspondence between glyph names and characters. The step of finding the corresponding character in the target mapping table based on the Unicode code point of each obfuscated character, and determining the character as the target text corresponding to the obfuscated text, includes:

[0021] Based on the Unicode code point corresponding to each of the obfuscated characters, the glyph name corresponding to each obfuscated character is found from the character mapping table;

[0022] Among the multiple glyph names found, those that meet the set aggregation conditions are aggregated to form a target glyph name, thus obtaining at least one target glyph name;

[0023] Based on the at least one target glyph name, the corresponding character is found from the glyph mapping table, and the found character is determined as the target text corresponding to the obfuscated text.

[0024] In one possible implementation, the step of aggregating the glyph names that meet the set aggregation conditions from the multiple found glyph names to form a target glyph name, thereby obtaining at least one target glyph name, includes:

[0025] s glyph names are sequentially extracted from the glyph name sequence, wherein the glyph name sequence is formed by arranging the found glyph names in order, and s is a preset positive integer;

[0026] Aggregate the currently extracted glyph names to form a candidate glyph name, and look up the glyph mapping table based on the candidate glyph name;

[0027] If the candidate glyph name is found in the glyph mapping table, the candidate glyph name is determined as the target glyph name;

[0028] If no candidate glyph name is found in the glyph mapping table, s+i glyph names are sequentially taken from the glyph name sequence, and the process returns to the step of aggregating the currently taken glyph names into a candidate glyph name, until the target glyph name is determined, or the first candidate glyph name is not found in the glyph mapping table; wherein i increments from 1, and the first candidate glyph name is formed by aggregating the glyph names corresponding to all the obfuscated characters.

[0029] Secondly, this application provides an obfuscated text recognition device, the device comprising:

[0030] The address extraction module is used to extract the URL address of the target font file from the content of the webpage to be identified;

[0031] The file acquisition module is used to acquire the target font file pointed to by the URL address of the target font file, and to acquire the target mapping table corresponding to the target font file. The target mapping table is used to map the Unicode code points of the obfuscated text to the corresponding characters in order to decode the obfuscated text.

[0032] The text decoding module is used to decode the obfuscated text in the target font file according to the target mapping table to obtain the target text corresponding to the obfuscated text.

[0033] In one possible implementation, the address extraction module is specifically used for:

[0034] Based on predefined regular expression matching rules, extract the target font file URL address from the content of the web page to be identified.

[0035] In one possible implementation, the file acquisition module includes:

[0036] The file search unit is used to search the cache area based on the target font file URL address, wherein the cache area includes the correspondence between font file URL addresses and font files;

[0037] The file determination unit is used to determine the font file corresponding to the target font file URL address as the target font file when the target font file URL address is found in the cache area.

[0038] The file download unit is used to download the target font file using the target font file URL address when the target font file URL address is not found in the cache area.

[0039] In one possible implementation, the file download unit is further configured to:

[0040] The target font file is stored in the cache area, and a correspondence between the target font file URL address and the target font file is established in the cache area.

[0041] In one possible implementation, the text decoding module includes:

[0042] The code point acquisition unit is used to acquire the Unicode code point of each obfuscated character in the obfuscated text in the target font file;

[0043] The target text determination unit is used to find the corresponding character from the target mapping table based on the Unicode code point of each obfuscated character, and determine the character as the target text corresponding to the obfuscated text.

[0044] In one possible implementation, the target mapping table includes a character mapping table and a glyph mapping table. The character mapping table includes a correspondence between Unicode code points and glyph names, and the glyph mapping table includes a correspondence between glyph names and characters. The target text determination unit further includes:

[0045] The lookup subunit is used to find the glyph name corresponding to each of the obfuscated characters from the character mapping table based on the Unicode code point corresponding to each of the obfuscated characters.

[0046] The aggregation subunit is used to aggregate the multiple found glyph names that meet the set aggregation conditions to form a target glyph name, thereby obtaining at least one target glyph name;

[0047] The target text determination subunit is used to find the corresponding character from the glyph mapping table based on the at least one target glyph name, and determine the found character as the target text corresponding to the obfuscated text.

[0048] In one possible implementation, the polymer subunit is specifically used for:

[0049] s glyph names are sequentially extracted from the glyph name sequence, wherein the glyph name sequence is formed by arranging the found glyph names in order, and s is a preset positive integer;

[0050] Aggregate the currently extracted glyph names to form a candidate glyph name, and look up the glyph mapping table based on the candidate glyph name;

[0051] If the candidate glyph name is found in the glyph mapping table, the candidate glyph name is determined as the target glyph name;

[0052] If no candidate glyph name is found in the glyph mapping table, s+i glyph names are sequentially taken from the glyph name sequence, and the process returns to the step of aggregating the currently taken glyph names into a candidate glyph name, until the target glyph name is determined, or the first candidate glyph name is not found in the glyph mapping table; wherein i increments from 1, and the first candidate glyph name is formed by aggregating the glyph names corresponding to all the obfuscated characters.

[0053] Thirdly, this application provides an electronic device, comprising: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus, wherein the processor is configured to:

[0054] Extract the URL address of the target font file from the content of the webpage to be identified;

[0055] Obtain the target font file pointed to by the URL address of the target font file, and obtain the target mapping table corresponding to the target font file. The target mapping table is used to map the Unicode code points of the obfuscated text to the corresponding characters in order to decode the obfuscated text.

[0056] For the obfuscated text in the target font file, the obfuscated text is decoded according to the target mapping table to obtain the target text corresponding to the obfuscated text.

[0057] Fourthly, this application also provides a computer storage medium storing computer-executable instructions for performing the obfuscated text recognition method described in any of the preceding claims.

[0058] Compared with the prior art, the technical solution provided in this application has the following advantages: The method provided in this application extracts the target font file URL address from the content of the webpage to be identified, obtains the target font file pointed to by the target font file URL address, and obtains the target mapping table corresponding to the target font file. For the obfuscated text in the target font file, the obfuscated text is decoded according to the target mapping table to obtain the target text corresponding to the obfuscated text. This method can dynamically track and obtain font files from different platforms and perform real-time identification of obfuscated text in font files, while improving the efficiency of obfuscated text identification and reducing hardware costs. Attached Figure Description

[0059] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0060] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0061] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0062] Figure 1 A flowchart illustrating an embodiment of an obfuscated text recognition method provided in this application;

[0063] Figure 2 A flowchart illustrating another embodiment of the obfuscated text recognition method provided in this application;

[0064] Figure 3 A flowchart illustrating another embodiment of the obfuscated text recognition method provided in this application;

[0065] Figure 4 A block diagram of an obfuscated text recognition device provided in an embodiment of this application;

[0066] Figure 5This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0068] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0069] To address the technical problem that existing technologies cannot handle the failure of obfuscated text recognition due to dynamic changes in font files and URLs on web platforms, the method provided in this application extracts the URL address of the target font file from the content of the webpage to be recognized, obtains the target font file pointed to by the URL address, and obtains the target mapping table corresponding to the target font file. For the obfuscated text in the target font file, the obfuscated text is decoded according to the target mapping table to obtain the target text corresponding to the obfuscated text. This method can dynamically track and obtain font files from different platforms and perform real-time recognition of obfuscated text within font files, while improving the efficiency of obfuscated text recognition.

[0070] Figure 1 A flowchart illustrating an embodiment of an obfuscated text recognition method provided in this application is shown below. Figure 1 As shown, the main steps include the following:

[0071] Step 101: Extract the URL address of the target font file from the content of the webpage to be identified.

[0072] The webpage to be identified refers to a target webpage that uses font obfuscation technology to protect critical data, such as the product details page, price page, and sales statistics page of an e-commerce platform. The characteristic of the content of the aforementioned webpage to be identified is that the visible text (e.g., "#") is obfuscated characters, requiring decoding to obtain the actual characters (e.g., "1"). The obfuscated text recognition method provided in this application can decode the obfuscated text in the content of the webpage to be identified, obtaining the actual text corresponding to the obfuscated text. For example, this application provides a method that can decode the obfuscated text "#%¥*#@" in the content of the webpage to be identified into the actual target text "123". The specific decoding process will be described in detail in the relevant embodiments below.

[0073] In one embodiment, the URL address of the target font file is extracted from the content of the webpage to be identified. The URL address points to a font file, and the corresponding target font file can be obtained through this URL address. Specifically, the URL address of the target font file can be located and extracted from the source code of the webpage to be identified using preset regular expression matching rules, so as to obtain the font file based on the URL address and then decode the font file.

[0074] Furthermore, different web platforms use different URL formats, and the font files pointed to by these URLs also have different formats. Therefore, different regular expression matching rules can be set for different platforms. Even if the URL in the webpage source code changes, the changed URL can still be extracted according to the preset regular expression matching rules, automatically obtaining the latest and most accurate font file without manual intervention.

[0075] Step 102: Obtain the target font file pointed to by the URL address of the target font file, and obtain the target mapping table corresponding to the target font file. The target mapping table is used to map the Unicode code points of the obfuscated text to the corresponding characters in order to decode the obfuscated text.

[0076] A font file is an electronic file that stores font glyph information and character mapping relationships. It is a technical tool used by some online platforms to obfuscate key data (such as price and sales volume).

[0077] The aforementioned target mapping table includes a character mapping table and a glyph mapping table. The character mapping table records the mapping relationship between Unicode code points and glyph names, while the glyph mapping table records the mapping relationship between glyph names and characters. The glyph mapping table can be a predefined mapping table adapted to the current network platform. The aforementioned target mapping table is used to map the Unicode code points of the obfuscated text to corresponding characters, thereby decoding the obfuscated text.

[0078] Unicode code points are unique numerical identifiers assigned to each character for representing and processing text in a computer. Each code point is typically represented in hexadecimal form and serves as the underlying identifier for a character within the computer's digital system. In short, Unicode code points are numerical identifiers for characters within a computer, used for storage and transmission.

[0079] In one embodiment, the target font file pointed to by the URL address of the target font file is obtained, and the target mapping table corresponding to the target font file is also obtained. Specifically, the font file can be obtained by sending an HTTP (Hypertext Transfer Protocol Request) request. Once the target font file is determined, a font parsing tool is used to parse the target font file to obtain the target mapping table.

[0080] An HTTP request is a message sent by a client (such as a browser or web crawler) to a server to request resources (such as web pages, images, font files, etc.). Its core function is to establish a data transmission channel between the client and the server, enabling the acquisition and interaction of resources.

[0081] For example, the font file is obtained by sending an HTTP request to the URL address of the font file. After obtaining the font file, a target font file is determined based on the existing multiple font files. Then, a font parsing tool is used to parse the target font file to obtain a character mapping table. The final target mapping table is determined based on a predefined adapted glyph mapping table and the aforementioned character mapping table. This is just an example; the specific method of determining the target font file will be explained in detail in the relevant embodiments below.

[0082] Step 103: For the obfuscated text in the target font file, decode the obfuscated text according to the target mapping table to obtain the target text corresponding to the obfuscated text.

[0083] Obfuscated text refers to the use of custom font files by online platforms to display fake characters after converting real characters into actual text, in order to protect critical data (such as prices and sales figures). Its core purpose is to prevent regular web crawlers from directly identifying and extracting real data, making it a crucial anti-crawler technique used by online platforms. Web crawlers, also known as web scrapers, are programs or scripts that automatically retrieve information from the internet according to specific rules. Their core function is to simulate human web browsing behavior, accessing webpage URLs, parsing webpage content, and extracting and storing target data. They are widely used in search engine data collection, data analysis, and content monitoring.

[0084] In one embodiment, for obfuscated text in the target font file, the obfuscated text is decoded according to the target mapping table to obtain the target text corresponding to the obfuscated text. Specifically, each character in the obfuscated text is traversed to obtain the Unicode code point corresponding to each character. Then, the Unicode code point is decoded according to the target mapping table corresponding to the current font file to finally obtain the actual character corresponding to the Unicode code point, which is the real text corresponding to the obfuscated text.

[0085] For example, suppose the obfuscated text in the target font file is "&". Traversing this obfuscated text, we obtain its corresponding Unicode code point "U+3042". Then, by looking up the target mapping table corresponding to the target font file, we find that the character corresponding to the code point "U+3042" is "1". This indicates that the target text corresponding to the obfuscated text is 1, thus completing the decoding of the obfuscated text. This decoding method ensures that the data obtained by the crawler is the actual text, facilitating further analysis and summarization of the data collected using crawling technology, thereby obtaining more accurate data analysis and summary results.

[0086] Therefore, it can be seen that the method provided by the above embodiments can use web crawling technology to collect real text data from various online platforms, rather than obfuscated text data.

[0087] The method provided in this application extracts the URL address of the target font file from the content of the webpage to be identified, obtains the target font file pointed to by the URL address, and obtains the target mapping table corresponding to the target font file. For the obfuscated text in the target font file, the obfuscated text is decoded according to the target mapping table to obtain the target text corresponding to the obfuscated text. This method can dynamically track and obtain font files from different platforms and perform real-time identification of obfuscated text in font files, while improving the efficiency of obfuscated text identification and reducing operation and maintenance costs.

[0088] Figure 2 See the flowchart of another embodiment of the obfuscated text recognition method provided in this application. Figure 2 As shown, this mainly describes how to determine the target font file, including the following steps:

[0089] Step 201: Extract the URL address of the target font file from the content of the webpage to be identified.

[0090] For step 201 above, see [link / reference]. Figure 1 The relevant description of the illustrated embodiment.

[0091] Step 202: Locate the cache based on the target font file URL address. The cache includes the correspondence between font file URL addresses and font files. If the target font file URL address is found in the cache, proceed to step 203. If the target font file URL address is not found in the cache, proceed to step 204.

[0092] Step 203: Identify the font file corresponding to the URL address of the found target font file as the target font file, and obtain the target mapping table corresponding to the target font file.

[0093] Step 204: Download the target font file using the target font file URL address, store the target font file in the cache area, establish the correspondence between the target font file URL address and the target font file in the cache area, and obtain the target mapping table corresponding to the target font file.

[0094] The following is a unified explanation of steps 202-204 above:

[0095] The cache area is used to store downloaded font files and target mapping tables generated through parsing. Cacheing previously downloaded font files and parsed target mapping tables in this area allows for subsequent reuse of these files and tables to decode obfuscated text, avoiding multiple downloads and parsings and improving the efficiency of obfuscated text recognition. Additionally, the cache area stores the mapping between font file URLs and font files, facilitating the matching and retrieval of target font files within the cache based on this mapping.

[0096] In one embodiment, a cache is searched based on the target font file URL address. The cache includes a mapping between font file URL addresses and font files. If the target font file URL address is found in the cache, the font file corresponding to the found URL address is identified as the target font file, and a target mapping table corresponding to the target font file is obtained. If the target font file URL address is not found in the cache, the target font file is downloaded using the URL address, stored in the cache, and a mapping between the target font file URL address and the target font file is established in the cache. A target mapping table corresponding to the target font file is then obtained.

[0097] Specifically, the URL address of the target font file extracted from the content of the webpage to be identified is obtained. The cache is searched based on the URL address. If the URL address is found, the font file corresponding to the URL address is identified as the target font file, and the target mapping table corresponding to the font file is searched in the cache. If the URL address is not found in the cache, the target font file pointed to by the above target font file URL address is downloaded, and the target font file is parsed to obtain the corresponding target mapping table. At the same time, the downloaded target font file and the parsed target mapping table are cached in the corresponding cache for reuse.

[0098] In another embodiment, based on the target font file URL address and the target font file hash value, it is checked whether there is a font file in the cache that matches the current target font file. If it exists, the font file in the cache and the target mapping table are used directly to decode the obfuscated text. If it does not exist, the target font file pointed to by the above target font file URL address is downloaded directly, and the target font file is parsed to obtain its corresponding target mapping table. At the same time, the target font file and the target mapping table are cached in the corresponding cache area.

[0099] Specifically, first, the URL address of the current target font file is matched against the URL address in the cache. If the URL address matches successfully, the calculated hash value of the current target font file is matched against the hash value of the font file in the cache. If the hash value matches successfully, it means that the URL address in the cache points to the same font file as the target font file pointed to by the current target font file's URL address. In this case, the font file in the cache is directly used to decode the obfuscated text. If either the URL address of the target font file or the hash value of the target font file fails to match, it means that the font file in the cache does not match the font file pointed to by the current target font file's URL address. In this case, the font file pointed to by the target font file's URL address is directly downloaded and parsed to obtain the target font file and the target mapping table.

[0100] Furthermore, it can be further verified whether the target mapping table corresponding to the target font file pointed to by the current target font file URL address is consistent with the target mapping table corresponding to the font file pointed to by the font file URL address in the cache. If the target mapping tables are consistent, the font file in the cache and its corresponding target mapping table are used directly. If the target mapping tables are inconsistent, the target font file pointed to by the current target font file URL address is downloaded and parsed directly to obtain the target font file and the target mapping table.

[0101] The above embodiments can handle situations where the URL address of the current network platform has not changed, but the font file or mapping table corresponding to the URL address has changed. Through the above multi-level matching verification mechanism, the continuity of the current crawler process and the accuracy of obfuscated text recognition can be ensured.

[0102] Step 205: For the obfuscated text in the target font file, decode the obfuscated text according to the target mapping table to obtain the target text corresponding to the obfuscated text.

[0103] For step 205 above, please refer to the above. Figure 1 The relevant description of the illustrated embodiment.

[0104] pass Figure 2 The illustrated embodiment uses a caching mechanism to cache font files and their corresponding target mapping tables in a designated cache area. This avoids users repeatedly downloading and parsing the font files and target mapping tables, significantly reducing memory and network bandwidth usage. Furthermore, in high-frequency web crawler data collection scenarios, the caching mechanism can improve the decoding speed of subsequent data collection while significantly reducing hardware investment and maintenance costs.

[0105] Figure 3 See the flowchart of another embodiment of the obfuscated text recognition method provided in this application. Figure 3 As shown, this mainly describes how to decode the obfuscated text in the target font file to obtain the real target text, including the following steps:

[0106] Step 301: Extract the target font file URL from the content of the webpage to be identified.

[0107] Step 302: Obtain the target font file pointed to by the URL address of the target font file, and obtain the target mapping table corresponding to the target font file. The target mapping table is used to map the Unicode code points of the obfuscated text to the corresponding characters in order to decode the obfuscated text.

[0108] For steps 301-302 above, please refer to the description of the relevant embodiments above.

[0109] Step 303: For the obfuscated text in the target font file, obtain the Unicode code point of each obfuscated character in the obfuscated text.

[0110] In one embodiment, for the obfuscated text in the target font file, the Unicode code point corresponding to each obfuscated character in the obfuscated text is obtained. Specifically, the target font file is parsed using a font parsing tool to obtain the obfuscated text in the target font file, and relevant functions are used to traverse each obfuscated character in the obfuscated text to obtain the Unicode code point corresponding to each obfuscated character.

[0111] For example, the `ord(char)` function can be used to iterate through each obfuscated character in the obfuscated text to obtain its corresponding Unicode code point. The `ord(char)` function is a built-in function in Python used to obtain the Unicode code point of a single character. This is merely an example; other related functions can also be used to iterate through obfuscated characters and obtain their corresponding Unicode code points, and this application does not limit this approach.

[0112] Step 304: Based on the Unicode code point corresponding to each obfuscated character, find the glyph name corresponding to each obfuscated character from the character mapping table.

[0113] according to Figure 1 As described in the related embodiments, the target mapping table in this application includes a character mapping table and a glyph mapping table. The character mapping table contains the mapping relationship between Unicode code points and glyph names, while the glyph mapping table contains the mapping relationship between characters of glyph names. These are used to decode the Unicode code points in the obfuscated text into the corresponding characters, thereby obtaining the actual target text corresponding to the obfuscated text. Furthermore, the aforementioned glyph mapping table can be a user-defined mapping table adapted to the current network platform.

[0114] In one embodiment, the glyph name corresponding to each obfuscated character is retrieved from the character mapping table based on the Unicode code point corresponding to each obfuscated character. Specifically, based on the Unicode code point corresponding to each obfuscated character in the obfuscated text obtained in step 303 above, the character mapping table is searched to obtain the glyph name corresponding to each Unicode code point.

[0115] For example, suppose the obfuscated text is "#¥*", which consists of three different obfuscated characters. The three corresponding Unicode code points are "U+3042", "U+0026", and "U+0031". Looking up the character mapping table based on these three Unicode code points yields the corresponding glyph names: "a", "b", and "c". This is merely illustrative; the representation of glyph names and the number of obfuscated characters in the obfuscated text are not limited in this embodiment. The example above assumes the obfuscated text contains three obfuscated characters, but in reality, obfuscated text can contain at least one obfuscated character; the number of obfuscated characters is not limited.

[0116] Step 305: Among the multiple glyph names found, those that meet the set aggregation conditions are aggregated to form a target glyph name, thus obtaining at least one target glyph name.

[0117] Step 306: Based on at least one target glyph name, find the corresponding character from the glyph mapping table, and determine the found character as the target text corresponding to the obfuscated text.

[0118] The following is a unified explanation of steps 305-306:

[0119] One of the technical measures implemented by major online platforms to prevent web crawling is to split the real character "1" into multiple obfuscated characters "#¥@". For example, if the price of a product on an e-commerce platform is "123", the platform will map the actual price to multiple obfuscated characters "@#¥!". @! ☆¥¥”, then, decoding the multiple obfuscated characters corresponding to the above product price will yield multiple glyph names “onetwothree”. Therefore, it is necessary to aggregate the above multiple glyph names into at least one target glyph name that can be found in the glyph mapping table, so as to find the corresponding actual character based on the target glyph name.

[0120] In one embodiment, the glyph names that satisfy the set aggregation conditions from the multiple found glyph names are aggregated to form a target glyph name, resulting in at least one target glyph name. Specifically, the multiple glyph names found in step 304 are aggregated to form a target glyph name that can be found in the glyph mapping table, resulting in at least one target glyph name. Here, satisfying the set aggregation conditions means that multiple glyph names can be aggregated to form a target glyph name that can be found in the glyph mapping table; therefore, these multiple glyph names are glyph names that satisfy the aggregation conditions. If there are glyph names among the multiple glyph names that satisfy the set aggregation conditions, these satisfying glyph names are aggregated into a target glyph name that can be found in the glyph mapping table, ultimately resulting in at least one target glyph name.

[0121] For example, suppose that decoding the obfuscated text yields multiple glyph names "onetwothree". If multiple glyph names are found to satisfy the aggregation condition, these glyph names are aggregated into at least one target glyph name: "one", "two", "three". Furthermore, these target glyph names can be found in the glyph mapping table.

[0122] In one embodiment, at least one target glyph name can be obtained by aggregating glyph names that meet the set aggregation conditions from multiple found glyph names in the following manner: s glyph names are sequentially extracted from the glyph name sequence, where the glyph name sequence is formed by arranging the multiple found glyph names in order, and s is a preset positive integer; the currently extracted glyph names are aggregated to form a candidate glyph name, and a glyph mapping table is searched based on the candidate glyph name; if a candidate glyph name is found in the glyph mapping table, the candidate glyph name is determined as the target glyph name; if no candidate glyph name is found in the glyph mapping table, s+i glyph names are sequentially extracted from the glyph name sequence, and the step of aggregating the currently extracted glyph names to form a candidate glyph name is returned to be executed until the target glyph name is determined, or the first candidate glyph name is not found in the glyph mapping table; where i increments from 1, and the first candidate glyph name is formed by aggregating the glyph names corresponding to all the obfuscated characters.

[0123] Specifically, decoding the obfuscated text yields multiple glyph names, which are combined into a glyph name sequence. From this sequence, s glyph names are sequentially extracted according to their order, where s is a preset positive integer. These s extracted glyph names are aggregated into a candidate glyph name, which consists of s glyph names. The glyph mapping table is then searched based on this candidate glyph name. If the candidate glyph name is found in the glyph mapping table, it is designated as the target glyph name. If no candidate glyph name is found in the glyph mapping table, s+i glyph names are taken from the glyph name sequence according to the order of the glyph name sequence, and these s+i glyph names are aggregated to form a candidate glyph name. The candidate glyph name is searched for again in the glyph mapping table until a candidate glyph name is found in the glyph mapping table. If a candidate glyph name is found in the glyph mapping table, then this candidate glyph name is determined as the target glyph name. Alternatively, all glyph names in the glyph name sequence are aggregated into a first candidate glyph name. If the first candidate glyph name is not found in the glyph mapping table, then the search stops.

[0124] For example, suppose the sequence of glyph names corresponding to the obfuscated text is "onetwo". First, take s glyph names from this sequence and aggregate them into a candidate glyph name. Let's say s is 3. Then the candidate glyph name is "one". Search the glyph mapping table based on the candidate glyph name "one". If the candidate glyph name "one" is found in the glyph mapping table, then the candidate glyph name is determined as the target glyph name. Then, take s (3) glyph names from the above sequence and aggregate them into a candidate glyph name "two". Search the glyph mapping table based on this candidate glyph name. If the candidate glyph name is found in the glyph mapping table, then the candidate glyph name is determined as the target glyph name.

[0125] For example, suppose the sequence of glyph names corresponding to the obfuscated text is "onnt". First, s glyph names (assuming s is 3) are taken from this sequence and aggregated into a candidate glyph name "onn". If the candidate glyph name is not found in the glyph mapping table, then s+i glyph names (assuming i is 1) are taken from the sequence and aggregated into a candidate glyph name "onnt". At this point, the candidate glyph name contains all the glyph names in the sequence and is the first candidate glyph name. So the search stops, and the final target text is "onnt".

[0126] In one embodiment, based on at least one target glyph name, the corresponding character is found in the glyph mapping table, and the found character is determined as the target text corresponding to the obfuscated text. Specifically, assuming there are two target glyph names "one" and "two", and the characters corresponding to the above target glyph names are found in the glyph mapping table as "1" and "2" respectively, then the above characters are determined as the target text corresponding to the obfuscated text, that is, the target text is "12".

[0127] pass Figure 3 The process shown utilizes the glyph mapping table and character mapping table in the target mapping table to decode the obfuscated text, obtaining the decoded target text. This method ensures that the data collected by the crawler is real text data, not obfuscated text data. It can also be used in conjunction with the aforementioned... Figure 1 and Figure 2 The dynamic URL extraction and caching mechanism shown in the process dynamically establishes the correspondence between Unicode code points and characters of the obfuscated text based on the glyph mapping table and character mapping table when the network platform frequently updates the URL or font file, thereby realizing real-time decoding of the obfuscated text.

[0128] Figure 4 This is a structural block diagram of an obfuscated text recognition device provided in an embodiment of this application. Figure 4The device includes:

[0129] Address extraction module 41 is used to extract the URL address of the target font file from the content of the webpage to be identified;

[0130] The file acquisition module 42 is used to acquire the target font file pointed to by the URL address of the target font file, and to acquire the target mapping table corresponding to the target font file. The target mapping table is used to map the Unicode code points of the obfuscated text to the corresponding characters in order to decode the obfuscated text.

[0131] The text decoding module 43 is used to decode the obfuscated text in the target font file according to the target mapping table to obtain the target text corresponding to the obfuscated text.

[0132] In one possible implementation, the address extraction module 41 is specifically used for:

[0133] Based on predefined regular expression matching rules, extract the target font file URL address from the content of the web page to be identified.

[0134] In one possible implementation, the file acquisition module 42 includes:

[0135] The file search unit is used to search the cache area based on the target font file URL address, wherein the cache area includes the correspondence between font file URL addresses and font files;

[0136] The file determination unit is used to determine the font file corresponding to the target font file URL address as the target font file when the target font file URL address is found in the cache area.

[0137] The file download unit is used to download the target font file using the target font file URL address when the target font file URL address is not found in the cache area.

[0138] In one possible implementation, the file download unit is further configured to:

[0139] The target font file is stored in the cache area, and a correspondence between the target font file URL address and the target font file is established in the cache area.

[0140] In one possible implementation, the text decoding module 43 includes:

[0141] The code point acquisition unit is used to acquire the Unicode code point of each obfuscated character in the obfuscated text in the target font file;

[0142] The target text determination unit is used to find the corresponding character from the target mapping table based on the Unicode code point of each obfuscated character, and determine the character as the target text corresponding to the obfuscated text.

[0143] In one possible implementation, the target mapping table includes a character mapping table and a glyph mapping table. The character mapping table includes a correspondence between Unicode code points and glyph names, and the glyph mapping table includes a correspondence between glyph names and characters. The target text determination unit further includes:

[0144] The lookup subunit is used to find the glyph name corresponding to each of the obfuscated characters from the character mapping table based on the Unicode code point corresponding to each of the obfuscated characters.

[0145] The aggregation subunit is used to aggregate the multiple found glyph names that meet the set aggregation conditions to form a target glyph name, thereby obtaining at least one target glyph name;

[0146] The target text determination subunit is used to find the corresponding character from the glyph mapping table based on the at least one target glyph name, and determine the found character as the target text corresponding to the obfuscated text.

[0147] In one possible implementation, the polymer subunit is specifically used for:

[0148] s glyph names are sequentially extracted from the glyph name sequence, wherein the glyph name sequence is formed by arranging the found glyph names in order, and s is a preset positive integer;

[0149] Aggregate the currently extracted glyph names to form a candidate glyph name, and look up the glyph mapping table based on the candidate glyph name;

[0150] If the candidate glyph name is found in the glyph mapping table, the candidate glyph name is determined as the target glyph name;

[0151] If no candidate glyph name is found in the glyph mapping table, s+i glyph names are sequentially taken from the glyph name sequence, and the process returns to the step of aggregating the currently taken glyph names into a candidate glyph name, until the target glyph name is determined, or the first candidate glyph name is not found in the glyph mapping table; wherein i increments from 1, and the first candidate glyph name is formed by aggregating the glyph names corresponding to all the obfuscated characters.

[0152] like Figure 5As shown in the figure, this application provides an electronic device, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.

[0153] Memory 113 is used to store computer programs;

[0154] In one embodiment of this application, when the processor 111 executes the program stored in the memory 113, it implements the obfuscated text recognition method provided in any of the foregoing method embodiments, including:

[0155] Extract the URL address of the target font file from the content of the webpage to be identified;

[0156] Obtain the target font file pointed to by the URL address of the target font file, and obtain the target mapping table corresponding to the target font file. The target mapping table is used to map the Unicode code points of the obfuscated text to the corresponding characters in order to decode the obfuscated text.

[0157] For the obfuscated text in the target font file, the obfuscated text is decoded according to the target mapping table to obtain the target text corresponding to the obfuscated text.

[0158] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the obfuscated text recognition method provided in any of the foregoing method embodiments.

[0159] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0160] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0161] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also include the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.

[0162] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for recognizing obfuscated text, characterized in that, The method includes: Extract the URL address of the target font file from the content of the webpage to be identified; Obtain the target font file pointed to by the URL address of the target font file, and obtain the target mapping table corresponding to the target font file. The target mapping table is used to map the Unicode code points of the obfuscated text to the corresponding characters in order to decode the obfuscated text. For the obfuscated text in the target font file, the obfuscated text is decoded according to the target mapping table to obtain the target text corresponding to the obfuscated text.

2. The method according to claim 1, characterized in that, The step of extracting the target font file URL address from the content of the webpage to be identified includes: Based on predefined regular expression matching rules, extract the target font file URL address from the content of the web page to be identified.

3. The method according to claim 1, characterized in that, The step of obtaining the target font file pointed to by the URL address of the target font file includes: The cache area is located based on the target font file URL address, and the cache area includes the correspondence between font file URL addresses and font files; If the target font file URL address is found in the cache area, the font file corresponding to the found target font file URL address is determined as the target font file; If the target font file URL address is not found in the cache, the target font file is downloaded using the target font file URL address.

4. The method according to claim 3, characterized in that, After downloading the target font file using the target font file URL address, the process further includes: The target font file is stored in the cache area, and a correspondence between the target font file URL address and the target font file is established in the cache area.

5. The method according to claim 1, characterized in that, The step of decoding the obfuscated text in the target font file according to the target mapping table to obtain the target text corresponding to the obfuscated text includes: For the obfuscated text in the target font file, obtain the Unicode code point of each obfuscated character in the obfuscated text; Based on the Unicode code point of each obfuscated character, the corresponding character is found in the target mapping table, and the character is determined as the target text corresponding to the obfuscated text.

6. The method according to claim 5, characterized in that, The target mapping table includes a character mapping table and a glyph mapping table. The character mapping table includes the correspondence between Unicode code points and glyph names, and the glyph mapping table includes the correspondence between glyph names and characters. The step of finding the corresponding character in the target mapping table based on the Unicode code point of each obfuscated character, and determining the character as the target text corresponding to the obfuscated text, includes: Based on the Unicode code point corresponding to each of the obfuscated characters, the glyph name corresponding to each obfuscated character is found from the character mapping table; Among the multiple glyph names found, those that meet the set aggregation conditions are aggregated to form a target glyph name, thus obtaining at least one target glyph name; Based on the at least one target glyph name, the corresponding character is found from the glyph mapping table, and the found character is determined as the target text corresponding to the obfuscated text.

7. The method according to claim 6, characterized in that, The process of aggregating multiple found glyph names that meet the set aggregation conditions to form a target glyph name, thereby obtaining at least one target glyph name, includes: s glyph names are sequentially extracted from the glyph name sequence, wherein the glyph name sequence is formed by arranging the found glyph names in order, and s is a preset positive integer; Aggregate the currently extracted glyph names to form a candidate glyph name, and look up the glyph mapping table based on the candidate glyph name; If the candidate glyph name is found in the glyph mapping table, the candidate glyph name is determined as the target glyph name; If no candidate glyph name is found in the glyph mapping table, s+i glyph names are sequentially taken from the glyph name sequence, and the process returns to the step of aggregating the currently taken glyph names into a candidate glyph name, until the target glyph name is determined, or the first candidate glyph name is not found in the glyph mapping table; wherein i increments from 1, and the first candidate glyph name is formed by aggregating the glyph names corresponding to all the obfuscated characters.

8. A device for recognizing obfuscated text, characterized in that, The device includes: The address extraction module is used to extract the URL address of the target font file from the content of the webpage to be identified; The file acquisition module is used to acquire the target font file pointed to by the URL address of the target font file, and to acquire the target mapping table corresponding to the target font file. The target mapping table is used to map the Unicode code points of the obfuscated text to the corresponding characters in order to decode the obfuscated text. The text decoding module is used to decode the obfuscated text in the target font file according to the target mapping table to obtain the target text corresponding to the obfuscated text.

9. An electronic device, characterized in that, include: A processor and a memory, the processor being configured to execute a scrambled text recognition program stored in the memory to implement the scrambled text recognition method according to any one of claims 1-7.

10. A storage medium, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the obfuscated text recognition method according to any one of claims 1-7.

Citation Information

Cited By

  • Business data retrieval, profile construction method and apparatus, device, and medium

    CN122548017A