Method, System and Electronic Device for Identifying External Links in Web Pages
By deploying the web page external link recognition method on the IPv6 server, and using multi-pattern matching algorithm and grammatical rules to recognize and convert external links in the web page, the problems of inaccurate external link recognition and untimely conversion in the existing technology are solved, and the user experience and access success rate are improved.
Patent Information
- Application Number
- CN202210622679.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-01
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-06-01
AI Technical Summary
The prior art cannot quickly and accurately identify external links in web pages, and cannot convert external link addresses before the IPv6 client clicks the external link, resulting in the IPv6 client failing to access external links.
Deploy a method for identifying external links in web pages on the IPv6 server. By obtaining the top-level domain name string in the HTTP response message, extracting context information, and determining whether it is a legal external link based on the preset syntax rules. Use multi-pattern matching algorithm and dictionary tree to quickly find and match domain name strings to improve the efficiency and accuracy of external link searches.
It realizes fast and accurate identification of external links in web pages, and converts formats before the IPv6 client clicks on external links, ensuring that the IPv6 client can successfully access the proxy server that supports IPv6.
Smart Images

Figure CN115022284B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and in particular, to a method, a system, and an electronic device for identifying external links in a web page. Background Art
[0002] After a website on the IPv4 network is upgraded to a website on the IPv6 network, there will be links to other websites on the website page, that is, external links. If the external link website does not have an IPv6 network, then when an IPv6 client clicks on the external link in the upgraded website, the access will fail. Therefore, before returning the web page to the client, the IPv6 website needs to identify the external links in the page in order to perform a certain format conversion on the external link domain name. The existing technical methods cannot achieve a quick search for external links, nor can they accurately determine whether the found link is an external link. Summary of the Invention
[0003] In view of this, an object of the present invention is to provide a method, a system, and an electronic device for identifying external links in a web page, so as to alleviate the technical problems in the prior art that cannot achieve a quick search for external links and cannot accurately determine whether the found link is an external link.
[0004] In a first aspect, an embodiment of the present invention provides a method for identifying external links in a web page, which is applied to an IPv6 server. The IPv6 server is deployed between an IPv6 client and an IPv4 HTTP site. The method includes: obtaining a top-level domain name string in an HTTP response message in the web page; wherein, the HTTP response message is obtained from the HTTP response message responded by the IPv4 HTTP site after the IPv4 HTTP site sends an HTTP request to the IPv6 client; extracting context information from the found top-level domain name string, and obtaining the complete domain name string where it is located; judging the context information based on a preset syntax rule to determine whether the domain name string is a legal external link.
[0005] Further, the step of obtaining the top-level domain name string in the HTTP response message in the web page includes: quickly finding the top-level domain name string in the HTTP response message through a string matching algorithm, and matching the HTTP response message with a preset trie tree through a multi-pattern matching algorithm to determine the top-level domain name string in the HTTP response message; wherein, the trie tree is pre-constructed based on common top-level domain names as keywords.
[0006] Here, the trie tree is pre-constructed by the IPv6 server using common top-level domain names, including: ".com", ".cn", ".net", and ".org", etc.
[0007] Further, the step of extracting the context information of the top-level domain name string from the HTTP response message includes: extracting the pre-domain name string of the top-level domain name string and the string after the top-level domain name string; determining the pre-domain name string and the string after the top-level domain name string as the context information of the top-level domain name string.
[0008] Further, the step of determining whether the domain name string is a legal external link based on a preset syntax rule for the context information includes: determining whether the characters in the pre-domain name string belong to legal domain name characters. If not, it is determined that the domain name string is not a legal external link; if so, the pre-domain name string is determined as the pre-legal domain name string; determining whether the pre-legal domain name string contains a specified character. If so, it is determined that the domain name string is not a legal external link; the specified characters include: the "@" character, consecutive "." characters, and consecutive "-" characters; if not, the domain name string is determined as a suspected external link, and it is determined whether the suspected external link is a suspected legal external link based on the string after the top-level domain name string.
[0009] Further, the step of determining whether the suspected external link is a suspected legal external link based on the string after the top-level domain name string includes: S1: determining whether the first character in the string after the top-level domain name string is the ":" character. If so, execute S2; if not, execute S3; S2: determining whether there are consecutive numeric characters after the ":" character and the value of the consecutive numeric characters is between 0 and 65535. If so, execute S3; if not, it is determined that the suspected external link is not a suspected legal external link; S3: determining whether the next character after the consecutive numeric characters is the "?" character or the " / " character or the "#" character. If so, it is determined that the suspected external link is a suspected legal external link; if not, execute S4; S4: determining the first non-space character after the consecutive numeric characters, and determining whether the non-space character is the """" character, or the ""'" character, or the "`" character. If so, it is determined that the suspected external link is a suspected legal external link; if not, it is determined that the suspected external link is not a suspected legal external link.
[0010] Further, after the step of determining the first non-space character after the consecutive numeric characters and determining whether the non-space character is the """" character, or the ""'" character, or the "`" character, and if so, determining that the suspected external link is a suspected legal external link, the method further includes: obtaining the domain name of the suspected legal external link; determining whether the domain name is a web domain name; if not, it is determined that the suspected legal external link is a legal external link.
[0011] Further, the method further includes: if an access request from an IPv6 client for an HTTP external link is received, encapsulating the HTTP response message into an IPv6 response message and sending it to the IPv6 client.
[0012] Second aspect, an external link recognition system in a web page provided by an embodiment of the present invention is applied to an IPv6 server. The IPv6 server is deployed between an IPv6 client and an IPv4 HTTP site. The system includes: an acquisition module, configured to acquire a top-level domain name string in an HTTP response message in a web page; where the HTTP response message is obtained from an HTTP response message responded by the IPv4 HTTP site after the IPv6 client sends an HTTP request to the IPv4 HTTP site; a context extraction module, configured to extract context information from the found top-level domain name string and obtain the complete domain name string where it is located; a judgment module, configured to judge the context information based on a preset syntax rule to determine whether the domain name string is a legal external link.
[0013] Third aspect, an embodiment of the present invention provides an electronic system, including a processing device and a storage device; a computer program is stored on the storage device, and when the computer program is run by the processing device, it executes the external link recognition method of any one of the above.
[0014] Fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processing device, it executes the steps of the external link recognition method as described in any one of the above.
[0015] An embodiment of the present invention provides an external link recognition method, system and electronic device in a web page, which are applied to an IPv6 server. The IPv6 server is deployed between an IPv6 client and an IPv4 HTTP site, and includes: acquiring a top-level domain name string in an HTTP response message in a web page; extracting context information from the found top-level domain name string and obtaining the complete domain name string where it is located; judging the context information based on a preset syntax rule to determine whether the domain name string is a legal external link. In this way, based on the domain name characteristics in the external link and the basic logical structure in the web page, the external links in the web page are searched by a multi-pattern matching algorithm to improve the efficiency and accuracy of external link search, thereby improving the user experience.
[0016] Other features and advantages of the present invention will be described in the following specification, and part of them will become obvious from the specification, or will be understood by implementing the present invention.
[0017] To make the above objects, features and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0018] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0019] Figure 1 Flowchart of the external link recognition method in a web page provided in the first embodiment of the present invention;
[0020] Figure 2 Flowchart for determining whether a suspected external link is a suspected legal external link based on the string after the top-level domain name string provided in the first embodiment of the present invention;
[0021] Figure 3 Schematic diagram of the external link recognition system in a web page provided in the second embodiment of the present invention.
[0022] Icons: 1 - Acquisition module; 2 - Context extraction module; 3 - Judgment module. Specific embodiments
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention with reference to the drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0024] To facilitate the understanding of this embodiment, the following will introduce the embodiments of the present invention in detail.
[0025] Embodiment 1:
[0026] After the website of the IPv4 network is upgraded to the website of the IPv6 network, there will be external links on its website pages. If the external link website does not have an IPv6 network, then when an IPv6 client clicks on this external link in the upgraded website, the access will fail. Currently, the existing external link recognition methods in web pages cannot quickly and accurately identify external links, nor can they convert the external link address before the IPv6 client clicks on the external link.
[0027] To solve the above problems, an idea provided by the present invention is that before the IPv6 server returns the web page to the client in the IPv6 website, it identifies the external links in the web page and converts the format of the external link domain names. When the IPv6 client clicks on the converted external link domain name, it will access the proxy server that supports IPv6, thus achieving successful access.
[0028] To describe the technical method of the present invention in detail, an embodiment of the present invention provides a method for identifying external links in a web page. Referring to Figure 1 , this method can be applied to an IPv6 server, which is deployed between an IPv6 client and an IPv4 HTTP site. The method includes the following steps:
[0029] Step S101, obtain the top-level domain name string in the HTTP response message in the web page; wherein, the HTTP response message is obtained from the HTTP response message responded by the IPv4 HTTP site after the IPv6 client sends an HTTP request to the IPv4 HTTP site;
[0030] Step S102, extract the context information from the found top-level domain name string and obtain the complete domain name string where it is located;
[0031] Step S103, judge the context information based on a preset syntax rule to determine whether the domain name string is a legal external link.
[0032] Here, the HTTP response message is the request response message generated by the IPv4 server after the IPv6 client sends a source station resource request message to the IPv4 server. After obtaining the domain name string in the HTTP response message, the top-level domain name string in the web page is found based on the characteristics of the domain name, and it is judged whether the context of the top-level domain name string is a domain name or a code segment according to the web page logic, so as to identify the legal external link.
[0033] Further, the step of obtaining the top-level domain name string in the HTTP response message in the web page includes:
[0034] Quickly find the top-level domain name string in the HTTP response message through a string matching algorithm, and match the HTTP response message with a preset trie through a multi-pattern matching algorithm to determine the top-level domain name string in the HTTP response message; wherein, the trie is pre-constructed based on common top-level domain name nouns as keywords.
[0035] Specifically, using the multi-pattern matching algorithm, sequentially match the domain name strings in the web page with the trie. If the current domain name string matches the trie successfully, it is determined that the current domain name string is a common top-level domain name noun; if the current domain name string fails to match the trie, it is determined that the current domain name string is not a common top-level domain name noun. If the current domain name string is not the last domain name string in the page, match the next domain name string with the trie until there are no unmatched domain name strings in the page.
[0036] Further, the step of extracting the context information of the top-level domain name string from the HTTP response message includes:
[0037] Extract the pre - domain name string before the top - level domain name string and the string after the top - level domain name string;
[0038] Determine the context information of the top - level domain name string by using the pre - domain name string and the string after the top - level domain name string.
[0039] Here, the pre - domain name string of the top - level domain name string refers to the string composed of all consecutive domain names before the top - level domain name string.
[0040] Further, the step of judging whether the domain name string is a legal external link based on a preset syntax rule includes:
[0041] Judge whether the characters in the pre - domain name string belong to legal domain name characters. If not, determine that the domain name string is not a legal external link;
[0042] If so, determine the pre - domain name string as the pre - legal domain name string;
[0043] Judge whether the pre - legal domain name string contains specified characters. If so, determine that the domain name string is not a legal external link; the specified characters include: the "@" character, consecutive "." characters, and consecutive "-" characters;
[0044] If not, determine the domain name string as a suspected external link, and determine whether the suspected external link is a suspected legal external link based on the string after the top - level domain name string.
[0045] Specifically, after finding the top - level domain name string, judge the characters in the pre - domain name string one by one to obtain a suspected external link. If the characters in the pre - domain name string are legal domain name characters, that is, numeric characters, English letter characters, "-" characters, and "." characters, then the current character belongs to the characters of the suspected external link until the first character that is not a legal domain name character is encountered, and then the judgment ends, and a suspected external link is obtained. If the first illegal domain name character is the "@" character, it means that the string is an email domain name and is not an external link either.
[0046] During the process of obtaining the suspected external link character by character, the "-" character and the "." character cannot exist adjacent to each other, such as: "-.", and the "-" character or the "." character cannot have the same adjacent characters, such as: "--" or "..". If the above conditions are not met, the current domain name string is not a suspected external link.
[0047] Further, referring to Figure 2The flowchart for determining whether a suspected external link is a suspected legal external link based on the string after the top-level domain name string. The steps for determining whether a suspected external link is a legal external link based on the second non-domain name string can be implemented through the following steps S1 - S4:
[0048] S1: Determine whether the first character in the string after the top-level domain name string is a ":" character. If it is, execute S2; if not, execute S3;
[0049] S2: Determine whether there are consecutive numeric characters after the ":" character, and the value of the consecutive numeric characters is between 0 - 65535. If it is, execute S3; if not, determine that the suspected external link is not a suspected legal external link;
[0050] S3: Determine whether the next character after the consecutive numeric characters is a "?" character or a " / " character or a "#" character. If it is, determine that the suspected external link is a suspected legal external link; if not, execute S4;
[0051] S4: Determine the first non-whitespace character after the consecutive numeric characters, and determine whether the non-whitespace character is a "" character, or a ''' character, or a ` character. If it is, determine that the suspected external link is a suspected legal external link; if not, determine that the suspected external link is not a suspected legal external link.
[0052] Specifically, if the current domain name string is determined to be a suspected external link, then determine whether the suspected external link is a legal external link based on the string after the top-level domain name string. Obtain the characters in the string after the top-level domain name string in sequence, and judge whether each character meets the requirements of the html (HyperText Markup Language) for the domain name format, and judge whether there is a port number in the string after the top-level domain name string and whether the format of the port number is correct. If the first character in the string after the top-level domain name string is a ":" character, according to the URL (Uniform Resource Locator) rule, there is a port number in the string after the top-level domain name string. Obtain the characters after the ":" character, and judge whether there are consecutive numeric characters among them, and the value of the consecutive numeric characters is between 0 - 65535. If not, the suspected external link is not a suspected legal external link.
[0053] If there are consecutive numeric characters after the ":" character, and the value of the consecutive numeric characters is between 0 - 65535, then judge whether the next character is one of the "?" character or the " / " character or the "#" character. If it is, it means that the suspected external link is a suspected legal external link.
[0054] If not, starting from the current character, judge whether the subsequent characters are whitespace characters one by one, such as the " " character, the "\n" character, the "\r" character, the "\t" character, until the first non-whitespace character is found. If the non-whitespace character is the " character or the ' character or the ` character, the suspected external link is a suspected legal external link; otherwise, the suspected external link is not a suspected legal external link.
[0055] Further, after determining the first non-whitespace character after the consecutive digital characters, judge whether the non-whitespace character is the " character, or the ' character, or the ` character. If so, after the step of determining that the suspected external link is a suspected legal external link, the method further includes:
[0056] Obtain the domain name of the suspected legal external link;
[0057] Judge whether the domain name is a web domain name;
[0058] If not, determine that the suspected legal external link is a legal external link.
[0059] Here, if the domain name of the current legal external link is not the domain name of this website, determine that the current suspected legal external link is a legal external link.
[0060] Further, the method for identifying external links in a web page further includes:
[0061] If an access request from an IPv6 client for an HTTP external link is received, encapsulate the HTTP response message into an IPv6 response message and send it to the IPv6 client.
[0062] Here, according to the search result of the HTTPS external link, convert the format of the external link domain name. When the IPv6 client clicks on the converted external link domain name, it will access the proxy server that supports IPv6, and thus the access is successful.
[0063] An embodiment of the present invention provides a method for identifying external links in a web page, which is applied to an IPv6 server. The IPv6 server is deployed between an IPv6 client and an IPv4 HTTP site. The method includes: obtaining the top-level domain name string in the HTTP response message in the web page; extracting context information from the found top-level domain name string and obtaining the complete domain name string where it is located; judging the context information based on a preset syntax rule to determine whether the domain name string is a legal external link. In this way, based on the domain name characteristics in the external link and the basic logical structure in the web page, the top-level domain name string in the web page is identified through a multi-pattern matching algorithm, and the context of the top-level domain name string is judged based on a preset syntax rule, so as to search for external links, improve the efficiency and accuracy of external link search, and thus improve the user experience.
[0064] Embodiment Two:
[0065] Figure 3 Schematic diagram of the external link recognition system in a web page provided in the second embodiment of the present invention.
[0066] Refer to Figure 3 , the external link recognition system in the web page is applied to an IPv6 server, and the IPv6 server is deployed between an IPv6 client and an IPv4 HTTP site. The system includes:
[0067] An obtaining module 1, configured to obtain the top-level domain name string in the HTTP response message in the web page; wherein, the HTTP response message is obtained from the HTTP response message responded by the IPv4 HTTP site after the IPv6 client sends an HTTP request to the IPv4 HTTP site;
[0068] A context extraction module 2, configured to extract context information from the found top-level domain name string and obtain the complete domain name string where it is located;
[0069] A judgment module 3, configured to judge the context information based on a preset syntax rule to determine whether the domain name string is a legal external link.
[0070] The embodiment of the present invention provides an external link recognition system in a web page, which is applied to an IPv6 server. The IPv6 server is deployed between an IPv6 client and an IPv4 HTTP site. The system includes: obtaining the top-level domain name string in the HTTP response message in the web page; extracting context information from the found top-level domain name string and obtaining the complete domain name string where it is located; judging the context information based on a preset syntax rule to determine whether the domain name string is a legal external link. In this system, the external links in the web page are searched through a multi-pattern matching algorithm to improve the efficiency and accuracy of external link search, thereby improving the user experience.
[0071] The embodiment of the present invention further provides an electronic system, including a processing device and a storage device; a computer program is stored on the storage device, and when the computer program is run by the processing device, it implements the steps of the external link recognition method provided in the above embodiment.
[0072] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processing device, it executes the steps of the external link recognition method in the above embodiment.
[0073] The computer program product provided by the embodiment of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the foregoing method embodiment. For specific implementation, reference can be made to the method embodiment, which will not be elaborated here.
[0074] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0075] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "install", "connect", and "couple" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0076] If the above-mentioned functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0077] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0078] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the technical field of the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention.
Claims
1. A method for identifying external links in a web page, characterized in that, it is applied to an IPv6 server, and the IPv6 server is deployed between an IPv6 client and an IPv4 HTTP site. The method includes: Obtaining the top-level domain name string in the HTTP response message of the web page; wherein, the HTTP response message is a request response message generated by the IPv4 HTTP site after the IPv6 client sends an HTTP request to the IPv4 HTTP site; Extracting context information from the found top-level domain name string and obtaining the complete domain name string where it is located; Judging the context information based on a preset syntax rule to determine whether the domain name string is a legal external link; The step of extracting the context information of the top-level domain name string from the HTTP response message includes: Extracting the pre-order domain name string of the top-level domain name string and the string after the top-level domain name string; Determining the pre-order domain name string and the string after the top-level domain name string as the context information of the top-level domain name string.
2. The method for identifying external links in a web page according to claim 1, characterized in that, the step of obtaining the top-level domain name string in the HTTP response message of the web page includes: Quickly finding the top-level domain name string in the HTTP response message through a string matching algorithm, and matching the HTTP response message with a preset trie through a multi-pattern matching algorithm to determine the top-level domain name string in the HTTP response message; wherein, the trie is pre-constructed based on common top-level domain names as keywords.
3. The method for identifying external links in a web page according to claim 1, characterized in that, the step of judging the context information based on a preset syntax rule to determine whether the domain name string is a legal external link includes: Judging whether the characters in the pre-order domain name string belong to legal domain name characters. If not, determining that the domain name string is not a legal external link; If so, determining the pre-order domain name string as a pre-order legal domain name string; Judging whether the pre-order legal domain name string contains specified characters. If so, determining that the domain name string is not a legal external link; the specified characters include: the "@” character, consecutive "." characters, and consecutive "-” characters; If not, determining the domain name string as a suspected external link and determining whether the suspected external link is a suspected legal external link based on the string after the top-level domain name string.
4. The method for identifying external links in a web page according to claim 3, characterized in that, the step of determining whether the suspected external link is a suspected legal external link based on the string after the top-level domain name string includes: S1: Judging whether the first character in the string after the top-level domain name string is a ":" character. If so, execute S2; if not, execute S3; S2: Determine whether there are consecutive numeric characters after the ":" character, and whether the value of the consecutive numeric characters is between 0 and 65535. If so, execute S3; if not, determine that the suspected external link is not the suspected legitimate external link; S3: Determine whether the next character after the consecutive numeric characters is a "?", " / ", or "#" character. If so, determine that the suspected external link is the suspected legitimate external link; if not, execute S4; S4: Determine the first non-space character after the consecutive numeric characters, and determine whether the non-space character is a """, "'", or "`" character. If so, determine that the suspected external link is the suspected legitimate external link; if not, determine that the suspected external link is not the suspected legitimate external link.
5. The method for identifying external links in a web page according to claim 4, wherein, after the step of determining the first non-space character after the consecutive numeric characters and determining whether the non-space character is a """, "'", or "`" character, and if so, determining that the suspected external link is the suspected legitimate external link, the method further includes: Obtain the domain name of the suspected legitimate external link; Determine whether the domain name is a web page domain name; If not, determine that the suspected legitimate external link is the legitimate external link.
6. The method for identifying external links in a web page according to claim 5, wherein, the method further includes: If an access request from the IPv6 client for the HTTP external link is received, encapsulate the HTTP response message into an IPv6 response message and send it to the IPv6 client.
7. A system for identifying external links in a web page, wherein, applied to an IPv6 server, the IPv6 server is deployed between an IPv6 client and an IPv4 HTTP site, and the system includes: An acquisition module, configured to acquire the top-level domain name string in the HTTP response message in the web page; wherein, the HTTP response message is a request response message generated by the IPv4 HTTP site after the IPv6 client sends an HTTP request to the IPv4 HTTP site; A context extraction module, configured to extract context information from the found top-level domain name string and obtain the complete domain name string where it is located; A judgment module, configured to judge the context information based on a preset syntax rule to determine whether the domain name string is a legitimate external link; The context extraction module is further configured to extract the pre-order domain name string of the top-level domain name string and the string after the top-level domain name string; and determine the pre-order domain name string and the string after the top-level domain name string as the context information of the top-level domain name string.
8. An electronic system, wherein, the electronic system includes: a processing device and a storage device; The storage device stores a computer program, and the computer program, when run by the processing device, executes the external link identification method according to any one of claims 1 to 6.
9. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is run by a processing device, it executes the steps of the external link recognition method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Matching method, client side, server and matching equipment
CN106161352A
IPv4 / IPv6 address conversion system
CN109451097A
Data processing method, device and system
CN112866439A