Web filtering system

The web filtering system allows control of data transmission based on pre-registered keywords using hashing and rolling hash methods, addressing the inefficiencies of existing systems by enabling precise and efficient prevention of inappropriate content.

JP2025103318AActive Publication Date: 2025-07-09NETABTAR
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2023220640
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-07-09
Estimated Expiration
2043-12-27

AI Technical Summary

Technical Problem

Existing web filtering systems lack the ability to control access based on the content of data transmitted to a website, leading to the need to prohibit entire websites or functions to prevent inappropriate information, which is inefficient and unsatisfactory for diverse user needs.

Method used

A web filtering system that includes a storage unit for pre-registered keywords and a determination unit to check if these keywords are included in the request body data, using hashing and rolling hash methods for efficient matching.

Benefits of technology

Enables precise control of data transmission based on user-defined keywords, preventing inappropriate information while improving processing speed and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025103318000001_ABST
    Figure 2025103318000001_ABST
Patent Text Reader

Abstract

To provide a technique for controlling an access based on the content of data to be transmitted to a website.SOLUTION: Individual keywords registered by a user and hashed keywords obtained by hashing the keywords are previously stored in DB. A filtering server sequentially hashes division data generated by dividing body data of a request so as to confirm whether or not individual keywords are included in the body data (S63 and S71), matches the hashed division data with the hashed keywords (S64), and determines a regulation (S65), when the hashed division data and the hashed keyword are matched (S64: Yes). Matching can be performed at high speed by hashing. When the second and subsequent hashed division data is considered as a target of matching, the tail part of the immediately preceding hashed division data is connected to the top of the hashed division data (S72), and accordingly the regulation can be determined without missing a keyword existing over the two division data.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to control of access to a website, and more particularly to a web filtering system that controls access based on the content of data to be transmitted to the website.

Background Art

[0002] Conventionally, a technique for controlling communication between a computer and a web server to which the computer attempts to access in accordance with predetermined rules is known. For example, in Patent Document 1, the administrator has previously set whether browsing is permitted or prohibited for each category, and when a browsing request for a website is received, the filtering server is queried about the category of the website, and browsing of the website is prohibited or permitted based on the category information transmitted from the filtering server. Further, Patent Document 2 describes that access restriction conditions for determining whether or not to restrict access from a first computer to a second computer may include settings for permission or prohibition in units of individual functions (login, mail, writing, upload, etc.) included in the service provided by the second computer.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Patent Document 3

Summary of the Invention

Problems to be Solved by the Invention

[0004] According to the former prior art, access to a website can be controlled based on the category of the website specified by the URL. According to the latter prior art, it is considered that the control of one website can be subdivided according to functions, such as permitting only the use of specific functions provided by the website while prohibiting the use of other functions. However, in any of the prior arts, since they do not involve the content of the information transmitted to the website, in order to prevent inappropriate information from being transmitted, even if there is no problem with browsing, it is necessary to prohibit the use of the entire website or all the functions provided by the website.

[0005] On the other hand, in an electronic bulletin board system, when a posted bulletin article contains a posting-prohibited term inappropriate for posting on the electronic bulletin board, the posting to the electronic bulletin board is automatically rejected. For bulletin articles that do not contain posting-prohibited terms but contain content that slanders an individual or personal information, there is also a system that deletes them after the administrator checks (see, for example, Patent Document 3). However, websites with such functions are only a very small part. Also, since the information considered inappropriate varies among users (such as companies and schools), even if all websites have such functions, it is not always possible to achieve satisfactory operation for users. In view of these points, in order to prevent information considered inappropriate by users from being transmitted, some control is required on the transmission side.

[0006] Therefore, an object of the present invention is to provide a technology for controlling access based on the content of data to be transmitted to a website.

Means for Solving the Problems

[0007] To solve the above problems, the present invention employs the following web filtering system. The following explanatory notes in parentheses are merely examples, and the present invention is not limited thereto.

[0008] That is, the web filtering system of the present invention includes a storage unit that stores one or more keywords pre-registered by a user, and a determination unit that, when receiving an inquiry regarding the regulation of a request from a browser to a web server, makes a determination based on whether the keyword is included in the request body data and returns the determination result to the inquiry source.

[0009] The keyword used for determination in the web filtering system is pre-registered by the user. Therefore, according to the web filtering system, it is possible to make a determination on whether to regulate in accordance with the user's intention, and it is possible to reliably regulate the transmission of data including keywords that the user considers inappropriate.

[0010] Preferably, in the web filtering system of the above-described aspect, the determination unit confirms whether the keyword is included in the body data by performing matching between the hashed data obtained by hashing the body data and the hashed keyword obtained by hashing the keyword.

[0011] According to the web filtering system of this aspect, since matching is performed between the hashed data and the hashed keyword, it is possible to execute the matching, and thus the confirmation of whether the keyword is included in the body data, at a higher speed compared to the case where matching is performed without hashing.

[0012] More preferably, in the web filtering system of any of the above-described aspects, the storage unit stores the keyword and the corresponding hashed keyword in advance.

[0013] According to the web filtering system of this aspect, since the hashed keyword stored in the storage unit can be used as it is during matching, it is not necessary to hash each keyword every time, and the efficiency of matching can be improved.

[0014] More preferably, in the web filtering system according to any of the above-described aspects, the determination unit divides the body data into predetermined sizes, performs matching between the hashed divided data obtained by hashing the divided data and the hashed keyword, and when the second and subsequent hashed divided data are the targets of matching, at the head of the hashed divided data, hash data corresponding to the number of characters of the longest keyword registered by the user, which forms the tail end of the previous hashed divided data, is concatenated.

[0015] According to the web filtering system of this aspect, since matching with the hashed keyword is performed on the hashed data formed by concatenating the tail end of the previous hashed divided data to the head of the second and subsequent hashed divided data, it is possible to make a regulation determination without missing a keyword existing across two divided data, and it becomes possible to surely regulate the transmission of data.

Advantages of the Invention

[0016] As described above, according to the present invention, access can be controlled based on the content of the data to be transmitted to the website.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Mode for Carrying Out the Invention

[0018] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the following embodiments are preferred examples, and the present invention is not limited to this example.

[0019] FIG. 1 is a block diagram showing a web filtering system 1 according to an embodiment. The web filtering system 1 is a system that controls access from a terminal in a user environment (for example, in an environment such as a school or a company) to various web servers WS existing on the Internet based on pre - registered information. The program for determining permission or restriction of access is implemented in a filtering server 10 arranged on the cloud, and the CPU 11 of the filtering server 10 is the execution entity of this program. In addition, a database (DB) 12 is provided in the filtering server 10, and various information used for determination by the program is stored therein.

[0020] In addition to the filtering server 10, a proxy server 20 and a redirect server 30 are arranged on the cloud. Also, a browser 40 is installed on the terminal in the user environment, and an application 50 and a browser extension 60 are further installed as necessary. When attempting to access the web server WS from the browser 40 of the terminal, a query regarding whether to restrict the request from the browser 40 is made to the filtering server 10 via any one of the proxy server 20, the redirect server 30, and the application 50, and communication is controlled based on the result.

[0021] In other words, in the web filtering system 1, three forms for controlling the communication between the browser 40 and the web server WS are provided, and a form suitable for the terminal is selected. Specifically, a proxy type that controls communication via the proxy server 20 using the proxy settings of the terminal, an ICAP type that controls communication via the application 50 without using the proxy settings of the terminal, and a redirect server type that controls communication via the browser extension 60 using the redirect server 30 are available.

[0022] FIG. 2 shows the proxy type communication control. In FIG. 2, (A) is a flowchart showing the flow of proxy type control, and (B) in FIG. 2 is a block diagram showing the configuration related to proxy type control. Note that the arrows in the block diagram of (B) indicate the main direction in which data flows, the step numbers attached to the arrows correspond to the step numbers in the flowchart of (A), and the black arrows indicate restrictions. Also, although the illustration of the Internet is omitted in the block diagram of (B), the communication between the terminal and the server is made via the Internet (the same applies in FIGS. 3 and 4). The following will be described along the flow.

[0023] Steps S11, S12: When the proxy server 20 receives a request from the browser 40 for the web server WS (step S11), it inquiries the filtering server 10 about whether to restrict this request (step S12).

[0024] Step S13: In response to an inquiry from the proxy server 20, the filtering server 10 executes a filtering determination process and returns the result to the proxy server 20. The details of the filtering determination process will be further described later with reference to another drawing.

[0025] Steps S14, S15: When the proxy server 20 receives "restriction" from the filtering server 10 (Step S14: Yes), it returns a restriction response to the browser 40 (Step S15). As a result, a screen indicating that access is restricted is displayed on the browser 40. On the other hand, when the proxy server 20 receives "permission" from the filtering server 10 (Step S14: No), it proceeds to Step S16.

[0026] Steps S16 - S18: The proxy server 20 sends a request to the web server WS (Step S16), receives a response from the web server WS in response thereto (Step S17), and sends it to the browser 40 (Step S18). As a result, the requested page is displayed on the browser 40.

[0027] FIG. 3 shows ICAP - type communication control. (A) in FIG. 3 is a flowchart showing the flow of ICAP - type control, and (B) in FIG. 3 is a block diagram showing the configuration related to ICAP - type control. The following will be described along the flow.

[0028] Steps S21, S22: The application 50 hooks a request from the browser 40 to the web server WS (Step S21) and inquires of the filtering server 10 whether the hooked request is restricted or not (Step S22).

[0029] Step S23: In response to an inquiry from the application 50, the filtering server 10 executes a filtering determination process and returns the result to the application 50.

[0030] Steps S24, S25: When the application 50 receives "Restriction" from the filtering server 10 (Step S24: Yes), it returns a restriction response to the browser 40 (Step S25). As a result, a screen indicating that access is restricted is displayed on the browser 40. On the other hand, when "Permission" is returned from the filtering server 10 (Step S24: No), the application 50 proceeds to Step S26.

[0031] Steps S26 - S28: The application 50 sends a request to the web server WS (Step S26), receives a response from the web server WS in response to this (Step S27), and sends it to the browser 40 (Step S28). As a result, the requested page is displayed on the browser 40.

[0032] Figure 4 shows communication control of the redirect server type. In Figure 4(A), it is a flowchart showing the control flow of the redirect server type, and in Figure 4(B), it is a block diagram showing the configuration related to the control of the redirect server type. The following will be described along the flow.

[0033] Step S31: The browser extension 60 hooks a request from the browser 40 to the web server WS and transfers it to the redirect server 30.

[0034] Step S32: The redirect server 30 inquires of the filtering server 10 whether the request transferred from the browser extension 60 is restricted.

[0035] Step S33: In response to the inquiry from the redirect server 30, the filtering server 10 executes filtering determination processing and returns the result to the redirect server 30.

[0036] Steps S34 - S36: When the redirect server 30 receives "Restriction" from the filtering server 10 (Step S34: Yes), it returns a redirect response to the restriction screen to the browser extension 60 (Step S35). As a result, the browser 40 redirects to a screen indicating that the access is restricted. On the other hand, when "Permission" is returned from the filtering server 10 (Step S34: No), the redirect server 30 returns a redirect response to the web server WS to the browser extension 60 (Step S36).

[0037] Steps S37, S38: The browser extension 60 sends a request to the web server WS (Step S37) and receives a response from the web server WS in response (Step S38). As a result, the requested page is displayed on the browser 40.

[0038] Figure 5 is a flowchart showing an example of the procedure of the filtering determination process. The filtering determination process is executed by the CPU 11 of the filtering server 10 that has received an inquiry regarding whether to restrict a request in the process of controlling communication between the browser 40 and the web server WS (Step S13 in FIG. 2, Step S23 in FIG. 3, Step S33 in FIG. 4). Hereinafter, the main processes executed in the filtering determination process will be described according to the example of the procedure.

[0039] Step S51: The CPU 11 executes a URL determination process. In the URL determination process, the CPU 11 determines whether to permit access to the website specified by the URL of the received request based on the domain information and category information pre - stored in the database 12, and returns a determination result of "Permission" or "Restriction". Note that since the specific determination method is the same as that described in, for example, Japanese Patent No. 6259175, the description is omitted here.

[0040] Step S52: If the return value of the URL determination process is "permitted" (Step S52: Yes), the CPU 11 proceeds to Step S53. On the other hand, if the return value is "restricted" (Step S52: No), the CPU 11 proceeds to Step S56.

[0041] Step S53: Next, the CPU 11 executes the keyword determination process. In the keyword determination process, the CPU 11 determines whether to permit access based on whether the data (request body data) about to be sent from the browser 40 to the web server WS contains keywords pre-registered by the user (such as a school, a company, etc.), and returns a determination result of "permitted" or "restricted". The detailed content of the keyword determination process will be described later in detail with reference to another drawing.

[0042] Step S54: If the return value of the keyword determination process is "permitted" (Step S54: Yes), the CPU 11 proceeds to Step S55. On the other hand, if the return value is "restricted" (Step S54: No), the CPU 11 proceeds to Step S56.

[0043] Steps S55, S56: If the return values of both the URL determination process and the keyword determination process are "permitted", the CPU 11 returns "permitted" to the inquiry sender (Step S55). In contrast, if the return value of the URL determination process or the keyword determination process is "restricted", the CPU 11 returns "restricted" to the inquiry sender (Step S56). After finishing the above procedure, the CPU 11 ends the filtering determination process.

[0044] Note that the above procedure example is merely an example and can be changed as appropriate. For example, in the above procedure example, the keyword determination process is executed when the return value of the URL determination process is "permitted". Instead, the keyword determination process may be executed first, and if the return value is "permitted", the URL determination process may be executed. In addition to the URL determination process and the keyword determination process, further determination processes may be combined and executed.

[0045] FIG. 6 is a flowchart showing an example of the procedure of the keyword determination process. The keyword determination process is executed by the CPU 11 of the filtering server 10 in the process of the filtering determination process (step S53 in FIG. 5). Hereinafter, it will be described along with an example of the procedure.

[0046] Step S61: The CPU 11 divides the body data of the request into predetermined sizes to generate N pieces of divided data. Here, the CPU 11 further divides the data sent after being divided from the browser 40 into predetermined sizes that are easy to process. As a result, the entire body data of the request is divided into N pieces. Note that when the size of the body data is less than the predetermined size, the number of divided data is one (N = 1).

[0047] Steps S62, S63: The CPU 11 sets 1 to a position counter c indicating the current position in the entire body data (step S62), hashes the first divided data, i.e., divided data 1, to obtain hashed divided data 1 (step S63).

[0048] The hashed and segmented data is an array H with the number of elements (number of characters + 1) corresponding to the number of characters constituting the original segmented data. Each element of the array H is set with a value calculated based on an integer value of 0 or more (hash value) obtained by hashing each character constituting the segmented data. Specifically, a fixed value "1" is set in H[0], and in H[k], a value based on the value obtained by accumulating while securing a predetermined number of bits (for example, 8 bits) for the hash values of each character from the first character to the k-th character of the segmented data. More specifically, the remainder obtained by dividing the sum of the value obtained by multiplying H[k - 1] by a predetermined value (for example, 256 (= the maximum value of 8 bits)) and the hash value of the k-th character of the segmented data by a sufficiently large divisor is set. By arranging the segmented data in such a manner, when performing the matching described later, the hash value of an arbitrary-length substring at an arbitrary position included in the segmented data can be easily calculated by a simple operation. Note that the hashing may be performed by an algorithm (hash function) developed independently or by using a generally known hash function.

[0049] Subsequently, for the hashed and segmented data 1, matching with individual keywords for restricting data transmission is performed. FIG. 7 shows, as an example of a keyword list, a partial extraction of the keyword list registered by a user named School A.

[0050] The database 12 stores a keyword list consisting of one or more keywords registered in advance by the user. In the keyword list of School A, various keywords (for example, words related to drugs and crimes, violent words, words that slander or discriminate against others, etc.) that School A has determined should block posts by students and teachers to SNS, bulletin boards, etc. are registered.

[0051] In order to perform matching efficiently, the database 12 further stores a list of hashed keywords consisting of hashed keywords represented by integers greater than or equal to 0 obtained by hashing individual keywords. In FIG. 7, although the specific numerical values of individual hashed keywords are not shown and are abbreviated as "......", different keywords result in different hashed keywords.

[0052] Also, as shown in FIG. 7, for Japanese keywords, hashed keywords hashed with three types of character codes (Shift_JIS, EUC, UTF-8) are stored respectively. Thereby, when performing matching, it is not necessary to hash individual keywords each time. Instead, the hashed keyword corresponding to the character code of the request body data (charset attribute of the request header) can be obtained from the database 12 and directly used for matching.

[0053] [Refer to FIG. 6] Step S64: The CPU 11 selects one keyword corresponding to the character code of the body data from the user's list of hashed keywords, and performs matching between this hashed keyword and the hashed split data obtained in the previous step. In the matching, based on the premise that if the character strings are equal, their hash values are also equal, using the rolling hash method, a search is made for a character string having the same hash value as the selected hashed keyword, targeting the hashed data obtained in the previous step (step S63 or step S72). If a location that matches the hashed keyword is found in the matching, it means that the keyword corresponding to the hashed keyword is included in the data before being hashed. Specific examples of the matching will be further described later with reference to another drawing.

[0054] Steps S65 to S67: If the selected hashed keyword matches (step S65: Yes), the CPU 11 returns "Restricted" (step S66), ends the keyword determination process, and returns to the calling filtering determination process.

[0055] On the other hand, if the selected hashed keyword does not match (step S65: No), the CPU 11 checks whether it has performed matching with all the keywords included in the hashed keyword list (in the case of Japanese keywords, the hashed keyword corresponding to the character code of the body data) (step S67). If there are still keywords for which matching has not been performed (step S67: No), the CPU 11 returns to step S64, selects a keyword for which matching has not been performed yet, performs the matching, and executes the subsequent steps again.

[0056] In contrast, when matching has been performed with all the keywords included in the hashed keyword list (step S67: Yes), the CPU 11 proceeds to step S68. Steps S68, S69: The CPU 11 checks whether the value of the position counter c is less than the number N of divided data (c < N). If c < N, that is, if there is still divided data for which matching has not been performed yet (step S68: Yes), the CPU 11 proceeds to step S70. On the other hand, if c = N, that is, if matching for all the divided data has been completed (step S68: No), the CPU 11 returns "Permitted" (step S69), ends the keyword determination process, and returns to the calling filtering determination process.

[0057] Steps S70, S71: The CPU 11 adds 1 to the position counter c (step S70), hashes the divided data (c), and sets this as the hashed divided data (c) (step S71). For example, if c = 2 in step S70, then in step S71, the second divided data, divided data 2, is hashed to generate the hashed divided data 2.

[0058] Step S72: The CPU 11 checks the number of characters of the longest keyword in the user's keyword list (hereinafter referred to as the "longest number of characters"), and obtains the hash data of the longest number of characters from the end of the hashed split data (c-1), that is, the previous hashed split data, and concatenates it to the beginning of the hashed split data (c). For example, when c = 2, the hash data of the longest number of characters at the end of the hashed split data 1 is concatenated to the beginning of the hashed split data 2. Then, the CPU 11 repeats the procedures after step S64 for the hashed data thus obtained.

[0059] As described above, in the keyword determination process, the position to be processed in the entire request body data is gradually moved from the front to the back, and matching is performed with the hashed keywords corresponding to the individual keywords registered in advance by the user. When a match is found with a hashed keyword, "restriction" is returned at that time, and "permission" is returned only when matching is performed up to the last position and no match is found with any hashed keyword.

[0060] Note that the seed value of the hash function is changed every time the filtering server 10 is started. Correspondingly, the hashed keyword list stored in the database 12 is also updated every time the filtering server 10 is started.

[0061] FIG. 8 is a diagram for explaining the hashed data to be matched at each stage of the keyword determination process.

[0062] (A) in FIG. 8 shows an example of the request body data from the browser 40. In the illustrated example, the body data is divided into three parts and split data 1 to 3 are generated. The "~~~" in the figure simply shows the arrangement of the characters constituting the split data.

[0063] In (B) of FIG. 8, in the keyword determination process for the body data composed of the divided data 1 to 3 shown in (A) of FIG. 8, the hashed data to be matched at each stage is shown. The "......" in the figure simply indicates the hashed data.

[0064] First, when the position counter c = 1, the hashed divided data 1 is the object of matching. Next, when c = 2, the hashed data obtained by concatenating the hashed data for the longest number of characters at the end of the hashed divided data 1 and the hashed divided data 2 is the object of matching. Then, when c = 3, the hashed data obtained by concatenating the character string for the longest number of characters at the end of the hashed divided data 2 and the hashed divided data 3 is the object of matching.

[0065] For example, assume that the longest keyword in the keyword list of school A shown in FIG. 7 is 10 characters. In this case, when c = 2, the hashed data obtained by concatenating the 10-character hashed data forming the end part of the hashed divided data 1 to the head of the hashed divided data 2 is the object of matching, and when c = 3, the hashed data obtained by concatenating the 10-character hashed data forming the end part of the hashed divided data 2 to the head of the hashed divided data 3 is the object of matching.

[0066] FIG. 9 shows an example when the keyword is included in the divided data 1. In FIG. 9, the data location corresponding to the keyword in the hashed divided data is shaded (the same applies in FIG. 10).

[0067] For example, assume that the character code of the request body data is "Shift_JIS", and among the keyword lists of school A shown in FIG. 7, the keyword "papa activity" is included in the divided data 1. In the process of the keyword determination process (FIG. 6), when the position counter c = 1, matching is performed between the hashed divided data 1 and the hashed keyword of "Shift_JIS" corresponding to the keyword "papa activity" (step S64 in FIG. 6).

[0068] Specifically, from the array H of the hashed and segmented data 1, two elements located at positions separated by the number of characters of the keyword "papa activity", that is, two elements separated by three positions, are used to calculate a hash value for three characters, and it is checked whether it matches the hashed keyword. First, a hash value for the first three characters is calculated using the values of H[0] and H[3], and it is checked whether it matches the hashed keyword. If it does not match, a hash value for three characters starting from the second character is calculated using the values of H[1] and H[4], and it is checked whether it matches the hashed keyword. If it does not match, a hash value for three characters starting from the third character is calculated using the values of H[2] and H[5], and it is checked whether it matches the hashed keyword, and so on. In this way, the matching for three characters is sequentially performed by shifting the position one by one backward.

[0069] And in the process of matching against the hashed and segmented data 1, since it matches the hashed keyword (step S65: Yes in FIG. 6), "restriction" is returned as the determination result (step S66 in FIG. 6). As a result, access from the browser 40 to the web server WS is restricted.

[0070] FIG. 10 is a diagram showing an example in the case where the keyword straddles the segmented data 1 to 2. For example, assume that the character code of the request body data is "EUC", and among the keyword lists of School A shown in FIG. 7, the keyword "cocaine" straddles the segmented data 1 to 2. In the process of the keyword determination process (FIG. 6), matching is performed between the hashed and segmented data and the hashed keyword of "EUC" corresponding to the keyword "cocaine" (step S64 in FIG. 6).

[0071] When the position counter c = 1, two elements located at a distance equal to the number of characters of the keyword "cocaine" from the array H of the hashed split data 1, that is, two elements four positions apart, are used to calculate a hash value for four characters, and it is checked whether it matches the hashed keyword. First, a hash value for the first four characters is calculated using the values of H[0] and H[4], and it is checked whether it matches the hashed keyword. If it does not match, a hash value for four characters starting from the second character is calculated using the values of H[1] and H[5], and it is checked whether it matches the hashed keyword. If it does not match, a hash value for four characters starting from the third character is calculated using the values of H[2] and H[6], and it is checked whether it matches the hashed keyword, and so on. In this way, the position is shifted one by one backward, and four-character matching is performed successively. As shown in FIG. 10, the hashed split data 1 contains only a part of the hashed keyword "EUC" corresponding to the keyword "cocaine" (the hash data corresponding to the first "co"). Therefore, when c = 1, it does not match the hashed keyword.

[0072] When c = 2, the hashed data obtained by concatenating the end part of the hashed split data 1 and the hashed split data 2 is targeted, and a hash value for four characters is calculated using two elements four positions apart from its array H, and four-character matching is performed successively in the manner described above. As shown in FIG. 10, the end part of the hashed split data 1 contains a part of the hashed keyword (the hash data corresponding to the first "co"), and the start part of the subsequent hashed split data 2 contains the remaining part (the hash data corresponding to "caine").

[0073] Therefore, since it matches the hashed keyword in the process of matching performed when c = 2 (step S65: Yes in FIG. 6), "Restricted" is returned as the determination result (step S66 in FIG. 6). As a result, access from the browser 40 to the web server WS will be restricted.

[0074] [Advantages of the Present Invention] As described above, according to the web filtering system of the above-described embodiment, the following effects can be obtained.

[0075] (1) Since it is determined whether or not to restrict access based on whether or not a pre-registered keyword is included in the request body data, it is possible to prevent inappropriate information from being transmitted from a terminal in the user environment.

[0076] (2) Since the keyword list used for the determination is registered by the user himself / herself, it is possible to make a determination of "restriction" or "permission" in accordance with the user's intention, and it is possible to surely restrict the transmission of information including keywords that an individual user considers inappropriate (keywords that the user does not want to be transmitted from the user environment).

[0077] (3) Since the matching between the hashed divided data generated by dividing the request body data and the hashed keyword obtained by hashing the pre-registered keyword is performed using the rolling hash method, the matching can be executed at high speed, and the processing speed can be increased as compared with the case of performing the matching with a non-hashed character string.

[0078] (4) In the matching between the hashed divided data 1 to N and each hashed keyword, when the second and subsequent hashed divided data are the targets of the matching, at the head of the hashed divided data (for example, the hashed divided data 2), hash data corresponding to the number of characters of the longest keyword obtained from the end of the previous hashed divided data (for example, the hashed divided data 1) is concatenated, and the end part of the previous hashed divided data is also included in the target of the matching. Therefore, even when a keyword exists across two divided data, it is possible to surely restrict the transmission of that information.

[0079] Since the hashed keyword list corresponding to the keyword list is pre-stored in the database 12, it is not necessary to hash each keyword every time matching is performed, so the matching efficiency can be improved compared to the case where the keyword is hashed each time.

[0080] The present invention can be variously modified and implemented without being restricted by the above-described embodiments.

[0081] In the above-described embodiment, the database 12 is provided in the filtering server 10, but the installation location of the database is not limited to the filtering server 10. For example, it may be provided on another server to which the filtering server 10 can be connected, or a plurality of databases provided on different servers may be selectively used according to the application. Alternatively, the keyword list or the hashed keyword list may be stored in a configuration file or the like on the filtering server 10 instead of the database 12.

[0082] In the above-described embodiment, groups such as schools and companies are assumed as users of the filtering server 10, but the usage form is not limited to this. For example, when an individual accesses the web server WS from a terminal in the individual's home environment as a user of the filtering server 10, communication control may be performed based on a keyword list set by the individual himself / herself. Thereby, it is possible to prevent the individual's family members from transmitting inappropriate information.

[0083] In addition, the configurations, numerical values, etc. mentioned in the process of describing the web filtering system 1 of the embodiment are merely examples, and it goes without saying that they can be appropriately modified when implementing the present invention.

Description of Reference Numerals

[0084] 1 Web filtering system 10 Filtering server 11 CPU (Determination unit) 12 Database (Memory Unit) 20 Proxy Server 30 Redirect Server 40 Browser 50 Application 60 Browser Extension

Claims

1. a storage unit storing one or more keywords registered in advance by a user; a determination unit that, when receiving an inquiry regarding restriction of a request from a browser to a web server, makes a determination based on whether the keyword is included in the body data of the request, and returns the determination result to the inquiry source A web filtering system comprising:

2. In the web filtering system according to claim 1, the determination unit confirms whether the keyword is included in the body data by performing matching between the hashed data obtained by hashing the body data and the hashed keyword obtained by hashing the keyword. A web filtering system characterized by this.

3. In the web filtering system according to claim 2, the storage unit is characterized in that the keyword and the corresponding hashed keyword are stored in advance. A web filtering system characterized by this.

4. In the web filtering system according to claim 2 or 3, the determination unit divides the body data into predetermined sizes, performs matching between the hashed divided data obtained by hashing the divided data and the hashed keyword, and when the second and subsequent hashed divided data are the targets of matching, at the head of the hashed divided data, hash data corresponding to the number of characters of the longest keyword registered by the user, which forms the end part of the previous hashed divided data, is concatenated. A web filtering system characterized by this.

Citation Information

Patent Citations

  • Method for rapidly filtering key word

    CN103186669A

  • Method and apparatus for handling http request

    CN107040606A

  • Contributed data evaluation system

    JP2006268303A

  • Electronic bulletin board system

    JP2007257346A

  • Access control device, screen generation device, program, access control method, and screen generation method

    JP2015141609A