Uniform resource locator (URL) detection system for network content access control

A processor-based system with a Machine Learning Model addresses false positives in network content access control by analyzing URLs and user identity, ensuring efficient filtering of inappropriate content.

WO2025253269A1PCT designated stage Publication Date: 2025-12-11NETSWEEPER BARBADOS
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/055665
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-03
Filing Date
2025-06-02
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Network content access control services based on keyword detection suffer from excessive false positive determinations, leading to inefficient content access and wastage of processing resources and bandwidth due to unnecessary blocking of valid content.

Method used

A processor-based system that utilizes a Machine Learning Model (MLM) to analyze URLs for category determination, combining initial heuristic keyword matching with TF-IDF vectorization and category-specific weights to reduce false positives, and considers user identity information for approval decisions.

Benefits of technology

Effectively detects undesirable content without generating false positives, enhancing content access efficiency and reducing resource waste by accurately filtering out inappropriate material.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025055665_11122025_PF_FP_ABST
    Figure IB2025055665_11122025_PF_FP_ABST
Patent Text Reader

Abstract

A content access control service comprising a processor configured to receive a content request that includes identity information and a requested Uniform Resource Locator (URL). The processor determines whether the request pertains to content related to a category by identifying as a candidate word a consecutive string of characters within the requested URL matching a predetermined word in a list of predetermined words indicative of the category, and, when the candidate word is identified, providing the requested URL to a Machine Learning Model (MLM) to determine whether the requested URL contains a match word indicative of the category to calculate a confidence score. The processor determines whether to approve the content request based on the category determination and the identity information and directs a client agent of the computing device to the requested content or redirects the client agent to pre-approved content when the content request is denied.
Need to check novelty before this filing date? Find Prior Art

Description

UNIFORM RESOURCE LOCATOR (URL) DETECTION SYSTEM FOR NETWORK CONTENT ACCESS CONTROL

[0001] This application claims priority to U.S. provisional patent application serial No.63 / 655,485, filed on June 3, 2024, which is hereby incorporated by reference.FIELD

[0002] This disclosure relates to computers and, more specifically, to network content access control services.BACKGROUND

[0003] Network content access control services can be used to prevent undesirable material from being retrieved by a computer. Such material can include adult-oriented material unsuitable for viewing by a child that has access to the computer, and which further may additionally include malicious code that detrimentally modifies the behavior of the retrieving computer.

[0004] A network content access control service based on detection of keywords in a Uniform Resource Locator (URL) may suffer from several drawbacks, for example, excessive false positive determinations, which may prevent the computer from retrieving significant amounts of content that should not have been blocked by the network content access control service.SUMMARY

[0005] An aspect of the specification provides a content access control service comprising a processor configured to receive a content request from a computing device. The content request includes identity information and a requested Uniform Resource Locator (URL). The processor determines whether the content request pertains to content related to a category by identifying as a candidate word a consecutive string of characters within the requested URL matching a predetermined word in a list of predetermined words indicative of the category and when the candidateword is identified, providing the requested URL to a Machine Learning Model (MLM) configured to determine whether the requested URL contains a match word indicative of the category to calculate a confidence score representative of whether the URL corresponds to the category. The processor further determines whether to approve the content request based on the category determination and the identity information; and directs a client agent of the computing device to the requested content when the content request is approved and redirects the client agent to pre-approved content when the content request is denied.

[0006] Another aspect of the specification provides a method of determining by a processor whether to approve a content request. The method comprises receiving the content request. The content request includes identity information and a requested URL. The method further comprises determining whether the content request pertains to content related to a category by identifying as a candidate word a consecutive string of characters within the requested URL matching a predetermined word in a list of predetermined words indicative of the category and when the candidate word is identified, providing the requested URL to a Machine Learning Model (MLM) configured to determine whether the requested URL contains a match word indicative of the category to calculate a confidence score representative of whether the URL corresponds to the category. The method further comprises determining whether to approve the content request based on the category determination and the identity information and directing a client agent of a computing device to the requested content when the content request is approved and redirecting the client agent to pre-approved content when the content request is denied.BRIEF DESCRIPTIONS OF THE DRAWINGS

[0007] Embodiments are described with reference to the following figures.

[0008] FIG. 1 depicts a schematic diagram of an example content exchange system.

[0009] FIG. 2 depicts a schematic diagram of a computing device of the example content exchange system, of FIG. 1 .

[0010] FIG. 3 depicts a schematic diagram of a server of the example content exchange system of FIG. 1.

[0011] FIG. 4 depicts a flowchart of an example method by which a processor may determine whether a content request is directed at content that falls within a predetermined category.

[0012] FIG. 5 depicts a flowchart of an example sub-method by which a processor may determine whether a content request is directed at content that falls within a predetermined category.DETAILED DESCRIPTION

[0013] Network content access control services based on detection of keywords can suffer from several drawbacks such as excessive false positive determinations, which can prevent a computer from retrieving significant amounts of content that should not have been blocked by the network content access control service, resulting in a network system with an overly limited content access functionality and poor content transmission efficiency, where a request for valid content from a client device to a content providing device (e.g. a content server) hosting the valid content may not result in the valid content being delivered to the client device through the network, which can translate into a significant waste of processing resources and bandwidth usage across the network system due to excessive blocked requests and / or intercepted transmissions of content. Example content exchange systems configured to perform content access control services that do not simply rely on keyword detection are described below, with references to the FIGS. These example content exchange systems may be useful in effectively detecting network requests to access undesirable material, for example, pornographic or violent content, by a computer without generating false positive determinations, thereby enabling a better content access experience and reducing the waste of processing resources and bandwidth across the network.

[0014] FIG. 1 depicts a schematic diagram of an example content exchange system 100. The system 100 includes computing devices (also referred to as client devices) 104-1 , 104-2, ... 104-N (collectively referred to as the computing devices 104 andspecifically referred to as a computing device 104, this nomenclature is used elsewhere in this specification) such as general-purpose computers, for example, desktop computers, laptop computers, tablets, smartphones, etc., or dedicated-purpose computers with web browsing capabilities, for example, e-book readers, smart TVs, gaming computers, etc.

[0015] The computing devices 104 are communicatively connected to a network device 108 such as a modem that provides access to a network 112 such as the internet. The network device 108 can represent a subsystem that includes one or more devices such as routers, switches, hubs, network cables, wireless access points, fiberoptic lines, and the like.

[0016] The network device 108 is further communicatively connected to the network 112 via a sub-network 1 14, which may include servers such as a gateway server, an account server, etc. For explanatory purposes, the sub-network 1 14 will be described as containing a gateway server 116 and an account server 120; however, it should be understood that more or fewer servers can be included, and that different processes and functions can be allotted to different servers in a myriad of ways. Functionality described herein with respect to several servers can be performed by fewer servers or even a single server, with any associated communications between physical servers described herein being configured instead as communications between processes. The gateway server 116 and the account server 120 may be provided and administered by an organization, such as an Internet service provider, a school system, a government, a company, or the like, that provides access to the network 112 for the computing devices 104. The gateway server 116 handles requests and responses to the computing devices 104 and maps the computing devices 104 to shared IP addresses, if required. The account server 120 is communicatively connected to the gateway server 116 and may be configured to store account information of users of the computing devices 104. Such information may include, for example, network access credentials, user identity information, personal information (e.g., name, address, etc.), billing information, etc.

[0017] The system 100 further includes content servers 124 communicatively connected to the network 112. The content servers 124 can include social mediawebsites, streaming media services such as video services, application servers, etc. As such, the content servers may store and make media available to the computing devices 104, said media may include files available for streaming and / or download, such as video, audio, text, databases, programs, websites, etc. The content servers 124 operate at one or more host names (e.g., www.example.com). Furthermore, the content servers 124 may also include various computing devices (such as the computing devices 104) that supply media files or portions thereof via peer-to-peer or decentralized file distribution techniques (e.g., torrents). In such cases, the content servers 124 may include one or more servers that have links to initiate downloading of peer-to-peer or decentralized files and tracker computers that assist in coordinating downloads.

[0018] The system 100 further includes a content access control service 126. The content access control service 126 can include various components such as , a policy server, an administration server, a log server, and a category server.

[0019] The system 100 further includes a client monitor application 127 that may be installed and run at computing devices 104 (for example, at computing device 104-N as shown in FIG. 1 ) or at another device within the system 100, for example, within subnetwork 114 (for example, at gateway server 116 as shown in FIG. 1 ). The client monitor application 127 may be considered part of the content access control service 126 or a separate element from it.

[0020] For explanatory purposes, the content access control service 126 will be described in terms of an administration server 128 and a policy server 132; however, it should be understood that more or fewer servers can form the content access control service 126, and that different processes and functions can be allotted to different servers in a myriad of ways. Functionality described herein with respect to several servers can be performed by fewer servers or even a single server, with any associated communications between physical servers described herein being configured instead as communications between processes. For example, the administration server 128 and the policy server 132 can be implemented on different servers, the same server, as a process running on a server of the sub-network 114 or on a computing device 104. In this embodiment, the content access control service 126 may be out-of-band or in-linewith network requests and responses between the computing devices 104 and the network 112.

[0021] The content access control service 126 is configured to apply resource access policy to restrict access to content, such that for each remote content request made by the computing devices 104 a policy decision is requested from the content access control service 126. The content access control service 126 can determine a whether the requested content may fall within (i.e. may be related to) a predetermined category, such as, for example, an Interactive Advertising Bureau (IAB) category such as careers, law, government and politics, shopping, gambling, games, etc., or a category outside the scope of the IAB such as pornographic content, violent content, etc., based on the result of the determination and at least one of a location of the content, such as a Uniform Resource Locator (URL), hostname, domain, etc., and user identity information received with the request, the content access control service 126 may approve or deny the request. If the request is denied, the computing device 104 that generated the request may be redirected to a redirection page. The redirection page can be a predefined page hosted by the content access control service 126, for example, in the administration server 128. Alternatively, the redirection page may be a pre-approved page hosted at a content server 124, a default page locally stored at the computing device 104, etc. The computing device 104 is redirected to the redirection page (predefined page) as a redirect from the requested host / domain. The response speed of the content access control service 126 is configured to be faster than the actual response of the requested host / domain.

[0022] The identity information may be unique to the user and comprise information such as name or identification number, or may be a broader group-based characteristic, and comprise information such as the user's age group, sex, organizational role (e.g., minor student, student at age of majority, teacher, parent, etc.), country, or legal jurisdiction (e.g., state, province, territory, city, special economic zone, etc.). The identity information may further include any combination of group-based characteristics and information unique to the user. The identity information can be provided by the user by way of, for example, a login credential that is stored at the computing devices 104 (e.g., in a cookie). The identity information can also be determined by policies set in thecontent access control service 126 based on the content of requests. For example, a request may include an IP address that can be mapped to a country or legal jurisdiction.

[0023] The policy server 132 of the content access control service 126 can be configured to determine whether a restrictive content access policy applies to the requested content. Restrictive policy may be based on the identity information of the user and the requested content, may be based on the requested content without regard to the identity information of the user (e.g., all users are subject to the same policy), or may be group-based. The policy server 132 may store a policy database.

[0024] FIG. 2 depicts a schematic diagram of an example computing device 104 of the system 100. The computing device 104 includes a processor 200 that may be implemented as a plurality of processors or one or more multi-core processors. The processor 108 may include one or more of a Central Processing Unit (CPU), a microcontroller, a microprocessor, a processing core, a Field-Programmable Gate Array (FPGA) or the like, and combinations thereof. The processor 200 may additionally include a built-in Graphics Processing Unit (GPU), however, in other embodiments a GPU may be provided separately from the processor 200.

[0025] The computing device 104 further includes Input / Output (I / O) devices 202 to which the processor 200 is communicatively connected to, and through which the user may interact with the computing device 104. The I / O devices 202 may include input devices such as, for example, a keyboard, a mouse, a sensor such as an accelerometer, etc., and may include output devices such as, for example, a display, a speaker, a set of headphones, a haptic device, etc.

[0026] The computing device 104 further includes a network interface 204 through which the processor 200 is communicatively connected to the network 112.

[0027] The computing device 104 further includes one or more memory units, including a volatile memory 208 and a non-volatile memory 212, communicatively coupled to the processor 200. The volatile memory 208 is based on any random-access memory (RAM) technology. For example, the volatile memory 208 can be based on a Double Data Rate (DDR) Synchronous Dynamic Random-Access Memory (SDRAM). Other types of volatile memory 208 are contemplated.

[0028] The non-volatile memory 212 can be based on any persistent memory technology, such as an Erasable Electronic Programmable Read Only Memory (“EEPROM”), flash memory, solid-state hard disk (SSD), other type of hard-disk, or combinations of them. The non-volatile memory 212 may also be described as a non- transitory computer readable media. Also, more than one type of non-volatile memory 212 may be provided.

[0029] Programming instructions in the form of applications such as the client monitor application 127 and applications 216 are typically maintained, persistently, in non-volatile memory 212 and used by the processor 200 which reads from and writes to the volatile memory 208 during the execution of the applications 127, 216. Various methods discussed herein can be coded as one or more applications 127, 216, such as, for example, client agents such as web browsers that retrieve content from the network 112. One or more tables or databases 220 are maintained in the non-volatile memory 212 for use by the applications 127, 216.

[0030] FIG. 3 depicts a schematic diagram of an example server 112, 116, 124, 128, or 132 of the system 100.

[0031] The server 112, 116, 124, 128, or 132 includes a processor 300 that may be implemented as a plurality of processors or one or more multi-core processors. The processor 300 may include one or more of a Central Processing Unit (CPU), a microcontroller, a microprocessor, a processing core, a Field-Programmable Gate Array (FPGA) or the like, and combinations thereof. The processor 300 may additionally include a built-in Graphics Processing Unit (GPU), however, in other embodiments a GPU may be provided separately from the processor 300.

[0032] The server 112, 116, 124, 128, or 132 further includes a network interface 304 through which the processor 300 is communicatively connected to the network 1 12.

[0033] The server 112, 116, 124, 128, or 132 further includes one or more memory units, including a volatile memory 308 and a non-volatile memory 312, communicatively coupled to the processor 300. The volatile memory 308 is based on any random-access memory (RAM) technology. For example, the volatile memory 308 can be based on aDouble Data Rate (DDR) Synchronous Dynamic Random-Access Memory (SDRAM). Other types of volatile memory 308 are contemplated.

[0034] The non-volatile memory 312 can be based on any persistent memory technology, such as an Erasable Electronic Programmable Read Only Memory (“EEPROM”), flash memory, solid-state hard disk (SSD), other type of hard-disk, or combinations of them. The non-volatile memory 312 may also be described as a non- transitory computer readable media. Also, more than one type of non-volatile memory 312 may be provided.

[0035] Programming instructions in the form of applications 316 are typically maintained, persistently, in non-volatile memory 312 and used by the processor 300 which reads from and writes to the volatile memory 308 during the execution of the applications 316. Various methods discussed herein can be coded as one or more applications 316, such as, for example, a policy decision application. One or more tables or databases 320, such as, for example, a policy database, are maintained in the non-volatile memory 312 for use by the applications 316.

[0036] A processor 300 of a device of the content access control service 126 may execute an application 316 to determine whether a content request by a computing device 104 may be directed at content that falls within a predetermined category, such as, for example, pornographic content. FIG. 4 illustrates an example method 400 by which the processor 300 may perform said determination.

[0037] Starting at block 405, the processor 300 receives a content request from a client agent, such as an internet browser, of a computing device 104. The content request may be directed to the processor 300 by the client monitor application 127 running at the computing device 104 or at a different device within the system 100, for example, at the gateway server 116. The content request comprises a requested URL and identity information. As discussed above, the identity information may be unique to the user and comprise information such as name or identification number, may be a broader group-based characteristic and comprise information such as the user's age group, sex, organizational role, country or legal jurisdiction, or may include any combination of group-based characteristics and / or information unique to the user. Theidentity information can be provided by the user by way of, for example, a login credential that is stored at the computing device 104 (e.g., in a cookie) and gets sent by the client agent along with the requested URL. Alternatively, the identity information may be added to the content request by client monitor application 127, for example, running at the computing device 104 or at a different device within the system 100, such as at the gateway server 116.

[0038] After receiving the content request, the processor 300 may proceed to block 410, where the processor 300 may categorize the identity information from the content request. The identity information may be categorized by the processor 300 by matching the identity information or additional information within the content request to a user database 320 that stores, for example, unique user IDs by age-groups, sex, organizational role, etc. Alternatively, or additionally, the identity information may also be categorized by, for example, mapping the IP address of the computing device 104 making the content request to a country or legal jurisdiction from an IP address database 320. As a further alternative, the identity information received by the processor 300 may include a category indication provided by, for example, the client agent or the gateway server 116.

[0039] After categorizing the identity information, the processor 300 proceeds to block 415, where the processor 300 determines whether the requested content falls within (is related to) a determined category, for example, a pornographic content category. FIG. 5 depicts an example sub-method 500 that can be performed as block 415.

[0040] Starting at block 505, as an initial heuristic step, the processor 300 may analyze the characters of requested URL to identify all consecutive strings of characters within the requested URL matching a list of predetermined words that are indicative of the determined category, such as, for example “porn”, “porno”, “pornographic”, “sex”, “nude”, etc. indicative of a pornographic category. The list of predetermined words may be contained in a database 320. If more than one category is to be determined, the processor 300 may analyze the characters of the requested URL against more than one list of predetermined words. Each set of consecutive strings of characters within therequester URL matching a word in the at least one list may be identified as a candidate (or suspicious) word (or keyword) present in the requested URL and temporarily stored in the volatile memory 308.

[0041] After analyzing the characters of the requested URL at block 505, the processor 300 proceeds to block 510, where the processor 300 may determine (identify) whether at least one candidate word is present in the requested URL, for example, by determining whether at least one candidate word was temporarily stored in the volatile memory 308.

[0042] If no candidate words are identified in the requested URL at block 510, the processor 300 proceeds to block 515 where it is determined that the requested content does not fall within the determined category.

[0043] If at least one candidate word is identified in the requested URL at block 510, the processor 300 proceeds to block 520, where the processor 300 provides the requested (suspicious) URL to a Machine Learning Model (MLM) configured to determine whether the requested URL falls within the determined category. The MLM can include, for example, a Recurrent Neural Network (RNN), a Convolutional Neural Network (CNN), a Conditional Random Field (CRF), a Hidden Markov Model (HMM), a conditional Generative Adversarial Network (cGAN), or a Transformer Model, such as, for example, a Bidirectional Encoder Representations from Transformers (BERT) model. The MLM can be adapted to handle character-level inputs by using a BERT tokenizer to tokenize the characters into sub-words or characters themselves to break down unknown words into known sub-words whilst using special characters to indicate their original order.

[0044] The MLM is configured to determine whether the URL contains a match word that actually belongs to, and is indicative of the determined category, whether there is(are) any candidate word(s) that is(are) actually one of one or more sub-words within the requested URL that are unrelated to the category and whether the requested URL contains any match word(s) indicative of the URL being unrelated to the category. Based on these determinations, the MLM model may calculate and output a numerical value representative of a level of confidence of the URL corresponding to the category(also referred to as a confidence score), so that when the numerical value is equal to or higher than a threshold value (also referred to as a threshold score), it may be determined that the URL corresponds to the category. For example, if the candidate word (also referred to as token) “sex” was identified within the requested URL “essex.com” at block 505, the MLM would identify that the URL contains no actual match words indicative of a pornographic category by identifying that the token “sex” is a sub-word of the word “Essex”. As a further example, if the candidate word “porno” was identified within the requested URL “vapornoodles.com” at block 505, the MLM would identify that the URL contains no actual match words indicative of a pornographic category by identifying that the token “porno” is actually comprised of sub-words “por” and “no”, belonging to non-match words “vapor” and “noodles”, respectively. Furthermore, the MLM can further determine whether the URL contains any word(s) that is / are indicative of the URL not falling within the determined category; for example, if the candidate word “sex” was identified within the requested URL “sexaddictionhelpline.com” at block 505, the MLM would identify that the URL does contain the match word “sex” indicative of a pornographic category, and the MLM would also identify that the URL contains the match word “helpline” that is unrelated to the pornographic category and that is indicative of the URL not being actually within the pornographic category.

[0045] The determination by the MLM can include performing a Term Frequency- Inverse Document Frequency (TF-IDF) vectorization, which converts textual data into numerical vectors to represent the importance of each word within a requested URL relative to a collection of webpages that are representative of the determined category, by calculating a TF-IDF score:TF-IDF = TF * IDF,TF = NTerm I TotNTerm, andIDF = log_e (TotNPages / NPagesTerm) where NTerm is the number of times a term or known word appears in the requested URL, TotNTerm is the total number of terms or known words in the requested URL, TotNPages is the total number of pages in the category and NPagesTerm is the numberof pages in the category containing the term or known word. By integrating BERT tokenization to split the URL string into tokens and processing the tokens through TF- IDF vectorization, word concatenation (e.g. “workshopornament”, “glassexchange”) can be handled by the MLM.

[0046] The calculated TF-IDF vector for the requested URL can then be multiplied by a series of category-specific weights learned by the MLM during the training process to calculate the confidence score so that, if the result of the vector multiplication is above the threshold score, it may be determined by the processor 300 that the requested URL falls within the determined category, for example, with a sufficient degree of confidence. In order to increase the effectiveness of the MLM, the MLM may be trained and deployed exclusively on URLs containing match words related to and indicative of the category (e.g. “porn”, “sex”, etc.) and on URLs containing both match words related and indicative of the category and match words that are indicative of the URL not falling within the category (e.g. “information”, “help”, “education” etc.).

[0047] After block 520, the processor 300 proceeds to block 525 where the processor 300 determines whether the requested URL falls within the determined category based on the output from the MLM (i.e. whether the MLM is sufficiently confident that the suspicious URL is related to the category).

[0048] If a negative determination is made by the processor 300 at block 525, the processor 300 continues to block 515 where the processor 300 determines that the requested content does not fall within the determined category.

[0049] If a positive determination is made by the processor 300 at block 525, the processor 300 continues to block 530 where the processor determines that the requested content falls within the determined category.

[0050] Through identifying the top n words of the weights for a category that the model most strongly associates with the category, the list of match words related to and indicative of the category used at block 505 may be generated. Additionally, to identify the words which led the model to a categorization, the TF-IDF vector for a given requested URL may be multiplied by the learned weights associated with a category by using the Hadamard product and the resultant vector may be analyzed with reference tothe vectorizer’s vocabulary. This functionality can optionally be performed by the processor 300 and recorded in a log for model effectiveness diagnosis.

[0051] Going back to FIG. 4, if the determination at block 415 is negative, the processor 300 proceeds to block 425 where the processor 300 may direct the client agent of the computing device 104 to the requested content, for example, to a page hosted by a content server 124.

[0052] If the determination at block 415 is affirmative, the processor 300 proceeds to block 420 where the processor 300 determines whether the user and / or the computing device 104 is approved to access the requested content. The processor 300 may perform the determination based on the category within which the requested URL was identified, on the identity information and on predetermined policies applicable to specific users and / or groups of users and / or specific computing devices 104 and / or groups of computing devices 104 that may be stored in a database 320.

[0053] If a positive determination is made by the processor 300 at block 420, the processor 300 proceeds to block 425 where the processor 300 may direct the client agent to the requested content.

[0054] If a negative determination is made by the processor 300 at block 420, the processor 300 proceeds to block 430 where the processor 300 redirects the client agent of the computing device 104 away from the requested content, for example, to a predetermined (predefined) page as a redirect from the requested host / domain, for example, hosted by administration server 128. The response speed of the content access control service 126 is configured to be faster than the actual response of the requested host / domain.

[0055] The example method 400 and the example sub-method 500 or variations thereof can be useful for a processor executing it for effectively detecting network requests to access undesirable material, for example, pornographic or violent content without generating false positive determinations, thereby enabling a better content access experience and reducing waste of processing resources and bandwidth across a network. The initial heuristic step 505 can be useful in constraining the semantic space which needs to be covered by the MLM, which can make the determination at block 420or a variation thereof more computationally efficient. Additionally, the MLM at block 520 can account for false word matches which may have occurred due to sub-words, wordconcatenation and so on, making the example determination sub-method 500 or variations thereof more functional by reducing the number of false positive results. Furthermore, the MLM at block 520 can also account for URLs which contain one or more match words that are indicative or suggestive of the URL being related to the category but that also contain one or more match words that are indicative of the URL not being actually related to the category.

[0056] The present invention has been described by way of examples. Modifications and variations to the above-described examples are possible and may occur to those skilled in the art. All such modifications and variations are believed to be within the scope of the present invention, as defined by the claims. For example, the example method 400 and the example sub-method 500 have been described as performed by a processor 300 of the content access control service 126 (for example, by policy server 132); however, the example method 400 and the example sub-method 500, or modifications thereof, may be programmed into any computing device with network access capabilities such as, for example, by one or more processors of another device within a system like the system 100 such as, another server such as, for example, gateway server 116, by computing devices 104, etc.

[0057] Furthermore, while the example sub-method 500 has been described as performing a determination block based on the requested URL, the determination block may be optionally performed by analyzing alternatively or additionally to the requested URL, the content at the requested URL or parts thereof, for example, by analyzing specific parts of a given Hypertext Markup Language (HTML) document at the requested URL such as the document “description” and “keyword” tags, which are generally associated with Search Engine Optimization (SEO) information. While making the determination based on analyzing the requested URL may be more computationally efficient than further basing the determination on additional content or a section of content at the requested URL, a modified sub-method 500 that uses the additional content for the determination may improve accuracy of the determination, for example, when the URL contains a candidate word identified at heuristic step 505 or amodification thereof that is then positively identified as indicative of a determined category but the additional content provides additional context (i.e. includes words matching to match words that are indicative of the URL not falling within the category) to determine that the requested content at the URL does not fall within the determined category. For example, for a URL such as “sexandu.ca” that contains the word “sex” indicative of a determined pornographic category and that does not contain additional words that indicate that the content at the URL does not fall in the determined category, but that has the following description in the header of the HTML document at the requested URL “Trusted Canadian sexual health information, including consent, contraception, and STI prevention” the example method 400 and the example submethod 500 may accurately determine that the URL does not correspond to the pornographic category based on the weights of the additional words “and”, “u” and “.ca” of the URL countering the weight of the word “sex” and alternatively, a modified submethod 500 that is configured to use the description in the header of the HTML document at the requested URL as input to a modified MLM may also identify that the content at the requested URL does not fall within the category based on the weights of the words in the description, which includes words such as “health”, “information” and “prevention” that are indicative of the URL not corresponding to the category.

[0058] Those skilled in the art will appreciate that in some embodiments, the functionality of the example methods or variations thereof may be implemented using pre-programmed hardware or firmware elements (e.g., application specific integrated circuits (ASICs), electrically erasable programmable read-only memories (EEPROMs), etc.), or other related components.

[0059] The scope of the claims should not be limited by the embodiments set forth in the above examples but should be given the broadest interpretation consistent with the description as a whole.

Claims

CLAIMS1 . A content access control service comprising: a processor configured to: receive a content request from a computing device, the content request including identity information and a requested Uniform Resource Locator (URL); determine whether the content request pertains to content related to a category by: identifying as a candidate word a consecutive string of characters within the requested URL matching a predetermined word in a list of predetermined words indicative of the category; and when the candidate word is identified, providing the requested URL to a Machine Learning Model (MLM) configured to determine whether the requested URL contains a match word indicative of the category to calculate a confidence score representative of whether the URL corresponds to the category; determine whether to approve the content request based on the category determination and the identity information; and direct a client agent of the computing device to the requested content when the content request is approved and redirect the client agent to preapproved content when the content request is denied.

2. The system of claim 1 wherein the MLM is configured to determine whether the requested URL contains an additional match word indicative of the URL being unrelated to the category.

3. The system of claim 1 wherein the processor is further configured to determine whether the content request pertains to content related to the category by further: providing the requested URL to a Bidirectional Encoder Representations from Transformers (BERT) tokenizer configured to tokenize characters of the requested URL by breaking down unknown words into known sub-words and maintaining a record of their original order; andproviding the tokenized characters of the requested URL to the MLM.

4. The system of claim 1 wherein the category is one of a pornographic content category and a violent content category.

5. A method of determining by a processor whether to approve a content request, the method comprising: receiving the content request, the content request including identity information and a requested Uniform Resource Locator (URL); determining whether the content request pertains to content related to a category by: identifying as a candidate word a consecutive string of characters within the requested URL matching a predetermined word in a list of predetermined words indicative of the category; and when the candidate word is identified, providing the requested URL to a Machine Learning Model (MLM) configured to determine whether the requested URL contains a match word indicative of the category to calculate a confidence score representative of whether the URL corresponds to the category; determining whether to approve the content request based on the category determination and the identity information; and directing a client agent of a computing device to the requested content when the content request is approved and redirecting the client agent to pre-approved content when the content request is denied.

6. The method of claim 5 wherein the MLM is configured to determine whether the requested URL contains an additional match word indicative of the URL being unrelated to the category.

7. The method of claim 5 wherein the determination of whether the content request pertains to content related to the category further comprises:tokenizing characters by breaking down unknown words into known sub-words and maintaining a record of their original order; calculating a Term Frequency-Inverse Document Frequency (TF-IDF) vector with the known sub-words; and multiplying the TF-IDF vector by a series of category-specific weights.

8. The method of claim 5 wherein the category is one of a pornographic content category and a violent content category.

Citation Information

Patent Citations

  • Utilizing Machine Learning for dynamic content classification of URL content

    US20220067581A1

  • URL-based content categorization

    US8078625B1

  • Age verification and content filtering systems and methods

    US8131763B2

  • Methods and systems for web site categorization and filtering

    US8539329B2

  • US202463655485P