Privacy protection method for extracting webpage information for CSS (Cascading Style Sheet) selector
By performing form-sensitive field filtering and HTTP response header checking on CSS selectors under the Scrapy framework, the problem of personal identity information leakage in web page data extraction is solved, and higher data security is achieved.
Patent Information
- Application Number
- CN202510473598.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-25
AI Technical Summary
When using the Scrapy framework and CSS selector for web page data extraction, there is a risk that unauthorized personal identity information (such as ID number, name, address, phone number, email address, etc.) will be leaked, resulting in data leakage.
Ensure compliance with the security policies of the website by filtering form-sensitive fields with selectors under the Scrapy framework, using regular expressions to detect and remove data items containing sensitive information, and check security policies in HTTP response headers such as X-Frame-Options, Strict-Transport-Security (HSTS), and Content-Security-Policy (CSP).
Effectively prevent the leakage of personal identity and other privacy data, and improve data security during the extraction of web page information.
Smart Images

Figure CN120372683A_ABST
Abstract
Description
Technical Field
[0001] The present invention discloses a privacy protection method for web page information extraction oriented to CSS selectors, which relates to the technical field of data protection. Background Art
[0002] In the field of computer applications, using the Scrapy framework and CSS selectors for web page data extraction is an efficient, flexible and applicable web page data extraction method for various types of websites. The Scrapy framework is an open-source web crawler framework written in Python, which can quickly capture website data; CSS selectors are a way to select HTML document elements in web development. However, in the process of web page information extraction, there are problems of collecting unauthorized personal identity information such as user ID numbers, names, addresses, phone numbers, and email addresses, resulting in the risk of data leakage. Summary of the Invention
[0003] Aiming at the problems of the prior art, the present invention provides a privacy protection method for web page information extraction oriented to CSS selectors. By selector filtering, adding data filtering rules, and response header analysis, the data security is enhanced, ensuring that personal identity or other privacy data will not be leaked during the information extraction process, and realizing the protection of user privacy.
[0004] The specific solution proposed by the present invention is as follows:
[0005] The present invention provides a privacy protection method for web page information extraction oriented to CSS selectors, including:
[0006] Step 1: Under the Scrapy framework, use selectors to filter sensitive form fields and exclude elements related to privacy information.
[0007] Step 2: Add data filtering rules before parsing the page content, and use regular expressions to detect and remove data items containing sensitive information.
[0008] Step 3: Obtain and check the HTTP response header returned by the server to ensure compliance with the website's security policy: check the header information related to privacy protection such as X-Frame-Options, Strict-Transport-Security (HSTS), and Content-Security-Policy (CSP). Among them, the X-Frame-Options header is used to control whether the page is <iframe>Shown; Strict-Transport-Security (HSTS) restricts browsers to access the site only via HTTPS connections; the Content-Security-Policy (CSP) header defines the resources that can be loaded and where to load the resources.
[0009] Further, in step 1 of the privacy protection method for web page information extraction oriented to CSS selectors, filtering form sensitive fields using selectors includes:
[0010] Using XPath expressions to find matching elements in the HTML document, selecting all elements with the class attribute of item,
[0011] Extracting data from the elements and returning the extracted data in dictionary form for subsequent processing by Scrapy.
[0012] Further, in step 2 of the privacy protection method for web page information extraction oriented to CSS selectors, defining the filter_sensitive_data function, based on the filter_sensitive_data function, defining a list of sensitive information regular expressions, the list of sensitive information regular expressions contains a regular expression for matching ID card numbers, checking whether the input string data contains ID card number sensitive information,
[0013] Using item.css to extract the content of elements from HTML, and each time extracting, calling the filter_sensitive_data function to judge whether there is ID card number sensitive information in the extracted content, and if so, filtering the corresponding content.
[0014] Further, in step 3 of the privacy protection method for web page information extraction oriented to CSS selectors, using the parse method to process the HTTP response and detect the response header information, where first defining the list headers_to_check, which contains the names of HTTP response headers to be checked: X-Frame-Options, Strict-Transport-Security (HSTS), Content-Security-Policy (CSP),
[0015] Traversing and checking each response header name in the headers_to_check list.
[0016] The present invention also provides a privacy protection device for web page information extraction oriented to CSS selectors, including a field filtering module, a sensitive information filtering module, and a response header checking module.
[0017] The step field filtering module filters sensitive form fields using a selector under the Scrapy framework to exclude elements related to privacy information.
[0018] The sensitive information filtering module adds data filtering rules before parsing the page content, and uses regular expressions to detect and remove data items containing sensitive information.
[0019] The response header checking module obtains and checks the HTTP response header returned by the server to ensure compliance with the website's security policy: checking the header information related to privacy protection such as X-Frame-Options, Strict-Transport-Security (HSTS), and Content-Security-Policy (CSP). Among them, the X-Frame-Options header is used to control whether the page is displayed in an <iframe>; Strict-Transport-Security (HSTS) restricts the browser to access the site only through HTTPS connections; the Content-Security-Policy (CSP) header defines the resources that can be loaded and where to load the resources.
[0020] Furthermore, the field filtering module of the privacy protection device for web page information extraction oriented to CSS selectors filters sensitive form fields using a selector, including:
[0021] Using an XPath expression to find matching elements in the HTML document, and selecting all elements with the class attribute of item.
[0022] Extracting data from the elements and returning the extracted data in dictionary form for subsequent processing by Scrapy.
[0023] Further, the sensitive information filtering module of the privacy protection device for web page information extraction oriented to CSS selectors defines the filter_sensitive_data function. Based on the filter_sensitive_data function, a list of sensitive information regular expressions is defined. The list of sensitive information regular expressions contains a regular expression for matching ID card numbers to check whether the input string data contains ID card number sensitive information.
[0024] Use item.css to extract the content of elements from HTML. Each time extraction is performed, call the filter_sensitive_data function to determine whether there is ID card number sensitive information in the extracted content. If so, filter the corresponding content.
[0025] Further, the response header checking module of the privacy protection device for web page information extraction oriented to CSS selectors uses the parse method to process HTTP responses and detect response header information. First, define the list headers_to_check, which contains the names of HTTP response headers to be checked: X-Frame-Options, Strict-Transport-Security (HSTS), Content-Security-Policy (CSP).
[0026] Traverse and check each response header name in the headers_to_check list.
[0027] The beneficial effects of the present invention are:
[0028] The privacy protection method for web page information extraction oriented to CSS selectors under the Scrapy framework filters and screens relevant selectors for CSS selectors, removes the elements involving privacy information selected, adds filtering rules, detects whether the corresponding data items contain the corresponding sensitive information according to regular expressions. If so, remove the data items, and check the HTTP response headers to prevent the leakage of personal identity and other privacy data, and improve the security of data in the process of web page information extraction.BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a schematic diagram of the method flow of the present invention.Detailed Implementation Modes
[0030] The following further describes the present invention in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the described embodiments are not intended to limit the present invention.
[0031] Embodiment 1
[0032] The present invention provides a privacy protection method for web page information extraction oriented to CSS selectors, including:
[0033] Step 1: Under the Scrapy framework, use selectors to filter sensitive form fields and exclude elements related to privacy information. Using selectors to filter sensitive form fields includes:
[0034] Use XPath expressions to find matching elements in the HTML document, and select all elements with the class attribute of item,
[0035] Extract data from the elements and return the extracted data in dictionary form for subsequent processing by Scrapy. The reference code is as follows:
[0036]
[0037] Step 2: Add data filtering rules before parsing the page content, and use regular expressions to detect and remove data items containing sensitive information. Define the filter_sensitive_data function. Based on the filter_sensitive_data function, define a list of regular expressions for sensitive information. The list of regular expressions for sensitive information contains a regular expression for matching ID card numbers, and check whether the input string data contains ID card number sensitive information,
[0038] Use item.css to extract the content of elements from the HTML. Each time extraction is performed, call the filter_sensitive_data function to determine whether there is ID card number sensitive information in the extracted content. If so, filter the corresponding content.The code can be referred to as follows:
[0039]
[0040]
[0041] Step 3: Obtain and check the HTTP response headers returned by the server to ensure compliance with the website's security policy: Check the header information related to X-Frame-Options, Strict-Transport-Security (HSTS), Content-Security-Policy (CSP), and privacy protection. Among them, the X-Frame-Options header is used to control whether the page is displayed in an <iframe>; Strict-Transport-Security (HSTS) restricts the browser to access the site only through HTTPS connections; the Content-Security-Policy (CSP) header defines the resources that can be loaded and where to load the resources.
[0042] Among them, the parse method is used to process the HTTP response and detect the response header information. First, a list headers_to_check is defined, which contains the names of the HTTP response headers to be checked: X-Frame-Options, Strict-Transport-Security (HSTS), Content-Security-Policy (CSP).
[0043] Traverse and check each response header name in the headers_to_check list.The code can be referred to as follows:
[0044]
[0045] Example 2
[0046] The present invention also provides a privacy protection device for web page information extraction oriented to CSS selectors, including a field filtering module, a sensitive information filtering module, and a response header checking module.
[0047] The step field filtering module, under the Scrapy framework, uses selectors to filter sensitive form fields and excludes elements related to privacy information.
[0048] The sensitive information filtering module adds data filtering rules before parsing the page content, and uses regular expressions to detect and remove data items containing sensitive information.
[0049] The response header checking module obtains and checks the HTTP response headers returned by the server to ensure compliance with the website's security policy: checks the X-Frame-Options, Strict-Transport-Security (HSTS), Content-Security-Policy (CSP) and other privacy protection-related header information. Among them, the X-Frame-Options header is used to control whether the page is displayed in an <iframe>; Strict-Transport-Security (HSTS) restricts the browser to access the site only through HTTPS connections; the Content-Security-Policy (CSP) header defines the resources that can be loaded and where to load the resources.
[0050] Regarding the information interaction and execution process among the above-mentioned modules in the device, since they are based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention and will not be elaborated here.
[0051] Similarly, the privacy protection method for web page information extraction of the device of the present invention oriented to CSS selectors under the Scrapy framework filters and screens the CSS selectors with relevant selectors, removes the elements related to privacy information selected, adds filtering rules, detects whether the corresponding data items contain the corresponding sensitive information according to regular expressions, and if so, removes the data items, and checks the HTTP response headers to prevent the leakage of personal identity and other privacy data, and improve the security of data in the process of web page information extraction.
[0052] It should be noted that not all steps and modules in the above processes and device structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted as needed. The system structures described in the above embodiments can be physical structures or logical structures, that is, some modules may be implemented by the same physical entity, or some modules may be implemented separately by multiple physical entities, or some components in multiple independent devices can be jointly implemented.
[0053] The above-described embodiments are merely preferred embodiments cited to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are within the protection scope of the present invention. The protection scope of the present invention is subject to the claims.< / iframe>
Claims
1. A privacy protection method for web page information extraction oriented to CSS selectors, characterized in that Including: Step 1: Under the Scrapy framework, use a selector to filter sensitive form fields and exclude elements related to privacy information. Step 2: Add data filtering rules before parsing the page content, and use regular expressions to detect and remove data items containing sensitive information. Step 3: Obtain and check the HTTP response headers returned by the server to ensure compliance with the website's security policy: Check the headers related to privacy protection such as X-Frame-Options, Strict-Transport-Security (HSTS), and Content-Security-Policy (CSP). Among them, the X-Frame-Options header is used to control whether the page is <iframe>Shown; Strict-Transport-Security (HSTS) restricts the browser to access the site only through HTTPS connections; the Content-Security-Policy (CSP) header defines the resources that can be loaded and where to load the resources.
2. A privacy protection method for web page information extraction oriented to CSS selectors according to claim 1, characterized in that in step 1, the form sensitive fields are filtered using selectors, including: Use XPath expressions to find matching elements in the HTML document, select all elements with the class attribute of item, Extract data from the element and return the extracted data in dictionary form for subsequent processing by Scrapy.
3. A privacy protection method for web page information extraction oriented to CSS selectors according to claim 1, characterized in that in step 2, define the filter_sensitive_data function. Based on the filter_sensitive_data function, define a list of sensitive information regular expressions. The list of sensitive information regular expressions contains a regular expression for matching ID card numbers, and check whether the input string data contains ID card number sensitive information. Use item.css to extract the content of elements from HTML. Each time extraction is performed, call the filter_sensitive_data function to determine whether there is ID card number sensitive information in the extracted content. If it exists, filter the corresponding content.
4. The privacy protection method for web page information extraction oriented to CSS selectors according to claim 1, characterized in that in step 3, the parse method is used to process the HTTP response and detect the response header information, where a list headers_to_check is first defined, which contains the names of HTTP response headers to be checked: X-Frame-Options, Strict-Transport-Security (HSTS), Content-Security-Policy (CSP),traverse and check each response header name in the headers_to_check list.
5. A privacy protection device for web page information extraction oriented to CSS selectors, characterized by including a field filtering module, a sensitive information filtering module, and a response header checking module.The step field filtering module, under the Scrapy framework, filters sensitive form fields using selectors to exclude elements related to privacy information.The sensitive information filtering module adds data filtering rules before parsing the page content, and uses regular expressions to detect and remove data items containing sensitive information.The response header checking module obtains and checks the HTTP response headers returned by the server to ensure compliance with the website's security policy: checks the header information related to privacy protection such as X-Frame-Options, Strict-Transport-Security (HSTS), and Content-Security-Policy (CSP). Among them, the X-Frame-Options header is used to control whether the page is displayed in an <iframe>; Strict-Transport-Security (HSTS) restricts the browser to access the site only through HTTPS connections; the Content-Security-Policy (CSP) header defines the resources that can be loaded and where to load the resources.
6. A privacy protection device for web page information extraction oriented to CSS selectors, characterized in that the field filtering module filters sensitive fields of forms using selectors, including: Use XPath expressions to find matching elements in the HTML document, and select all elements with the class attribute of item. Extract data from the elements and return the extracted data in dictionary form for subsequent processing by Scrapy.
7. A privacy protection device for web page information extraction oriented to CSS selectors, characterized in that the sensitive information filtering module defines a filter_sensitive_data function. Based on the filter_sensitive_data function, a list of sensitive information regular expressions is defined. The list of sensitive information regular expressions contains a regular expression for matching ID card numbers to check whether the input string data contains sensitive ID card number information. Use item.css to extract the content of elements from HTML. Each time extraction is performed, call the filter_sensitive_data function to determine whether there is sensitive ID card number information in the extracted content. If it exists, filter the corresponding content.
8. A privacy protection device for web page information extraction oriented to CSS selectors, characterized in that the response header check module uses the parse method to process the HTTP response and detect the response header information, where a list headers_to_check is first defined, which contains the names of the HTTP response headers to be checked: X-Frame-Options, Strict-Transport-Security (HSTS), Content-Security-Policy (CSP),traverse and check each response header name in the headers_to_check list..< / iframe>