Network data extraction method and device, equipment and storage medium
By using a pre-defined browser to load visual representations and perform data recognition and transformation within the TOR network, the problem of unreliable data extraction caused by the high latency of the TOR network is solved, achieving efficient and reliable data extraction and reducing maintenance costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING KNOWNSEC INFORMATION TECHNOLOGY CO LTD
- Filing Date
- 2025-12-18
- Publication Date
- 2026-04-10
AI Technical Summary
The high latency of TOR networks causes large-scale data extraction tasks to fail, increasing the cost and unreliability of data extraction.
By obtaining the request address of the target website, loading the visual representation using a preset browser, and performing data identification and positioning through a visual analysis model, the data is converted into readable structured data storage, avoiding the encryption and relay operations of the TOR network and improving the reliability of data extraction.
It reduces long-term maintenance costs, improves the success rate and efficiency of data extraction, and solves the problem of unreliable data extraction caused by high latency in TOR networks.
Smart Images

Figure CN121834073A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of network information collection, in particular to a network data extraction method and device, equipment and storage medium. BACKGROUND
[0002] TOR (The Second Generation Onion Router) is an anonymous communication network based on distributed nodes, and its core function is to hide the network identity and browsing behavior of users. In order to achieve anonymity, the TOR protocol will perform multi-layer encryption and relay on a node network distributed globally and run by volunteers, which will significantly increase network delay. When large-scale data extraction is required from TOR, such high delay is easy to cause the data extraction task to fail, thereby increasing the cost of data extraction. SUMMARY
[0003] Therefore, the present application aims to provide a network data extraction method, device, equipment and storage medium to solve the problem that the existing data extraction method is easy to fail due to high delay.
[0004] To achieve the above-mentioned purpose, the technical solutions adopted by the embodiments of the present application are as follows: In a first aspect, the embodiments of the present application provide a network data extraction method, comprising: obtaining a request address corresponding to a user request from a target website; loading the request address by using a preset browser to obtain a visual representation corresponding to the request address; performing data recognition and positioning on the visual representation by using a visual analysis model to obtain target data in the visual representation; converting the target data into readable structured data and storing the same.
[0005] In an optional implementation, the step of obtaining a request address corresponding to a user request from a target website comprises: obtaining the user request, forwarding the user request to the target website through a proxy service; obtaining a request address corresponding to the user request from the target website.
[0006] In an optional implementation, the step of forwarding the user request to the target website through a proxy service comprises: if the user request forwarding fails, requesting a new address through the proxy service; forwarding the user request to the target website through the new address.
[0007] In an optional implementation, the preset browser comprises a headless browser, and the step of loading the request address by using the preset browser to obtain the visual representation corresponding to the request address comprises: In the headless browser loads the request address and renders the headless browser according to the request address to obtain the visual representation corresponding to the request address.
[0008] In an optional implementation, the method further comprises: generating a corresponding document object model tree according to the visual representation; corresponding the visual representation with the document object model tree to obtain a corresponding relationship between the visual representation and the document object model tree.
[0009] In an optional implementation, the step of performing data recognition and positioning on the visual representation by using a visual analysis model to obtain target data in the visual representation comprises: obtaining layout information of the visual representation, and dividing the visual representation into a plurality of visual regions according to the layout information; performing data recognition and positioning on the visual regions to obtain a target region in the visual regions; obtaining a target model tree node corresponding to the target region from the document object model tree by using a mapping algorithm and the corresponding relationship, and obtaining the target data from the target model tree node.
[0010] In an optional implementation, the step of converting the target data into readable structured data and storing the target data comprises: extracting original text from the target model tree node; performing data cleaning and structured processing on the original text to obtain the target data.
[0011] In a second aspect, an embodiment of the present application provides a network data extraction device, comprising: a request address acquisition module configured to acquire a request address corresponding to a user request from a target website; a request address loading module configured to load the request address by using a preset browser to obtain a visual representation corresponding to the request address; a target data acquisition module configured to perform data recognition and positioning on the visual representation by using a visual analysis model to obtain target data in the visual representation; a target data storage module configured to convert the target data into readable structured data and store the target data.
[0012] Thirdly, embodiments of the present invention provide an electronic device, including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor can execute the machine-executable instructions to implement the network data extraction method described in the first aspect.
[0013] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the network data extraction method described in the first aspect.
[0014] The present invention provides a network data extraction method, apparatus, device, and storage medium. By obtaining the request address corresponding to the user's request from the target website and loading it in a preset browser, the visual representation of the target website can be reproduced in the preset browser. Then, the target data is extracted from the visual representation and stored, which solves the technical problem of unreliable data extraction and greatly reduces long-term maintenance costs.
[0015] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A schematic diagram of the structure of an electronic device provided in an embodiment of the present invention is shown; Figure 2 A flowchart illustrating a network data extraction method provided by an embodiment of the present invention is shown; Figure 3 A flowchart illustrating a vision-to-target model tree node mapping algorithm provided by an embodiment of the present invention is shown. Figure 4 The diagram shows a functional block diagram of a network data extraction device provided in an embodiment of the present invention.
[0018] icon: 100 - Electronic device; 110 - Memory; 120 - Processor; 130 - Communication module; 400 - Network data extraction device; 401 - Request address acquisition module; 402 - Request address loading module; 403 - Target data acquisition module; 404 - Target data storage module. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0020] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0021] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0022] The relevant terms in this embodiment are explained as follows: Dark Web: World Wide Web content that exists on dark networks or overlay networks and can only be accessed with special software, special authorization, or special computer settings.
[0023] TOR (The Second Generation Onion Router): Also known as the Onion Network, it is software used for anonymous communication. The name comes from the acronym of the original software project name "The Onion Router". The Tor Network consists of more than 7,000 relay nodes, each of which is provided free of charge by volunteers around the world. Through layers of relay nodes, the user's real address is hidden and network monitoring and traffic analysis are avoided.
[0024] Visual representation refers to the pixel-based two-dimensional image that a webpage ultimately presents to the human user after it has been fully rendered by the browser engine (including JavaScript execution). This representation can be a screenshot in a bitmap format (such as PNG or JPEG) or a pixel buffer in memory. The key is that it reflects the final visual layout of the page, rather than its source code structure.
[0025] Please refer to Figure 1 , Figure 1 This is a block diagram of an electronic device 100 provided in this embodiment. The electronic device 100 includes a memory 110, a processor 120, and a communication module 130. The memory 110, processor 120, and communication module 130 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0026] The memory 110 is used to store programs or data. The memory 110 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0027] The processor 120 is used to read / write data or programs stored in the memory 110 and to perform corresponding functions.
[0028] The communication module 130 is used to establish a communication connection between the electronic device 100 and other communication terminals through the network, and to send and receive data through the network.
[0029] It should be understood that, Figure 1 The structure shown is only a schematic diagram of the electronic device 100. The electronic device 100 may also include components that are larger than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown. Figure 1 The components shown can be implemented using hardware, software, or a combination thereof.
[0030] Please refer to Figure 2 ,Figure 2 This is a flowchart illustrating a network data extraction method provided in this embodiment. The method can... Figure 1 The method is performed by the electronic device shown, and includes: S201. Obtain the request address corresponding to the user's request from the target website.
[0031] When a user wants to scrape data from certain websites, a user request is generated. However, when accessing certain special websites such as the dark web, for information security reasons, the user will access the target website on the dark web through the TOR network. Since the TOR network is usually anonymous, and considering that some websites will block anonymous access, the user request can be proxied through some request forwarding or proxy services, and then the request address corresponding to the user request can be obtained from the target website.
[0032] S202. Load the request address using a preset browser to obtain a visual representation corresponding to the request address.
[0033] The target website sends its data response to the proxy service based on the user's request. Since directly extracting data from the target website can lead to high latency due to the encryption and relay operations of the TOR network, or even interruptions in data extraction for various reasons, this is not feasible. Therefore, the data response obtained from the target website can be a request URL. This request URL is then loaded and rendered in a default browser to reproduce the visual representation of the target website. The default browser can be the user's client browser. Because the visual representation is reloaded and rendered in the default browser, users accessing the default browser through their client and extracting data from the visual representation do not need to undergo encryption and relay operations, and the data extraction operation is less prone to interruption.
[0034] S203. The visual representation is identified and located using a visual analysis model to obtain the target data in the visual representation.
[0035] A visual representation can be a complete page image. After the browser loads the request address and renders the visual representation, the desired data, i.e. the target data, can be extracted from the visual representation through a visual analysis model.
[0036] S204. Convert the target data into readable structured data and store it.
[0037] Finally, the extracted target data undergoes format conversion and other operations to obtain readable structured data, which is then stored for subsequent real-time retrieval of the target data.
[0038] This embodiment obtains the request address corresponding to the user's request from the target website and loads it in a preset browser, so that the visual representation of the target website can be reproduced in the preset browser. Then, the target data is extracted from the visual representation and stored, which solves the technical problem of unreliable data extraction and greatly reduces long-term maintenance costs.
[0039] In one implementation, the step of obtaining the request address corresponding to the user request from the target website includes: The user request is obtained, and the user request is forwarded to the target website through a proxy service; Obtain the request address corresponding to the user's request from the target website.
[0040] To achieve anonymity, users can use the TOR network to encrypt and relay their requests. The user request is encrypted sequentially with three different keys, each layer corresponding to a relay node. Each relay node can be randomly selected from the network, and each node only decrypts the data corresponding to its own layer, unable to view the complete content or the user's identity. The final exit node then sends the data to the target website.
[0041] To enhance security, some websites deploy various anti-scraping tools to block anonymous automated access. Furthermore, because the list of exit nodes for the TOR network is public, many websites directly block traffic from known TOR IP addresses. Additionally, the IP addresses of TOR exit nodes are shared and often have low credibility due to their potential for malicious activity. This makes requests from the TOR network more likely to trigger CAPTCHAs or other bot-detection challenges, which are extremely difficult or even impossible for automated systems to resolve.
[0042] Therefore, all requests can be routed using a proxy service, such as a SOCKS5 proxy. That is, by using a SOCKS5 proxy as the front-end entry point, user requests are first sent to the SOCKS5 proxy server, which then forwards the request to the entry node of the TOR network, thus achieving anonymity for the user request. To enable access to the target website, the SOCKS5 proxy can generate a corresponding virtual IP address based on the user request and forward the virtual IP address along with the user request to the TOR network. Then, access to the target website is achieved through this virtual IP address, preventing the user request from being intercepted by the target website.
[0043] If the user request forwarding fails, a new address is requested through the proxy service, and the user request is forwarded to the target website through the new address. The new address is different from the previously blocked virtual IP addresses in order to improve the success rate of the user request.
[0044] Once the user request is successfully forwarded, the SOCKS5 proxy retrieves the corresponding data response from the TOR network and forwards it to the user client.
[0045] This embodiment uses a proxy service to forward user requests and generates a virtual IP address based on the user request. When the request forwarding fails, a new address is generated to re-request, which can effectively prevent user requests from being blocked by the target website and improve the success rate of requests and data extraction.
[0046] In one implementation, the preset browser includes a headless browser, and the step of loading the request address using the preset browser to obtain a visual representation corresponding to the request address includes: The headless browser loads the request address and renders the headless browser according to the request address to obtain a visual representation corresponding to the request address.
[0047] The headless browser loads the requested address and generates a rendered visual representation, such as an image in PNG or JPEG format. At this point, a final, stable interface is obtained, and data extraction is no longer affected by the target website.
[0048] Headless browsers can be based on frameworks such as Playwright, Puppeteer, or Selenium. A headless browser is a browser without a graphical user interface. It does not display visual elements such as windows and buttons as in traditional browsers. Therefore, by loading a request address through a headless browser, the original visual representation of the request address can be completely restored. Moreover, headless browsers have low memory usage at runtime, which can save system resources.
[0049] In one embodiment, the method further includes: Generate a corresponding document object model tree based on the visual representation; The visual representation is mapped to the document object model tree to obtain the correspondence between the visual representation and the document object model tree.
[0050] After loading the requested URL and obtaining the visual representation of the webpage corresponding to the URL, the browser parses the HTML code of the webpage into a tree structure composed of nodes, namely the Document Object Model (DOM) tree. The root node is at the top, and other elements, text, etc., are distributed below the root node as branches and leaf nodes. Different visual representations correspond to different nodes in the DOM tree.
[0051] In one implementation, the step of performing data recognition and localization on the visual representation using a visual analysis model to obtain target data in the visual representation includes: Obtain the layout information of the visual representation, and divide the visual representation into multiple visual regions based on the layout information; Data recognition and localization are performed on the visual region to obtain the target region within the visual region; Using the mapping algorithm and the correspondence, the target model tree node corresponding to the target region is obtained from the document object model tree, and the target data is obtained from the target model tree node.
[0052] Please refer to Figure 3 , Figure 3 This is a flowchart illustrating a vision-to-target model tree node mapping algorithm provided in this embodiment.
[0053] After acquiring the visual regions, the predicted visual bounding boxes for each visual region are determined. All elements in the target model tree (DOM) nodes are rendered and traversed to calculate the rendered bounding box for each DOM element. Then, the correspondence between the target region and the target model tree nodes is determined based on the interaction ratio score between the predicted visual bounding boxes and the rendered bounding boxes.
[0054] A visual analysis model can be an object detection model. The object detection model divides the visual representation, such as a webpage screenshot, into multiple regions. Then, it performs corresponding detection on the fields contained in each region, such as keyword extraction. Regions containing keyword fields are selected as candidate regions, and each candidate region is assigned a confidence score. The final target region is determined based on the confidence score.
[0055] For example, when extracting data from certain dark web ransomware, multiple bounding boxes are generated using a target detection model. Each bounding box is used to select text in different areas of a webpage screenshot. Then, the text area containing keywords, such as the victim's name, is used as the target visual bounding box, and the position of the target text in the image is recorded.
[0056] After determining the location of the target text in the image, the tree node corresponding to the target text is determined from the document object model tree according to the mapping algorithm. Then, the original text is extracted from the target model tree node, and the original text is cleaned and structured, such as being converted into a JSON object, to obtain the target data. Visual analytics models can also be multimodal Transformer models, such as LayoutLMv3, which can simultaneously process visual representations, text extracted via OCR, and the layout and position information of the text. This allows it to understand not only the content but also the structure and relationships between elements. For example, this model can learn that the text to the right of the label "victim:" is the victim's name. This is used to extract key-value pairs commonly found on ransomware websites on the dark web.
[0057] This embodiment uses different visual analysis models to analyze and locate visual representations, thereby quickly and accurately extracting target data from visual representations and improving the efficiency of data extraction.
[0058] To perform the corresponding steps in the above embodiments and various possible methods, an implementation of a network data extraction device is given below. Please refer to [link / reference]. Figure 4 , Figure 4 This is a functional block diagram of a network data extraction device provided in an embodiment of the present invention. It should be noted that the basic principle and technical effects of the network data extraction device provided in this embodiment are the same as those in the above embodiments. For the sake of brevity, any parts not mentioned in this embodiment can be referred to the corresponding content in the above embodiments. The network data extraction device 400 includes: The request address acquisition module 401 is used to obtain the request address corresponding to the user's request from the target website; The request address loading module 402 is used to load the request address using a preset browser to obtain a visual representation corresponding to the request address; The target data acquisition module 403 is used to perform data recognition and positioning on the visual representation through a visual analysis model to obtain the target data in the visual representation. The target data storage module 404 is used to convert the target data into readable structured data and store it.
[0059] Optionally, the above modules can be stored in the form of software or firmware. Figure 1 The memory shown may be stored in or embedded in the operating system (OS) of the network data extraction device, and may be generated by... Figure 1 The processor executes the commands. Meanwhile, the data and program code required to execute these modules can be stored in memory.
[0060] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative; for example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0061] In addition, the functional modules in the various embodiments of the present invention can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0062] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0063] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for extracting network data, characterized in that, include: Obtain the request address corresponding to the user's request from the target website; The request address is loaded using a preset browser to obtain a visual representation corresponding to the request address; The visual representation is identified and located using a visual analysis model to obtain the target data in the visual representation. The target data is converted into readable structured data and stored.
2. The network data extraction method according to claim 1, characterized in that, The step of obtaining the request address corresponding to the user's request from the target website includes: The user request is obtained, and the user request is forwarded to the target website through a proxy service; Obtain the request address corresponding to the user's request from the target website.
3. The network data extraction method according to claim 2, characterized in that, The step of forwarding the user request to the target website through a proxy service includes: If the user request forwarding fails, a new address is requested through the proxy service; The user request is forwarded to the target website via the new address.
4. The network data extraction method according to claim 1, characterized in that, The preset browser includes a headless browser. The step of loading the request address using the preset browser to obtain a visual representation corresponding to the request address includes: The headless browser loads the request address and renders the headless browser according to the request address to obtain a visual representation corresponding to the request address.
5. The network data extraction method according to claim 4, characterized in that, The method further includes: Generate a corresponding document object model tree based on the visual representation; The visual representation is mapped to the document object model tree to obtain the correspondence between the visual representation and the document object model tree.
6. The network data extraction method according to claim 5, characterized in that, The step of performing data recognition and localization on the visual representation using a visual analysis model to obtain the target data in the visual representation includes: Obtain the layout information of the visual representation, and divide the visual representation into multiple visual regions based on the layout information; Data recognition and localization are performed on the visual region to obtain the target region within the visual region; Using the mapping algorithm and the correspondence, the target model tree node corresponding to the target region is obtained from the document object model tree, and the target data is obtained from the target model tree node.
7. The network data extraction method according to claim 6, characterized in that, The step of converting the target data into readable structured data and storing it includes: Extract the original text from the target model tree nodes; The original text is cleaned and structured to obtain the target data.
8. A network data extraction device, characterized in that, include: The request address acquisition module is used to obtain the request address corresponding to the user's request from the target website; The request address loading module is used to load the request address using a preset browser and obtain a visual representation corresponding to the request address; The target data acquisition module is used to perform data recognition and localization on the visual representation through a visual analysis model to obtain the target data in the visual representation. The target data storage module is used to convert the target data into readable structured data and store it.
9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor to implement the network data extraction method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the network data extraction method as described in any one of claims 1-7.