A web crawler detection system based on application scenarios

By using a web crawler detection system based on application scenarios, and leveraging user analysis, human-machine recognition, and identity verification, the system achieves triple identification of malicious crawlers. This solves the problem that conventional detection methods cannot identify new types of malicious crawlers, and improves detection flexibility and user experience.

CN115525813BActive Publication Date: 2026-05-26WUHAN JIYI NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN JIYI NETWORK TECH CO LTD
Filing Date
2022-09-30
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

When faced with novel malicious crawler designs, existing technologies and conventional detection methods are unable to effectively screen out malicious crawlers, leading to the encroachment of enterprise resources and the impact on network operations.

Method used

A web crawler detection system based on application scenarios is adopted. It performs triple malicious crawler identification through user analysis unit, human-machine recognition unit, secondary verification unit and space development unit. Combined with the purpose classification of enterprise network and user behavior analysis, it uses CAPTCHA verification, identity information verification and application scenario correlation analysis to achieve comprehensive screening.

Benefits of technology

It improves the flexibility and accuracy of malicious crawler detection, reduces false positives on real users, lowers enterprise network operating costs, and provides powerful operational management support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115525813B_ABST
    Figure CN115525813B_ABST
Patent Text Reader

Abstract

This invention discloses a web crawler detection system based on application scenarios, including a crawler detection platform. The crawler detection platform comprises a user analysis unit, a human-machine recognition unit, a secondary verification unit, a space development unit, and a reality integration unit. This invention relates to the field of web crawler detection technology. This application scenario-based web crawler detection system divides the enterprise network into different application scenarios according to their uses. Based on these application scenarios, it performs triple malicious crawler identification, achieving a more comprehensive screening of malicious crawlers. Simultaneously, it provides registered users with personal spaces, offering convenience and effectively avoiding false positives on genuine users, thus improving user experience. Furthermore, by combining network information related to the application scenarios to manage personal spaces, it greatly enhances the flexibility of malicious crawler detection while effectively locating abnormal parts of the enterprise network, providing powerful assistance for the operation and management of the enterprise network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of web crawler detection technology, specifically a web crawler detection system based on application scenarios. Background Technology

[0002] A web crawler, also known as a web spider or web robot, is a program or script used to automatically browse the World Wide Web. Crawlers can verify hyperlinks and HTML code, and are used for web scraping. Web search engines and other sites use crawler software to update their own website content or their indexes of other websites. The process of a crawler accessing a website consumes the target system's resources, so when accessing a large number of pages, crawlers need to consider issues such as planning and load.

[0003] In real-world enterprise networks, malicious web crawlers often infiltrate and steal core enterprise resources and documents, causing losses and disrupting normal network operations. Conventional enterprise network anti-crawler designs typically detect malicious crawlers by setting up IP blacklists / whitelists or using user agent checks. Newer malicious crawler designs include setting up massive IP proxy pools and controlling browsers through emulators. Conventional detection methods are often ineffective in filtering out these new malicious crawlers. Summary of the Invention

[0004] (a) Technical problems to be solved

[0005] To address the shortcomings of existing technologies, this invention provides a web crawler detection system based on application scenarios, which solves the problem that conventional detection methods often fail to effectively screen out malicious crawlers in novel malicious crawler designs.

[0006] (II) Technical Solution

[0007] To achieve the above objectives, the present invention provides the following technical solution: a web crawler detection system based on application scenarios, comprising a crawler detection platform, wherein the crawler detection platform includes a user analysis unit, a human-machine recognition unit, a secondary verification unit, a spatial development unit, and a reality integration unit. The user analysis unit is used to divide the enterprise network into several application scenarios according to their uses, and simultaneously record the browsing behavior of test users in the test version, generating user usage specifications. The user analysis unit interfaces with the human-machine recognition unit, which performs human-machine recognition to achieve the first layer of crawler identification. The human-machine recognition unit interfaces with the secondary verification unit, wherein the secondary verification unit sets an alarm threshold for the number of daily verification markings in the same application scenario. When the actual number of markings in the application scenario exceeds the alarm threshold, the registered user is marked as abnormal, and human-machine recognition verification is triggered simultaneously. The second layer of crawler identification involves the secondary verification unit, which interfaces with the space development unit, and the space development unit, which in turn interfaces with the human-machine identification unit. The space development unit sets a daily limit for the number of repeated browsing sessions in the same application scenario. When this limit is reached, basic identity verification is initiated for the registered user. After successful verification, the registered user's personal space is opened within the enterprise network for downloading application scenario information. The space development unit interfaces with the reality integration unit, which in turn interfaces with the secondary verification unit. The reality integration unit synchronizes the updates of application scenario information on the enterprise network with the registered user's personal space information. It also detects official website information related to the browsing direction of the application scenario and performs correlation analysis before crawler identification of the registered user and their opened personal space, thus achieving the third layer of crawler identification.

[0008] By adopting the above technical solution, the enterprise network is divided into different application scenarios according to its purpose. Based on the application scenario, a three-dimensional malicious crawler identification is performed, which more comprehensively achieves the screening of malicious crawlers. At the same time, personal spaces are opened for registered users to store application scenario information, which provides convenience for registered users and effectively avoids false positives to real users, improves user experience, and combines network information related to application scenarios to manage personal spaces. This greatly improves the flexibility of malicious crawler detection and effectively locates abnormal parts of the enterprise network, providing strong support for the operation and management of the enterprise network.

[0009] The present invention is further configured such that: the user analysis unit includes a big data recording module, a data analysis module, a specification setting module, and a setting verification module;

[0010] The big data recording module is used to divide the enterprise network into several application scenarios according to its purpose, and at the same time record the browsing of data by test users in different application scenarios of the enterprise network in the test version and the opening ratio of the test users' personal space within a limited time.

[0011] The data analysis module is used to analyze the browsing behavior of test users in the test version and summarize the similarities in user usage.

[0012] The specification setting module is used to refine enterprise requirements into a framework based on the summarized user usage similarities, forming the enterprise network usage specifications as the initial user usage specifications.

[0013] The setting verification module is used to put the initial user usage specifications into the test version for testing and adjustment, so as to obtain the revised user usage specifications.

[0014] By adopting the above technical solution, the framework of user usage specifications is formulated using the big data generated by the enterprise network test version. After adding enterprise requirements to the framework, tests and adjustments are made to ensure that the use of the enterprise network is truly and effectively in line with the user group, while realizing the enterprise's needs and providing a comparison standard for the third layer of identification of malicious crawlers.

[0015] The present invention is further configured such that: the human-machine recognition unit includes a verification code testing module, a human-machine recognition module, and an initial identification module;

[0016] The verification code testing module is used to send random verification codes to registered users on the enterprise network, perform human-machine identification verification, and send identity verification information to registered users;

[0017] The human-machine recognition module is used to receive feedback from registered users on verification codes and identity verification information, and to verify the correctness of the feedback on verification codes and identity verification information.

[0018] The initial identification module is used to trace the registered user and mark the crawler when the returned verification code is incorrect, and to trace the registered user and mark the crawler when the returned identity verification information is incorrect.

[0019] By adopting the above technical solution and using CAPTCHA to perform human-machine identification, the first layer of malicious web crawler identification is achieved, and feedback verification of identity information is provided. This assists in the opening of personal space and the third layer of malicious web crawler identification, and deepens the coordination of the three layers of malicious web crawler identification.

[0020] The present invention is further configured such that: the secondary verification unit includes a usage setting module, a threshold setting module, an anomaly triggering module, and a crawler marking module;

[0021] The setting module is used to set the number of times a registered user's identity information is verified per day in enterprise network application scenarios, and to permanently mark the number of times a registered user is verified per day.

[0022] The threshold setting module is used to set alarm thresholds for the number of times a verification is performed in a single day in the same application scenario, and to generate a monthly report of permanent number of verifications for registered users.

[0023] The anomaly triggering module is used to mark the registered user as abnormal when the actual number of markings in the application scenario exceeds the alarm threshold, and at the same time trigger human-machine recognition verification.

[0024] The crawler tagging module is used to trace the registered user back to the network and tag the user as a crawler when the human-machine identification verification fails.

[0025] By adopting the above technical solution, combined with the human-machine recognition unit for a second layer of malicious crawler identification, and providing enterprises with statistics on the number of daily verifications of registered users in the form of monthly reports, research data is provided for enterprises' progress in opening up personal spaces.

[0026] The present invention is further configured such that: the space development unit includes a duplicate verification module, an identity information verification module, and a personal space development module, wherein the duplicate verification module is connected to the identity information verification module, and the identity information verification module is connected to the personal space development module.

[0027] The present invention is further configured such that: the duplicate verification module is used to record the number of times a registered user repeatedly browses the same application scenario, and to set the maximum number of times a single application scenario is repeatedly browsed in a single day;

[0028] The identity information verification module is used to send basic identity information verification to the registered user through a secondary verification unit when the number of repeated browsings in the same application scenario reaches an extreme value in a single day.

[0029] The personal space creation module is used to open up personal space for downloading application scenario information after a registered user has verified their identity information.

[0030] By adopting the above technical solution, the extreme number of times registered users browse application scenarios is limited. This ensures that registered users who have passed the first layer of malicious crawler screening can browse normally, while also assisting in the determination of application scenarios that registered users are interested in. This avoids the creation of a large number of personal spaces corresponding to application scenarios, reduces the operating costs of the enterprise network, and provides a second layer of identification and screening for malicious crawlers, effectively ensuring the normal operation of the enterprise network.

[0031] The present invention is further configured such that: the reality integration unit includes a space management module, a reality correlation analysis module, an anomaly marking module, and a space closure module; the space management module is connected to the reality correlation analysis module, the reality correlation analysis module is connected to the anomaly marking module, and the anomaly marking module is connected to the space closure module.

[0032] The present invention is further configured such that: the space management module is used to download information according to the application scenario of the registered user, encrypt it with a key, and update the application scenario to the personal space after the application scenario is updated in the enterprise network;

[0033] The actual correlation analysis module is used to record the open time period of personal space, calculate the open percentage of personal space within the limited time, compare it with the open percentage of the test version, and when it exceeds the open percentage of the test version, it summarizes and analyzes the application scenarios browsed by registered users, determines the browsing direction, detects the official website information related to the browsing direction, and performs correlation analysis.

[0034] When the official website related to the application scenario has not published relevant information, the anomaly marking module marks the registered user with spatial anomalies and sends a deep-level identity verification to the registered user.

[0035] The space closing module is used to close the personal space when a registered user fails the verification request, and to add a crawler tag after tracing the registered user.

[0036] By adopting the above technical solutions, personal spaces are managed and controlled. Combined with big data analysis in the user analysis unit, the percentage of personal spaces open within a limited time is compared to directly locate abnormal data, narrowing the scope of malicious crawlers. At the same time, by combining the usage of the personal space with the application scenarios corresponding to the browsing direction, correlation analysis of official website information is conducted to determine the cause of the anomaly, thereby helping to identify the third layer of malicious crawlers.

[0037] (III) Beneficial Effects

[0038] This invention provides a web crawler detection system based on application scenarios. It has the following beneficial effects:

[0039] (1) This application scenario-based web crawler detection system divides the enterprise network into different application scenarios according to their uses, and performs triple malicious crawler identification based on the application scenarios. This system can more comprehensively screen malicious crawlers while opening personal spaces for registered users to store application scenario information and provide convenience for registered users. This effectively avoids harming real users and improves user experience. Furthermore, it combines network information related to application scenarios to manage personal spaces, which greatly improves the flexibility of malicious crawler detection and effectively locates abnormal parts of the enterprise network, providing strong support for the operation and management of the enterprise network.

[0040] (2) This application scenario-based web crawler detection system uses big data generated by enterprise network test versions to formulate a framework for user usage specifications, and adds enterprise requirements to the framework for testing and adjustment. This effectively ensures that the use of enterprise networks is truly and effectively in line with the user group, while realizing the enterprise's needs and providing a comparison standard for the third layer of malicious crawler identification.

[0041] (3) This web crawler detection system based on application scenarios uses CAPTCHA to perform human-machine identification, thereby achieving the first layer of malicious crawler identification and providing feedback verification of identity information, which assists in the opening of personal space and the third layer of malicious crawler identification, and deepens the coordination of the three layers of malicious crawler identification.

[0042] (4) This web crawler detection system based on application scenarios performs a second layer of malicious crawler identification by working with a human-machine recognition unit, and provides enterprises with statistics on the number of daily verifications of registered users in the form of monthly reports, providing research data for the progress of enterprises in opening up personal spaces.

[0043] (5) This application scenario-based web crawler detection system limits the maximum number of times registered users browse application scenarios, ensuring that registered users who have passed the first layer of malicious crawler screening can browse normally, while assisting in the judgment of application scenarios that registered users are interested in. This avoids the creation of a large number of personal spaces corresponding to application scenarios, reduces the operating costs of the enterprise network, and at the same time performs a second layer of identification and screening of malicious crawlers, effectively ensuring the normal operation of the enterprise network.

[0044] (6) This application scenario-based web crawler detection system manages personal spaces, combines big data analysis in the user analysis unit, compares the openness ratio of personal spaces within a limited time, directly locates abnormal data, narrows the scope of malicious crawlers, and combines the use of the application scenario corresponding to the personal space to conduct correlation analysis of the official website information related to the browsing direction, thereby determining the cause of the anomaly and thus helping to identify the third layer of malicious crawlers. Attached Figure Description

[0045] Figure 1 This is a system principle block diagram of the present invention;

[0046] Figure 2 This is a system principle block diagram of the user analysis unit of the present invention;

[0047] Figure 3 This is a system principle block diagram of the human-machine recognition unit of the present invention;

[0048] Figure 4 This is a system principle block diagram of the secondary verification unit of the present invention;

[0049] Figure 5 This is a system principle block diagram of the space development unit of the present invention;

[0050] Figure 6 This is a system principle block diagram of the actual integration unit of the present invention.

[0051] In the diagram, 1. Crawler Detection Platform; 2. User Analysis Unit; 3. Human-Machine Recognition Unit; 4. Secondary Verification Unit; 5. Space Development Unit; 6. Reality Integration Unit; 7. Big Data Recording Module; 8. Data Analysis Module; 9. Specification Setting Module; 10. Verification Setting Module; 11. CAPTCHA Testing Module; 12. Human-Machine Recognition Module; 13. Initial Identification Module; 14. Usage Setting Module; 15. Threshold Setting Module; 16. Anomaly Trigger Module; 17. Crawler Marking Module; 18. Duplicate Verification Module; 19. Identity Information Verification Module; 20. Personal Space Development Module; 21. Space Management Module; 22. Reality-Related Analysis Module; 23. Anomaly Marking Module; 24. Space Closure Module. Detailed Implementation

[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0053] Please see Figure 1-6 This invention provides a technical solution: a web crawler detection system based on application scenarios, as shown in the attached figure. Figure 1 As shown, it includes a web crawler detection platform 1, which includes a user analysis unit 2, a human-machine recognition unit 3, a secondary verification unit 4, a spatial development unit 5, and a reality integration unit 6.

[0054] As a preferred embodiment, User Analysis Unit 2 is used to categorize the enterprise network into several application scenarios based on their purpose, while simultaneously recording the browsing behavior of test users in the test version, and generating user usage guidelines, as detailed in the appendix. Figure 2 As shown, the user analysis unit 2 includes a big data recording module 7, a data analysis module 8, a specification setting module 9, and a setting verification module 10;

[0055] The big data recording module 7 is used to divide the enterprise network into several application scenarios according to its purpose, and at the same time record the browsing of data by test users in different application scenarios of the enterprise network in the test version and the opening ratio of the test users' personal space within a limited time.

[0056] Data analysis module 8 is used to analyze the browsing behavior of test users in the test version and summarize the similarities in user usage;

[0057] The specification setting module 9 is used to refine enterprise requirements into the framework based on the summarized user usage similarities, forming the enterprise network usage specifications as the initial user usage specifications.

[0058] To ensure that the user usage guidelines better match the actual use of registered users, a verification module 10 is set up to put the initial user usage guidelines into the test version for testing. Based on the feedback from users in the test version, the initial user usage guidelines are adjusted to obtain the revised user usage guidelines.

[0059] As a preferred embodiment, user analysis unit 2 is interfaced with human-machine recognition unit 3, which is used for human-machine recognition, as detailed in the attached document. Figure 3 As shown, the human-machine recognition unit 3 includes a verification code testing module 11, a human-machine recognition module 12, and an initial identification module 13;

[0060] The verification code testing module 11 is used to send random verification codes to registered users on the enterprise network, perform human-machine recognition verification, and send identity verification information to registered users;

[0061] The human-machine recognition module 12 is used to receive feedback from registered users on verification codes and identity verification information, and to verify the correctness of the feedback on verification codes and identity verification information.

[0062] The initial identification module 13 is used to trace the registered user and mark the crawler when errors occur in the returned verification code and identity verification information.

[0063] As explained in detail, when the verification code returned by the initial identification module 13 is incorrect, it traces the registered user and marks the crawler, thus achieving the first layer of malicious crawler identification. When the basic identity verification information returned by the initial identification module 13 is incorrect, it traces the registered user and marks the crawler, thus achieving the second layer of malicious crawler identification. The basic identity verification information includes, but is not limited to, name, gender, and birthday. When the deeper identity verification information returned by the initial identification module 13 is incorrect, it traces the registered user and marks the crawler, thus achieving the third layer of malicious crawler identification. The deeper identity information includes, but is not limited to, facial recognition and ID card number, thus providing convenience for users while effectively protecting user privacy.

[0064] As a preferred embodiment, the human-machine recognition unit 3 is connected to the secondary verification unit 4. The secondary verification unit 4 is used to set an alarm threshold for the number of markings per day in the same application scenario. When the actual number of markings in the application scenario exceeds the alarm threshold, the registered user is marked as abnormal, and human-machine recognition verification is triggered simultaneously. Details are as follows (see attached diagram). Figure 4 As shown, the secondary verification unit 4 includes a setting module 14, a threshold setting module 15, an anomaly triggering module 16, and a crawler marking module 17.

[0065] The setting module 14 is used to set the number of times registered user identity information is verified per day in enterprise network application scenarios, and to permanently mark the number of times registered users are verified per day.

[0066] The threshold setting module 15 is used to set alarm thresholds for the number of times a single day is verified in the same application scenario, and to generate a monthly report of the number of times registered users are permanently marked, providing research data for the progress of enterprises opening up personal spaces;

[0067] The anomaly triggering module 16 is used to mark the registered user as abnormal when the actual number of markings in the application scenario exceeds the alarm threshold, and at the same time trigger human-machine recognition verification.

[0068] The crawler tagging module 17 is used to trace the registered user back to the network and tag the crawler when the human-machine identification verification fails.

[0069] As a preferred embodiment, the secondary verification unit 4 is connected to the space development unit 5, and the space development unit 5 is connected to the human-machine recognition unit 3. The space development unit 5 is used to set the maximum number of repeated browsing times for a single application scenario per day. When the maximum number of repeated browsing times for the same application scenario per day is reached, basic identity information verification is initiated to the registered user. After successful verification, the registered user's personal space is opened in the enterprise network for downloading application scenario information. Details are as follows (see attached). Figure 5As shown, the space development unit 5 includes a duplicate verification module 18, an identity information verification module 19, and a personal space development module 20. The duplicate verification module 18 is used to record the number of times a registered user repeatedly browses the same application scenario and to set the maximum number of times a user repeatedly browses the application scenario in a single day.

[0070] The duplicate verification module 18 is connected to the identity information verification module 19. The identity information verification module 19 is used to send basic identity information verification to the registered user when the number of repeated browsings in the same application scenario reaches the extreme value in a single day.

[0071] The identity information verification module 19 is connected to the personal space development module 20. The personal space development module 20 is used to open the personal space after the registered user passes the identity information verification, so as to download application scenario information.

[0072] As a preferred solution, the space development unit 5 interfaces with the reality integration unit 6, and the reality integration unit 6 interfaces with the secondary verification unit 4. The reality integration unit 6 is used to synchronize the updates of enterprise network application scenarios and registered users' personal space information. Based on the browsing direction of the registered user's application scenario, it detects official website publication information related to the browsing direction, performs correlation analysis, and then performs web crawler identification of registered users and their opened personal spaces. Specifically, see the attached diagram. Figure 6 As shown, the reality integration unit 6 includes a space management module 21, a reality-related analysis module 22, an anomaly marking module 23, and a space closing module 24. The space management module 21 is used to download information according to the application scenario of the registered user, encrypt it with a key, and update the application scenario to the personal space after the application scenario is updated in the enterprise network.

[0073] The space management module 21 is connected to the actual correlation analysis module 22. The actual correlation analysis module 22 is used to record the open time period of personal space, calculate the open percentage of personal space within the limited time, and compare it with the open percentage of the test version. When it exceeds the open percentage of the test version, the application scenarios browsed by registered users are summarized and analyzed. After determining the browsing direction, the official website information related to the browsing direction is detected and correlation analysis is performed. As a detailed explanation, when the application scenario is CET-4 and CET-6 English learning content, the relevant official website information includes but is not limited to information related to CET-4 and CET-6 English exams.

[0074] The actual relevant analysis module 22 is connected to the anomaly marking module 23. When the official website related to the application scenario has not published relevant information, the anomaly marking module 23 marks the registered user as having spatial anomalies and sends a deeper identity information verification to the registered user.

[0075] The anomaly marking module 23 interfaces with the space closing module 24. The space closing module 24 is used to close the personal space when a registered user fails the verification request and add a crawler mark after tracing the registered user.

[0076] First layer of malicious crawler identification: The verification code testing module 11 sends a random verification code to the registered users of the enterprise network for human-machine identification verification. The human-machine identification module 12 receives the feedback of the registered users on the verification code and verifies the correctness of the verification code feedback. When the feedback verification code is incorrect, the initial identification module 13 traces the registered user and marks it as a crawler. When the actual number of markings in the application scenario exceeds the alarm threshold, the anomaly triggering module 16 marks the registered user as an anomaly and triggers human-machine identification verification at the same time. When human-machine identification verification fails, the crawler marking module 17 traces the registered user to the network and marks it as a crawler.

[0077] The second layer of malicious crawler identification: The duplicate verification module 18 records the number of times a registered user repeatedly browses the same application scenario and sets an extreme value for the number of repeated browsings in a single application scenario per day. When the number of repeated browsings in the same application scenario per day reaches the extreme value, the identity information verification module 19 sends basic identity information verification to the registered user, the verification code testing module 11 sends basic identity information verification to the registered user on the enterprise network, the human-machine recognition module 12 receives feedback from the registered user on the basic identity information verification and verifies the correctness of the feedback. When an error occurs in the basic identity information verification, the initial identification module 13 traces the registered user and marks the crawler.

[0078] The third layer of malicious crawler identification: Record the open time period of personal space, calculate the open percentage of personal space within the limited time, and compare it with the open percentage of the test version. When it exceeds the open percentage of the test version, the actual related analysis module 22 summarizes and analyzes the application scenarios browsed by the registered user, determines the browsing direction, detects the official website information related to the browsing direction, and performs correlation analysis. When the official website has not published relevant information related to the application scenario, the anomaly marking module 23 marks the registered user's space as abnormal and sends a deeper identity information verification to the registered user. The human-machine recognition module 12 receives the feedback of the registered user on the deeper identity information verification and verifies the correctness of the feedback. When the deeper identity information verification fails, the initial identification module 13 traces the registered user and marks the crawler.

[0079] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A web crawler detection system based on application scenarios, comprising a crawler detection platform (1), characterized in that: The crawler detection platform (1) includes a user analysis unit (2), a human-machine recognition unit (3), a secondary verification unit (4), a space development unit (5), and a reality integration unit (6). The user analysis unit (2) is used to divide the enterprise network into several application scenarios according to its purpose, and at the same time record the browsing situation of test users in the test version and generate user usage specifications. The user analysis unit (2) is connected to the human-machine recognition unit (3). The human-machine recognition unit (3) is used to perform human-machine recognition. The human-machine recognition unit (3) is connected to the secondary verification unit (4). The secondary verification unit (4) is used to set an alarm threshold for the number of daily verification markings in the same application scenario. When the actual number of markings in the application scenario exceeds the alarm threshold, the registered user is marked abnormally, and human-machine recognition verification is triggered at the same time. The secondary verification unit (4) is connected to the space development unit (5). Unit (5) is connected, and the space development unit (5) is connected to the human-machine recognition unit (3). The space development unit (5) is used to set the extreme value of the number of repeated browsing times of the application scenario in a single day. When the number of repeated browsing times in the same application scenario reaches the extreme value, basic identity information verification is initiated to the registered user. After verification, the registered user's personal space is opened in the enterprise network for downloading application scenario information. The space development unit (5) is connected to the reality integration unit (6), and the reality integration unit (6) is connected to the secondary verification unit (4). The reality integration unit (6) is used to realize the synchronization of the update of enterprise network application scenario and registered user personal space information. According to the browsing direction of the registered user browsing the application scenario, the official website release information related to the browsing direction is detected. After performing correlation analysis, the registered user and the personal space opened by the registered user are crawled and identified. The human-machine recognition unit (3) includes a verification code testing module (11), a human-machine recognition module (12), and an initial identification module (13). The verification code testing module (11) is used to send random verification codes to registered users of the enterprise network, perform human-machine identification verification, and send identity verification information to registered users; The human-machine recognition module (12) is used to receive feedback from registered users on verification codes and identity verification information, and to verify the correctness of the feedback on verification codes and identity verification information. The initial identification module (13) is used to trace the registered user and mark the crawler when the verification code and identity information returned are incorrect. The reality integration unit (6) includes a space management module (21), a reality correlation analysis module (22), an anomaly marking module (23), and a space closure module (24). The space management module (21) is connected to the reality correlation analysis module (22), the reality correlation analysis module (22) is connected to the anomaly marking module (23), and the anomaly marking module (23) is connected to the space closure module (24). The space management module (21) is used to download information according to the application scenario of the registered user, encrypt it with a key, and update the application scenario to the personal space after the application scenario is updated in the enterprise network. The actual correlation analysis module (22) is used to record the open time period of personal space, calculate the open ratio of personal space in the limited time, compare it with the open ratio of the test version, and when it exceeds the open ratio of the test version, summarize and analyze the application scenarios browsed by registered users, determine the browsing direction, detect the official website release information related to the browsing direction, and perform correlation analysis. When the official website related to the application scenario has not published relevant information, the anomaly marking module (23) marks the registered user as having spatial anomalies and sends a deep-level identity information verification to the registered user. The space closing module (24) is used to close the personal space when the registered user fails the verification request, and add a crawler mark after tracing the registered user.

2. The web crawler detection system based on application scenarios according to claim 1, characterized in that: The user analysis unit (2) includes a big data recording module (7), a data analysis module (8), a specification setting module (9), and a setting verification module (10). The big data recording module (7) is used to divide the enterprise network into several application scenarios according to its purpose, and at the same time record the browsing of data by test users in different application scenarios of the enterprise network and the opening ratio of the test users' personal space within a limited time. The data analysis module (8) is used to analyze the browsing behavior of test users in the test version and summarize the similarities used by users. The specification setting module (9) is used to refine the enterprise's needs into the framework based on the summarized user usage similarities, and to form the enterprise network usage specifications as the initial user usage specifications. The setting verification module (10) is used to put the initial user usage specifications into the test version for testing and adjustment, so as to obtain the revised user usage specifications.

3. The web crawler detection system based on application scenarios according to claim 1, characterized in that: The secondary verification unit (4) includes a usage setting module (14), a threshold setting module (15), an anomaly triggering module (16), and a crawler marking module (17). The setting module (14) is used to set the number of times the registered user's identity information is verified per day in the enterprise network application scenario, and to permanently mark the number of times the registered user is verified per day. The threshold setting module (15) is used to set an alarm threshold for the number of times the verification is performed in a single day in the same application scenario, and to generate a monthly report of the number of times the registered user is permanently marked. The abnormal triggering module (16) is used to mark the registered user abnormally when the actual number of markings in the application scenario exceeds the alarm threshold, and at the same time trigger human-machine recognition verification. The crawler marking module (17) is used to trace the registered user through the network and mark the crawler when the human-machine identification verification fails.

4. The web crawler detection system based on application scenarios according to claim 1, characterized in that: The space development unit (5) includes a duplicate verification module (18), an identity information verification module (19), and a personal space development module (20). The duplicate verification module (18) is connected to the identity information verification module (19), and the identity information verification module (19) is connected to the personal space development module (20).

5. The web crawler detection system based on application scenarios according to claim 4, characterized in that: The duplicate verification module (18) is used to record the number of times a registered user repeatedly browses the same application scenario, and to set the maximum number of times a single application scenario is repeatedly browsed in a single day. The identity information verification module (19) is used to send basic identity information verification to the registered user when the number of repeated browsings in the same application scenario reaches an extreme value. The personal space development module (20) is used to open the personal space after the registered user has verified their identity information, so as to download application scenario information.