Data anti-crawling method and system based on large model
By building a multi-dimensional data model and dynamic strategy, combining machine learning algorithms to identify and defend against network crawlers, the problem that existing technology is difficult to effectively defend against complex crawling methods is solved, and data security and integrity protection is achieved.
Patent Information
- Application Number
- CN202510201322.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-30
AI Technical Summary
The existing technology is difficult to fundamentally improve data security, and traditional anti-crawling methods are difficult to cope with complex crawling methods. The existing anti-crawling strategies based on big data analysis have failed to effectively defend against network crawlers.
By building multi-dimensional data models and dynamic adjustment strategies, accurate identification and defense of network crawlers can be achieved, machine learning algorithms can be used to identify crawler behaviors, and simulated data can be generated to deceive crawlers.
It realizes accurate identification and effective defense of network crawlers, protects data security and integrity, destroys the accuracy of crawling data, and achieves the purpose of preventing crawling.
Smart Images

Figure CN120068059A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data security and protection, and more particularly to a data anti-crawling method and system based on a large model. Background Art
[0002] With the rapid development of Internet technology, web crawler technology has been widely used in data collection and analysis. However, this technology is often misused by criminals to illegally collect and use data in commercial systems through automated means, especially commercial systems provided in the form of SaaS services, causing huge losses to data owners. Traditional anti-crawling methods, such as simple IP blocking and CAPTCHA verification, are no longer sufficient to cope with increasingly sophisticated crawling techniques.
[0003] In the prior art, although there are some anti-crawling strategies based on big data analysis, most of them focus on the identification and temporary blocking of crawling behaviors, rather than fundamentally enhancing data security. Therefore, developing a data anti-crawling method based on a large model has become an urgent problem to be solved in the current data security field. Summary of the Invention
[0004] The technical task of the present invention is to address the above deficiencies by providing a data anti-crawling method and system based on a large model, which can accurately identify and effectively defend against web crawlers by constructing a multi-dimensional data model and a dynamic adjustment strategy, thereby protecting the security and integrity of data.
[0005] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0006] A data anti-crawling method based on a large model, the implementation of which includes:
[0007] Constructing and optimizing a multi-dimensional data model based on historical crawling behavior data;
[0008] Accurately identifying crawling behaviors based on machine learning algorithms;
[0009] Invoking the data model to generate simulation data and providing it to the crawler to achieve a deception effect.
[0010] This method constructs a data model that can accurately identify web crawler behaviors. After identifying a web crawler, it constructs a data model that can analyze the crawled data and generate simulation data, and provides the simulation data to the web crawler, thereby destroying the accuracy of the data to achieve the purpose of anti-crawling.
[0011] Furthermore, the specific implementation steps of this method are as follows:
[0012] 1) Collect and analyze historical crawling behavior data;
[0013] 2) Based on big data analysis technology, construct a multi-dimensional data model;
[0014] 3) Construct a model and train it using the real business data in the local database, so that the model can recognize the business data and generate corresponding simulation data according to the real data;
[0015] 4) Monitor network access behavior in real time and use the model to identify the access behavior;
[0016] 5) After identifying the crawling behavior, update the crawler information to the crawler registration library;
[0017] 6) The business system sets up a unified crawler interceptor. After intercepting the crawling requests of the corresponding crawlers according to the crawler information in the crawler registration library, call the model to generate simulation data according to the real business data returned by the system, and then provide the simulation data to the crawlers.
[0018] Further, the historical crawling behavior data includes key information such as request frequency, request time, request source, request content, etc.
[0019] Further, the multi-dimensional data model includes user behavior model, request feature model, access pattern model, etc.
[0020] The present invention also claims to protect a data anti-crawling system based on a large model, including:
[0021] A data analysis module for collecting and analyzing historical crawling behavior data;
[0022] A data model construction module for constructing a multi-dimensional data model;
[0023] A crawling behavior recognition module for accurately recognizing crawling behavior based on machine learning algorithms;
[0024] A simulation data generation module for generating simulation data according to real business data and providing it to the crawlers to achieve a deception effect.
[0025] Further, the process of the system realizing data anti-crawling is as follows:
[0026] 1) Collect and analyze historical crawling behavior data;
[0027] 2) Based on big data analysis technology, construct a multi-dimensional data model;
[0028] 3) Construct a model and train it using the real business data in the local database, so that the model can recognize the business data and generate corresponding simulation data according to the real data;
[0029] 4) Monitor network access behavior in real time and use a model to identify access behavior;
[0030] 5) After identifying the crawling behavior, update the crawler information to the crawler registration library;
[0031] 6) The business system sets up a unified crawler interceptor. After intercepting the crawling requests of the corresponding crawlers according to the crawler information in the crawler registration library, it calls the model to generate simulation data based on the real business data returned by the system, and then provides the simulation data to the crawler.
[0032] Further, the historical crawling behavior data includes key information such as request frequency, request time, request source, and request content.
[0033] Further, the multi-dimensional data model includes a user behavior model, a request feature model, an access pattern model, etc.
[0034] The present invention also claims a data anti-crawling device based on a large model, including at least one memory and at least one processor;
[0035] The at least one memory is used to store machine-readable programs;
[0036] The at least one processor is used to call the machine-readable program to implement the above method.
[0037] The present invention also claims a computer-readable medium, on which computer instructions are stored, and when the computer instructions are executed by a processor, the processor is enabled to implement the above method.
[0038] A data anti-crawling method and system based on a large model of the present invention have the following beneficial effects compared with the prior art:
[0039] The present invention realizes the accurate identification and effective defense of network crawlers by constructing a multi-dimensional data model and a dynamic adjustment strategy. By constructing a data model that can accurately identify network crawler behaviors, and after identifying the network crawlers, by constructing a data model that can analyze the crawled data and generate simulation data, and providing the simulation data to the network crawlers, thereby destroying the accuracy of the data to achieve the purpose of anti-crawling and protecting the security and integrity of the data. Description of the Drawings
[0040] Figure 1 is a flowchart of a data anti-crawling method based on a large model provided by an embodiment of the present invention. Detailed Embodiments
[0041] The present invention will be further described below with reference to specific embodiments.
[0042] An embodiment of the present invention provides a method for preventing data crawling based on a large model. The implementation of the method includes:
[0043] Constructing and optimizing a multi-dimensional data model, constructing a multi-dimensional data model based on historical crawling behavior data;
[0044] Accurately identifying crawling behaviors based on machine learning algorithms;
[0045] Invoking the data model to generate simulation data and providing it to the crawler to achieve a deception effect.
[0046] Combined with the attached Figure 1 As shown, the specific implementation process of this method is as follows:
[0047] 1. Collect and analyze historical crawling behavior data, including key information such as request frequency, request time, request source, and request content.
[0048] 2. Based on big data analysis technology, construct a multi-dimensional data model, including user behavior models, request feature models, access pattern models, etc.
[0049] 3. Construct a model and train it using real business data in the local database, so that the model can identify business data and generate corresponding simulation data according to the real data.
[0050] 4. Real-time monitor network access behaviors and use the model to identify access behaviors.
[0051] 5. After identifying the crawling behavior, update the crawler information to the crawler registration library.
[0052] 6. The business system sets up a unified crawler interceptor. After intercepting the crawling requests of the corresponding crawlers according to the crawler information in the crawler registration library, call the model to generate simulation data according to the real business data returned by the system, and then provide the simulation data to the crawler.
[0053] This method effectively prevents the illegal acquisition of sensitive or key data by network crawlers by constructing a complex data model and a dynamic defense mechanism, protecting the rights and interests of data owners.
[0054] An embodiment of the present invention also provides a system for preventing data crawling based on a large model, including:
[0055] A data analysis module for collecting and analyzing historical crawling behavior data;
[0056] A data model construction module for constructing a multi-dimensional data model;
[0057] A crawling behavior identification module for accurately identifying crawling behaviors based on machine learning algorithms;
[0058] A simulation data generation module, which is used to generate simulation data according to real business data and provide it to the crawler to achieve a deception effect.
[0059] The process of the system for preventing data from being crawled is as follows:
[0060] 1. The data analysis module collects and analyzes historical crawling behavior data, including key information such as request frequency, request time, request source, and request content.
[0061] 2. Based on big data analysis technology, construct multi-dimensional data models, including user behavior models, request feature models, access pattern models, etc.
[0062] 3. Construct a model and use the real business data in the local database for training, so that the model can identify business data and generate corresponding simulation data according to the real data.
[0063] 4. Real-time monitor network access behavior and use the model to identify the access behavior.
[0064] 5. After identifying the crawling behavior, update the crawler information to the crawler registration library.
[0065] 6. The business system sets up a unified crawler interceptor. After intercepting the crawling requests of the corresponding crawlers according to the crawler information in the crawler registration library, call the model to generate simulation data according to the real business data returned by the system, and then provide the simulation data to the crawler.
[0066] An embodiment of the present invention also provides a data anti-crawling device based on a large model, including at least one memory and at least one processor;
[0067] The at least one memory is used to store machine-readable programs;
[0068] The at least one processor is used to call the machine-readable program to implement the data anti-crawling method based on the large model described in the above embodiment.
[0069] An embodiment of the present invention also provides a computer-readable medium. Computer instructions are stored on the computer-readable medium. When the computer instructions are executed by a processor, the processor executes the data anti-crawling method based on the large model described in the above embodiment. Specifically, a system or device equipped with a storage medium can be provided. Software program codes for implementing the functions of any one of the above embodiments are stored on the storage medium, and the computer (or CPU or MPU) of the system or device reads and executes the program codes stored on the storage medium.
[0070] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments. Therefore, the program code and the storage medium storing the program code constitute a part of the present invention.
[0071] Examples of the storage medium for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as a CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer via a communication network.
[0072] Furthermore, it should be clear that not only can the functions of any one of the above embodiments be realized by executing the program code read by the computer, but also by causing an operating system or the like operating on the computer based on the instructions of the program code to complete part or all of the actual operations.
[0073] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in an expansion board inserted into the computer or into the memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU or the like installed on the expansion board or the expansion unit is caused to execute part or all of the actual operations, thereby realizing the functions of any one of the above embodiments.
[0074] The present invention has been described in detail above with reference to the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above-mentioned multiple embodiments, those skilled in the art can know that more embodiments of the present invention can be obtained by combining the code review means in the above different embodiments, and these embodiments are also within the protection scope of the present invention.
Claims
1. A data anti-crawling method based on a large model, characterized in that: The implementation of this method includes: Construction and optimization of multi-dimensional data models, building multi-dimensional data models based on historical crawling behavior data; Accurately identify crawling behaviors based on machine learning algorithms; Call the data model to generate simulation data and provide it to the crawler to achieve a deceptive effect.
2. According to the large model-based data anti-crawling method of claim 1, it is characterized in that: The specific implementation steps of this method are as follows: 1) Collect and analyze historical crawling behavior data; 2) Build a multi-dimensional data model based on big data analysis technology; 3) Build a model and use real business data from the local database for training, so that the model can recognize business data and generate corresponding simulation data based on real data; 4) Monitor network access behavior in real time and use models to identify access behavior; 5) After identifying the crawling behavior, update the crawler information to the crawler registration library; 6) The business system sets up a unified crawler interceptor. After intercepting the crawling request of the corresponding crawler according to the crawler information in the crawler registration library, the model is called to generate simulation data according to the real business data returned by the system, and then the simulation data is provided to the crawler.
3. A data anti-crawling method based on a large model according to claim 2, characterized in that: The historical crawling behavior data includes request frequency, request time, request source, and request content information.
4. According to the large model-based data anti-crawling method of claim 2, it is characterized in that: The multi-dimensional data model includes a user behavior model, a request feature model, and an access pattern model.
5. A data anti-crawling system based on a large model, characterized in that: include: Data analysis module, used to collect and analyze historical crawling behavior data; Data model building module, used to build multi-dimensional data models; Crawling behavior identification module, which accurately identifies crawling behavior based on machine learning algorithms; The simulation data generation module is used to generate simulation data based on real business data and provide it to the crawler to achieve a deceptive effect.
6. A data anti-crawling system based on a large model according to claim 5, characterized in that: The process of the system to achieve data anti-crawling is as follows: 1) Collect and analyze historical crawling behavior data; 2) Build a multi-dimensional data model based on big data analysis technology; 3) Build a model and use real business data from the local database for training, so that the model can recognize business data and generate corresponding simulation data based on real data; 4) Monitor network access behavior in real time and use models to identify access behavior; 5) After identifying the crawling behavior, update the crawler information to the crawler registration library; 6) The business system sets up a unified crawler interceptor. After intercepting the crawling request of the corresponding crawler according to the crawler information in the crawler registration library, the model is called to generate simulation data according to the real business data returned by the system, and then the simulation data is provided to the crawler.
7. A data anti-crawling system based on a large model according to claim 6, characterized in that: The historical crawling behavior data includes request frequency, request time, request source, and request content information.
8. The data anti-crawling system based on a large model according to claim 6, characterized in that: The multi-dimensional data model includes a user behavior model, a request feature model, and an access pattern model.
9. A data anti-crawling device based on a large model, characterized in that: comprising at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is used to call the machine-readable program to implement the method described in any one of claims 1 to 4.
10. A computer-readable medium, characterized in that The computer readable medium stores computer instructions, which, when executed by a processor, enable the processor to implement the method according to any one of claims 1 to 4.