Method and apparatus for anti-tampering of asynchronous interfaces
By receiving and analyzing the asynchronous interface request logs of the browser, the system identifies and prevents advanced crawlers from illegally acquiring data, thus addressing the shortcomings of existing anti-scraping technologies and achieving security protection for asynchronous interfaces and improved user experience.
Patent Information
- Application Number
- CN202211481745.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-24
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2042-11-24
AI Technical Summary
Existing technologies cannot effectively prevent advanced web crawlers from illegally obtaining core data of network products through asynchronous interfaces, and frequency control based on user IP addresses can easily harm legitimate users.
By receiving asynchronous interface data requests from the browser and front-end request logs, recording back-end request logs, determining a risk score based on the front-end and back-end logs, and outputting a verification code box for human verification when the risk score is higher than the threshold, the data request is stopped if the verification fails.
It effectively identifies and protects against advanced web crawlers, reduces false positives for legitimate users, and improves user experience and data security.
Smart Images

Figure CN116436626B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the field of computers, in particular to the field of network security, and specifically to a method and device for preventing asynchronous interface scraping. BACKGROUND
[0002] An asynchronous interface refers to sending a request without waiting for a return, and at any time a next request can be sent, i.e. without waiting. In network products, in order to improve the loading speed of a webpage, data is often returned to the browser page through an asynchronous interface. Black production simulates normal users to initiate an HTTP request to the asynchronous interface through a web crawler and the like, and illegally obtains core data of the network product.
[0003] If no intervention is made, the asynchronous interface data is leaked by black production scraping. The prior art simply limits and blocks primary crawlers, but advanced crawlers can easily obtain data in the asynchronous interface. If frequency control is performed on the user IP dimension, normal users of public export IPs such as schools and airports are very easy to be mistaken. SUMMARY
[0004] The present disclosure provides a method and device for preventing asynchronous interface scraping, equipment, a storage medium and a computer program product.
[0005] According to a first aspect of the present disclosure, a method for preventing asynchronous interface scraping is provided, comprising: receiving an asynchronous interface data request and a front-end request log from a browser; recording a back-end request log according to the asynchronous interface data request before sending a response message of the asynchronous interface data request to the browser; determining a risk score based on the front-end request log and the back-end request log; outputting a verification code box for human-computer verification if the risk score is higher than a predetermined threshold; and stopping sending the response message of the asynchronous interface data request to the browser if the human-computer verification fails.
[0006] According to a second aspect of the present disclosure, a device for preventing asynchronous interface scraping is provided, comprising: a receiving unit configured to receive an asynchronous interface data request and a front-end request log from a browser; a recording unit configured to record a back-end request log according to the asynchronous interface data request before sending a response message of the asynchronous interface data request to the browser; an identification unit configured to determine a risk score based on the front-end request log and the back-end request log; a verification unit configured to output a verification code box for human-computer verification if the risk score is higher than a predetermined threshold; and a disabling unit configured to stop sending the response message of the asynchronous interface data request to the browser if the human-computer verification fails.
[0007] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of the first aspect.
[0008] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the method of the first aspect.
[0009] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method of the first aspect.
[0010] Embodiments of the present disclosure provide a method and device for preventing grabbing of an asynchronous interface, which replays user behavior based on front-end and back-end logs requested by the user. For a normal user, a front-end and a back-end each generate a log when the user browses a webpage once. By aggregating the logs, crawler traffic analysis is performed to obtain a risk score. For a request with high risk, human-computer verification is performed to identify whether it is black production. Only a browser that passes the human-computer verification can normally obtain data of the asynchronous interface.
[0011] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:
[0013] Figure 1 is an exemplary system architecture diagram to which an embodiment of the present disclosure can be applied;
[0014] Figure 2 is a flowchart of an embodiment of the method for preventing grabbing of an asynchronous interface according to the present disclosure;
[0015] Figure 3 is a schematic diagram of an application scenario of the method for preventing grabbing of an asynchronous interface according to the present disclosure;
[0016] Figure 4 is a flowchart of another embodiment of the method for preventing grabbing of an asynchronous interface according to the present disclosure;
[0017] Figure 5 is a schematic diagram of another application scenario of the method for preventing grabbing of an asynchronous interface according to the present disclosure;
[0018] Figure 6 is a structural diagram of one embodiment of the anti-grabbing device according to the asynchronous interface of the present disclosure;
[0019] Figure 7 is a structural diagram of a computer system of an electronic device suitable for implementing embodiments of the present disclosure. DETAILED DESCRIPTION
[0020] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, which should be considered in their context only. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, descriptions of well-known functions and structures are omitted in the following description for the sake of clarity and conciseness.
[0021] Figure 1 An exemplary system architecture 100 to which the anti-grabbing method of the asynchronous interface or the anti-grabbing device of the asynchronous interface according to the present disclosure can be applied is shown.
[0022] As shown in Figure 1 , the system architecture 100 can include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is a medium to provide a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0023] A user can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the terminal devices 101, 102, 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0024] The terminal devices 101, 102, and 103 can be hardware or software. When the terminal devices 101, 102, and 103 are hardware, they can be various electronic devices with a display screen and supporting web browsing, including but not limited to a smart phone, a tablet computer, an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, a desktop computer, and the like. When the terminal devices 101, 102, and 103 are software, they can be installed in the above-listed electronic devices. They can be implemented as multiple software or software modules (for example, to provide distributed services) or as a single software or software module. No specific limitation is made herein.
[0025] The server 105 can be a server providing various services, for example, a background web server supporting a web page displayed on the terminal devices 101, 102, and 103. The background web server can analyze and process received data such as an asynchronous interface data request, and feed back the processing result (for example, asynchronous interface data) to the terminal device. If the server 105 analyzes that the asynchronous interface data request is simulated by black production rather than manually issued through a browser, no asynchronous interface data is returned.
[0026] It should be noted that the server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules (for example, multiple software or software modules for providing distributed services) or as a single software or software module. No specific limitation is made herein. The server can also be a server of a distributed system or a server combined with a blockchain. The server can also be a cloud server or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0027] It should be noted that the method for preventing asynchronous interface grabbing provided by the embodiments of the present disclosure is generally executed by the server 105, and accordingly, the device for preventing asynchronous interface grabbing is generally arranged in the server 105.
[0028] It should be understood that Figure 1 The number of terminal devices, networks, and servers in
[0029] With reference back to Figure 2, shows a flow 200 of one embodiment of the anti-scraping method of the asynchronous interface according to the present disclosure. The anti-scraping method of the asynchronous interface comprises the following steps:
[0030] Step 201, receiving an asynchronous interface data request and a front-end request log from a browser.
[0031] In this embodiment, the execution subject of the anti-scraping method of the asynchronous interface (for example Figure 1 The server shown can receive an asynchronous interface data request from a browser, in addition to which the browser will send a log request, i.e. a front-end request log, to the crawler identification service of the server at the same time as initiating the asynchronous interface data request. The server runs multiple tasks, including the crawler identification service, the verification code service, etc. The crawler identification service mainly performs the task of identifying crawlers, analyzes by receiving the front-end request log and the back-end request log to determine whether the user is a crawler. Then the user suspected to be a crawler is handed over to the verification code service for further verification, and if it is a crawler, it is not allowed to obtain asynchronous interface data.
[0032] One of the prerequisites of the present application is that the browser can send a front-end request log at the same time as sending an asynchronous interface data request. The front-end request log can include some attribute information of the browser, such as IP address, Cookie, UserAgent, Referer, etc. The front-end request log can also include time information.
[0033] Even if the browser does not support the function of sending a front-end request log at the same time as sending an asynchronous interface data request, it will not cause a false judgment. Instead, it will analyze the risk score according to the back-end request log, and finally it can also be verified through human-computer interaction. As long as it is a manual operation, it will not be mistaken for a crawler. If the browser supports the function of sending a front-end request log at the same time as sending an asynchronous interface data request, the probability of human-computer interaction can be reduced, and the user experience can be improved.
[0034] Step 202, recording a back-end request log according to the asynchronous interface data request before sending a response message to the browser.
[0035] In this embodiment, the server forwards the request to the crawler identification service, i.e. the back-end request log, when processing the request in the traffic access layer after the browser initiates the asynchronous interface data request on the page. The information recorded in the back-end request log is the same as that recorded in the front-end request log, only with a time difference.
[0036] Step 203, determining a risk score based on the front-end request log and the back-end request log.
[0037] In this embodiment, based on the front-end and back-end logs of user requests, the user behavior is played back. For a normal user, a web page is browsed once, and the front-end and back-end each generates a log. Through aggregation of the logs, crawler traffic analysis is performed, a risk model identification operator is executed, and finally a risk score is obtained, which is issued to the business layer. The identification operator of the risk model can include browsing duration, browsing frequency, browsing time period, etc. The scores of each identification operator are weighted to obtain the final risk score. For example, the longer the browsing duration, the higher the risk score. The higher the browsing frequency, the higher the risk score. The more the browsing time period does not conform to the work and rest time of ordinary people (for example, 2-5 a.m.), the higher the risk score.
[0038] In step 204, if the risk score is higher than a predetermined threshold, a verification code box is output for human-computer verification.
[0039] In this embodiment, if the risk score is too high, further human-computer interaction is required to verify whether it is a human operation. The existing technology common verification code can be output to let the user operate according to the requirements to verify whether it is a human operation. For example, an askew picture is output, and the user is required to slide the operation block to rotate the picture to the correct angle. The conventional treatment means is to jump to an independent page for human-computer verification, and after verification, the source page is returned. Since the asynchronous interface is only responsible for initiating a data request, and the specific rendering work is done by the synchronous page. If the verification is passed and the asynchronous interface is returned, the user cannot normally browse the web page, affecting the user experience.
[0040] Therefore, for the treatment means of the asynchronous interface, after the crawler request is identified, the front-end displays a mask effect and pops up a verification code box, requiring the user to perform human-computer verification. Before verification, other operations will be prohibited.
[0041] In step 205, if the human-computer verification fails, the sending of the response message of the asynchronous interface data request to the browser is stopped.
[0042] In this embodiment, if the human-computer verification fails, it means that a crawler is detected, and the sending of the response message of the asynchronous interface data request to the browser is not allowed.
[0043] The conventional network product can only judge whether the request is a crawler through single information such as Cookie, UserAgent, Referer, etc. This is a common crawler identification means based on the back-end request. In this case, normal users are easily mistaken, affecting the user experience. Therefore, the present application proposes an innovative solution based on the front-end and back-end logs of user requests to play back the user behavior. For a normal user, a web page is browsed once, and the front-end and back-end each generates a log. Through aggregation of the logs, crawler traffic analysis is performed.
[0044] In some optional implementations of the present embodiment, the determining of the risk score based on the front-end request log and the back-end request log comprises: if the difference between the number of the back-end request log and the number of the front-end request log is greater than a predetermined value, setting the risk score as a highest risk value, wherein the highest risk value is greater than the predetermined threshold. A real browser sends a front-end request log, so the number of the corresponding front-end request log and the number of the back-end request log should be consistent, and even if there is data loss, the difference should not be large. However, when a crawler sends an asynchronous interface data request, no front-end request log is simulated or it is impossible to completely simulate the sending of a front-end request log, which may result in a large difference between the number of the front-end request log and the number of the back-end request log. If the difference between the number of the back-end request log and the number of the front-end request log is greater than a predetermined value, it means that the request is most likely sent by a crawler, so the highest risk value is set. The number of front-end request logs and back-end request logs can quickly identify crawlers, reduce processing delay, and improve user experience.
[0045] In some optional implementations of the present embodiment, the determining of the risk score based on the front-end request log and the back-end request log comprises: determining the duration and frequency of user access to the website based on the front-end request log and the back-end request log; if the duration is greater than a predetermined duration threshold or the frequency is greater than a predetermined frequency threshold, setting the risk score as a score greater than the predetermined threshold. According to the time of the front-end request log and the back-end request log, the duration of user access to the website can be determined. According to the number of the front-end request log and the back-end request log, the frequency of user access to the website can be calculated. If the duration is greater than a predetermined duration threshold or the frequency is greater than a predetermined frequency threshold, it means that the page browsing is not performed by a human being, but by a crawler. Crawlers can be quickly and accurately identified without false positives.
[0046] In some optional implementations of the present embodiment, the determining of the risk score based on the front-end request log and the back-end request log comprises: determining the duration, time period, and frequency of user access to the website based on the front-end request log and the back-end request log; and inputting the duration, time period, and frequency into a pre-trained risk prediction model to obtain the risk score. The training samples can include front-end request logs and back-end request logs sent by crawlers and front-end request logs and back-end request logs sent by humans. Through the risk prediction model, the characteristics of the duration, time period, and frequency of crawler access to the website are learned, and the risk prediction model is trained. When applied, the risk prediction model comprehensively scores the duration, time period, and frequency of user access to the website to obtain the risk score. The risk prediction model can analyze crawlers in cases where the duration, time period, and frequency alone cannot clearly identify crawlers, thereby improving the accuracy and efficiency of identification.
[0047] In some optional implementations of the present embodiment, the method further comprises: if the human-computer verification succeeds, adding the browser to a white list and setting an exempt time, and no longer performing human-computer verification within the exempt time. In this way, the number of human-computer verifications can be reduced, the browsing efficiency can be improved, and the user experience can be improved.
[0048] Continuing to refer to Figure 3 , Figure 3 is a schematic diagram of an application scenario of the method for preventing grabbing of an asynchronous interface according to the present embodiment. In the application scenario of Figure 3 , the following implementation steps are implemented:
[0049] 1. A user browses a page, and the page initiates a data request of an asynchronous interface. At the same time when the data request of the asynchronous interface is initiated, a log request is sent to a crawler identification service, that is, a front-end request log.
[0050] 2. After the server page initiates the data request of the asynchronous interface, the request is forwarded to the crawler identification service at a traffic access layer, that is, a back-end request log.
[0051] 3. The crawler identification service aggregates the front-end request log and the back-end request log, and sends a risk score to a business layer after analysis.
[0052] 4. The business layer perceives that the risk score exceeds a set threshold, renders a verification code component, and prompts the user to perform human-computer verification.
[0053] 5. If the human-computer verification passes, the user is added to a white list, and exempted from verification for a period of time.
[0054] 6. If the human-computer verification fails, the human-computer verification is repeatedly performed until the human-computer verification passes, and then the data of the asynchronous interface is returned.
[0055] Further reference is made to Figure 4 , which shows a flow 400 of still another embodiment of the method for preventing grabbing of an asynchronous interface. The flow 400 of the method for preventing grabbing of an asynchronous interface includes the following steps:
[0056] Step 401, in response to detecting an access request of a browser, sending an admission identifier to the browser to make the browser report attribute information.
[0057] In the present embodiment, an execution subject (for example, a server as shown in Figure 1 may receive an access request from a browser. The first access request can be a synchronous interface data request, for example, a synchronous request for accessing a webpage through a url. An admission identifier is sent to the browser to make the browser report attribute information. As Figure 5The flow of authenticating the browser before the browser sends the asynchronous interface data request is shown.
[0058] At step 402, in response to receiving the attribute information sent by the browser, a communication identifier is generated according to the attribute information and returned to the browser, for the browser to carry when sending the asynchronous interface data request.
[0059] In this embodiment, the attribute information can include IP, device identifier, Cookie, UserAgent, Referer, etc. Different browsers can be assigned a communication identifier, which can be randomly assigned a non-repeating communication identifier, or generated according to a predetermined encoding rule based on the attribute information, for example, the communication identifier can be: IP+device identifier+Cookie. The server records the correspondence between the communication identifier already assigned and the browser, and when receiving the asynchronous interface data request sent by the browser, checks whether the browser carries the communication identifier.
[0060] At step 403, the asynchronous interface data request and the front-end request log from the browser are received.
[0061] Step 403 is basically the same as step 201, and thus will not be described again. The difference is that this time the asynchronous interface data request should carry the communication identifier assigned in step 402.
[0062] At step 404, before sending the response message to the browser for the asynchronous interface data request, the back-end request log is recorded according to the asynchronous interface data request.
[0063] Step 404 is basically the same as step 202, and thus will not be described again.
[0064] At step 405, if the browser does not carry the communication identifier when sending the asynchronous interface data request, the risk score is set to the highest risk value, otherwise the risk score is determined based on the front-end request log and the back-end request log.
[0065] In this embodiment, the crawler is first identified according to the communication identifier, because the asynchronous interface data request sent by the crawler cannot carry the communication identifier, while the asynchronous interface data request sent by the browser is carried by the communication identifier. If the asynchronous interface data request does not carry the communication identifier, it proves that the request may be fake by the crawler. The risk score can be directly determined to be the highest risk value based on the front-end request log and the back-end request log. If the communication identifier is carried, the risk score needs to be determined according to the same process as step 203.
[0066] At step 406, if the risk score is higher than a predetermined threshold, an authentication code box is output for human-computer verification.
[0067] Step 407, if the man-machine verification fails, stop sending the response message of the asynchronous interface data request to the browser.
[0068] Steps 406-407 are basically the same as steps 204-205, and thus will not be described again.
[0069] As can be seen from Figure 4 , compared with the embodiment corresponding to Figure 2 , the flow 400 of the anti-scraping method of the asynchronous interface in the embodiment embodies the step of access authentication of the browser. Thus, the scheme described in the embodiment can identify the crawler before analyzing the front-end request log and the back-end request log, thereby reducing the identification process, reducing the time delay of data request, improving the accuracy of identification, and improving the user experience.
[0070] Further referring to Figure 6 , as an implementation of the method shown in the above figures, the disclosure provides an embodiment of an anti-scraping device of an asynchronous interface, which corresponds to the method embodiment shown in Figure 2 , and the device can be specifically applied to various electronic devices.
[0071] As shown in Figure 6 , the anti-scraping device 600 of the asynchronous interface in the embodiment includes a receiving unit 601, a recording unit 602, an identification unit 603, a verification unit 604, and a disabling unit 605. The receiving unit 601 is configured to receive an asynchronous interface data request and a front-end request log from a browser; the recording unit 602 is configured to record a back-end request log according to the asynchronous interface data request before sending a response message of the asynchronous interface data request to the browser; the identification unit 603 is configured to determine a risk score based on the front-end request log and the back-end request log; the verification unit 604 is configured to output a verification code box for man-machine verification if the risk score is higher than a predetermined threshold; and the disabling unit 605 is configured to stop sending the response message of the asynchronous interface data request to the browser if the man-machine verification fails.
[0072] In the embodiment, the specific processing of the receiving unit 601, the recording unit 602, the identification unit 603, the verification unit 604, and the disabling unit 605 of the anti-scraping device 600 of the asynchronous interface can refer to steps 201, 202, 203, 204, and 205 in the corresponding embodiment. Figure 2
[0073] In some optional implementation manners of the embodiment, the identification unit 603 is further configured to set the risk score to a highest risk value if the difference between the number of the back-end request log and the front-end request log is greater than a predetermined value, wherein the highest risk value is greater than the predetermined threshold.
[0074] In some optional implementations of the embodiment, the identifying unit 603 is further configured to determine the duration and frequency of the user accessing the website based on the front-end request log and the back-end request log; and set the risk score to a score greater than the predetermined threshold if the duration is greater than a predetermined duration threshold or the frequency is greater than a predetermined frequency threshold.
[0075] In some optional implementations of the embodiment, the identifying unit 603 is further configured to determine the duration, time period and frequency of the user accessing the website based on the front-end request log and the back-end request log; and input the duration, time period and frequency into a pre-trained risk prediction model to obtain the risk score.
[0076] In some optional implementations of the embodiment, the verifying unit 604 is further configured to add the browser to a white list and set an exempt time during which the bot check is no longer performed if the bot check is successful.
[0077] In some optional implementations of the embodiment, the apparatus 600 further includes an admission unit (not shown in the figure) configured to send an admission identifier to the browser to make the browser report attribute information in response to detecting the access request of the browser; and generate a communication identifier according to the attribute information and return the communication identifier to the browser in response to receiving the attribute information sent by the browser, so as to be carried by the browser when sending the asynchronous interface data request.
[0078] In some optional implementations of the embodiment, the identifying unit 603 is further configured to set the risk score to a highest risk value if the communication identifier is not carried by the browser when sending the asynchronous interface data request.
[0079] In the technical solution of the disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solution comply with relevant laws and regulations and do not violate public order and good customs.
[0080] According to the embodiments of the disclosure, the disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0081] An electronic device includes at least one processor and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the flow 200 or 400.
[0082] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to perform the method of flow 200 or 400.
[0083] A computer program product comprising a computer program which, when executed by a processor, implements the method of root flow 200 or 400.
[0084] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present disclosure described and / or claimed in this document.
[0085] As shown in Figure 7 The device 700 includes a computing unit 701 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 702 or a computer program loaded into a random access memory (RAM) 703 from a storage unit 708. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0086] Various components in the device 700 are connected to the I / O interface 705, including an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0087] The computing unit 701 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs various methods and processes described above, such as the method of asynchronous interface anti-pinch. For example, in some embodiments, the method of asynchronous interface anti-pinch can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded onto the RAM 703 and executed by the computing unit 701, one or more steps of the method of asynchronous interface anti-pinch described above can be performed. Alternatively, in other embodiments, the computing unit 701 can be configured to perform the method of asynchronous interface anti-pinch by any other suitable means, such as by means of firmware.
[0088] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0089] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0090] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0091] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0092] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0093] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0094] It should be understood that the various forms of flow shown above can be used to reorder, add, or remove steps. For example, the steps described in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure are achieved, which is not limited herein.
[0095] The specific implementation described above does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for anti-scraping of an asynchronous interface, comprising: receiving an asynchronous interface data request and a front-end request log from a browser, wherein the browser is capable of sending one more front-end request logs while sending the asynchronous interface data request; recording a back-end request log according to the asynchronous interface data request before sending a response message of the asynchronous interface data request to the browser; determining a risk score based on the front-end request log and the back-end request log, comprising: performing a crawler traffic analysis by aggregating the front-end request log and the back-end request log, executing a risk model identification operator to obtain the risk score, wherein the risk model identification operator comprises at least one of: browsing duration, browsing frequency, browsing time period; if the risk score is higher than a predetermined threshold, outputting a verification code box for human-computer verification; if the human-computer verification fails, stopping sending the response message of the asynchronous interface data request to the browser.
2. The method of claim 1, wherein, The determining of the risk score based on the front-end request log and the back-end request log comprises: if a difference between a number of the back-end request log and a number of the front-end request log is greater than a predetermined value, setting the risk score as a highest risk value, wherein the highest risk value is greater than the predetermined threshold.
3. The method of claim 1, wherein, The determining of the risk score based on the front-end request log and the back-end request log comprises: determining a duration and a frequency of a user accessing a website based on the front-end request log and the back-end request log; if the duration is greater than a predetermined duration threshold or the frequency is greater than a predetermined frequency threshold, setting the risk score as a score greater than the predetermined threshold.
4. The method of claim 1, wherein, The determining of the risk score based on the front-end request log and the back-end request log comprises: determining a duration, a time period, and a frequency of a user accessing a website based on the front-end request log and the back-end request log; inputting the duration, the time period, and the frequency into a pre-trained risk prediction model to obtain the risk score.
5. The method of claim 1, wherein, The method further comprises: if the human-computer verification succeeds, adding the browser to a white list and setting an inspection-free time, and not performing the human-computer verification again within the inspection-free time.
6. The method of claim 1, wherein, The method further comprises: in response to detecting an access request of the browser, sending an access identification to the browser to make the browser report attribute information; in response to receiving the attribute information sent by the browser, generating a communication identification according to the attribute information and returning the communication identification to the browser, so as to be carried by the browser when sending the asynchronous interface data request.
7. The method of claim 6, wherein, The determining of the risk score based on the front-end request log and the back-end request log comprises: if the communication identification is not carried by the browser when sending the asynchronous interface data request, setting the risk score as a highest risk value. 8.An apparatus for anti-scraping of an asynchronous interface, comprising: a receiving unit configured to receive an asynchronous interface data request and a front-end request log from a browser, wherein the browser is capable of sending one more front-end request logs while sending the asynchronous interface data request; a recording unit configured to record a back-end request log according to the asynchronous interface data request before sending a response message of the asynchronous interface data request to the browser; The identifying unit is configured to determine a risk score based on the front-end request log and the back-end request log, including: performing a crawler traffic analysis by aggregating the front-end request log and the back-end request log, executing a risk model identification operator to obtain a risk score, wherein the risk model identification operator includes at least one of: a browsing duration, a browsing frequency, a browsing time period; The verification unit is configured to output a verification code box for human-computer verification if the risk score is higher than a predetermined threshold value; The disabling unit is configured to stop sending a response message of an asynchronous interface data request to the browser if the human-computer verification fails.
9. The apparatus of claim 8, wherein, The identifying unit is further configured to: If the difference between the number of the back-end request log and the front-end request log is greater than a predetermined value, set the risk score to a highest risk value, wherein the highest risk value is greater than the predetermined threshold value.
10. The apparatus of claim 8, wherein, The identifying unit is further configured to: Determine a duration and a frequency of a user accessing a website based on the front-end request log and the back-end request log; If the duration is greater than a predetermined duration threshold value or the frequency is greater than a predetermined frequency threshold value, set the risk score to a score greater than the predetermined threshold value.
11. The apparatus of claim 8, wherein, The identifying unit is further configured to: Determine a duration, a time period, and a frequency of a user accessing a website based on the front-end request log and the back-end request log; Input the duration, the time period, and the frequency into a pre-trained risk prediction model to obtain a risk score.
12. The apparatus of claim 8, wherein, The verification unit is further configured to: If the human-computer verification is successful, add the browser to a white list and set an exempt time, and no longer perform human-computer verification within the exempt time.
13. The apparatus of claim 8, wherein, The device further includes an admission unit configured to: In response to detecting an access request of the browser, send an admission identifier to the browser to make the browser report attribute information; In response to receiving the attribute information sent by the browser, generate a communication identifier according to the attribute information and return it to the browser, so as to be carried when the browser sends an asynchronous interface data request.
14. The apparatus of claim 13, wherein, The identifying unit is further configured to: If the communication identifier is not carried when the browser sends an asynchronous interface data request, set the risk score to a highest risk value.
15. An electronic device, comprising: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.
16. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to make the computer execute the method according to any one of claims 1-7.
17. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-7.
17. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-7.
Citation Information
Patent Citations
Crawler recognition method and system based on user behavior burial point
CN108712426A
Sensitive data interface crawler identification method and device
CN113821754A