High-concurrency data request processing method, device, electronic device and storage medium

By using the traffic timing prediction model and the sentinel flow control framework, combined with heuristic interception rules, the problem of insufficient adaptability and robustness of high concurrent data processing in the existing technology is solved, and more efficient and secure data processing is achieved.

CN116094945BActive Publication Date: 2025-05-13GRG BANKING IT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310115575.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-13
Publication Date
2025-05-13
Estimated Expiration
2043-02-13

AI Technical Summary

Technical Problem

The prior art lacks a preprocessing mechanism when processing large amounts of data with high concurrency, resulting in poor adaptability and robustness of the system and inability to effectively process large amounts of data with high concurrency.

Method used

By obtaining the traffic timing prediction model to predict the data request volume at the next moment, the concurrent data request threshold of the sentinel traffic control framework is determined, and unsafe user data requests are intercepted through heuristic interception rules.

Benefits of technology

Improves the adaptability and robustness of high concurrent data processing, while improving the security of data processing, preventing the system from crashing when facing large traffic shocks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116094945B_ABST
    Figure CN116094945B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, electronic device and storage medium for processing high concurrent data requests, which belongs to the field of data processing. The processing method includes: obtaining the predicted data request amount at the next moment output by the traffic timing prediction model based on the user data request amount at the current moment; determining the concurrent data request threshold of the sentinel traffic control framework at the next moment based on the predicted data request amount at the next moment; receiving and responding to the concurrent data request of the trusted user at the next moment when it is determined that the concurrent data request amount of the trusted user at the next moment is less than or equal to the concurrent data request threshold of the sentinel traffic control framework at the next moment. This method can enhance the adaptability, security and robustness of high concurrent data processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular to a method, device, electronic device and storage medium for processing high-concurrency data requests. Background Art

[0002] With the widespread dissemination of distributed systems, computer systems are increasingly faced with large amounts of data requests. Currently, high-concurrency processing technology is usually used on computer systems to cope with large amounts of data requests.

[0003] At present, large amounts of highly concurrent data are processed by filtering or restricting access through the Zuul gateway to prevent system crashes. These processing methods lack a preprocessing mechanism, and the system's adaptability and robustness are poor, making it impossible to effectively process large amounts of highly concurrent data. Summary of the invention

[0004] The present application aims to solve at least one of the technical problems existing in the prior art. To this end, the present application proposes a method, device, electronic device and storage medium for processing high-concurrency data requests, which can improve the security and robustness of data processing.

[0005] In a first aspect, the present application provides a method for processing high-concurrency data requests, the method comprising:

[0006] Obtain the traffic time series prediction model based on the user data request volume at the current moment, and output the predicted data request volume at the next moment;

[0007] Based on the predicted data request amount at the next moment, determining a concurrent data request threshold of the sentinel flow control framework at the next moment;

[0008] In the case where it is determined that the concurrent data request amount of the trusted user at the next moment is less than or equal to the concurrent data request threshold of the sentinel flow control framework at the next moment, receiving and responding to the concurrent data request of the trusted user at the next moment;

[0009] The traffic time series prediction model is trained based on a historical data set, wherein the historical data set includes historical time data and historical user data requests, and the trusted user is a user whose identity is confirmed by a heuristic interception rule.

[0010] According to the method for processing high-concurrency data requests provided in the embodiment of the present application, the predicted data request volume at the next moment is predicted by using a traffic timing prediction model, and the predicted data request volume is used to determine the threshold of the sentinel traffic control framework, so as to ensure that the sentinel traffic control framework will not exceed the limit when processing data at the next moment, thereby improving the adaptability and robustness of high-concurrency data processing, and at the same time, intercepting unsafe user data requests through heuristic interception rules to improve the security of data processing.

[0011] According to an embodiment of the present application, after determining the concurrent data request threshold of the sentinel flow control framework at the next moment, the method further includes:

[0012] In the case where it is determined that the amount of concurrent data requests of the trusted user at the next moment is greater than the concurrent data request threshold of the sentinel flow control framework at the next moment, receiving and responding to the first concurrent data request of the trusted user at the next moment, the concurrent data request of the trusted user at the next moment including the first concurrent data request and the second concurrent data request, the second concurrent data request being a data request that exceeds the concurrent data request threshold of the sentinel flow control framework at the next moment;

[0013] If it is determined that the concurrent data request amount of the second concurrent data request is less than or equal to the queue length threshold of the Kafka message queue at the next moment, the second concurrent data request is cached in the Kafka message queue

[0014] Among them, the queue length threshold of the Kafka message queue at the next moment is determined based on the predicted data request amount at the next moment.

[0015] According to an embodiment of the present application, after receiving and responding to the first concurrent data request of the trusted user at the next moment, the method further includes:

[0016] When it is determined that the concurrent data request amount of the second concurrent data request is greater than the queue length threshold of the Kafka message queue at the next moment, the concurrent data request of the trusted user at the next moment is rejected, and a prompt message is output, wherein the prompt message is used to prompt the trusted user to access later.

[0017] According to one embodiment of the present application, the traffic time series prediction model is trained by the following steps:

[0018] Determining historical traffic time series data based on the historical moment data and the historical user data request;

[0019] Based on the historical traffic time series data, the BiGRU-Dual Attention network structure is trained to obtain the traffic time series prediction model.

[0020] According to one embodiment of the present application, the BiGRU-Dual Attention network structure includes a bidirectional gating layer, a first attention layer, a second attention layer and an output layer. The BiGRU-Dual Attention network structure is trained based on the historical traffic time series data to obtain the traffic time series prediction model, including:

[0021] Inputting the historical traffic time series data into the bidirectional gating layer to obtain the historical traffic time series features of the historical traffic time series data output by the bidirectional gating layer;

[0022] Inputting the historical traffic time series feature into the first attention layer to obtain a first attention weight output by the first attention layer;

[0023] Inputting the first attention weight and the historical traffic time series data into the second attention layer to obtain a second attention weight output by the second attention layer;

[0024] Outputting the second attention weight to the output layer to obtain the historical prediction data request amount output by the output layer;

[0025] Based on the historical predicted data request volume, the parameters of the BiGRU-Dual Attention network structure are updated to obtain the traffic time series prediction model.

[0026] According to one embodiment of the present application, the heuristic interception rules include intercepting users of hostile IPs, intercepting users of IPs whose user data request sending frequency is greater than a target threshold, intercepting users of IPs whose user data requests for a first target number are for the same type of pages, intercepting users of IPs whose user data requests for a second target number are for the same page access path and whose average stay time on each page does not exceed a first target duration, and intercepting at least one of users who repeatedly switch between multiple IPs within a second target duration.

[0027] In a second aspect, the present application provides a high-concurrency data request processing device, the device comprising:

[0028] An acquisition module is used to obtain the predicted data request volume at the next moment output by the traffic time series prediction model based on the user data request volume at the current moment;

[0029] A first processing module, configured to determine a concurrent data request threshold of a sentinel flow control framework at the next moment based on the predicted data request amount at the next moment;

[0030] The second processing module is used to receive and respond to the concurrent data request of the trusted user at the next moment when it is determined that the concurrent data request amount of the trusted user at the next moment is less than or equal to the concurrent data request threshold of the sentinel flow control framework at the next moment;

[0031] The traffic time series prediction model is trained based on a historical data set, wherein the historical data set includes historical time data and historical user data requests, and the trusted user is a user whose identity is confirmed by a heuristic interception rule.

[0032] According to the high-concurrency data request processing device provided by the embodiment of the present application, the predicted data request amount at the next moment is predicted by using a traffic timing prediction model, and the predicted data request amount is used to determine the threshold of the sentinel traffic control framework, so as to ensure that the sentinel traffic control framework will not exceed the limit when processing data at the next moment, thereby improving the adaptability and robustness of high-concurrency data processing, and at the same time, intercepting unsafe user data requests through heuristic interception rules to improve the security of data processing.

[0033] In a third aspect, the present application provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the processor implements the method for processing high-concurrency data requests as described in the first aspect above.

[0034] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method for processing high-concurrency data requests as described in the first aspect above is implemented.

[0035] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the method for processing high-concurrency data requests as described in the first aspect above.

[0036] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0038] Figure 1This is one of the flowcharts of the method for processing high-concurrency data requests provided in the embodiment of the present application;

[0039] Figure 2 This is the second flow chart of the method for processing high-concurrency data requests provided in the embodiment of the present application;

[0040] Figure 3 This is a flowchart of a method for processing high-concurrency data requests provided in an embodiment of the present application;

[0041] Figure 4 It is a structural diagram of the BiGRU-Dual Attention network structure provided in an embodiment of the present application;

[0042] Figure 5 is a schematic diagram of the structure of a bidirectional gating layer provided in an embodiment of the present application;

[0043] Figure 6 It is a structural diagram of the GRU module provided in an embodiment of the present application;

[0044] Figure 7 is a schematic diagram of the structure of the attention module provided in an embodiment of the present application;

[0045] Figure 8 This is a fourth flow chart of a method for processing high-concurrency data requests provided in an embodiment of the present application;

[0046] Fig. 9 It is a structural diagram of a device for processing high-concurrency data requests provided in an embodiment of the present application;

[0047] Fig.10 It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0048] The following will be combined with the drawings in the embodiments of the present application to clearly describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.

[0049] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0050] Combine the following Figure 1-Figure 10 , through specific embodiments and their application scenarios, the high-concurrency data request processing method, high-concurrency data request processing device, electronic device and readable storage medium provided in the embodiments of the present application are described in detail.

[0051] The method for processing high-concurrency data requests may be applied to a terminal, and may be specifically executed by hardware or software in the terminal.

[0052] The terminal includes, but is not limited to, a portable communication device such as a mobile phone or tablet computer with a touch-sensitive surface (e.g., a touch screen display and / or a touch pad). It should also be understood that in some embodiments, the terminal may not be a portable communication device, but a desktop computer with a touch-sensitive surface (e.g., a touch screen display and / or a touch pad).

[0053] In the following various embodiments, a terminal including a display and a touch-sensitive surface is described. However, it should be understood that the terminal may include one or more other physical user interface devices such as a physical keyboard, a mouse and a joystick.

[0054] like Figure 1 As shown, the method for processing high-concurrency data requests includes steps 110 to 130.

[0055] Step 110: Obtain the predicted data request volume at the next moment output by the traffic time series prediction model based on the trusted user data request volume at the current moment.

[0056] It is understandable that there is a certain time interval between the current moment and the next moment. For example, when the current moment is 5:59:01, the next moment may be 5:59:02, so the time length between the current moment and the next moment is 1 second.

[0057] The time length between the current moment and the next moment can be 1 second or 2 seconds, and there is no limitation on the time length between the current moment and the next moment.

[0058] The traffic time series prediction model is used to output the predicted data request volume at the next moment based on the input trusted user data request volume at the current moment.

[0059] It can be understood that by adjusting the length of time between the current moment and the next moment, changes in the amount of trusted user data requests can be addressed.

[0060] For example, when the amount of trusted user data requests changes greatly, the length of the current moment can be set to 0.5 seconds. At this time, the traffic timing prediction model can update the predicted data request amount every 0.5 seconds, thereby effectively responding to changes in the amount of trusted user data requests and enhancing the robustness of the system.

[0061] Step 120: Based on the predicted data request volume at the next moment, determine the concurrent data request threshold of the sentinel flow control framework at the next moment.

[0062] Among them, the sentinel traffic control framework is a high-availability traffic protection component for distributed service architecture. The concurrent data request threshold refers to the maximum amount of user data requests that the sentinel traffic control framework can process at the same time.

[0063] In this embodiment, when dealing with high concurrency, the sentinel traffic control framework can implement peak shaving and valley filling according to the amount of user data requests, and can also implement real-time monitoring of traffic. The amount of received user data requests can be seen in the console of the sentinel traffic control framework.

[0064] The sentinel flow control framework is also used to adjust the sending data of network packets and adjust user data requests according to the system processing capacity.

[0065] For example, the sentinel traffic control framework can autonomously determine whether to accept, adjust, or discard randomly arriving user data requests, thereby obtaining requests that are suitable for the system to receive.

[0066] In this embodiment, the sentinel flow control framework is provided with a concurrent data request threshold, which is the maximum value of the data request amount processed by the sentinel flow control framework. The concurrent data request threshold can be determined based on the predicted data request amount at the next moment predicted by the flow timing prediction model.

[0067] For example, if the traffic time series prediction model predicts that the predicted data request volume at the next moment is 50,000, then 50,000 can be used as the maximum value of the data request volume processed by the sentinel traffic control framework.

[0068] For another example, if the traffic time series prediction model predicts that the predicted data request volume at the next moment is 50,000, 50,000+n can also be used as the maximum value of the data request volume processed by the sentinel traffic control framework.

[0069] Here, n represents the fault tolerance of the flow control framework, and n can be an integer, for example, 1, 2, 3, 4, 5.

[0070] In actual implementation, when receiving a large amount of user data requests, the sentinel flow control framework may receive very few user data requests that exceed the concurrent data request threshold when it is close to the concurrent data request threshold. Setting a reasonable fault tolerance can effectively avoid problems such as non-response and over-limit in the sentinel flow control framework.

[0071] In actual implementation, the concurrent data request threshold of the sentinel traffic control framework can be set to the predicted data request amount at the next moment output by the traffic timing prediction model, that is, the concurrent data request threshold is equal to the predicted data request amount at the next moment.

[0072] Step 130: When it is determined that the concurrent data request amount of the trusted user at the next moment is less than or equal to the concurrent data request threshold of the sentinel flow control framework at the next moment, receive and respond to the concurrent data request of the trusted user at the next moment.

[0073] Among them, the traffic time series prediction model is trained based on the historical data set, which includes historical moment data and historical user data requests. The trusted user is the user whose identity is confirmed by the heuristic interception rules.

[0074] In this embodiment, whether the user is a trusted user is determined based on heuristic interception rules.

[0075] Among them, the trusted user is a user whose IP is identified by the heuristic interception rules and is not identified as a malicious IP by the heuristic interception.

[0076] In this embodiment, receiving and responding to the concurrent data request of the trusted user at the next moment refers to receiving the concurrent data request of the trusted user and providing the data resources required in the concurrent data request of the trusted user.

[0077] For example, by entering the keyword "down jacket" in the search box of the Taobao APP, a user data request representing down jackets can be sent to the Taobao system, and relevant page data resources about down jackets can be displayed on the Taobao APP page.

[0078] Heuristic interception rules are rules established based on feature value scanning technology, which are used to intercept malicious data requests sent by malicious IP addresses and obtain user data requests from trusted users.

[0079] In the related technology, large amounts of highly concurrent data are processed by filtering or restricting access through the Zuul gateway to prevent system crashes. These processing methods lack a preprocessing mechanism, and the system's adaptability and robustness are poor, making it impossible to effectively process large amounts of highly concurrent data.

[0080] In an embodiment of the present application, a traffic timing prediction model is used to output the predicted data request amount at the next moment based on the current moment, and the request amount is used as the threshold of the sentinel traffic control framework to implement a preprocessing mechanism, which effectively ensures that the data processed by the sentinel traffic control framework at the next moment will not exceed the limit. When receiving the user data request at the current moment, the user data request is intercepted by a heuristic interception rule, and the malicious data requests sent by malicious IPs are eliminated to process high-concurrency data requests.

[0081] According to the method for processing high-concurrency data requests provided in the embodiment of the present application, the predicted data request volume at the next moment is predicted by using a traffic timing prediction model, and the predicted data request volume is used to determine the threshold of the sentinel traffic control framework, so as to ensure that the sentinel traffic control framework will not exceed the limit when processing data at the next moment, thereby improving the adaptability and robustness of high-concurrency data processing, and at the same time, intercepting unsafe user data requests through heuristic interception rules to improve the security of data processing.

[0082] In some embodiments, after determining the concurrent data request threshold of the sentinel flow control framework at the next moment, the method further includes:

[0083] In the case where it is determined that the amount of concurrent data requests of the trusted user at the next moment is greater than the concurrent data request threshold of the sentinel flow control framework at the next moment, receiving and responding to the first concurrent data request of the trusted user at the next moment, the concurrent data request of the trusted user at the next moment includes the first concurrent data request and the second concurrent data request, and the second concurrent data request is a data request that exceeds the concurrent data request threshold of the sentinel flow control framework at the next moment;

[0084] When it is determined that the concurrent data request amount of the second concurrent data request is less than or equal to the queue length threshold of the Kafka message queue at the next moment, the second concurrent data request is cached in the Kafka message queue;

[0085] Among them, the queue length threshold of the Kafka message queue at the next moment is determined based on the predicted data request volume at the next moment.

[0086] Among them, the Kafka message queue is an event stream platform. The Kafka message queue can maintain stable performance after storing a large number of messages, ensuring zero downtime and zero data loss, and the queue length of the Kafka message queue can be adjusted.

[0087] For example, when the amount of user data requests is small, the queue length of the Kafka message queue can be shortened to release the memory space occupied by the Kafka message queue, thereby improving the performance of the Kafka message queue in processing user data requests.

[0088] When the amount of user data requests is large, the queue length of the Kafka message queue can be expanded to store enough user data requests while still maintaining stable performance.

[0089] Among them, the sentinel flow control box is used to process concurrent user data requests, and the Kafka message queue is used to store user data requests that exceed the concurrent data request threshold.

[0090] In this embodiment, after receiving the user data request, the amount of user data requests received at the next moment may exceed the concurrent data request threshold of the sentinel flow control framework at the next moment. At this time, the concurrent data request at the next moment is divided into two parts, including a first concurrent data request and a second concurrent data request.

[0091] The first concurrent data request amount is equal to the concurrent data request threshold of the sentinel flow control framework. When the concurrent data request at the next moment is received, the first concurrent data request is input into the sentinel flow control framework.

[0092] The second concurrent data volume is data requests that exceed the concurrent data request threshold of the sentinel flow control framework at the next moment. This part of user data requests can be marked as data requests to be queued.

[0093] When it is determined that the concurrent data request amount of the data request to be queued is less than or equal to the queue length threshold of the Kafka message queue at the next moment, the data request to be queued is cached in the Kafka message queue.

[0094] For example, the concurrent data request threshold of the sentinel flow control framework at the next moment is P. The first P concurrent data requests of the concurrent data requests at the next moment are input into the sentinel flow control framework, and the P+1th concurrent data request is marked as a data request to be queued and cached in the Kafka message queue.

[0095] Among them, the queue length threshold of the Kafka message queue at the next moment is obtained based on the traffic timing prediction model. The user data request amount at the current moment is input into the traffic timing prediction model. The traffic timing prediction model outputs the predicted data request amount at the next moment, and uses the predicted data request amount to update the queue length threshold of the Kafka queue.

[0096] For example, when the predicted data request volume output by the traffic time series prediction model is 50,000, the queue length threshold of the Kafka message queue can be set to 50,000 based on the predicted data request volume.

[0097] Similar to the sentinel flow control framework, the queue length threshold of the Kafka message queue can also be set to 50000+t based on the predicted data request volume, where t is an integer and can be 1, 2, 3, 4, or 5.

[0098] By introducing a fault tolerance t, it can be ensured that when the Kafka message queue encounters a large flow of traffic and receives more data requests than the Kafka message queue's own queue length threshold due to a momentary error, the Kafka message queue can maintain a normal working state.

[0099] By introducing the Kafka message queue, when the concurrency exceeds the receiving threshold of the Sentinel flow control framework, a buffer space can be reserved, which greatly improves the security and robustness of the system and prevents the system from crashing when it is hit by large traffic.

[0100] In some embodiments, after receiving and responding to a first concurrent data request from a trusted user at a next moment, the method further includes:

[0101] When it is determined that the concurrent data request amount of the second concurrent data request is greater than the queue length threshold of the Kafka message queue at the next moment, the concurrent data request of the trusted user at the next moment is rejected, and a prompt message is output, where the prompt message is used to prompt the trusted user to access later.

[0102] It is understandable that if the number of concurrent data requests at the next moment is too large, after the sentinel flow control framework receives and processes the first concurrent data request, the second concurrent data request may still exceed the queue length threshold of the Kafka message queue at the next moment.

[0103] In this embodiment, when the second concurrent data request amount at the next moment is greater than the queue length threshold of the Kafka message queue at the next moment, the queued data requests with a value equal to the Kafka message queue length threshold at the next moment are cached in the Kafka message queue, and the concurrent data requests that exceed the Kafka message queue length threshold at the next moment are rejected, and a prompt message is output to the user.

[0104] The prompt information includes explaining that the current user access channel is crowded and prompting the user to access later.

[0105] In some embodiments, the traffic time series prediction model is trained by the following steps:

[0106] Determine historical traffic time series data based on historical time data and historical user data request volume;

[0107] Based on historical traffic time series data, the BiGRU-Dual Attention network structure is trained to obtain the traffic time series prediction model.

[0108] In this embodiment, the historical moment data includes and the historical user data requests may be the amount of user data requests received by the system at historical moments and the corresponding moment data.

[0109] Each time point data corresponds to a user data request amount. For example, the user data request amount corresponding to 9:59:01 is 2000.

[0110] The historical traffic time series data includes a moment data and a corresponding user data request amount. The BiGRU-DualAttention network structure is a neural network structure with bidirectional gating and dual attention layers.

[0111] In this embodiment, the historical time data and the historical user data request volume are used to train the BiGRU-DualAttention network structure, update the BiGRU-Dual Attention network structure, and obtain the traffic time series prediction model.

[0112] During the training process, the historical traffic time series data is input into the traffic time series prediction model to obtain the output of the traffic time series prediction model and update the model parameters.

[0113] In some embodiments, Figure 4 As shown in the figure, the BiGRU-Dual Attention network structure includes a bidirectional gating layer, a first attention layer, a second attention layer, and an output layer. Based on the historical traffic time series data, the BiGRU-Dual Attention network structure is trained to obtain a traffic time series prediction model, including:

[0114] Input the historical traffic time series data into the bidirectional gating layer to obtain the historical traffic time series features of the historical traffic time series data output by the bidirectional gating layer;

[0115] Input the historical traffic time series features into the first attention layer, and obtain the first attention weight output by the first attention layer;

[0116] Input the first attention weight and the historical traffic time series data into the second attention layer to obtain the second attention weight output by the second attention layer;

[0117] Output the second attention weight to the output layer to obtain the historical prediction data request amount output by the output layer;

[0118] Based on the historical forecast data request volume, the parameters of the BiGRU-Dual Attention network structure are updated to obtain the traffic time series prediction model.

[0119] In this embodiment, based on historical traffic time series data, the BiGRU-Dual Attention network structure is trained to obtain a traffic time series prediction model.

[0120] Among them, before entering the model, you also need to apply the formula

[0121]

[0122] Among them, X norm is the normalized data, X is the original data, and X max is the maximum value X of the historical traffic time series data min It is the minimum value of the historical traffic time series data.

[0123] Perform normalization preprocessing on historical traffic time series data.

[0124] The preprocessed historical traffic time series data is input into the bidirectional gating layer, which is used to mine the time series characteristics and trend characteristics of the historical time series data and output the time series characteristics and trend characteristics.

[0125] like Figure 5 As shown, the bidirectional gating layer includes a reverse layer, a forward layer, and an input layer.

[0126] In actual implementation, the hidden state of the bidirectional gating layer is composed of the current historical traffic time series data and the output of the forward hidden state at time (t-1). and reverse output The three parts jointly decide.

[0127] In this embodiment, the GRU module unit determines the hidden state output of the bidirectional gating layer, such as Figure 6As shown, r represents the reset gate and z represents the update gate.

[0128] Among them, the hidden layer state output application formula

[0129] h i:i+L =BiGRU(X i ,X i+1 ,...,X i+L )

[0130] Among them, X i Represents the historical traffic time series data at time i.

[0131] The hidden layer state output is the historical traffic time series characteristics of the historical traffic time series data.

[0132] The update gate controls the influence of the hidden layer output at the previous moment on the current hidden layer. The larger the value of z, the greater the influence of the hidden state at the previous moment on the learning of the current hidden layer. The reset gate controls the degree to which the hidden layer information at the previous moment is ignored. The smaller the value of r, the smaller the positive effect of the forward information on the training of the current hidden layer, and the more it needs to be ignored.

[0133] tanh is the activation function, σ is the sigmoid function, the memory gate uses the tanh activation function, and each element of its output vector is in [-1, 1]. The reset gate and the update gate use the sigmoid function, and each element of their output vector is in [0, 1]. The dimension of the memory gate output vector is equal to the dimension of the reset gate and the update gate output vector.

[0134] Among them, the reset gate and update gate are used to control the amount of information flowing through each dimension, and the memory gate is used to store information flowing through the GRU module.

[0135] Apply the following formula

[0136]

[0137]

[0138] r=σ(W r h t-1 +U r x t )

[0139] z=σ(W z h t-1 +U z x t )

[0140] Among them, z is the update gate, r is the reset gate, c is the memory gate, tanh() is the activation function, σ() is the sigmoid function, and x tis the input of the GRU module.

[0141] Determine the hidden state h at time t t .

[0142] In this embodiment, if Figure 7 As shown, the input of the first attention layer can be the historical traffic time series features. In this layer, the introduction of the first attention weight can assign different attention weights to the historical traffic time series features at different times, highlight the key information at recent times, and reduce the influence of long-term historical features. The first attention layer outputs the first attention weight.

[0143] Among them, the application formula

[0144]

[0145] u i:i+L =tanh(W w h i:i+L +b w )

[0146] E(h i:i+L )=∑α i:i+L h i:i+L

[0147] Among them, E(h i:i+L ) is the hidden state weight of the historical traffic time series feature, h i:i+L is the hidden layer state output, and L is the historical moment.

[0148] Determine the hidden layer state weights of the historical traffic time series characteristics at each historical moment.

[0149] In this embodiment, the input of the second attention layer is the first attention weight and historical traffic time series data. The second attention layer is used to extract scattered features existing in the long-term history of the historical traffic time series data, and obtain the weight distribution of different scattered features in the scattered features. The second attention layer outputs the second attention weight.

[0150] In this embodiment, the output layer is used to process the second attention weight of the input of the second attention layer, and the second attention weight is in the fully connected layer.

[0151] Apply the formula

[0152] P=∑WX+b

[0153] Among them, W is the output of the second attention layer, X is the input of the second attention layer, and b is the bias parameter.

[0154] Calculate the historical forecast data request volume.

[0155] The historical predicted data request volume is used to update the parameters of the BiGRU-Dual Attention network structure to obtain the traffic time series prediction model.

[0156] In some embodiments, the heuristic interception rules include intercepting users of hostile IPs, intercepting users of IPs whose user data request sending frequency is greater than a target threshold, intercepting users of IPs whose user data requests are for the same type of pages with a first target number, intercepting users of IPs whose user data requests are for the same page access path with a second target number and whose average stay time on each page does not exceed a first target duration, and intercepting at least one of users who repeatedly switch between multiple IPs within a second target duration.

[0157] Among them, the heuristic interception rules are used to target malicious data request behaviors of malicious IPs, reduce the interference of malicious data request behaviors of malicious IPs on the system, and within half an hour after being identified as a malicious IP by the heuristic interception rules, the IP's access to the system will be restricted.

[0158] After passing the heuristic interception rules, users who are not identified as malicious IPs by the heuristic interception rules are trusted users.

[0159] The heuristic interception rules may include a variety of judgment criteria for intercepting malicious IPs, including at least the following five situations.

[0160] First, hostile IP.

[0161] The IP of the user who sends the data request can be detected. If the user's IP has a hostile relationship with this system, the user data request sent by the user will be intercepted, and the user will be prohibited from accessing this system within half an hour.

[0162] Second, the sending frequency.

[0163] The sending frequency of the user IP can be calculated. The sending frequency can be obtained by dividing the requests sent within a certain period of time by the time. If the sending frequency exceeds the target threshold, the user data request sent by the user will be intercepted, and the user will be prohibited from accessing the system within half an hour.

[0164] Third, the same request.

[0165] The first target number of user data requests for the same user can be detected. If the user's recent first target number of user data requests all point to the same type of page, it means that the user may have assembled the URL and page ID by himself through a web crawler for access. This type of access is an invalid access and brings a lot of server pressure to the system.

[0166] In this case, the user data request sent by the user is intercepted, and the user is prohibited from accessing the system within half an hour.

[0167] Fourth, page dwell time.

[0168] The second target number of user data requests for the same user and subsequent operations after responding to the user data requests can be detected. If the access paths of the user's recent second target number of user data requests are basically consistent, and the average time the user stays on each page after receiving the response does not exceed the first target time.

[0169] Then the user may actually simulate the behavior of browsing web pages through a web crawler, which will also bring a lot of server pressure to the system, and the access will also be an invalid access.

[0170] In this case, the user data request sent by the user is also intercepted, and the user is prohibited from accessing the system within half an hour.

[0171] Fifth, IP switching.

[0172] The IP of the same user can be detected within the second target time. If the IP of the user switches repeatedly within the second target time, it means that the IP of the user may come from the IP proxy pool and try to deceive the system by switching the IP.

[0173] In this case, the user data request sent by the user is also intercepted, and the user is prohibited from accessing the system within half an hour.

[0174] A specific embodiment is introduced below to describe the specific process of the high-concurrency data request processing method.

[0175] like Figure 2 and Figure 3 As shown, during the historical data collection phase, the historical traffic time series data in the background work log can be obtained by consulting the system work log.

[0176] In the traffic time series prediction model construction stage, the traffic time series prediction model is constructed through the BiGRU-Dual Attention network structure, and the historical traffic time series data is input into the traffic time series prediction model, the predicted traffic time series data is output, and the model is updated according to the output predicted traffic time series data.

[0177] In the model application stage, the current time t is input into the traffic time series prediction model to obtain the output result p of the traffic time series prediction model, where p represents the predicted user data request volume at time t+1.

[0178] It should be noted that the time t+1 represents the next time of the current time. For example, when the selected prediction time length is 1 second, the current time is 59 minutes and 06 seconds, then the next time is 59 minutes and 07 seconds;

[0179] For another example, when the selected predicted time length is 2 seconds, the current time is 59 minutes and 6 seconds, and the next time is 59 minutes and 8 seconds.

[0180] After obtaining the predicted traffic time series data for the next moment, Figure 8 As shown in the figure, after the traffic prediction model outputs the predicted traffic time series data, the concurrent data request threshold of the sentinel traffic control framework and the kafka message queue is immediately updated to the predicted traffic time series data.

[0181] In the heuristic interception rule setting stage, different heuristic interception rules can be set for different malicious IPs and different malicious user data requests. After setting different heuristic interception rules, high-concurrency services are started.

[0182] In the stage of receiving user requests, the heuristic interception rules are used to determine whether the user IP is credible. If it is not credible, the user data request is intercepted, and the user IP is set as a malicious IP, prohibiting the user from accessing the system within half an hour.

[0183] If it is credible, the user is regarded as a credible user, and the data request sent by the user is received. It is determined whether the concurrency at that moment is greater than the concurrent data request threshold of the sentinel flow control framework. If the concurrency at that moment is less than the concurrent data request threshold of the sentinel, then all user data requests at that moment are responded to.

[0184] If the concurrency at that moment is greater than the concurrent data request threshold of Sentinel, then respond to the amount of trusted user data requests equal to the concurrent data request threshold at that moment, and compare the excess trusted user data requests with the queue length of the Kafka message queue.

[0185] If the amount of data requests from the redundant trusted users is less than the queue length of the message queue, all data requests from the redundant trusted users are cached in the message queue.

[0186] If the excess amount of trusted user data requests is greater than the queue length of the message queue, then the data requests of trusted users equal to the queue length of the message queue will be cached in the message queue. The data requests of trusted users cached in the message queue will wait for the data requests of the previous trusted users to end. After the previous data request ends, the data request of the trusted user will be responded to.

[0187] The data requests of the redundant part of the trusted users are intercepted, and the trusted users who sent the data requests are prompted to access the data later.

[0188] The method for processing high-concurrency data requests provided in the embodiments of the present application may be executed by an electronic device or a functional module or functional entity in the electronic device that can implement the method for processing high-concurrency data requests. The electronic devices mentioned in the embodiments of the present application include but are not limited to mobile phones, tablet computers, computers, cameras, and wearable devices. The method for processing high-concurrency data requests provided in the embodiments of the present application is described below using an electronic device as an example of an execution subject.

[0189] The processing method for high-concurrency data requests provided in the embodiment of the present application can be executed by a processing device for high-concurrency data requests. In the embodiment of the present application, the processing method for high-concurrency data requests executed by a processing device for high-concurrency data requests is taken as an example to illustrate the processing device for high-concurrency data requests provided in the embodiment of the present application.

[0190] The embodiment of the present application also provides a device for processing high-concurrency data requests.

[0191] like Fig. 9 As shown, the processing device for high-concurrency data requests includes:

[0192] An acquisition module 910 is used to acquire the predicted data request volume at the next moment output by the traffic time series prediction model based on the user data request volume at the current moment;

[0193] The first processing module 920 is used to determine the concurrent data request threshold of the sentinel flow control framework at the next moment based on the predicted data request amount at the next moment;

[0194] The second processing module 930 is used to receive and respond to the concurrent data request of the trusted user at the next moment when it is determined that the concurrent data request amount of the trusted user at the next moment is less than or equal to the concurrent data request threshold of the sentinel flow control framework at the next moment;

[0195] Among them, the traffic time series prediction model is trained based on the historical data set, which includes historical moment data and historical user data requests. The trusted user is the user whose identity is confirmed by the heuristic interception rules.

[0196] According to the high-concurrency data request processing device provided by the embodiment of the present application, the predicted data request amount at the next moment is predicted by using a traffic timing prediction model, and the predicted data request amount is used to determine the threshold of the sentinel traffic control framework, so as to ensure that the sentinel traffic control framework will not exceed the limit when processing data at the next moment, thereby improving the adaptability and robustness of high-concurrency data processing, and at the same time, intercepting unsafe user data requests through heuristic interception rules to improve the security of data processing.

[0197] In some embodiments, the first processing module 920 is used to receive and respond to the first concurrent data request of the trusted user at the next moment when it is determined that the concurrent data request amount of the trusted user at the next moment is greater than the concurrent data request threshold of the sentinel flow control framework at the next moment, the concurrent data request of the trusted user at the next moment includes the first concurrent data request and the second concurrent data request, and the second concurrent data request is a data request that exceeds the concurrent data request threshold of the sentinel flow control framework at the next moment;

[0198] The first processing module 920 is further configured to cache the second concurrent data request to the Kafka message queue when it is determined that the concurrent data request amount of the second concurrent data request is less than or equal to the queue length threshold of the Kafka message queue at the next moment;

[0199] Among them, the queue length threshold of the Kafka message queue at the next moment is determined based on the predicted data request volume at the next moment.

[0200] In some embodiments, the first processing module 920 is also used to reject the concurrent data request of the trusted user at the next moment when it is determined that the concurrent data request amount of the second concurrent data request is greater than the queue length threshold of the Kafka message queue at the next moment, and output a prompt message, where the prompt message is used to prompt the trusted user to access later.

[0201] In some embodiments, based on the historical moment data and the historical user data requests, historical traffic time series data is determined;

[0202] Based on historical traffic time series data, the BiGRU-Dual Attention network structure is trained to obtain the traffic time series prediction model.

[0203] In some embodiments, the BiGRU-Dual Attention network structure includes a bidirectional gating layer, a first attention layer, a second attention layer, and an output layer. Based on historical traffic time series data, the BiGRU-Dual Attention network structure is trained to obtain a traffic time series prediction model, including:

[0204] Input the historical traffic time series data into the bidirectional gating layer to obtain the historical traffic time series features of the historical traffic time series data output by the bidirectional gating layer;

[0205] Input the historical traffic time series features into the first attention layer, and obtain the first attention weight output by the first attention layer;

[0206] Input the first attention weight and the historical traffic time series data into the second attention layer to obtain the second attention weight output by the second attention layer;

[0207] Output the second attention weight to the output layer to obtain the historical prediction data request amount output by the output layer;

[0208] Based on the historical forecast data request volume, the parameters of the BiGRU-Dual Attention network structure are updated to obtain the traffic time series prediction model.

[0209] In some embodiments, the heuristic interception rules include intercepting users of hostile IPs, intercepting users of IPs whose user data request sending frequency is greater than a target threshold, intercepting users of IPs whose user data requests are for the same type of pages with a first target number, intercepting users of IPs whose user data requests are for the same page access path with a second target number and whose average stay time on each page does not exceed the first target duration, and intercepting at least one of users who repeatedly switch between multiple IPs within the second target duration.

[0210] The processing device for high concurrent data requests in the embodiment of the present application can be an electronic device, or a component in the electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal, or it can be other devices other than the terminal. Exemplarily, the electronic device can be a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted electronic device, a mobile Internet device (Mobile Internet Device, MID), an augmented reality (augmented reality, AR) / virtual reality (virtual reality, VR) device, a robot, a wearable device, an ultra-mobile personal computer (ultra-mobile personal computer, UMPC), a netbook or a personal digital assistant (personal digital assistant, PDA), etc., and can also be a server, a network attached storage (Network Attached Storage, NAS), a personal computer (personal computer, PC), a television (television, TV), a teller machine or a self-service machine, etc., which is not specifically limited in the embodiment of the present application.

[0211] The high concurrent data request processing device in the embodiment of the present application may be a device having an operating system. The operating system may be an Android operating system, an IOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.

[0212] The high-concurrency data request processing device provided in the embodiment of the present application can achieve Figures 1 to 8 To avoid repetition, the various processes implemented by the method embodiment are not described here.

[0213] In some embodiments, Fig.10 As shown, an embodiment of the present application also provides an electronic device 1000, including a processor 1001, a memory 1002, and a computer program stored in the memory 1002 and executable on the processor 1001. When the program is executed by the processor 1001, each process of the above-mentioned high-concurrency data request processing method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0214] It should be noted that the electronic devices in the embodiments of the present application include the mobile electronic devices and non-mobile electronic devices mentioned above.

[0215] An embodiment of the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned high-concurrency data request processing method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0216] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0217] An embodiment of the present application also provides a computer program product, including a computer program, which implements the above-mentioned method for processing high-concurrency data requests when executed by a processor.

[0218] The processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory ROM, a random access memory RAM, a magnetic disk or an optical disk.

[0219] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.

[0220] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a computer software product, which is stored in a storage medium (such as ROM / RAM, a disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0221] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.

[0222] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.

[0223] Although the embodiments of the present application have been shown and described, those skilled in the art will appreciate that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present application, and that the scope of the present application is defined by the claims and their equivalents.

Claims

1. A method for processing high-concurrency data requests, characterized in that: include: Obtain the traffic time series prediction model based on the user data request volume at the current moment, and output the predicted data request volume at the next moment; Based on the predicted data request amount at the next moment, determining a concurrent data request threshold of the sentinel flow control framework at the next moment; In the case where it is determined that the concurrent data request amount of the trusted user at the next moment is less than or equal to the concurrent data request threshold of the sentinel flow control framework at the next moment, receiving and responding to the concurrent data request of the trusted user at the next moment; The traffic time series prediction model is obtained by training based on a historical data set, the historical data set includes historical time data and historical user data requests, and the trusted user is a user whose identity is confirmed by a heuristic interception rule; After determining the concurrent data request threshold of the sentinel flow control framework at the next moment, the method further includes: In the case where it is determined that the amount of concurrent data requests of the trusted user at the next moment is greater than the concurrent data request threshold of the sentinel flow control framework at the next moment, receiving and responding to the first concurrent data request of the trusted user at the next moment, the concurrent data request of the trusted user at the next moment including the first concurrent data request and the second concurrent data request, the second concurrent data request being a data request that exceeds the concurrent data request threshold of the sentinel flow control framework at the next moment; When it is determined that the concurrent data request amount of the second concurrent data request is less than or equal to the queue length threshold of the Kafka message queue at the next moment, cache the second concurrent data request to the Kafka message queue; Among them, the queue length threshold of the Kafka message queue at the next moment is determined based on the predicted data request amount at the next moment.

2. The method for processing high concurrent data requests according to claim 1, characterized in that: After receiving and responding to the first concurrent data request of the trusted user at the next moment, the method further includes: When it is determined that the concurrent data request amount of the second concurrent data request is greater than the queue length threshold of the Kafka message queue at the next moment, the concurrent data request of the trusted user at the next moment is rejected, and a prompt message is output, wherein the prompt message is used to prompt the trusted user to access later.

3. The method for processing high concurrent data requests according to claim 1, characterized in that: The traffic time series prediction model is trained through the following steps: Determining historical traffic time series data based on the historical moment data and the historical user data request; Based on the historical traffic time series data, the BiGRU-Dual Attention network structure is trained to obtain the traffic time series prediction model.

4. The method for processing high concurrent data requests according to claim 3, characterized in that: The BiGRU-DualAttention network structure includes a bidirectional gating layer, a first attention layer, a second attention layer and an output layer. The BiGRU-DualAttention network structure is trained based on the historical traffic time series data to obtain the traffic time series prediction model, including: Inputting the historical traffic time series data into the bidirectional gating layer to obtain the historical traffic time series features of the historical traffic time series data output by the bidirectional gating layer; Inputting the historical traffic time series feature into the first attention layer to obtain a first attention weight output by the first attention layer; Inputting the first attention weight and the historical traffic time series data into the second attention layer to obtain a second attention weight output by the second attention layer; Outputting the second attention weight to the output layer to obtain the historical prediction data request amount output by the output layer; Based on the historical predicted data request volume, the parameters of the BiGRU-Dual Attention network structure are updated to obtain the traffic time series prediction model.

5. The method for processing high-concurrency data requests according to any one of claims 1 to 4, characterized in that: The heuristic interception rules include intercepting users of hostile IPs, intercepting users of IPs whose user data request sending frequency is greater than a target threshold, intercepting users of IPs whose user data requests of a first target number are for the same type of pages, intercepting users of IPs whose user data requests of a second target number are for the same page access path and whose average stay time on each page does not exceed the first target duration, and intercepting at least one of users who repeatedly switch between multiple IPs within the second target duration.

6. A device for processing high concurrent data requests, characterized in that: include: An acquisition module is used to obtain the predicted data request volume at the next moment output by the traffic time series prediction model based on the user data request volume at the current moment; A first processing module, configured to determine a concurrent data request threshold of a sentinel flow control framework at the next moment based on the predicted data request amount at the next moment; The second processing module is used to receive and respond to the concurrent data request of the trusted user at the next moment when it is determined that the concurrent data request amount of the trusted user at the next moment is less than or equal to the concurrent data request threshold of the sentinel flow control framework at the next moment; The traffic time series prediction model is obtained by training based on a historical data set, the historical data set includes historical time data and historical user data requests, and the trusted user is a user whose identity is confirmed by a heuristic interception rule; The first processing module is used to receive and respond to the first concurrent data request of the trusted user at the next moment when it is determined that the concurrent data request amount of the trusted user at the next moment is greater than the concurrent data request threshold of the sentinel flow control framework at the next moment, wherein the concurrent data request of the trusted user at the next moment includes the first concurrent data request and the second concurrent data request, and the second concurrent data request is a data request that exceeds the concurrent data request threshold of the sentinel flow control framework at the next moment; The first processing module is further configured to cache the second concurrent data request to the Kafka message queue when it is determined that the concurrent data request amount of the second concurrent data request is less than or equal to the queue length threshold of the Kafka message queue at the next moment; Among them, the queue length threshold of the Kafka message queue at the next moment is determined based on the predicted data request amount at the next moment.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method for processing high-concurrency data requests as described in any one of claims 1-5 is implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for processing high-concurrency data requests as described in any one of claims 1 to 5 is implemented.

9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for processing high-concurrency data requests as described in any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Resource self-adaptive adjusting system and method of multiple virtual machines under single physical machine

    CN104283946A

  • High-concurrency automatic capacity expansion and contraction method and system, computer equipment and medium

    CN114358134A