Risk detection method, device and equipment
By analyzing and analyzing the stacked traffic data of the target server, determining whether there are malicious traffic characteristics and calculating the estimated data leakage amount, the problem of failure to effectively reduce the risk of data leakage in the existing technology is solved, real-time detection and prevention of data leakage incidents is realized, and data security is improved.
Patent Information
- Application Number
- CN202510331243.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-13
AI Technical Summary
The existing data leakage prevention technology is mainly aimed at emergency remedial measures after data leakage, which fails to effectively reduce the risk of data leakage, resulting in low data security.
By analyzing the stacked traffic data of the target server, determining whether there are malicious traffic characteristics, calculating the expected data leakage, and determining whether there is a data leakage incident based on the data set and threshold within the preset time.
Real-time data leakage incident detection of target servers is realized, and data leakage is avoided in a timely manner, thereby improving data security.
Smart Images

Figure CN120151047A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network security technology, and in particular, to a risk detection method, device, and equipment. Background Art
[0002] Under the background that the digital economy will continue to develop in the directions of informatization, digitization, intelligence, etc., data is an important cornerstone of economic development, and data security is the foundation of the development of the digital economy. Protecting data security is to escort the development of the digital economy.
[0003] In recent years, the trend of data leakage risk will continue to intensify, and the construction of data leakage prevention security has a long way to go. At present, data leakage prevention technologies such as data disaster recovery and data anti-ransomware are all emergency remedies after data leakage. In fact, they do not reduce the risk of data leakage, and the security of data is still relatively low. Summary of the Invention
[0004] This application provides a risk detection method, device, and equipment for improving the security of data.
[0005] In a first aspect, an embodiment of this application provides a risk detection method, including: parsing first traffic data from a target server to obtain first traffic characteristics, where the first traffic data is outgoing stack traffic data; determining whether the first traffic data meets a first condition according to the digital certificate, request information, and communication address in the first traffic characteristics, where the first condition is used to indicate that the first traffic characteristics include malicious traffic characteristics; if the first traffic data meets the first condition, determining the expected data leakage amount of the first traffic data according to the first traffic characteristics; determining a detection result of the target server according to the total expected data leakage amount of the first traffic data set within a preset duration and a first threshold, where the detection result is used to indicate whether there is a data leakage event in the target server, and the first traffic data set includes the first traffic data.
[0006] In the embodiment of this application, by determining whether the outgoing stack traffic data of the target server includes malicious traffic characteristics, it is determined whether there is a data leakage risk in the outgoing stack traffic data. Further, in the case where it is determined that the first traffic data has a data leakage risk, the expected data leakage amount of the first traffic data is determined, so that the total expected data leakage amount of the first traffic data set within a preset duration can be calculated. According to the total expected data leakage amount and the first threshold, it is determined whether there is a data leakage event in the target server. For example, when the total expected data leakage amount is greater than the first threshold, it is determined that there is a data leakage event in the target server. Through this detection method, data leakage events in the target server can be detected in real time to avoid data leakage in a timely manner, thereby improving the security of data.
[0007] In a possible implementation, the data transmission protocol of the first traffic data is the Hypertext Transfer Protocol http; the first condition includes at least one of the following: the digital certificate in the first traffic feature is a malicious digital certificate; the http request information in the first traffic feature includes malicious information; the communication address in the first traffic feature is a malicious communication address.
[0008] In a possible implementation, determining whether the first traffic data meets the first condition according to the digital certificate in the first traffic feature includes: extracting the certificate features of the digital certificate in the first traffic feature; determining whether the digital certificate is a malicious digital certificate according to the certificate features, a pre-stored reference certificate feature library, and a preset certificate rule; if the digital certificate is a malicious digital certificate, determining that the first traffic data meets the first condition.
[0009] In a possible implementation, the data transmission protocol of the first traffic data is the Hypertext Transfer Protocol http; determining whether the first traffic data meets the first condition according to the request information in the first traffic feature includes: extracting multiple parameter information in the http request information in the first traffic feature, where the multiple parameter information includes a request header, a Uniform Resource Locator URL, request parameters, and response body content; determining whether there is malicious information in the multiple parameter information according to the multiple parameter information, a pre-stored reference malicious information library, and a preset http rule; if there is malicious information, determining that the first traffic data meets the first condition.
[0010] In a possible implementation, the first traffic feature further includes the size of the data packet transmitted by the first traffic data; before determining the expected data leakage amount of the first traffic data according to the first traffic feature, the method further includes: if the first traffic data meets the first condition, inputting the traffic log, the data packet size, the communication address, and a preset intelligence value level coefficient of the first traffic data into a pre-trained data leakage risk index model to obtain a data leakage risk index, where the data leakage risk index is used to indicate the possibility of data leakage risk of the first traffic data; if it is determined that the data leakage risk index exceeds the warning value, determining the expected data leakage amount of the first traffic data.
[0011] In this embodiment, after determining that the first traffic data meets the first condition and before determining the predicted data leakage volume of the first traffic data, the data leakage risk index of the first traffic data may also be determined according to a pre-trained data leakage risk index model. Then, based on the magnitude of the data leakage risk index, the magnitude of the data leakage risk of the first traffic data can be judged. Further, it can be determined whether to include the predicted data leakage volume of the first traffic data in the total predicted data leakage volume for judging whether there is a data leakage event in the target server, thereby improving the accuracy of detecting data leakage events in the target server.
[0012] In a possible implementation, the method further includes: if the first traffic data does not meet the first condition and there is no traffic data in the second traffic data set that meets the first condition within a preset duration, then determine whether the second traffic data set meets a second condition, where the second condition is used to indicate that the data transmission volume and / or the number of requests for the data popped out of the second traffic data set are abnormal; if it is determined that the second traffic data set meets the second condition, then determine that there is a suspected data leakage event in the second traffic data set; generate a prompt message and send the prompt message to the user device, where the prompt message is used to indicate to confirm whether there is a data leakage event in the second traffic data set.
[0013] In this embodiment, when no traffic data that does not meet the first condition is detected within a period of time, it can be judged whether the data transmission volume and / or the number of requests for the data popped out of the second traffic data set within this period of time are abnormal. When the data transmission volume for the data popped out is abnormally large or the number of requests is abnormally large within a short period of time, it can be determined that there is a suspected data leakage event in the second traffic data set. At this time, the staff is prompted to further confirm, thereby improving the accuracy of detecting data leakage events in the target server and being beneficial to improving data security.
[0014] In a possible implementation, the second condition includes at least one of the following: the number of traffic data requested by the same user identifier in the second traffic data set exceeds a second threshold; the total amount of sensitive data transmitted by the second traffic data set exceeds a third threshold; there is abnormal traffic data in the second traffic data set with a packet size exceeding a fourth threshold.
[0015] In a possible implementation, after generating the prompt message and sending the prompt message to the user device, the method further includes: receiving a response message from the user device, where the response message is used to indicate whether there is a data leakage event in the second traffic data set and the abnormal traffic data in the second traffic data set; if there is a data leakage event in the second traffic data set, then analyze the abnormal traffic data in the second traffic data set, extract the traffic feature information in the abnormal traffic data and store it in the malicious traffic feature library.
[0016] In this embodiment, when the staff determines that there is a data leakage event in the second traffic data, the traffic feature information of the abnormal traffic data in the second traffic data set is extracted and stored in the malicious traffic feature library, which is beneficial to improving the accuracy of subsequent detection of abnormal traffic data.
[0017] In a second aspect, an embodiment of the present application provides a risk detection device, including: a parsing module, configured to parse the first traffic data from a target server to obtain a first traffic feature, where the first traffic data is outgoing stack traffic data; a determining module, configured to determine whether the first traffic data meets a first condition according to a digital certificate, a request message, and a communication address in the first traffic feature, where the first condition is used to indicate that the first traffic feature includes a malicious traffic feature; the determining module is further configured to, if the first traffic data meets the first condition, determine an expected data leakage amount of the first traffic data according to the first traffic feature; a detection module, configured to determine a detection result of the target server according to a total expected data leakage amount of the first traffic data set within a preset duration and a first threshold, where the detection result is used to indicate whether there is a data leakage event in the target server, and the first traffic data set includes the first traffic data.
[0018] In a possible implementation manner, a data transmission protocol of the first traffic data is a hypertext transfer protocol http; the first condition includes at least one of the following: the digital certificate in the first traffic feature is a malicious digital certificate; the http request message in the first traffic feature includes malicious information; the communication address in the first traffic feature is a malicious communication address.
[0019] In a possible implementation manner, the determining module is specifically configured to: extract a certificate feature of the digital certificate in the first traffic feature; determine whether the digital certificate is a malicious digital certificate according to the certificate feature, a pre-stored reference certificate feature library, and a preset certificate rule; if the digital certificate is a malicious digital certificate, determine that the first traffic data meets the first condition.
[0020] In a possible implementation manner, a data transmission protocol of the first traffic data is a hypertext transfer protocol http; the determining module is specifically configured to: extract multiple parameter information in the http request message in the first traffic feature, where the multiple parameter information includes a request header, a uniform resource locator URL, request parameters, and response body content; determine whether there is malicious information in the multiple parameter information according to the multiple parameter information, a pre-stored reference malicious information library, and a preset http rule; if there is malicious information, determine that the first traffic data meets the first condition.
[0021] In a possible implementation, the first traffic feature further includes the packet size of the first traffic data transmission; the determining module is further configured to, before determining the estimated data leakage amount of the first traffic data according to the first traffic feature, if the first traffic data meets the first condition, input the traffic log of the first traffic data, the packet size, the communication address, and a preset intelligence value level coefficient into a pre-trained data leakage risk index model to obtain a data leakage risk index, where the data leakage risk index is used to indicate the possibility of data leakage risk existing in the first traffic data; if it is determined that the data leakage risk index exceeds the warning value, determine the estimated data leakage amount of the first traffic data.
[0022] In a possible implementation, the detection module is further configured to: if the first traffic data does not meet the first condition and there is no traffic data in the second traffic data set that meets the first condition within a preset time period, determine whether the second traffic data set meets a second condition, where the second condition is used to indicate that the outgoing data transmission volume and / or the number of requests in the second traffic data set are abnormal; if it is determined that the second traffic data set meets the second condition, determine that there is a suspected data leakage event in the second traffic data set; generate a prompt message and send the prompt message to the user device, where the prompt message is used to indicate to confirm whether there is a data leakage event in the second traffic data set.
[0023] In a possible implementation, the second condition includes at least one of the following: the number of traffic data requests by the same user identifier in the second traffic data set exceeds a second threshold; the total amount of sensitive data transmitted by the second traffic data set exceeds a third threshold; there is abnormal traffic data in the second traffic data set with a transmitted packet size exceeding a fourth threshold.
[0024] In a possible implementation, the detection module is further configured to, after generating the prompt message and sending the prompt message to the user device, receive a response message from the user device, where the response message is used to indicate whether there is a data leakage event in the second traffic data set and the abnormal traffic data in the second traffic data set; if there is a data leakage event in the second traffic data set, parse the abnormal traffic data in the second traffic data set, extract the traffic feature information in the abnormal traffic data, and store it in the malicious traffic feature library.
[0025] In a third aspect, an embodiment of the present application provides a risk detection device, including: at least one processor, and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the at least one processor realizes the method as described in the first aspect and any possible implementation manner by executing the instructions stored in the memory.
[0026] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer instructions, which, when running on a computer, cause the computer to execute the method as described in the first aspect and any possible implementation manner.
[0027] In a fifth aspect, an embodiment of the present application provides a computer program product containing computer instructions, which, when running on a computer, cause the method as described in the first aspect and any possible implementation manner to be realized.
[0028] For the beneficial effects of the second aspect to the fifth aspect, reference may be correspondingly made to the content described in the first aspect above, and details are not repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 FIG. is a schematic diagram of an application scenario of a risk detection method provided by an embodiment of the present application; Figure 2 FIG. is a schematic flowchart of a risk detection method provided by an embodiment of the present application; Figure 3 FIG. is a schematic structural diagram of a risk detection device provided by an embodiment of the present application; Figure 4 FIG. is a schematic structural diagram of a risk detection device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] To make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application. Without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other arbitrarily. And although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0031] In the description and claims of this application and the above-mentioned drawings, the terms "first" and "second" are used to distinguish different objects, rather than to describe a specific order. In addition, the term "comprising" and any variations thereof are intended to cover non-exclusive protection. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices. "Multiple" in this application may mean at least two, for example, it may be two, three or more, and the embodiments of this application do not make limitations.
[0032] The following describes exemplary embodiments of this application with reference to the accompanying drawings, including various details of the embodiments of this application to facilitate understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the disclosure of this application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description. It should be noted that in the embodiments of this application, some industry-existing solutions such as certain software, components, models, etc. may be mentioned, which should be considered exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of this application, but it does not mean that the applicant has already or necessarily used this solution.
[0033] In the technical solution of this application, the acquisition, transmission, storage, use, etc. of data all comply with the requirements of relevant national laws and regulations.
[0034] Please refer to Figure 1 , which is a schematic diagram of an application scenario of a risk detection method provided by an embodiment of this application. As Figure 1 shown, the schematic diagram includes a target server 110 and a risk detection device 120. Among them, a wired or wireless communication can be carried out between the target server 110 and the risk detection device 120.
[0035] The target server 110 refers to a server with data transmission requirements, such as any one of a single business server, a server cluster, or a distributed server of a cloud platform within an enterprise, institution, or organization. The embodiments of this application do not make limitations on this. The risk detection device 120 refers to a device or apparatus with data processing capabilities. The device is, for example, a terminal device or a server. The terminal device includes, but is not limited to, a mobile phone, a personal computer (PC), a tablet computer, a laptop computer, a handheld computer, a mobile internet device (MID), etc. When the risk detection device 120 is a device with data processing capabilities, it may specifically be a functional module in a device. The embodiments of this application do not make limitations on this.
[0036] Exemplarily, the target server 110 can receive access requests from the client in real time, process the access requests, and return response information to the client. The client can be, for example, a browser, a mobile application, etc. For example, the client can send an access request to the target server 110 through a terminal device, and the terminal device can refer to the content described above. The risk detection device 120 can collect the first traffic data between the target server 110 and the client in real time, and parse the first traffic data to obtain the first traffic characteristics. Furthermore, the risk detection device 120 determines whether the first traffic data meets the first condition according to the digital certificate, request information, and communication address in the first traffic characteristics. If the first condition is met, the expected data leakage amount of the first traffic data is further determined according to the first traffic characteristics. Finally, according to the total expected data leakage amount of the first traffic data set within a preset duration and the first threshold, it is determined whether there is a data leakage event in the target server. Among them, the specific implementation manner for the risk detection device 120 to determine whether there is a data leakage event in the target server will be described in detail below.
[0037] Please refer to Figure 2 , which is a schematic flowchart of a risk detection method provided by an embodiment of the present application. Among them, the following description is based on the risk detection device executing Figure 2 the steps shown. The risk detection device is, for example, Figure 1 the risk detection device 120 shown, Figure 2 and the target server involved is, for example, Figure 1 the target server 110 shown.
[0038] S201: Parse the first traffic data from the target server to obtain the first traffic characteristics, where the first traffic data is the outgoing traffic data.
[0039] The first traffic data is the network traffic data between the target server and the client, and usually includes all communication information between the client and the target server, such as request information, response information, digital certificate, etc. The first traffic characteristics include the digital certificate, data transmission protocol, and communication address in the first traffic data. Among them, the data transmission protocol is, for example, the hypertext transfer protocol (HTTP), the transmission control protocol (TCP), etc. The communication address includes the destination internet protocol (IP) address and the source IP address.
[0040] In a possible implementation manner, the first traffic characteristics further include information such as the request time of the first traffic data and the size of the transmitted data packet.
[0041] When the risk detection device obtains the first traffic data, it can parse features such as digital certificates, request details, and response details included in the first traffic data to obtain the first traffic feature.
[0042] S202. Determine whether the first traffic data meets the first condition according to the digital certificate, request information, and communication address in the first traffic feature, where the first condition is used to indicate that the first traffic feature includes malicious traffic features.
[0043] That is to say, the risk detection device determines whether the first traffic feature includes malicious traffic features. If it includes malicious traffic features, it determines that the first traffic data meets the first condition, and then executes S203. If it is determined that the first traffic feature does not include malicious traffic features, it determines that the first traffic data does not meet the first condition. In this case, the risk detection device regards the first traffic data as normal traffic data. Therefore, the detection of the first traffic data can be ended.
[0044] In a possible implementation manner, when the data transmission protocol of the first traffic data is http, the first condition may include at least one of the following: (1) The digital certificate in the first traffic feature is a malicious digital certificate; (2) The http request information in the first traffic feature includes malicious information; (3) The communication address in the first traffic feature is a malicious communication address.
[0045] Among them, when the first condition includes two or three of the above, if the first traffic feature meets any one, it is determined that the first traffic data meets the first condition. That is to say, when it is determined that the first traffic feature includes a malicious digital certificate, or malicious information, or a malicious communication address, it is determined that the first traffic data meets the first condition.
[0046] The following will separately describe the specific implementation manners for the risk detection device to determine whether the first traffic data meets the first condition according to the digital certificate, request information, and communication address in the first traffic feature.
[0047] 1. Digital certificate The risk detection device can extract the certificate features of the digital certificate in the first traffic feature. The certificate features can include issuer information, subject information, validity period, fingerprint, certificate chain, extension fields, etc. Further, the risk detection device can compare according to the certificate features, the pre-stored reference malicious certificate feature library, and the preset certificate rules to determine whether the digital certificate is a malicious digital certificate. Among them, the reference malicious certificate feature library can be extracted from the malicious digital certificates collected in advance, and the reference malicious certificate feature library can include untrusted issuer information, untrusted fingerprints, untrusted subject information, etc. The certificate rules can include certificate rules in multiple dimensions, such as dimensions of registrant, registration time, registration country, etc., which will not be elaborated here one by one. Among them, the certificate rules can be summarized and extracted by the staff from malicious digital certificates.
[0048] Specifically, the risk detection device can match the certificate features with the reference malicious certificate feature library to determine whether there are malicious certificate features included in the reference malicious certificate feature library in the certificate features. If included, it is determined that the digital certificate is a malicious digital certificate. If not included, it is determined that the digital certificate is a normal digital certificate. The risk detection device can also compare the certificate features with the certificate rules to determine whether the certificate features conform to the certificate rules. If not, it is determined that the digital certificate is a malicious certificate feature. If it conforms, it is determined that the digital certificate is a normal digital certificate. When it is determined that the digital certificate is a malicious digital certificate, it is determined that the first traffic data meets the first condition.
[0049] Exemplarily, taking the issuer in the certificate features as an example, when the risk detection device determines that the issuer of the digital certificate is consistent with the issuer information in the reference malicious certificate feature library, it determines that the digital certificate is a malicious digital certificate. Otherwise, it is not. Another example is taking the validity period in the certificate features as an example. If it is determined that the current usage time of the digital certificate has exceeded the specified validity period of the digital certificate, it is determined that the digital certificate is a malicious digital certificate. Otherwise, it is not.
[0050] In a possible implementation, when it is determined that the digital certificate is not a malicious digital certificate based on the reference malicious certificate feature library and certificate rules, it is also possible to further determine whether the digital certificate is a malicious digital certificate from the behavioral characteristics of the digital certificate to improve the accuracy of malicious digital certificate detection, wherein the behavioral characteristics may include certificate chain, certificate update frequency, certificate usage mode, etc. For example, the risk detection device may determine the integrity of the certificate chain to determine whether the digital certificate is a malicious digital certificate. Wherein, the certificate chain includes an intermediate certificate and a root certificate. Based on this, if it is determined that the certificate chain is incomplete, the digital certificate is determined to be a malicious digital certificate. For another example, if it is determined that the update frequency of the digital certificate is high, that is, it is frequently updated, the digital certificate is determined to be a malicious digital certificate. For another example, the risk detection device may extract the usage mode of the digital certificate from network mapping or communication session traffic, and detect whether the usage mode of the digital certificate is an abnormal certificate usage mode. If it is an abnormal usage mode, the digital certificate is determined to be a malicious digital certificate, for example, the digital certificate is associated with a malicious website.
[0051] It should be noted that in order to improve the accuracy of detecting malicious digital certificates, the digital certificate can be determined to be a malicious certificate only when it is detected that the digital certificate satisfies the above-mentioned multiple conditions at the same time. The specific setting can be based on actual needs, and the embodiment of the present application does not limit this.
[0052] In a possible implementation, in order to improve the efficiency of digital certificate recognition, malicious digital certificates can be identified by training a malicious digital certificate recognition model. The malicious digital certificate recognition model can be a machine learning model, such as a decision tree, a random forest, a support vector machine, a neural network model, etc. Specifically, the risk detection device can extract a malicious feature set from a pre-collected sample malicious digital certificate set, and extract a reference feature set from a sample normal digital certificate set. The reference feature set may include reference certificate features, reference behavior features, and other reference features. Among them, the malicious feature set may include malicious certificate features, behavior features, and other malicious features. Other malicious features include comparing the sample malicious digital certificate with other historical malicious digital certificates, and extracting compatible new malicious features from them. Exemplarily, by comparing the current certificate with the historical certificate, the similarities and differences between the two certificates are extracted, and new features are extracted from the current certificate, and it is checked whether the new feature is compatible with the malicious features in the historical certificate. If compatible, it means that the new feature is a malicious feature. The content of the malicious certificate feature may correspond to the content of the reference malicious certificate feature library mentioned above, and the content of the behavior feature may correspond to the content of the behavior feature mentioned above. The content of the reference feature set may correspond to the content of the malicious feature set mentioned above, and will not be repeated here.
[0053] The risk detection device can train the initial malicious digital certificate recognition model according to the extracted malicious feature set and normal feature set, combined with the cross-validation algorithm, until the trained malicious digital certificate recognition model is obtained. The trained malicious digital certificate model can identify the malicious features included in the malicious digital certificate, thereby determining that the traffic data where the malicious digital certificate is located does not meet the first condition.
[0054] In a possible implementation manner, the digital certificates (including normal digital certificates and malicious digital certificates) detected by the malicious digital certificate recognition model can be stored in the malicious feature set for updating and iterating the malicious certificate recognition model. To improve the accuracy of the malicious digital certificate recognition model in identifying malicious digital certificates, malicious digital certificates can also be obtained from other channels to expand the malicious feature set for updating the malicious digital certificate recognition model.
[0055] 2. Request information When the risk detection device determines that the data transmission protocol of the first traffic data is http, it can extract multiple parameter information of the http request information in the first traffic feature. Among them, the http request information specifically includes: http request detail information and http response detail information. The http request detail information includes request line, request headers, request body content, response headers, etc. The request line includes request method, request target, http version, etc. The request headers include request host, content type accepted by the client, data type, credentials, etc. The http response detail information includes response status, response headers, response body content. Among them, the response status includes http version, status code, status message, etc. The response headers include content type, length, service, cookies, etc. The multiple parameter information includes at least request headers, Uniform Resource Locator (URL), request parameters, and response body content.
[0056] The risk detection device can determine whether there is malicious information in the multiple parameter information according to the multiple parameter information, the pre-stored reference malicious information library, and the preset http rules. The reference malicious information library can be extracted from the pre-collected malicious http traffic. The reference malicious information library can include abnormal URLs, abnormal request parameters (such as parameter names in the blacklist), abnormal request headers, sensitive information, etc. The preset http rules can be summarized by the staff from the malicious http traffic. The http rules are used to judge whether the parameter information is abnormal. For example, check whether the parameter name conforms to the preset format, whether the logical relationship and dependency relationship of the parameter combination are reasonable, whether the type of the parameter value conforms to the preset format, whether the length of the parameter value conforms to the preset length, whether the parameter value is within the preset reasonable range, etc.
[0057] Specifically, multiple parameter information can be compared with a pre-stored reference malicious information database to determine whether there is malicious information in the reference malicious information database among the multiple parameter information. At the same time, according to the http rules, it is determined whether there is malicious information that does not conform to the http rules among the multiple parameter information. If malicious information is determined to exist in any of the above cases, it is determined that the first traffic data meets the first condition; otherwise, it does not meet the condition.
[0058] Exemplarily, taking the parameter information as the URL as an example, if it is determined that the URL is an abnormal URL in the reference malicious information database, it is determined that the URL belongs to malicious information. For another example, taking the parameter information as the request parameter as an example, if it is determined that the parameter name in the request parameter is the parameter name in the reference malicious information database, it is determined that the request parameter is malicious information, or if it is determined that the content in the request parameter involves sensitive information included in the reference malicious information database, it is determined that the request parameter is malicious information.
[0059] In a possible implementation manner, in order to improve the detection efficiency of http request information, a malicious information recognition model can be trained to recognize malicious information in the http request information. The malicious information recognition model can be a machine learning model, and the content of the machine learning model can be correspondingly referred to the content described above. Specifically, the risk detection device can extract a normal information feature set and a malicious information feature set from a pre-collected sample normal traffic data set and a sample malicious traffic data set. Among them, the content included in the malicious information feature set can be the malicious information in the previous reference malicious information database and the malicious information determined based on the http rules. The content of the normal information feature set is the request information extracted from the sample normal traffic data set, which will not be elaborated here.
[0060] The risk detection device can train the initial malicious information recognition model according to the extracted normal information feature set and malicious information feature set, combined with the cross-validation algorithm, until the trained malicious information recognition model is obtained. The trained malicious information recognition model can recognize the malicious information contained in the malicious traffic information, so as to determine that the malicious traffic information does not meet the first condition.
[0061] In a possible implementation manner, the malicious information detected by the malicious information recognition model can be stored in the malicious information feature set for updating and iterating the malicious information recognition model. In order to improve the accuracy of the malicious information recognition model in recognizing malicious information, malicious information can also be obtained from other channels to expand the malicious information feature set for updating the malicious information recognition model.
[0062] 3. Communication Address The risk detection device can extract the communication address of the first traffic data from the first traffic characteristics, and determine whether the communication address is a malicious communication address according to the communication address and a pre-stored reference malicious communication address library. The reference malicious communication address library is a pre-configured blacklist address library, which can be extracted from malicious traffic data. Specifically, if the communication address matches any malicious communication address in the reference malicious communication address library, it is determined that the communication address is a malicious communication address, and then it is determined that the first traffic data meets the first condition. If the communication address does not match each malicious communication address in the reference malicious communication address library, it is determined that the communication address is not a malicious communication address, and then it is determined that the first traffic data does not meet the first condition.
[0063] In a possible implementation manner, when the communication address is not a malicious communication address, it can be determined whether to determine the communication address as a malicious communication address according to whether the first traffic data includes a malicious digital certificate or malicious information. Specifically, if it is determined that the first traffic data includes a malicious digital certificate or malicious information, it can be determined that the first traffic data is abnormal traffic data. Therefore, the communication address in the first traffic characteristics can be stored in the malicious communication address library to enrich the malicious communication address library. If it is determined that the first traffic data does not include a malicious digital certificate or malicious information, the communication address in the first traffic data is stored in the whitelist address library.
[0064] When the first traffic data does not meet the first condition and there is no traffic data that meets the first condition in the second traffic data set within a preset duration, in order to avoid missed detection, it can be determined whether there is a data leakage event according to the outgoing data transfer volume and / or the number of requests in the second traffic data set. The second traffic data set includes the first traffic data. The preset duration is pre-configured in the risk detection device, and the specific value can be set according to actual needs, which is not limited in the embodiments of the present application.
[0065] In a possible implementation manner, if it is determined that the first traffic data does not meet the first condition and there is no traffic data that meets the first condition in the second traffic data set within a preset duration, then it is determined whether the second traffic data set meets the second condition. The second condition is used to indicate that the outgoing data transfer volume and / or the number of requests in the second traffic data set are abnormal. If it is determined that the second traffic data set meets the second condition, it is determined that there is a suspected data leakage event in the second traffic data set. If it is determined that the second traffic data set does not meet the second condition, it is determined that there is no suspected data leakage event in the second traffic data set.
[0066] That is to say, if no abnormal traffic data is detected within a period of time, but the outgoing data transfer volume of the second traffic data set within this period of time is abnormally large or the number of requests is abnormally large, there may also be a data leakage event.
[0067] In a possible implementation, the second condition may include at least one of the following: (1) The amount of traffic data requested by the same user identifier in the second traffic data set exceeds a second threshold.
[0068] That is to say, if the same user frequently initiates a large number of requests within a preset period of time, it can be determined that the number of requests from the user is abnormal and there is a suspected data leakage event. In other words, from the perspective of outbound, if the target server transmits a large amount of data to the same IP address, it is considered that there is a suspected data leakage event.
[0069] (2) The total amount of sensitive data transmitted by the second traffic data set exceeds the third threshold.
[0070] When the total amount of sensitive data transmitted by the second traffic data set exceeds the third threshold within the preset time period, it can be determined that there is a suspected data leakage event. Sensitive information includes, for example, user information, commercial data, etc.
[0071] (3) There is abnormal traffic data in the second traffic data set, in which the size of the transmitted data packet exceeds the fourth threshold.
[0072] When the second traffic data set stores traffic data with an abnormally large amount of transmitted data, it can also be determined that there is a suspected data leakage event.
[0073] Among them, when the second condition includes the above two or three items, if the second traffic data set meets any one of the items, it is determined that the second traffic data set meets the second condition.
[0074] Furthermore, in order to improve the accuracy of detecting data leakage events, when it is determined that there is a suspected data leakage event in the second traffic data set, prompt information can be generated and sent to the user device, and the prompt information is used to indicate whether the second traffic data set has a data leakage event. In other words, when it is determined that there is a suspected data leakage event in the second traffic data set, the staff is prompted to further determine whether there is a data leakage event in the second traffic data set, thereby improving the accuracy of detecting data leakage events.
[0075] In a possible implementation, the risk detection device may receive response information from a user device, where the response information is used to indicate whether there is a data leakage event in the second traffic data set and abnormal traffic data in the second traffic data set. If it is determined that there is a data leakage event in the second traffic data set, the abnormal traffic data in the second traffic data set is parsed, and the traffic feature information of the abnormal traffic data is extracted and stored in a malicious traffic feature library, which is used to update a malicious digital certificate recognition model and a malicious information recognition model, and update a reference malicious communication address library. Further, the risk detection device may also extract the traffic feature information of the normal traffic data in the second traffic data set, which is used to update the malicious digital certificate recognition model and the malicious information recognition model, and update the whitelist address library in the communication address. The whitelist address library can be used to filter normal traffic data, so as to improve the detection efficiency of the risk detection device for abnormal traffic data and reduce the misjudgment of normal traffic data.
[0076] S203. Determine the predicted data leakage amount of the first traffic data according to the first traffic feature.
[0077] Specifically, the risk detection device may determine the packet size of the first traffic data transmission from the first traffic feature. The risk detection device may further determine the size of the packet header in the packet, and determine the difference between the packet size and the size of the packet header in the packet as the predicted data leakage amount of the first traffic data.
[0078] In order to improve the accuracy of detecting data leakage events, the risk detection device may evaluate the data leakage risk of the first traffic data, so as to determine the possibility that the first traffic data has a data leakage risk, and filter out traffic data with a lower data leakage risk, reducing the impact of normal traffic data on subsequent detection results.
[0079] In a possible implementation, before the risk detection device determines the predicted data leakage amount of the first traffic data, the traffic log, packet size, communication address, and a preset intelligence value level coefficient of the first traffic data may be input into a pre-trained data leakage risk index model to obtain a data leakage risk index of the first traffic data, which is used to indicate the possibility that the first traffic data has a data leakage risk. Among them, the intelligence value level coefficient may be pre-configured in the risk detection device. The intelligence value level coefficient is related to the attack type, that is, different attack types correspond to different intelligence value level coefficients. The intelligence value level coefficients corresponding to different attack types may be set by the user, and the embodiments of the present application do not limit this.
[0080] If the data leakage risk index exceeds the warning value, it can be determined that the data leakage risk of the first traffic data is relatively high, and then the data leakage volume of the first traffic data can be determined. If the data leakage risk index does not exceed the warning value, it indicates that the possibility of data leakage risk for the first traffic data is relatively low. Therefore, it is not necessary to calculate the expected data leakage volume of the first traffic data. Among them, the warning value is preset in the risk detection device, and the specific value can be set according to actual needs or calculated based on historical abnormal traffic data. This application embodiment does not limit this.
[0081] Exemplarily, if the determined data leakage risk index is 85 and the preset warning value is 80, at this time 85>80, it can be determined that the data leakage risk of the first traffic data is relatively high, and then it is necessary to determine the data leakage volume of the first traffic data.
[0082] Among them, the pre-trained data leakage risk index model can be trained based on traffic logs, packet sizes, communication addresses, and intelligence value level coefficients in the sample traffic dataset.
[0083] S204. Determine the detection result of the target server according to the expected total data leakage amount of the first traffic dataset within the preset duration and the first threshold. The detection result is used to indicate whether there is a data leakage event in the target server. The first traffic dataset includes the first traffic data.
[0084] Among them, the first traffic dataset includes traffic data that has met the first condition and whose expected data leakage volume has been calculated in the target server within the preset duration. The specific implementation manners of determining whether each traffic data in the first traffic dataset meets the first condition and calculating the expected data leakage volume can respectively refer to the content of the first traffic data, which will not be elaborated here. The first threshold is pre-configured in the risk detection device, and the first threshold can be set according to actual needs, such as being set according to the experience of the staff. This application embodiment does not limit this.
[0085] Specifically, the risk detection device can determine whether the expected total data leakage amount of the first traffic dataset exceeds the first threshold. If it exceeds the first threshold, it indicates that the data leakage volume is abnormal, and it is determined that there is a data leakage event in the target server. If it does not exceed the first threshold, it indicates that the data leakage volume is normal, and it is determined that there is no data leakage event in the target server.
[0086] Exemplarily, a calculation formula for calculating the expected total data leakage amount is as follows: Expected total data leakage amount = sum of the packet sizes transmitted by the first traffic dataset - number of packets * packet header In a possible implementation, when it is determined that there is a data leakage event in the target server, a warning message is generated and sent to the user device so that the staff can process it in time after viewing the warning message. The warning message is used to indicate that there is a data leakage event in the target server.
[0087] In a possible implementation, the warning message is also used to indicate the malicious IP address associated with the data leakage event and the estimated total amount of data leakage.
[0088] Based on the same inventive concept, an embodiment of the present application provides a risk detection device, which is used to implement any of the above risk detection methods, for example Figure 2 the risk detection method shown, and the device can also implement the functions of the risk detection device in the foregoing.
[0089] Please refer to Figure 3 , which is a schematic structural diagram of a risk detection device provided by an embodiment of the present application. As Figure 3 shown, the risk detection device 300 includes a parsing module 301, a determination module 302, and a detection module 303.
[0090] Exemplarily, the parsing module 301 is configured to parse the first traffic data from the target server to obtain a first traffic feature, and the first traffic data is outgoing traffic data; the determination module 302 is configured to determine whether the first traffic data meets a first condition according to the digital certificate, request information, and communication address in the first traffic feature, and the first condition is used to indicate that the first traffic feature includes malicious traffic features; the determination module 302 is further configured to, if the first traffic data meets the first condition, determine the estimated data leakage amount of the first traffic data according to the first traffic feature; the detection module 303 is configured to determine the detection result of the target server according to the total estimated data leakage amount of the first traffic data set within a preset duration and a first threshold, and the detection result is used to indicate whether there is a data leakage event in the target server, and the first traffic data set includes the first traffic data.
[0091] In a possible implementation, the data transmission protocol of the first traffic data is the Hypertext Transfer Protocol http; the first condition includes at least one of the following: the digital certificate in the first traffic feature is a malicious digital certificate; the http request information in the first traffic feature includes malicious information; the communication address in the first traffic feature is a malicious communication address.
[0092] In a possible implementation, the determination module 302 is specifically configured to: extract the certificate feature of the digital certificate in the first traffic feature; determine whether the digital certificate is a malicious digital certificate according to the certificate feature, the pre-stored reference certificate feature library, and the preset certificate rule; if the digital certificate is a malicious digital certificate, determine that the first traffic data meets the first condition.
[0093] In a possible implementation, the data transmission protocol of the first traffic data is the Hypertext Transfer Protocol http; the determining module 302 is specifically configured to: extract multiple parameter information in the http request information of the first traffic feature, where the multiple parameter information includes a request header, a Uniform Resource Locator (URL), request parameters, and response body content; determine whether there is malicious information in the multiple parameter information according to the multiple parameter information, a pre-stored reference malicious information library, and a preset http rule; if there is malicious information, determine that the first traffic data meets the first condition.
[0094] In a possible implementation, the first traffic feature further includes the packet size of the first traffic data transmission; the determining module 302 is further configured to, before determining the predicted data leakage amount of the first traffic data according to the first traffic feature, if the first traffic data meets the first condition, input the traffic log, packet size, communication address, and a preset intelligence value level coefficient of the first traffic data into a pre-trained data leakage risk index model to obtain a data leakage risk index, where the data leakage risk index is used to indicate the possibility of data leakage risk existing in the first traffic data; if it is determined that the data leakage risk index exceeds the warning value, determine the predicted data leakage amount of the first traffic data.
[0095] In a possible implementation, the detection module 303 is further configured to: if the first traffic data does not meet the first condition and there is no traffic data that meets the first condition in the second traffic data set within a preset duration, determine whether the second traffic data set meets the second condition, where the second condition is used to indicate that the outgoing data transmission volume and / or the number of requests in the second traffic data set are abnormal; if it is determined that the second traffic data set meets the second condition, determine that there is a suspected data leakage event in the second traffic data set; generate a prompt message and send the prompt message to the user device, where the prompt message is used to indicate to confirm whether there is a data leakage event in the second traffic data set.
[0096] In a possible implementation, the second condition includes at least one of the following: the number of traffic data requested by the same user identifier in the second traffic data set exceeds a second threshold; the total amount of sensitive data transmitted by the second traffic data set exceeds a third threshold; there is abnormal traffic data in the second traffic data set whose transmitted packet size exceeds a fourth threshold.
[0097] In a possible implementation, the detection module 303 is further configured to, after generating a prompt message and sending the prompt message to the user device, receive response information from the user device, where the response information is used to indicate whether a data leakage event exists in the second traffic data set and abnormal traffic data in the second traffic data set; if a data leakage event exists in the second traffic data set, parse the abnormal traffic data in the second traffic data set, extract traffic feature information in the abnormal traffic data, and store the traffic feature information in the malicious traffic feature library.
[0098] Based on the same inventive concept, an embodiment of the present application provides a risk detection device, which is used to implement any of the above risk detection methods, for example, Figure 2 the risk detection method shown, and this device can also implement the functions of the risk detection device described above.
[0099] Please refer to Figure 4 , which is a schematic structural diagram of a risk detection device provided by an embodiment of the present application. As Figure 4 shown, the risk detection device 400 includes at least one processor 401 and a memory 402 communicatively connected to the at least one processor 401.
[0100] Among them, the processor 401 can be a general-purpose processor or a dedicated processor, etc. The processor 401 includes, for example, a baseband processor or a central processing unit, etc. The baseband processor can be used to process communication protocols and communication data. The central processing unit can be used to control the risk detection device 400, execute software programs, and / or process data. Different processors can be independent devices or can be provided in one or more processing circuits, for example, integrated on one or more application-specific integrated circuits.
[0101] In one embodiment, the memory 402 stores instructions executable by the at least one processor 401, and the at least one processor 401 realizes the functions of the risk detection device described above by executing the instructions stored in the memory 402. Correspondingly, the steps executed by the risk detection device described above can also be realized.
[0102] In this embodiment, the risk detection device 400 can also realize the functions of the risk detection device 300 described above, and at least one processor 401 in the risk detection device 400 can also realize the functions of the parsing module 301, the determination module 302, and the detection module 303 described above.
[0103] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, which stores computer instructions. When the computer instructions run on a computer, the computer is caused to execute any of the risk detection methods described above, for example, Figure 2 the risk detection method shown.
[0104] Based on the same inventive concept, an embodiment of the present application provides a computer program product, which includes computer instructions. When the computer instructions run on a computer, the risk detection method as described above in any of the foregoing, for example, Figure 2 the risk detection method shown is implemented.
[0105] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0106] The present application is described with reference to the flowcharts and / or block diagrams of the method, device (system), and computer program product according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0107] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0109] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these modifications and variations.
Claims
1. A risk detection method, characterized in that: include: Parsing first traffic data from a target server to obtain a first traffic feature, wherein the first traffic data is outbound traffic data; Determining whether the first traffic data satisfies a first condition according to the digital certificate, request information, and communication address in the first traffic feature, wherein the first condition is used to indicate that the first traffic feature includes a malicious traffic feature; If the first flow data satisfies the first condition, determining an estimated data leakage amount of the first flow data according to the first flow characteristics; The detection result of the target server is determined based on the estimated total amount of data leakage of the first traffic data set within a preset time period and a first threshold, and the detection result is used to indicate whether there is a data leakage incident on the target server. The first traffic data set includes the first traffic data.
2. The method according to claim 1, characterized in that The data transmission protocol of the first traffic data is the Hypertext Transfer Protocol (http); the first condition includes at least one of the following: The digital certificate in the first traffic feature is a malicious digital certificate; The http request information in the first traffic feature includes malicious information; The communication address in the first traffic feature is a malicious communication address.
3. The method according to claim 1, characterized in that Determining whether the first flow data satisfies a first condition according to the digital certificate in the first flow feature includes: Extracting the certificate feature of the digital certificate in the first traffic feature; Determining whether the digital certificate is a malicious digital certificate based on the certificate characteristics, a pre-stored reference certificate characteristics library, and a preset certificate rule; If the digital certificate is a malicious digital certificate, it is determined that the first traffic data satisfies the first condition.
4. The method according to claim 1, characterized in that: The data transmission protocol of the first flow data is the Hypertext Transfer Protocol (http); and determining whether the first flow data satisfies a first condition according to the request information in the first flow feature includes: Extracting multiple parameter information in the http request information in the first traffic feature, the multiple parameter information including a request header, a uniform resource locator URL, request parameters and response body content; Determine whether malicious information exists in the multiple parameter information according to the multiple parameter information, a pre-stored reference malicious information library and a preset http rule; If malicious information exists, it is determined that the first traffic data meets the first condition.
5. The method according to any one of claims 1 to 4, characterized in that: The first traffic characteristic also includes the data packet size of the first traffic data transmission; Before determining the estimated data leakage amount of the first traffic data according to the first traffic feature, the method further includes: If the first traffic data satisfies the first condition, the traffic log of the first traffic data, the data packet size, the communication address, and a preset intelligence value level coefficient are input into a pre-trained data leakage risk index model to obtain a data leakage risk index, where the data leakage risk index is used to indicate the possibility that the first traffic data has a data leakage risk; If it is determined that the data leakage risk index exceeds the warning value, the estimated data leakage amount of the first traffic data is determined.
6. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: If the first traffic data does not satisfy the first condition, and there is no traffic data satisfying the first condition in the second traffic data set within a preset time period, then determining whether the second traffic data set satisfies a second condition, where the second condition is used to indicate that the amount of data transmitted and / or the number of requests popped from the second traffic data set is abnormal; If it is determined that the second traffic data set meets the second condition, then determining that there is a suspected data leakage event in the second traffic data set; Generate prompt information and send the prompt information to the user equipment, where the prompt information is used to indicate whether there is a data leakage event in the second traffic data set.
7. The method according to claim 6, characterized in that The second condition includes at least one of the following: The amount of traffic data requested by the same user identifier in the second traffic data set exceeds a second threshold; The total amount of sensitive data transmitted by the second traffic data set exceeds a third threshold; The second traffic data set contains abnormal traffic data whose transmitted data packet size exceeds a fourth threshold.
8. The method according to claim 6, characterized in that After generating the prompt information and sending the prompt information to the user equipment, the method further includes: receiving response information from a user device, wherein the response information is used to indicate whether a data leakage event exists in the second traffic data set and abnormal traffic data in the second traffic data set; If there is a data leakage event in the second traffic data set, the abnormal traffic data in the second traffic data set is parsed, and the traffic feature information in the abnormal traffic data is extracted and stored in a malicious traffic feature library.
9. A risk detection device, characterized in that: include: A parsing module, used for parsing first flow data from a target server to obtain a first flow feature, wherein the first flow data is outbound flow data; a determination module, configured to determine whether the first traffic data satisfies a first condition according to the digital certificate, request information, and communication address in the first traffic feature, wherein the first condition is configured to indicate that the first traffic feature includes a malicious traffic feature; The determination module is further configured to determine an estimated data leakage amount of the first flow data according to the first flow feature if the first flow data satisfies the first condition; A detection module is used to determine the detection result of the target server based on the estimated total amount of data leakage of the first traffic data set within a preset time period and a first threshold, and the detection result is used to indicate whether there is a data leakage incident on the target server, and the first traffic data set includes the first traffic data.
10. A risk detection device, characterized in that: include: at least one processor, and a memory communicatively coupled to the at least one processor; The memory stores instructions that can be executed by the at least one processor, and the at least one processor implements the method according to any one of claims 1 to 8 by executing the instructions stored in the memory.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed on a computer, the computer is caused to execute the method according to any one of claims 1 to 8.
12. A computer program product, characterized in that The method comprises computer instructions, which, when executed on a computer, enable the method according to any one of claims 1 to 8 to be implemented.
Citation Information
Cited By
Food and drug safety intelligent supervision and analysis system and method based on big data
CN121350116A
Big data-based food and drug safety intelligent supervision and analysis system and method
CN121350116B