Risk detection method and device for configuration file, cluster, storage medium and program product

By performing dual detection of content and historical behavior on the configuration files on the cloud server, the problem of difficult to identify configuration files with no obvious violation words but with illegal content in the existing technology is solved, the accuracy and effectiveness of detection are improved, and the content security management of cloud service manufacturers is enhanced.

CN120074859APending Publication Date: 2025-05-30SHENZHEN HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411999699.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

It is difficult to identify the configuration files stored on cloud servers in the existing technology. Although there are no obvious violation words, there is actually a risk of violation content, which poses a threat to the content security of cloud service manufacturers.

Method used

By detecting the content dimension and historical behavior dimension of the configuration file, the first detection result and the second detection result are obtained, and the risk detection rate and accuracy of the detection are improved by combining the two.

Benefits of technology

It significantly improves the effectiveness and accuracy of risk detection of configuration files without obvious violation words, and enhances the content security management capabilities of cloud service manufacturers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120074859A_ABST
    Figure CN120074859A_ABST
Patent Text Reader

Abstract

The invention discloses a risk detection method and device for configuration files, a cluster, a storage medium and a program product, and relates to the technical field of computers. By detecting the configuration files without obvious violation words in different dimensions, the detection rate and the accuracy are remarkably improved. The method is applied to a cloud server, the cloud server stores a configuration file uploaded by a tenant, the configuration file is used for positioning resources, stored outside the cloud server, of the tenant, and the method comprises the steps that the configuration file is acquired; the configuration file comprises a domain name used for positioning resources; obtaining a first detection result of the configuration file; the first detection result is used for representing whether the configuration file is a candidate configuration file; the risk that the resource positioned by the candidate configuration file has illegal content is higher than a first risk threshold value; if the first detection result represents that the configuration file is the candidate configuration file, obtaining historical behavior portrait data of the tenant; according to the historical behavior portrait data, a second detection result of the candidate configuration file is determined, and the first risk threshold value is smaller than a second risk threshold value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a method, apparatus, cluster, storage medium, and program product for risk detection of configuration files. Background Art

[0002] With the rapid development and wide popularization of Internet services, Internet applications have also developed diversely. While the Internet has brought great convenience to the general public, the issues of network content security and compliance have become increasingly prominent.

[0003] For example, while major cloud service providers offer convenient storage, computing, and other services to tenants, they also need to detect and handle in a timely manner whether tenants misuse the corresponding services. In the past, some tenants directly stored illegal content on cloud servers. Correspondingly, by detecting whether there are illegal words in the public content of tenants, the providers could ensure a certain recall rate and precision rate.

[0004] However, with the escalation of the confrontation, some tenants store configuration files on cloud servers that do not contain obvious illegal words. These configuration files may only contain a domain name or an Internet Protocol (IP) address, etc. When an application or a web page on a terminal device runs, it will pull illegal content stored in other locations through the domain name in the configuration file. Traditional illegal word detection methods cannot effectively identify such configuration files without obvious illegal words, posing a serious threat to the content security of cloud service providers. Summary of the Invention

[0005] This application provides a method, apparatus, cluster, storage medium, and program product for risk detection of configuration files. Based on the detection results in terms of content dimension and historical behavior dimension for configuration files that do not contain obvious illegal words, the detection rate and accuracy of detecting the risk of whether there is illegal content in the resources located by the configuration files are significantly improved.

[0006] In a first aspect, the present application provides a method for detecting risks of a configuration file, which is applied to a cloud server. The cloud server stores configuration files uploaded by tenants, and the configuration files are used to locate resources stored by the tenants outside the cloud server. The method includes: obtaining a configuration file; the configuration file includes a domain name for locating resources; obtaining a first detection result of the configuration file; the first detection result is used to characterize whether the configuration file is a candidate configuration file; the risk that the resources located by the candidate configuration file have illegal content is higher than a first risk threshold; if the first detection result indicates that the configuration file is a candidate configuration file, obtaining historical behavior portrait data of the tenant; determining a second detection result of the candidate configuration file according to the historical behavior portrait data; the second detection result is used to characterize whether the risk that the resources located by the candidate configuration file have illegal content is higher than a second risk threshold; wherein, the first risk threshold is less than the second risk threshold.

[0007] It can be understood that through the first detection result of the content of the configuration file, the risk that the resources located by the configuration file have illegal content can be initially determined. Furthermore, when the first detection result indicates that the configuration file is a candidate configuration file, the risk that the resources located by the candidate configuration file have illegal content is detected again according to the historical behavior portrait data of the tenant to which the candidate configuration file belongs. Through the detection results of two different dimensions, the effectiveness of detecting configuration files without obvious illegal words can be significantly improved.

[0008] In a possible implementation manner, the historical behavior portrait data of the tenant includes: the number of files in each data bucket associated with the tenant on the cloud server, the life cycle of the configuration file stored by the tenant on the cloud server, one or more of the total number of historical accesses, and the life cycle of the configuration file is determined according to the first access time and the most recent access time of the configuration file.

[0009] It can be understood that when a tenant locates resources with illegal content through a configuration file, the behavior pattern is usually different from that of ordinary tenants. That is to say, by analyzing one or more pieces of historical behavior portrait data, the effectiveness of detecting the risks of the configuration file can be further improved.

[0010] In a possible implementation manner, determining a second detection result of a candidate configuration file based on historical behavior portrait data includes: obtaining one or more of a first ratio of data buckets with the number of files lower than a first target number threshold to the data buckets associated with a tenant, a second ratio of configuration files with a life cycle lower than a first target life cycle threshold to target configuration files, and a third ratio of configuration files with the total number of historical accesses lower than a first target number threshold to target configuration files, where the target configuration files are all public configuration files stored by the tenant on a cloud server; if one or more of the first ratio, the second ratio, and the third ratio are higher than the corresponding ratio thresholds, it is determined that the risk of the resource located by the candidate configuration file having illegal content is higher than a second risk threshold.

[0011] It can be understood that when a tenant locates a resource with illegal content through a configuration file, the tenant usually replaces the configuration file frequently in an attempt to reduce the probability of being detected and disposed of. That is to say, the life cycle of the configuration files stored by such a tenant in the data bucket is short, and the short life cycle also results in a small total number of historical accesses. Moreover, such a tenant will not store a large number of files in one data bucket, but will disperse the configuration files in multiple data buckets, and usually there is only one or a few configuration files in each data bucket. For these characteristics, by counting the proportions of the configuration files and data buckets that meet the corresponding characteristics, the accuracy of detecting the risk of the configuration file can be improved.

[0012] In a possible implementation manner, determining a second detection result of a candidate configuration file based on historical behavior portrait data includes: determining a second detection result of the candidate configuration file according to the historical behavior portrait data of the tenant for the candidate configuration file in the historical behavior portrait data; where the historical behavior portrait data of the tenant for the candidate configuration file includes one or more of the number of files in the data bucket where the candidate configuration file is located, the historical access times of the candidate configuration file, and the life cycle of the candidate configuration file, and the life cycle of the candidate configuration file is determined according to the first access time and the most recent access time of the candidate configuration file.

[0013] It can be understood that in some cases, the proportion of high-risk configuration files may not be high. For example, a certain configuration file of an ordinary tenant is tampered with or misstored, or the tenant intentionally stores a large number of compliant files in an attempt to bypass detection. In this way, through the historical behavior portrait data of the tenant for the candidate configuration file, risk detection can be carried out more specifically, thereby improving the accuracy of detecting the risk of the configuration file.

[0014] In a possible implementation, based on the historical behavior portrait data of the tenant for the candidate configuration file, determining the second detection result of the candidate configuration file includes: if the historical behavior portrait data meets one or more of the following conditions, determining that the risk of the resource having illegal content is higher than the second risk threshold, and the foregoing conditions include: the number of files in the data bucket where the candidate configuration file is located is lower than the second target quantity threshold; the historical access times of the candidate configuration file are lower than the second target times threshold; the life cycle of the candidate configuration file is lower than the second target life cycle threshold.

[0015] It can be understood that by comparing the historical behavior portrait data of the tenant for the candidate configuration file with the corresponding thresholds, it is possible to accurately determine whether there is data that does not meet the corresponding thresholds in the historical behavior portrait data, and thus improve the accuracy of detecting the risk of the configuration file.

[0016] In a possible implementation, if the configuration file also includes an Internet Protocol (IP) address for locating the resource, and if the first detection result indicates that the configuration file is a candidate configuration file, obtaining the historical behavior portrait of the tenant includes: if the IP address in the configuration file hits the preset risk IP list and the first detection result indicates that the configuration file is a candidate configuration file, obtaining the historical behavior portrait of the tenant.

[0017] It can be understood that by detecting other data (such as IP address) in the configuration file, it is possible to further broaden the detection dimension of high-risk configuration files, prevent the tenant from bypassing the detection by storing the IP address, and further improve the accuracy of detecting the risk of the configuration file.

[0018] In a possible implementation, the first detection result is determined by performing clustering or classification analysis on the features obtained by extracting the features of the domain name in the configuration file.

[0019] It can be understood that for a large number of domain names that can be accurately classified, the classification analysis method can be adopted; for the situation where the domain name classification is complex, the classification accuracy is low and it requires a large amount of manpower and material resources, the clustering analysis method can be adopted. Different categories are formed through the features extracted from the domain name, avoiding manual classification of domain names and improving efficiency.

[0020] In a possible implementation, obtaining the first detection result of the configuration file includes: inputting the domain name in the configuration file into a violation detection model to obtain the first detection result output by the violation detection model, and the violation detection model includes at least one neural network model trained according to sample domain names and the corresponding classification labels.

[0021] It can be understood that the neural network model trained according to the sample domain names and the corresponding classification labels of the sample domain names can detect the domain names in the input configuration file and obtain the corresponding first detection result.

[0022] In a possible implementation, the violation detection model is a Convolutional Neural Network-Long Short-Term Memory Network (CNN-LSTM) model.

[0023] It can be understood that the CNN-LSTM model can simultaneously extract the local semantic features and string sequence features of the domain name, and can accurately detect the risk that the domain name is used to locate resources containing violation content.

[0024] In a possible implementation, obtaining the first detection result of the configuration file includes: if the violation detection model includes two or more neural network models, determining the first detection result according to the weight coefficients corresponding to each neural network model and the output result of the violation detection model.

[0025] It can be understood that by performing weighted calculation on the output results of each violation detection model through the weight coefficients corresponding to each neural network model, the accuracy when determining the first detection result can be improved, which in turn helps to improve the accuracy of detecting the configuration file.

[0026] In a second aspect, the present application provides a risk detection device for a configuration file. The risk detection device for the configuration file is used to execute any one of the risk detection methods for the configuration file provided in the first aspect above.

[0027] In a possible implementation, the present application can divide the functional modules of the risk detection device for the configuration file according to the method provided in the first aspect above. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one module. Exemplarily, the present application can divide the risk detection device for the configuration file into a first acquisition module, a second acquisition module, a third acquisition module, a determination module, etc. according to functions. The descriptions of the possible technical solutions and beneficial effects executed by each of the above-divided functional modules can refer to the technical solutions provided in the first aspect or its corresponding possible implementation manners above, and will not be elaborated here.

[0028] In a third aspect, an embodiment of the present application provides a computing device. The computing device includes a processor and a memory, and the processor is coupled to the memory; the memory is used to store computer instructions, and the computer instructions are loaded and executed by the processor so that the computing device implements the risk detection method for the configuration file as described in the above aspects.

[0029] Fourthly, embodiments of the present application provide a computing device cluster, which includes at least one computing device. Each computing device includes a processor and a memory, and the processor is coupled to the memory. The processors of at least one computing device are configured to execute computer instructions stored in the memory of at least one computing device, so that the computing device cluster executes the risk detection method for the configuration file provided in various alternative implementations of the first aspect above.

[0030] Fifthly, embodiments of the present application provide a computer-readable storage medium, in which at least one computer program instruction is stored. The computer program instruction is loaded and executed by a processor to implement the risk detection method for the configuration file as described in the above aspects.

[0031] Sixthly, embodiments of the present application provide a computer program product, which includes computer instructions stored in a computer-readable storage medium. The processor of a computing device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computing device executes the risk detection method for the configuration file provided in various alternative implementations of the first aspect above.

[0032] For the specific descriptions of the second to sixth aspects and their various implementations in the present application, reference may be made to the detailed descriptions in the first aspect and its various implementations; and, for the beneficial effects of the second to sixth aspects and their various implementations, reference may be made to the analysis of the beneficial effects in the first aspect and its various implementations, which will not be elaborated here.

[0033] These aspects or other aspects of the present application will be more clearly understood in the following description. Description of the Drawings

[0034] Figure 1 A schematic diagram of a prior art provided by an embodiment of the present application;

[0035] Figure 2 A schematic diagram of another prior art provided by an embodiment of the present application;

[0036] Figure 3 A schematic diagram of another prior art provided by an embodiment of the present application;

[0037] Figure 4 A schematic diagram of the architecture of a risk detection system for a configuration file provided by an embodiment of the present application;

[0038] Figure 5 A schematic diagram of the structure of a computing device provided by an embodiment of the present application;

[0039] Figure 6Flowchart of a risk detection method for a configuration file provided by an embodiment of the present application;

[0040] Figure 7 For Figure 6 Schematic diagram of constructing and updating historical behavior portrait data involved in the illustrated embodiment;

[0041] Figure 8 Flowchart of a model training provided by an embodiment of the present application;

[0042] Figure 9 Schematic diagram of the structure of a risk detection device for a configuration file provided by an embodiment of the present application;

[0043] Figure 10 Schematic diagram of a computing device provided by an embodiment of the present application;

[0044] Figure 11 Schematic diagram of a computing device cluster provided by an embodiment of the present application;

[0045] Figure 12 Schematic diagram of the connection method between computing device clusters provided by an embodiment of the present application. Detailed implementation manners

[0046] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0047] As used herein, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0048] Moreover, in the description of the embodiments of the present application, unless otherwise specified, "a plurality of" means two or more than two. "At least one (item)" or similar expressions thereof refer to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c may be single or multiple.

[0049] In addition, to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and roles. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different. At the same time, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner for easy understanding.

[0050] First, an exemplary introduction to the application scenarios of the embodiments of the present application is provided.

[0051] To prevent the improper use of cloud services, each manufacturer usually needs to detect the public content stored by tenants in cloud servers in a timely manner to discover and process illegal content. Existing content security detection technologies generally include the following three types:

[0052] 1) Word matching-based method

[0053] As Figure 1 shown, Figure 1 is a schematic diagram of an existing technology provided by an embodiment of the present application. In this method, it is necessary to first construct a Figure 1 shown illegal word library (including multiple illegal words such as illegal word 1 and illegal word 2) based on expert experience, and then perform word segmentation on the given text to be tested (such as Figure 1 shown text content extracted from public content), to obtain multiple words such as Figure 1 shown word 1, word 2, word 3, and word 4, and then use a string matching algorithm to determine whether these words contain words in the illegal word library. Commonly used algorithms in the aforementioned word segmentation operation include character-based word segmentation, forward maximum matching, bidirectional maximum matching, machine learning-based word segmentation, etc. Commonly used string matching algorithms include the Rabin-Karp algorithm, the Knuth-Morris-Pratt operation (also known as the KMP algorithm), the Aho-Corasick automaton, etc.

[0054] 2) Rule-based method

[0055] As Figure 2 shown, Figure 2This is a schematic diagram of another prior art provided in the embodiment of the present application. This method requires experts to analyze a large amount of illegal content, summarize the corresponding patterns and rules, and then use regular expressions and other technologies to construct Figure 2 The rule base shown in the figure (including multiple rules such as Rule 1 and Rule 2) is then used to test the text (such as Figure 2 Compared with the aforementioned word matching-based method, this method can mine the deep semantic information of the text and obtain a better detection rate and accuracy.

[0056] 3) Machine learning-based methods

[0057] like Figure 3 As shown, Figure 3 A schematic diagram of another prior art provided for an embodiment of the present application, firstly, a machine learning model is trained through training samples, and the trained machine learning model is deployed on a cloud server, and the data to be tested is tested. After prediction by the machine learning model, it can be classified into compliance, violation or other types.

[0058] Generally speaking, machine learning-based methods mainly consist of three core parts: text vectorization, training models, and predicting data. The text vectorization stage mainly extracts some obvious feature points by performing text representation and data statistics on samples. For example, the text space vector can be represented by the term frequency-inverse document frequency (TF-IDF) model, and then these vectors can be used to train models such as support vector machines, random forests, decision trees, etc. to obtain a classifier, and then the classifier is used to predict the content of the new text to be detected.

[0059] In general, the above three methods are only applicable to the case where there are obvious illegal words in the public content stored by the tenant in the cloud server. For configuration files without obvious illegal words, for example, there may be only one domain name "a1b12c.com" in the configuration file, the above three methods cannot effectively identify them.

[0060] In view of this, an embodiment of the present application proposes a risk detection method for a configuration file, which significantly improves the effectiveness and efficiency of a computing device in identifying whether a configuration file is risky by comprehensively judging the content of the configuration file and the behavioral characteristics of the tenant to which the configuration file belongs.

[0061] In some feasible embodiments, the method is applied to a cloud server that stores configuration files uploaded by tenants. The configuration files are used to locate resources stored outside the cloud server. The method includes: obtaining a configuration file; the configuration file includes a domain name for locating resources; obtaining a first detection result of the configuration file; the first detection result is used to characterize whether the configuration file is a candidate configuration file; the risk that the resources located by the candidate configuration file have illegal content is higher than a first risk threshold; if the first detection result indicates that the configuration file is a candidate configuration file, obtaining historical behavior portrait data of the tenant; according to the historical behavior portrait data, determining a second detection result of the candidate configuration file; the second detection result is used to characterize whether the risk that the resources located by the candidate configuration file have illegal content is higher than a second risk threshold; wherein, the first risk threshold is less than the second risk threshold. Since the first detection result is a detection result of the content of the configuration file, and the second detection result is a detection result of the historical behavior portrait data of the tenant to which the configuration file belongs, that is to say, through the detection results of two different dimensions of the configuration file, it is possible to significantly improve the effectiveness of detecting such configuration files without obvious illegal words.

[0062] Secondly, an exemplary introduction to the system architecture of the embodiments of the present application is given.

[0063] Figure 4 It is a schematic diagram of the architecture of a risk detection system for a configuration file provided by an embodiment of the present application, wherein, Figure 4 In the risk detection system 1000 of the configuration file shown, at least includes: an application layer 1100 and a detection module 1110 therein. The detection module 1110 can be used to perform risk detection on the configuration file to be detected, including but not limited to detecting the content of the configuration file based on a trained neural network model, detecting the behavior of the tenant associated with the suspicious configuration file, and in the case where the risk of the configuration file exceeds the set threshold, performing alarm output according to the set alarm format, etc.

[0064] Optionally, Figure 4 The application layer 1100 shown also includes a review platform module 1120 and a disposal module 1130. The review platform module 1120 can be used to display model alarm data, receive the review results of operation and maintenance personnel and send the corresponding review results to the training module of the neural network model as feedback results, etc.; the disposal module 1130 can be used to dispose of the specified tenant resources based on the feedback of the operation and maintenance personnel, etc.

[0065] Optionally, Figure 4 The risk detection system 1000 of the configuration file shown may also include one or more of a model layer 1200, a data layer 1300, and a management layer 1400.

[0066] Specifically, the model layer 1200 is mainly used for preprocessing and training the data to be analyzed, and may include a data cleaning module 1210, a word segmentation module 1220, a model training module 1230, an automatic update module 1240, and a behavior baseline module 1250. Among them, the data cleaning module 1210 is mainly used to remove invalid data; the word segmentation module 1220 is used to convert the domain name into a character sequence or a gram sequence for model input; the model training module 1230 trains the neural network based on black and white samples; the automatic update module 1240 is used to re-learn and update the model based on the manual review results; the behavior baseline module 1250 is used to construct the behavior baseline of the tenant associated with each text file based on historical statistical data for identifying abnormal behaviors.

[0067] The data layer 1300 is mainly used to obtain the data required by other modules and provide it to the corresponding modules, and may include a log module 1310, a data collection module 1320, and a result collection module 1330. Among them, the log module 1310 can be used to collect the original access logs of the cloud server; the data collection module 1320 can be used to collect the public content of the tenant (including the configuration file to be detected) based on the original access logs of the cloud server, collect public data sets, the audit result data of the online prediction module 1110, and / or the operation and maintenance personnel, etc.; the result collection module 1330 can be used to collect the model parameter information obtained after the model training module 1230 completes training, the prediction results output by the model, and the alarm information sent by the online prediction module 1110, etc.

[0068] The management layer 1400 is mainly used to provide management and storage functions, and can be used to manage the data storage platform 1410, the model training platform 1420, the job scheduling module 1430, the data docking module 1440, etc. Among them, the data storage platform 1410 can be used to store various types of data involved in the data layer 1300; the model training platform 1420 can be used for the initial training, fine-tuning, etc. of the model; the job scheduling module 1430 can be used to schedule the job progress and schedule of model training tasks, model prediction tasks, model update tasks, and alarm tasks, etc. The data docking module 1440 can be used to perform data format conversion and data docking between various modules when needed.

[0069] Figure 5 It is a schematic structural diagram of a computing device provided by an embodiment of the present application. One or more components in the foregoing risk detection system 1000 of the configuration file can run on Figure 5 the computing device shown, Figure 5 The computing device 2000 shown at least includes: a memory 2010, a processor 2020, and a bus 2030. The computing device 2000 can at least be used to run Figure 4The detection module 1110 shown.

[0070] Among them, the processor 2020 can be used to obtain the configuration file publicly available to the tenant on the cloud server, obtain the first detection result of the configuration file, and if the first detection result indicates that the configuration file is a candidate configuration file, obtain the historical behavior portrait data of the tenant; according to the historical behavior portrait data, determine the second detection result of the candidate configuration file, and so on. The memory 2010 can be used to store the logic code corresponding to the risk detection method of the configuration file provided in the embodiments of the present application.

[0071] Optionally, the computing device 2000 can be various types of server devices such as a rack server, a whole cabinet server, etc., and can provide cloud services such as cloud computing and cloud storage for the tenant; the computing device 2000 can also be a terminal computing device such as a computer, a mobile phone terminal, a tablet computer, a laptop computer, a desktop computer, an all-in-one computer, a personal digital assistant (PDA), an ultra-mobile personal computer (UMPC), etc.

[0072] Optionally, the memory 2010 can include a random access memory (RAM), a read-only memory (ROM), etc. Among them, the memory 2010 can run the necessary operating system in its RAM, as well as modules such as the first acquisition module, the second acquisition module, the third acquisition module, and the determination module for executing the risk detection method of the configuration file provided in the present application.

[0073] Optionally, the processor 2020 can be a central processing unit (CPU) or other general-purpose processors, application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), digital signal processors (DSPs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0074] Optionally, the bus 2030 may be a Peripheral Component Interconnect (PCI) bus or the like. The present application does not limit the type of the bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 5 only one line is used to represent it in Figure 5 , but it does not mean that there is only one bus or one type of bus. The bus 2030 may include a path for transmitting information between various components of the computing device 2000 (for example, the memory 2010 and the processor 2020).

[0075] It should be noted that the system architecture and application scenarios described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art may know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0076] For the sake of easy understanding, the risk detection method for the configuration file provided in the embodiments of the present application is introduced exemplarily below in conjunction with the accompanying drawings. This method is applicable to Figure 5 the computing device shown.

[0077] Figure 6 FIG. 14 is a flowchart of a risk detection method for a configuration file provided in an embodiment of the present application. This method can be applied to a cloud server. The cloud server stores a configuration file uploaded by a tenant. The configuration file is used to locate resources stored by the tenant outside the cloud server. The method specifically includes the following steps:

[0078] S110, the computing device obtains the configuration file.

[0079] In this step, the computing device may first obtain the access traffic log, then extract the Uniform Resource Locator (URL) therein. Furthermore, the computing device may obtain the configuration file publicly available to the tenant by accessing these URLs. Some of the configuration files include domain names for locating resources.

[0080] In the foregoing obtaining process, considering that the access traffic is large, which may cause the step S110 to take a long time, in the embodiments of the present application, the computing device first screens the access traffic log through the access count threshold range set by the user, and only extracts the URLs in the access traffic log that meet the requirements, thereby improving the efficiency of obtaining the configuration file and helping to improve the detection efficiency.

[0081] In a possible implementation, after obtaining the configuration file, the computing device may associate the configuration file with the identifier of the data bucket where it is located, the identifier of the tenant to which the data bucket belongs, the creation time of the configuration file, the historical access times of the configuration file, and the time information of each access, etc. Further, the computing device may save the associated information, for example, save it to Figure 4 the data storage platform 1410 shown in the figure, so as to provide a data basis for subsequent detection processes and improve the detection efficiency.

[0082] In a possible implementation, based on the aforementioned associated information, the computing device may construct the historical behavior portrait data of the tenant, including: counting the number of files in the tenant's data bucket, the historical total access times of each file in the tenant's data bucket, the first access time and the most recent access time of each file in the tenant's data bucket, and calculating the life cycle of the file accordingly.

[0083] Further, the computing device may continuously update the historical behavior portrait data of the relevant tenant in the set update period.

[0084] In the embodiments of the present application, the data bucket may be an object storage service (OBS) bucket, etc., and the present application does not limit this.

[0085] Exemplarily, as Figure 7 shown, Figure 7 For Figure 6 a schematic diagram of constructing and updating historical behavior portrait data involved in the shown embodiment, where the update period set by the user is 1 day. Further, the computing device uses the latest traffic logs every day to obtain the public configuration files stored by the tenant on the cloud server through the URL access in the traffic logs, and associates the obtained configuration files with the identifier of the data bucket where they are located, the identifier of the tenant to which the data bucket belongs, the creation time of the configuration file, the historical access times of the configuration file, and the time information of each access, etc., so as to realize the construction of the tenant historical behavior portrait as Figure 7 shown, specifically including the number of files in the data bucket, the total historical access times of the configuration file, and the life cycle of the configuration file. After that, the computing device can repeat the aforementioned process every day, so as to continuously update the historical behavior portrait data of the tenant.

[0086] S120, the computing device obtains the first detection result of the configuration file.

[0087] Among them, the first detection result is used to characterize whether the configuration file is a candidate configuration file, and a candidate configuration file refers to a configuration file whose located resource has a risk of violating content higher than the first risk threshold.

[0088] In a possible implementation, the computing device may match the domain names and IP addresses in the configuration file with those in a preset violation list. If the match is successful, or in other words, the domain names and IP addresses in the configuration file hit the preset violation list, the first detection result obtained by the computing device can indicate that the configuration file is a candidate configuration file.

[0089] If the content in the configuration file is ciphertext, and the tenant to which the configuration file belongs authorizes decryption, the ciphertext is decrypted, and the domain names, IP addresses, and other content in the decrypted content are obtained in advance.

[0090] Through research, it is found that in the previously discovered violation configuration files, most of them use domain names to locate the violation content. Since these domain names have a high risk of being blocked, they are often generated quickly using a certain random generation algorithm. And compliant domain names often include strings such as brand names and company names that are easy for users to remember. That is to say, the domain names used to locate the violation content are often different from compliant domain names. For example, they lack readability and have a low proportion of vowel letters. In step S120, the computing device can mainly obtain the corresponding first detection result by detecting the domain names in the configuration file.

[0091] In a possible implementation, the first detection result can be determined by performing clustering or classification analysis on the features obtained after feature extraction of the domain names in the configuration file.

[0092] Specifically, the computing device can obtain the statistical features of compliant domain names by counting the length, digit ratio, vowel letter ratio, readability score, etc. of compliant domain names, and then measure the difference between the domain names in the configuration file and the foregoing statistical features. For example, if the digit ratio, vowel letter ratio, and readability score of the domain names in the configuration file are all significantly different from those of compliant domain names, the computing device can determine that the resources located by the domain names have a relatively high risk of containing violation content and should make a further judgment based on the corresponding historical behavior portrait data.

[0093] In a possible implementation, the computing device can also assist in judging the risk that the resources located by the domain name contain violation content by analyzing the domain name suffix (or top-level domain name) of the domain names in the configuration file. Common top-level domain names such as ".com" and ".cn" usually require higher costs. Therefore, compared with other less popular top-level domain names, the resources located by such domain names have a relatively lower risk of containing violation content. For example, the computing device can match the domain name suffix of the domain names in the configuration file with a preset list of top-level domain names and determine the risk that the resources located by the domain name contain violation content according to the matching result.

[0094] In addition to the aforementioned statistical method, when the computing device obtains the first detection result of the configuration file, it can also input the domain name in the configuration file into the violation detection model to obtain the first detection result output by the violation detection model. Among them, the violation detection model includes at least one neural network model trained according to the sample domain name and the classification label corresponding to the sample domain name. The training process of the neural network model will be introduced below in conjunction with the accompanying drawings and will not be elaborated here.

[0095] In the embodiment of the present application, the violation detection model can be a Convolutional Neural Network-Long Short-Term Memory (CNN-LSTM) model.

[0096] The CNN-LSTM model combines the advantages of two models, namely convolutional neural networks (CNN) and long short-term memory (LSTM), and can capture both the local semantic features of the domain name string and the sequential relationship features in the domain name string, thereby being able to more accurately detect whether the domain name in the configuration file is used to locate the illegal content.

[0097] Specifically, the CNN-LSTM model can generally be divided into an embedding layer, a convolutional layer, an LSTM layer, and an output layer, etc. Among them, the embedding layer is used to convert the input domain name string into a vector and input it into the convolutional layer. The convolutional layer in the embodiment of the present application can include multiple convolutional kernels of different sizes for capturing local semantic features of different dimensions, and the LSTM layer can be used to extract the character sequence features of the domain name string. An LSTM layer can be connected after each convolutional kernel of each size.

[0098] In a possible implementation manner, to further improve the precision and recall rate of the model when detecting domain names, the domain name string can be split into sequences of multiple lengths and used as the basic input unit of the violation detection model. Among them, the sequence length is positively correlated with the number of characters included. For example, a unigram sequence containing only a single character, a bigram sequence containing two characters, a trigram sequence containing three characters, and so on.

[0099] Furthermore, the computing device can determine the sequence length used when splitting the domain name string according to the length of the domain name string. For example, when the domain name string is less than or equal to 12 characters, unigram is used; when the domain name string is greater than 12 characters and less than or equal to 24 characters, bigram is used; when the domain name string is greater than 24 characters and less than or equal to 36 characters, trigram is used, and so on. The specific splitting strategy can also be set according to factors such as actual requirements and the detection effect of the model.

[0100] In the embodiments of the present application, the number of violation detection models is not limited. That is to say, when the computing device determines the first detection result, it can input the domain name in the configuration file into multiple trained neural network models, and then comprehensively determine the first detection result according to the output results of each neural network model and the corresponding weight coefficients, which can further improve the detection accuracy.

[0101] In some feasible embodiments, the violation detection model may further include a Convolutional Neural Network - Bidirectional Gated Recurrent Unit Network (CNN - BiGUR) model, a Generative Adversarial Network (GAN) model, and so on.

[0102] Exemplarily, in the configuration file obtained by the computing device, there is a domain name: "a3d6yx0d52ah1bx1h12c.vyn". Then the computing device determines that the length of this domain name is 20 (excluding the domain name suffix), the proportion of digits is 40%, the proportion of vowel letters is 10%, and the readability score is 0. After statistics, the length of compliant domain names rarely exceeds 12, the proportion of digits is usually less than 20%, the proportion of vowel letters is about 30% - 50%, and they have good readability. It can be seen that the domain names in the foregoing configuration file do not conform to the foregoing characteristics. Then the computing device determines that this configuration file is a candidate configuration file.

[0103] Alternatively, the computing device inputs the domain name in the foregoing configuration file into a pre - trained CNN - LSTM model and obtains the first detection result output by the model, and this first detection result is used to represent that this configuration file is a candidate configuration file.

[0104] Or, the first risk threshold is 80%. The computing device inputs the domain name in the foregoing configuration file into a pre - trained CNN - LSTM model and a GAN model, and respectively obtains their output results as: 95% and 85%. Among them, the foregoing percentages are the probabilities that the model determines that this configuration file is a candidate configuration file, or rather, there is a risk that the resource located by this configuration file has violation content. According to the corresponding weight coefficients (both are 0.5), it is comprehensively determined that the risk that the resource located by this configuration file has violation content is 90%, which is higher than the first risk threshold. Then the first detection result is determined as: this configuration file is a candidate configuration file.

[0105] S130. If the first detection result represents that the configuration file is a candidate configuration file, the computing device obtains the historical behavior portrait data of the tenant.

[0106] In this step, since the first detection result indicates that the configuration file is a candidate configuration file, that is, the risk of the resources located by this configuration file having illegal content is higher than the first risk threshold, the computing device further obtains the historical behavior portrait data of the tenant to which the candidate configuration file belongs.

[0107] In the embodiments of the present application, the historical behavior portrait data can be further divided into two levels, including:

[0108] The first level is the historical behavior portrait data of the tenant to which the candidate configuration file belongs for all the public configuration files stored on the cloud server by it.

[0109] Specifically, the data included in the historical behavior portrait data at the first level is more comprehensive, specifically including: the number of files in each data bucket associated by the tenant to which the candidate configuration file belongs on the cloud server, and one or more of the life cycle and total historical access times of each public configuration file stored by the tenant on the cloud server.

[0110] The second level is the historical behavior portrait data of the tenant to which the candidate configuration file belongs for this candidate configuration file.

[0111] Specifically, the data included in the historical behavior portrait data at the second level is more targeted, specifically including: the number of files in the data bucket where the candidate configuration file is located, and one or more of the life cycle and total historical access times of this candidate configuration file.

[0112] Among them, the life cycle of the candidate configuration file is determined according to the first access time and the most recent access time of this candidate configuration file.

[0113] It should be noted that in the embodiments of the present application, the historical behavior portrait data at any level is determined according to the public content stored by the tenant on the cloud server.

[0114] Exemplarily, the historical behavior portrait data of the tenant obtained by the computing device: the number of files in each data bucket associated by the tenant to which the candidate configuration file belongs on the cloud server is: the number of files in the No. 1 data bucket is 1, and the number of files in the No. 2 data bucket is 1; the life cycles of each public configuration file stored by the tenant on the cloud server are 2 hours and 1 hour respectively, and the total historical access times of the configuration files are 2 times and 1 time respectively.

[0115] Or, the historical behavior portrait data of the tenant obtained by the computing device for this candidate configuration file is: the number of files in the data bucket where the candidate configuration file is located is 1, and the life cycle of this candidate configuration file is 1 hour and the total historical access times is 1 time.

[0116] S140, the computing device determines a second detection result of the candidate configuration file according to the historical behavior portrait data.

[0117] Wherein, the second detection result is used to characterize whether the risk of the resource located by the candidate configuration file having illegal content is higher than the second risk threshold; wherein, the first risk threshold is less than the second risk threshold.

[0118] In this step, the computing device further detects the candidate configuration file according to the historical behavior portrait data obtained in step S130.

[0119] In a possible implementation manner, if the historical behavior portrait data obtained by the computing device in step S130 includes: the number of files in each data bucket associated with the tenant to which the candidate configuration file belongs on the cloud server, and the life cycle and total historical access times (first level) of each public configuration file stored by the tenant on the cloud server, then the computing device obtains the number of data buckets with the number of files lower than the first target number threshold by comparing the number of files in each data bucket with the first target number threshold; the life cycle of each configuration file with the first target life cycle threshold; and the historical access times of each configuration file with the first target number of times threshold, respectively, to obtain the number of data buckets with the number of files lower than the first target number threshold, the number of configuration files with the life cycle lower than the first target life cycle threshold, and the number of configuration files with the historical access times lower than the first target number of times threshold.

[0120] Moreover, the computing device further calculates the proportion (first proportion) of the aforementioned data buckets with the number of files lower than the first target number threshold in all the data buckets associated with the tenant, the proportion (second proportion) of the configuration files with the life cycle lower than the first target life cycle threshold in all the public configuration files stored by the tenant on the cloud server, and the proportion (third proportion) of the configuration files with the historical access times lower than the first target number of times threshold in all the public configuration files stored by the tenant on the cloud server, and determines the second detection result according to the values of the aforementioned three proportions.

[0121] Specifically, in the case that one or more of the aforementioned three proportions are higher than the corresponding proportion thresholds, the computing device can determine that the second detection result of the candidate configuration file is: the risk of the resource located by the candidate configuration file having illegal content is higher than the second risk threshold.

[0122] In a possible implementation, if the historical behavior portrait data obtained by the computing device in step S130 includes: the number of files in the data bucket where the candidate configuration file is located, the life cycle of the candidate configuration file, and the total number of historical accesses (second level), the computing device determines whether the number of files in the data bucket is lower than the second target quantity threshold, whether the life cycle of the candidate configuration file is lower than the second target life cycle threshold, and whether the historical access times of the candidate configuration file are lower than the second target times threshold. Furthermore, the computing device determines the second detection result of the candidate configuration file based on the comparison results described above.

[0123] Furthermore, if the number of files in the data bucket where the candidate configuration file is located is lower than the second target quantity threshold, and / or the life cycle of the candidate configuration file is lower than the second target life cycle threshold, and / or the historical access times of the candidate configuration file are lower than the second target times threshold. In other words, as long as any one of the above three conditions is met, the computing device can determine that the second detection result of the candidate configuration file is: the risk that the resource located by the candidate configuration file has illegal content is higher than the second risk threshold.

[0124] Optionally, when the second detection result indicates that the risk that the resource located by the configuration file has illegal content is higher than the second risk threshold, the computing device sends the detection result to the manual review platform for relevant operation and maintenance personnel to conduct a manual review. If the manual review result is consistent with the second detection result, the operation and maintenance personnel can take corresponding actions on the illegal configuration file and tenant; if the manual review result is inconsistent with the second detection result, the manual review result can be used as feedback for updating the neural network model and as a basis for modifying parameters such as the aforementioned target life cycle threshold, target times threshold, and target quantity threshold to improve the detection accuracy.

[0125] Taking the following data as an example, the historical behavior portrait data of the tenant for the candidate configuration file obtained by the computing device is: the number of files in the data bucket where the candidate configuration file is located is 1, the life cycle of the candidate configuration file is 1 hour, and the total number of historical accesses is 1 time; the second target quantity threshold is 2, the second target life cycle threshold is 6 hours, and the second target times threshold is 10 times; furthermore, the computing device can determine that each item in the historical behavior portrait data of the candidate configuration file does not meet the corresponding threshold by comparing the above data. Furthermore, the computing device determines that the second detection result of the candidate configuration file is: the risk that the resource located by the candidate configuration file has illegal content is higher than the second risk threshold, and the computing device sends the detection result to the Figure 4 review platform module 1120 as shown to enable relevant operation and maintenance personnel to conduct a manual review and take corresponding actions.

[0126] Further, if the aforementioned tenant deploys a large number of compliant configuration files on the cloud server to reduce the probability of being detected, since the aforementioned detection process is targeted, the risks of the candidate configuration files can be effectively detected; if the aforementioned tenant deploys a large number of similar candidate configuration files on the cloud server, the computing device can obtain the historical behavior portrait data of the first level, thereby further improving the accuracy of risk detection.

[0127] In some feasible embodiments, the computing device can first obtain the historical behavior portrait data of the first level, and when the second detection result determined according to the data is that the risk of the resource located by the candidate configuration file having illegal content is lower than the second risk threshold, then focus on the candidate configuration file, that is, obtain the historical behavior portrait data of the second level, and perform risk detection more specifically. By performing risk detection twice successively according to the historical behavior portrait data, the detection rate and accuracy of the configuration files with high risks can be further improved.

[0128] Through the above steps S110 - S140, the computing device can significantly improve the detection rate and accuracy when performing risk detection on the configuration file according to the detection results of at least two different dimensions of the configuration file (file content and the historical behavior portrait data of the tenant), and improve the work efficiency of the operation and maintenance personnel.

[0129] The following combines the attached Figure 8 to illustrate and explain the training process of the neural network model involved in the embodiments of the present application. Among them, Figure 8 is a flowchart of model training provided by the embodiments of the present application.

[0130] It should be noted that the computing device used to train the neural network model and the computing device that executes the risk detection method of the configuration file provided by the embodiments of the present application may be the same computing device or may not be the same computing device.

[0131] The following takes Figure 5 the computing device 2000 shown as an example of the execution subject. Correspondingly, the training process of the neural network model includes:

[0132] S210, the computing device obtains the original training samples.

[0133] In this step, the computing device uses the obtained dataset of illegal domain names, the public dataset, and the in-network domain name dataset as the original training samples. Among them, the original training samples include at least a large number of domain names. In some feasible embodiments, the original training samples also include classification labels corresponding to the domain names.

[0134] S220, the computing device performs data cleaning.

[0135] In this step, the computing device can perform data cleaning on the original training samples (including classification labels corresponding to domain names) obtained in step S210.

[0136] The computing device can first screen and divide the original training samples to roughly obtain the following two types of domain names: compliant domain names (white samples) and non-compliant domain names (black samples). Furthermore, the computing device can clean the aforementioned black samples based on the publicly available white domain name list on the network, clean the aforementioned white samples based on the self-inspection and self-examination results, use the training samples after data cleaning as the training set for training the neural network model, and data cleaning can significantly improve the quality of the training set, which helps to improve the detection accuracy of the neural network model.

[0137] S230, the computing device trains the neural network model.

[0138] In this step, the computing device can convert the black and white samples obtained in step S220 into character sequences, construct a dictionary set, and then convert each character in each training sample into its index in the dictionary, that is, each training sample is converted into a vector composed of the indices of the characters in the dictionary. Then, the word2vec technology is used to learn the semantic information of the samples, and each character is trained to obtain the word vector corresponding to each character.

[0139] Furthermore, the computing device uses the obtained word vectors to convert the training samples to obtain a word vector list corresponding to each training sample. Through the gradient descent algorithm, each word vector list is used as input to train the parameters in the neural network model until all word vector lists are trained to obtain a trained neural network model.

[0140] In some feasible embodiments, there are no labeled tags in the original training samples obtained by the computing device. In this way, unsupervised learning algorithms can also be used to train models such as GAN, and the trained neural network model can be used for risk monitoring of configuration files.

[0141] In a possible implementation manner, the computing device can also combine the supervised learning algorithm and automatic learning update based on the daily manual review results, so as to continuously perform online learning and tuning on the model based on the latest data, and then solve the problems of untimely update of traditional static models and low detection accuracy for the latest data, making the neural network model have stronger detection capabilities for continuously changing domain names.

[0142] Through the above steps S210 - S230, the computing device can construct a high - quality training set through data cleaning and train the neural network model, so that the neural network model has the ability to accurately detect domain names in the configuration file.

[0143] The above mainly introduces the solution of the embodiment of the present application from the perspective of the method. It can be understood that in order to implement the above functions, the risk detection device for the configuration file includes at least one of the corresponding hardware structures and software modules for executing each function. Those skilled in the art should easily realize that, combined with the units and algorithm steps of each example described in the embodiments disclosed in this article, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0144] The embodiments of the present application can divide the functional units of the risk detection device for the configuration file according to the above - mentioned method examples. For example, each functional unit can be divided corresponding to each function, or two or more functions can be integrated into one processing unit. The above - integrated unit can be implemented in the form of hardware or in the form of a software functional unit. It should be noted that for the units in the embodiments of the present application.

[0145] Exemplarily, Figure 9 FIG. 11 is a schematic structural diagram of a risk detection device for a configuration file provided by an embodiment of the present application. The risk detection device 800 for the configuration file is applied to a computing device, or the risk detection device 800 for the configuration file can be a computing device. The risk detection device 800 for the configuration file includes:

[0146] A first acquisition module 810, configured to acquire a configuration file; the configuration file includes a domain name for locating a resource.

[0147] A second acquisition module 820, configured to acquire a first detection result of the configuration file; the first detection result is used to characterize whether the configuration file is a candidate configuration file; the risk that the resource located by the candidate configuration file has illegal content is higher than a first risk threshold.

[0148] A third acquisition module 830, configured to, if the first detection result indicates that the configuration file is a candidate configuration file, acquire historical behavior portrait data of the tenant.

[0149] A determination module 840, configured to determine a second detection result of a candidate profile according to historical behavior portrait data; the second detection result is used to characterize whether the risk that the resource located by the candidate profile has illegal content is higher than a second risk threshold; wherein, the first risk threshold is less than the second risk threshold.

[0150] For example, in combination with Figure 4 , the first acquisition module 810 can be used to execute S110 as shown in Figure 6 , the second acquisition module 820 can be used to execute S120 as shown in Figure 6 , the third acquisition module 830 can be used to execute S130 as shown in Figure 6 , and the determination module 840 can be used to execute S140 as shown in Figure 6 .

[0151] In a possible implementation manner, the historical behavior portrait data of the tenant includes: the number of files in each data bucket associated by the tenant on the cloud server, the life cycle of the configuration files stored by the tenant on the cloud server, and one or more of the total number of historical accesses. The life cycle of the configuration file is determined according to the first access time and the most recent access time of the configuration file.

[0152] In a possible implementation manner, the determination module 840 is further configured to obtain one or more of: a first ratio of data buckets with the number of files lower than a first target number threshold to the data buckets associated by the tenant, a second ratio of configuration files with a life cycle lower than a first target life cycle threshold to the target configuration files, and a third ratio of configuration files with the total number of historical accesses lower than a first target number threshold to the target configuration files. The target configuration files are all public configuration files stored by the tenant on the cloud server; if one or more of the first ratio, the second ratio, and the third ratio are higher than the corresponding ratio thresholds, it is determined that the risk that the resource located by the candidate profile has the illegal content is higher than the second risk threshold.

[0153] In a possible implementation manner, the determination module 840 is further configured to determine a second detection result of the candidate profile according to the historical behavior portrait data of the tenant for the candidate profile; wherein, the historical behavior portrait data of the tenant for the candidate profile includes one or more of the number of files in the data bucket where the candidate profile is located, the historical access times of the candidate profile, and the life cycle of the candidate profile. The life cycle of the candidate profile is determined according to the first access time and the most recent access time of the candidate profile.

[0154] In a possible implementation, the determining module 840 is further configured to determine that the risk of the resource having the illegal content is higher than the second risk threshold if the historical behavior portrait data meets one or more of the following conditions, where the conditions include: the number of files in the data bucket where the candidate profile is located is lower than the target number threshold; the historical access times of the candidate profile are lower than the target times threshold; the life cycle of the candidate profile is lower than the target life cycle threshold.

[0155] In a possible implementation, if the profile further includes an Internet Protocol IP address for locating the resource, the third obtaining module 830 is further configured to obtain the historical behavior portrait of the tenant if the IP address in the profile hits a preset list of risky IPs and the first detection result indicates that the profile is the candidate profile.

[0156] In a possible implementation, the first detection result is determined by performing clustering or classification analysis on the features obtained by extracting features from the domain name in the profile.

[0157] In a possible implementation, the second obtaining module 820 is further configured to input the domain name in the profile into an illegal detection model to obtain the first detection result output by the illegal detection model, where the illegal detection model includes at least one neural network model trained according to sample domain names and the classification labels corresponding to the sample domain names.

[0158] In a possible implementation, the illegal detection model is a Convolutional Neural Network-Long Short-Term Memory Network CNN-LSTM model.

[0159] In a possible implementation, the second obtaining module 820 is further configured to determine the first detection result according to the weight coefficients corresponding to each neural network model and the output result of the illegal detection model if the illegal detection model includes two or more neural network models.

[0160] As a feasible example, the risk detection device 800 for the profile provided in this application is implemented by software modules. For example, the software modules can be provided to users for use through a cloud service subscription model, and users can select different subscription levels according to their needs; for another example, the software modules can also provide enterprise-level customization services with professional domain customization, interface personalization, and extended functions according to the needs of users or enterprises.

[0161] In addition, the risk detection device 800 for the configuration file provided in this application can also be provided to users as a value-added service, and this application does not limit this. When the risk detection device 800 for the configuration file is implemented through software modules, the risk detection device 800 for the configuration file can also be embedded in a cloud web application firewall software (web application firewall, WAF) or a situation awareness platform system.

[0162] An embodiment of this application also provides a computing device 100. As Figure 10 shown, the computing device 100 includes: a bus 102, a processor 104, a memory 106, and a communication interface 108. The processor 104, the memory 106, and the communication interface 108 communicate with each other through the bus 102. The computing device 100 can be a server or a terminal device. It should be understood that this application embodiment does not limit the number of processors and memories in the computing device 100.

[0163] The bus 102 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 10 only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus. The bus 102 can include a path for transmitting information between various components (for example, the memory 106, the processor 104, the communication interface 108) of the computing device 100.

[0164] The processor 104 can include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0165] The memory 106 may include volatile memory, such as random access memory (RAM). The processor 104 may also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0166] The executable program code is stored in the memory 106, and the processor 104 executes the executable program code to implement the functions of the foregoing first acquisition module, second acquisition module, third acquisition module, and determination module respectively, so as to implement the risk detection method for the configuration file. That is to say, instructions for executing the risk detection method for the configuration file are stored on the memory 106.

[0167] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.

[0168] As Figure 11 shown, the computing device cluster includes at least one computing device 100. Instructions for executing the risk detection method for the configuration file may be stored in the memory 106 in one or more of the computing devices 100 in the computing device cluster.

[0169] In some possible implementation manners, partial instructions for executing the risk detection method for the configuration file may also be stored separately in the memory 106 of one or more of the computing devices 100 in the computing device cluster. In other words, a combination of one or more computing devices 100 may jointly execute the instructions for executing the risk detection method for the configuration file.

[0170] It should be noted that the memories 106 in different computing devices 100 in the computing device cluster may store different instructions, which are respectively used to execute partial functions of the risk detection device for the configuration file. That is to say, the instructions stored in the memories 106 in different computing devices 100 may implement the functions of one or more of the first acquisition module, second acquisition module, third acquisition module, and determination module.

[0171] In some possible implementation manners, one or more computing devices in the computing device cluster may be connected through a network. Among them, the network may be a wide area network or a local area network, etc.Figure 12 A possible implementation is shown. As Figure 12 shown, two computing devices 100A and 100B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device. In this type of possible implementation, the instructions for executing the functions of the first acquisition module are stored in the memory 106 of the computing device 100A. At the same time, the instructions for executing the functions of the second acquisition module, the third acquisition module, and the determination module are stored in the memory 106 of the computing device 100B.

[0172] It should be understood that Figure 12 the functions of the computing device 100A shown in

[0173] This application embodiment also provides a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, it causes at least one computing device to execute the risk detection method of the configuration file.

[0174] This application embodiment also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive), etc. The computer-readable storage medium includes instructions that instruct the computing device to execute the risk detection method of the configuration file, or instruct the computing device to execute the risk detection method of the configuration file.

[0175] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the protection scope of the technical solutions of the embodiments of the present invention.

Claims

1. A risk detection method for a configuration file, characterized in that: The method is applied to a cloud server, the cloud server stores a configuration file uploaded by a tenant, the configuration file is used to locate resources stored by the tenant outside the cloud server, and the method includes: Obtaining the configuration file; the configuration file includes a domain name for locating the resource; Obtaining a first detection result of the configuration file; the first detection result is used to indicate whether the configuration file is a candidate configuration file; the risk of the resource located by the candidate configuration file having illegal content is higher than a first risk threshold; If the first detection result indicates that the configuration file is the candidate configuration file, obtaining historical behavior profile data of the tenant; Determine a second detection result of the candidate profile based on the historical behavior profile data; the second detection result is used to characterize whether the risk of the resource located by the candidate profile containing the illegal content is higher than a second risk threshold; wherein the first risk threshold is less than the second risk threshold.

2. The method according to claim 1, characterized in that The tenant's historical behavior profile data includes: One or more of the number of files in each data bucket associated with the tenant on the cloud server, the life cycle of the configuration file stored by the tenant on the cloud server, and the total number of historical access times, wherein the life cycle of the configuration file is determined based on the first access time and the most recent access time of the configuration file.

3. The method according to claim 2, characterized in that The determining, based on the historical behavior profile data, a second detection result of the candidate configuration file includes: Obtain one or more of a first ratio of the data buckets whose file quantity is lower than a first target quantity threshold to the data buckets associated with the tenant, a second ratio of the configuration files whose lifecycle is lower than a first target lifecycle threshold to the target configuration files, and a third ratio of the configuration files whose total number of historical accesses is lower than a first target number threshold to the target configuration files, wherein the target configuration files are all public configuration files stored by the tenant on the cloud server; If one or more of the first ratio, the second ratio, and the third ratio is higher than the corresponding ratio threshold, it is determined that the risk of the resource located by the candidate profile having the illegal content is higher than the second risk threshold.

4. The method according to claim 1, characterized in that The determining, based on the historical behavior profile data, a second detection result of the candidate configuration file includes: determining a second detection result of the candidate profile according to the historical behavior profile data of the tenant for the candidate profile in the historical behavior profile data; Among them, the tenant's historical behavior portrait data for the candidate profile includes one or more of the number of files in the data bucket where the candidate profile is located, the number of historical access times of the candidate profile, and the life cycle of the candidate profile. The life cycle of the candidate profile is determined based on the first access time and the most recent access time of the candidate profile.

5. The method according to claim 4, characterized in that The determining, according to the historical behavior profile data of the tenant for the candidate profile, a second detection result of the candidate profile, includes: If the historical behavior profile data satisfies one or more of the following conditions, it is determined that the risk of the resource containing the illegal content is higher than the second risk threshold, wherein the conditions include: The number of files in the data bucket where the candidate configuration file is located is lower than the target number threshold; The number of historical visits to the candidate profile is lower than a target number threshold; The life cycle of the candidate profile is lower than a target life cycle threshold.

6. The method according to any one of claims 1 to 5, characterized in that: If the configuration file further includes an Internet Protocol IP address for locating the resource, if the first detection result indicates that the configuration file is the candidate configuration file, obtaining the historical behavior profile of the tenant includes: If the IP address in the configuration file hits a preset risk IP list, and the first detection result indicates that the configuration file is the candidate configuration file, a historical behavior profile of the tenant is obtained.

7. The method according to any one of claims 1 to 6, characterized in that: The first detection result is determined by clustering or classifying the features obtained after feature extraction of the domain name in the configuration file.

8. The method according to any one of claims 1 to 7, characterized in that: The obtaining of the first detection result of the configuration file includes: The domain name in the configuration file is input into a violation detection model to obtain the first detection result output by the violation detection model, wherein the violation detection model includes at least one neural network model trained according to a sample domain name and a classification label corresponding to the sample domain name.

9. The method according to claim 8, characterized in that The violation detection model is a convolutional neural network-long short-term memory network CNN-LSTM model.

10. The method according to claim 8 or 9, characterized in that: The obtaining of the first detection result of the configuration file includes: If the violation detection model includes two or more neural network models, the first detection result is determined according to the weight coefficients corresponding to the neural network models and the output results of the violation detection model.

11. A risk detection device for a configuration file, characterized in that: The device comprises: A first acquisition module is used to acquire the configuration file; the configuration file includes a domain name for locating the resource; A second acquisition module is used to acquire a first detection result of the configuration file; the first detection result is used to indicate whether the configuration file is a candidate configuration file; the risk of the resource located by the candidate configuration file having illegal content is higher than a first risk threshold; A third acquisition module, configured to acquire historical behavior portrait data of the tenant if the first detection result indicates that the configuration file is the candidate configuration file; A determination module is used to determine a second detection result of the candidate configuration file based on the historical behavior portrait data; the second detection result is used to characterize whether the risk of the resource located by the candidate configuration file containing the illegal content is higher than a second risk threshold; wherein the first risk threshold is less than the second risk threshold.

12. The device according to claim 11, characterized in that The tenant's historical behavior profile data includes: One or more of the number of files in each data bucket associated with the tenant on the cloud server, the life cycle of the configuration file stored by the tenant on the cloud server, and the total number of historical access times, wherein the life cycle of the configuration file is determined based on the first access time and the most recent access time of the configuration file.

13. The device according to claim 12, characterized in that The determination module is further used to: Obtain one or more of a first ratio of the data buckets whose file quantity is lower than a first target quantity threshold to the data buckets associated with the tenant, a second ratio of the configuration files whose lifecycle is lower than a first target lifecycle threshold to the target configuration files, and a third ratio of the configuration files whose total number of historical accesses is lower than a first target number threshold to the target configuration files, wherein the target configuration files are all public configuration files stored by the tenant on the cloud server; If one or more of the first ratio, the second ratio, and the third ratio is higher than the corresponding ratio threshold, it is determined that the risk of the resource located by the candidate profile having the illegal content is higher than the second risk threshold.

14. The device according to claim 11, characterized in that The determination module is further used to: determining a second detection result of the candidate profile according to the historical behavior profile data of the tenant for the candidate profile in the historical behavior profile data; Among them, the tenant's historical behavior portrait data for the candidate profile includes one or more of the number of files in the data bucket where the candidate profile is located, the number of historical access times of the candidate profile, and the life cycle of the candidate profile. The life cycle of the candidate profile is determined based on the first access time and the most recent access time of the candidate profile.

15. The device according to claim 14, characterized in that The determination module is further used to: If the historical behavior profile data satisfies one or more of the following conditions, it is determined that the risk of the resource containing the illegal content is higher than the second risk threshold, wherein the conditions include: The number of files in the data bucket where the candidate configuration file is located is lower than the target number threshold; The number of historical visits to the candidate profile is lower than a target number threshold; The life cycle of the candidate profile is lower than a target life cycle threshold.

16. The device according to any one of claims 11 to 15, characterized in that: If the configuration file also includes an Internet Protocol IP address for locating the resource, the third acquisition module is further used to: If the IP address in the configuration file hits a preset risk IP list, and the first detection result indicates that the configuration file is the candidate configuration file, a historical behavior profile of the tenant is obtained.

17. The device according to any one of claims 11 to 16, characterized in that: The first detection result is determined by clustering or classifying the features obtained after feature extraction of the domain name in the configuration file.

18. The device according to any one of claims 11 to 17, characterized in that: The second acquisition module is further used to: The domain name in the configuration file is input into a violation detection model to obtain the first detection result output by the violation detection model, wherein the violation detection model includes at least one neural network model trained according to a sample domain name and a classification label corresponding to the sample domain name.

19. The device according to claim 18, characterized in that The violation detection model is a convolutional neural network-long short-term memory network CNN-LSTM model.

20. The device according to claim 18 or 19, characterized in that The second acquisition module is further used to: If the violation detection model includes two or more neural network models, the first detection result is determined according to the weight coefficients corresponding to the neural network models and the output results of the violation detection model.

21. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the risk detection method for the configuration file according to any one of claims 1 to 10.

22. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes computer instructions; when the computer instructions are executed in a computing device, the computing device executes the risk detection method for a configuration file according to any one of claims 1 to 10.

23. A computer program product, characterized in that When the computer program product is run in a computing device, the computing device executes the risk detection method for a configuration file according to any one of claims 1 to 10.