A data security protection method for application software system development
By using the decision tree as the base classifier in the SAMME algorithm and weight correction based on the similarity between iterative risk level and preset risk level, the problem of low risk level classification accuracy in the application software system development data is solved, and higher risk level evaluation accuracy and multi-classification effect are achieved.
Patent Information
- Application Number
- CN202510180579.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-19
AI Technical Summary
Directly using the traditional SAMME algorithm to uniformly allocate weights, it is easy to ignore the complex scenarios of application software system development data, resulting in low accuracy of risk level division.
By obtaining development data related to the user's operation of the application software, performing data quantification and preset risk level annotation, using the decision tree as the base classifier of the SAMME algorithm, the weight correction coefficient of the iterative analysis data is determined through the similarity between the iterative risk level and the preset risk level under different iterations, and the initial weight of the iterative analysis data in the original SAMME algorithm is corrected and adjusted.
The multi-classification effect and the accuracy of risk level evaluation of application software system development data is improved, making the overall risk level classification of development data more accurate and reliable.
Smart Images

Figure CN119670102B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data security protection, and in particular to a data security protection method for application software system development. Background Art
[0002] Application software system development data mainly includes data generated and used in the process of software development, testing, and deployment, mainly including source code, configuration files, user data, log files, connection databases, etc. Data security protection has the advantages of maintaining user privacy, preventing data tampering, and protecting corporate reputation. Therefore, it is very important to protect application software system development data. Since different data have different importance, it is necessary to manage the risk of development data. By identifying the risk level of data, potential threats can be better assessed, and corresponding preventive measures and emergency plans can be formulated to protect data. Determine the risk level of the data to determine whether there is a security threat, and take corresponding measures according to the risk level of the data.
[0003] In related technologies, the SAMME algorithm is often used to implement multi-classification of development data. However, the SAMME algorithm considers that the categories are independent of each other, that is, the algorithm only judges the correctness or error of the category classification. However, there is a correlation between the risk levels of actual development data, that is, there are also degrees of severity between recognition errors, and the misclassification is not equal in the risk level, so the direct use of the traditional SAMME algorithm to uniformly assign weights is likely to ignore the complex scenarios of the development data, resulting in low accuracy in risk level classification. Summary of the invention
[0004] In order to solve the technical problem that the method of uniformly allocating weights in the traditional SAMME algorithm in the related art easily ignores the complex scenarios of development data, resulting in low accuracy of risk level classification, the present invention provides a method for protecting application software system development data security, and the technical solution adopted is as follows:
[0005] The present invention proposes a method for protecting data security in application software system development, the method comprising:
[0006] Acquire development data related to the application software run by the user, quantify the development data and mark the preset risk level to obtain training data, wherein the training data has different attributes, and the attributes are key-value pairs;
[0007] The decision tree is used as the base classifier of SAMME, and the training data is input and processed with the same initial weights to output the iterative risk levels of different training data at different iteration times. Based on the similarity between the iterative risk levels at different iteration times and the preset risk levels of the training data, the training data with the correct risk level classification at each iteration is determined as the iterative analysis data.
[0008] According to the attribute distribution of the iterative analysis data under the same iteration, the attribute distribution significance feature of each attribute is determined; according to the attribute distribution significance feature, the target attribute is screened from the iterative analysis data under the same iteration; the similarity coefficient between the change of the attribute distribution significance feature of any target attribute under different risk levels and the change of the risk level is determined, and the weight correction coefficient of the iterative analysis data is determined according to the similarity coefficient of all target attributes in the same iterative analysis data;
[0009] The initial weight of the iterative analysis data of the same iteration in the original SAMME algorithm is corrected according to the weight correction coefficient, and the correction weight of different iterative analysis data under the same iteration is determined; according to the difference between the iterative risk level and the preset risk level, the correction weight is adjusted to obtain the adjusted weight of the iterative analysis data, and the SAMME algorithm is iteratively analyzed according to the adjusted weights of all iterative analysis data to obtain a security protection model.
[0010] Further, the similarity between the iterative risk level at different iteration times and the preset risk level of the training data is determined as the iterative analysis data with the correct risk level classification at each iteration, including:
[0011] Determine whether the iteration risk level of the training data in any iteration is the same as the preset risk level. If they are the same, use the corresponding training data as the iteration analysis data in the corresponding iteration.
[0012] Furthermore, the step of iteratively analyzing the attribute distribution of the data in the same iteration to determine the attribute distribution significance feature of each attribute includes:
[0013] Iteratively analyze the data under the same iteration and construct an empirical distribution function and a uniform distribution function for the attribute of the same attribute, use the KS test to determine the difference between the empirical distribution function and the uniform distribution function, and determine the test p value according to the difference;
[0014] The kurtosis value and the interquartile range of the empirical distribution function are determined, and the attribute distribution significance characteristics of the attribute are determined by combining the test p value, the kurtosis value and the interquartile range.
[0015] Furthermore, the attribute distribution significance feature is a numerical representation, and the attribute distribution significance feature of the attribute is determined by combining the test p value, the kurtosis value and the interquartile range. The corresponding calculation formula is:
[0016]
[0017] In the formula, f represents the attribute distribution significance characteristic of the attribute, ku represents the kurtosis value, IQR represents the interquartile range, and p represents the test p value.
[0018] Furthermore, the step of selecting the target attribute from the iterative analysis data in the same iteration according to the attribute distribution significance feature includes:
[0019] According to the distribution discreteness of the attribute distribution significance characteristics of the same attribute under different risk level classifications, the importance index of the attribute is determined;
[0020] The attribute whose value of the importance index is less than a preset importance threshold is taken as the target attribute.
[0021] Furthermore, the importance index of the attribute is determined according to the distribution discreteness of the attribute distribution significance characteristics of the same attribute under different risk level classifications, including:
[0022] In the same iteration, the standard deviation of the attribute distribution significance characteristics of the same attribute under different risk level classifications is determined, and the standard deviation is used as a significance distribution dispersion indicator;
[0023] The significance distribution discrete index is normalized to obtain the importance index of each attribute.
[0024] Furthermore, the step of determining the weight correction coefficient of the iterative analysis data according to the similarity coefficients of all target attributes in the same iterative analysis data includes:
[0025] The mean of the similarity coefficients of all target attributes in the same iterative analysis data is used as the weight correction coefficient of the iterative analysis data.
[0026] Further, the initial weight of the iterative analysis data of the same iteration in the original SAMME algorithm is corrected according to the weight correction coefficient to determine the correction weight of different iterative analysis data under the same iteration, including:
[0027] The ratio of the initial weight to the weight correction coefficient is used as the correction weight.
[0028] Furthermore, the adjusting the correction weight according to the difference between the iterative risk level and the preset risk level to obtain the adjusted weight of the iterative analysis data includes:
[0029] Using the difference between the preset risk level and the iterative risk level as an adjustment coefficient;
[0030] The product of the adjustment coefficient and the correction weight is normalized to obtain the adjustment weight.
[0031] Furthermore, the SAMME algorithm is iteratively analyzed according to the adjusted weights of all iterative analysis data to obtain a security protection model, including:
[0032] The SAMME algorithm is used to process the adjusted weights of the iterative analysis data under all iteration numbers until the change rate of the weights of all development data before and after the iteration is within 2%. The iteration is stopped and the security protection model is obtained.
[0033] The present invention has the following beneficial effects:
[0034] The present invention obtains development data related to application software run by users, quantifies the development data and labels the development data with preset risk levels, and obtains training data, wherein the training data has different attributes; a decision tree is used as a base classifier of SAMME, the training data is input, and the training data is processed through the same initial weight, and the iterative risk levels of different training data under different iteration times are output; based on the similarity between the iterative risk levels under different iteration times and the preset risk levels of the training data, the training data with correct risk level classification under each iteration is determined as iterative analysis data; according to the attribute distribution of the iterative analysis data under the same iteration, the attribute distribution significance feature of each attribute is determined; according to the attribute distribution significance feature, the iterative analysis data under the same iteration are obtained. The target attribute is obtained by screening from the iterative analysis data under different risk levels; the similarity coefficient between the change of the attribute distribution significance characteristics of any target attribute under different risk levels and the change of the risk level is determined, and the weight correction coefficient of the iterative analysis data is determined according to the similarity coefficient of all target attributes in the same iterative analysis data; the initial weight of the iterative analysis data of the same iteration in the original SAMME algorithm is corrected according to the weight correction coefficient, and the correction weight of different iterative analysis data under the same iteration is determined; according to the difference between the iterative risk level and the preset risk level, the correction weight is adjusted to obtain the adjustment weight of the iterative analysis data, and the SAMME algorithm is iteratively analyzed according to the adjustment weights of all iterative analysis data to obtain the security protection model. Since the adjustment weight is obtained adaptively according to the attribute performance of the development data under different iterations, the overall multi-classification effect of the application software system development data is better and the risk level assessment accuracy is higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0036] Figure 1 A flow chart of a data security protection method for application software system development provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0037] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the data security protection method for application software system development proposed by the present invention, its specific implementation method, structure, features and effects, in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.
[0038] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.
[0039] The following is a detailed description of a data security protection method for application software system development provided by the present invention in conjunction with the accompanying drawings.
[0040] See also Figure 1 , which shows a flow chart of a data security protection method for application software system development provided by an embodiment of the present invention, the method comprising:
[0041] S101: Acquire development data related to the user running the application software, quantify the development data and label the development data with preset risk levels to obtain training data, where the training data has different attributes, and the attributes are key-value pairs.
[0042] Through the analysis of the user's application of the software and the correlation between the data attributes and the software, the user's relevant data is obtained, such as information login, software operation, token acquisition, access frequency, location, browsing status, etc. The user's behavior generates corresponding development data and the data is stored in the database.
[0043] Among them, data quantization specifically involves performing operations such as feature extraction and dimensionality reduction on the data to obtain corresponding vectors, and using vectors to represent the development data of the application software, thereby completing data quantization.
[0044] Among them, the preset risk level is the risk level pre-labeled for different development data, which can be pre-labeled by relevant security testers. Security testers obtain the status of these development data, evaluate and test the existing risks, and obtain the preset risk level of each development data. No risk level is represented as 0, and the risk increases successively, represented as 1, 2, 3, etc., and the preset risk level is used as the label of the development data. Of course, corresponding tests can also be performed on all simulated data generated by software development to obtain the preset risk level of each data.
[0045] Among them, attributes are attribute features corresponding to training data. One training data may have multiple attributes. The attributes in the embodiment of the present invention are mainly embodied in the form of key-value pairs, so they are collectively referred to as attributes.
[0046] S102: Use the decision tree as the base classifier of SAMME, input the training data, process it with the same initial weights, and output the iterative risk levels of different training data under different iteration numbers; based on the similarity between the iterative risk levels under different iteration numbers and the preset risk levels of the training data, determine the training data with the correct risk level classification under each iteration as the iterative analysis data.
[0047] Since the judgment of the preset risk level by security testers is mainly obtained through corresponding data testing and subjective experience, the amount of data is relatively small in the development stage, and its overall reliability is low. If the application software is to be put on the shelves, it cannot only be tested on a small scale, otherwise the data generated by each user will be carefully judged by the security testers, which will lead to the consumption of a large amount of system resources and computing power, affecting the user experience and the long-term development of enterprise application software. Therefore, it is necessary to conduct supervised learning on the training data obtained in step S101, and build a suitable data model so that after the application software is put on the shelves, it can quickly and accurately judge the risk level corresponding to the data generated by the user, give the user corresponding prompts and feedback, and make corresponding processing measures and protection.
[0048] In the embodiment of the present invention, the construction of the data model is realized by the SAMME algorithm. The SAMME algorithm is a multi-classification AdaBoost algorithm, which is a multi-classification algorithm well known to those skilled in the art, and no further limitation or elaboration is made thereon.
[0049] First, a model capable of realizing multi-classification is selected as the base classifier of SAMME. The present invention uses a decision tree as the base classifier of SAMME. Then, data analysis is performed by inputting training data. The specific analysis process is described in the subsequent embodiments.
[0050] It is understandable that the traditional SAMME algorithm believes that development data of different categories (ie, risk levels) are independent of each other, that is, there is no connection between risk level 1 and risk level 2. Therefore, the wrongly classified development data are given the same weight value for adjustment. Since the initial weights are the same, the wrongly classified development data in the subsequent process will all be represented by the same weight.
[0051] It should be noted that, in fact, the risk situation of development data is more complicated. On the one hand, there is a certain correlation between different risk levels, and the risk levels are gradually transitioned. When performing specific classification, although some development data are classified incorrectly, they should also be given weight adjustments of different sizes to target the complex risk situations in the development data. On the other hand, the impact caused by the incorrect identification of risk level categories in development data is unequal, that is, the safety hazards caused by misidentification of low-risk development data as high-risk and misidentification of high-risk as low-risk are different. This also shows that in the face of the risk level of development data, different development data misclassifications should also be given different weight adjustments. Based on this, this solution provides a weight analysis method for implementing adaptive weight adjustments for complex risk situations in development data.
[0052] Different risk levels of development data focus on different data attributes (key, value). For example, high-risk data will pay more attention to the source code to prevent Trojans or viruses from causing security vulnerabilities in application software; while lower-risk data will focus on log information, etc. Development data is generated from historical data, and its value decreases over time. Therefore, development data of different risk levels focus on different attributes. By analyzing the features of development data that are correctly classified in each iteration, the weights of misclassified development data are adjusted based on these features, thereby improving the accuracy of the SAMME algorithm in judging the risks of unknown development data.
[0053] In the embodiment of the present invention, the training data is input, processed by the same initial weight, and the iterative risk level of different training data under different iteration times is output. The iterative risk level is the risk level obtained under the iteration of the SAMME algorithm. Since the effects of different iterations are different, the present invention analyzes the data under any iteration number in the subsequent analysis. That is, each iteration is analyzed once.
[0054] Furthermore, in some embodiments of the present invention, based on the similarity between the iterative risk level under different numbers of iterations and the preset risk level of the training data, the training data with correct risk level classification under each iteration is determined as the iterative analysis data, including: judging whether the iterative risk level of the training data under any iteration is the same as the preset risk level, and if they are the same, using the corresponding training data as the iterative analysis data under the corresponding iteration.
[0055] In any iteration, for example, the second iteration, determine whether the iterative risk level of the training data is the same as the preset risk level. If so, it means that the corresponding training data classification is correct. In the embodiment of the present invention, it is necessary to adjust the weight of the data with the correct risk level classification, thereby using it as the iterative analysis data in the second iteration.
[0056] S103: Determine the attribute distribution significance characteristics of each attribute according to the attribute distribution of the iterative analysis data under the same iteration; select the target attribute from the iterative analysis data under the same iteration according to the attribute distribution significance characteristics; determine the similarity coefficient between the change of the attribute distribution significance characteristics of any target attribute under different risk levels and the change of risk level, and determine the weight correction coefficient of the iterative analysis data according to the similarity coefficient of all target attributes in the same iterative analysis data.
[0057] The distribution of attributes can be used to characterize attribute features. In the embodiment of the present invention, attribute analysis is performed on iterative analysis data through the significant features of attribute distribution.
[0058] Furthermore, in some embodiments of the present invention, the attribute distribution significance characteristics of each attribute are determined based on the attribute distribution of iteratively analyzed data under the same iteration, including: constructing an empirical distribution function and a uniform distribution function for the attributes of the same attribute of iteratively analyzed data under the same iteration, using the KS test to determine the difference between the empirical distribution function and the uniform distribution function, and determining the test p-value based on the difference; determining the kurtosis value and interquartile range of the empirical distribution function, and determining the attribute distribution significance characteristics of the attribute by combining the test p-value, kurtosis value and interquartile range.
[0059] Select the iterative analysis data with the correct risk level classification under a certain iteration to form a data set, obtain the kth attribute in the data set, and obtain the attribute characteristics corresponding to the attribute. The attribute characteristics of the iterative analysis data are obtained through the distribution of the attribute value. The distribution of the attribute value is obtained by constructing the empirical distribution function F. Compare the empirical distribution function with the distribution function U of the uniform distribution. If there is a large similarity between the empirical distribution function F and U, it means that under the attribute corresponding to the risk category of the iterative analysis data, the attribute value distribution presents a high randomness, and thus has a small correlation with the category; on the contrary, if the difference is large, it means that the distribution of the attribute value is not a uniform distribution, and it presents a high distribution concentration in certain attribute value intervals; that is, the greater the difference between the distribution function F and the uniform distribution U, the more concentrated the attribute value distribution of the iterative analysis data under the attribute.
[0060] Therefore, in the embodiment of the present invention, the significant feature of the attribute distribution is represented as a numerical value, that is, it is analyzed as a specific numerical value, and the KS test is used to determine the difference between the empirical distribution function and the uniform distribution function, and the test p value is determined according to the difference.
[0061] The KS test is used to determine whether the attribute of the current iterative analysis data comes from a uniform distribution to obtain a test p value. The larger the test p value, the greater the possibility that the iterative analysis data of the attribute value comes from a uniform distribution, and vice versa. The KS test is a technology well known to those skilled in the art, and will not be further limited or elaborated on.
[0062] Furthermore, in some embodiments of the present invention, the attribute distribution significance characteristics of the attribute are determined by combining the test p-value, kurtosis value and interquartile range, and the corresponding calculation formula is:
[0063]
[0064] In the formula, f represents the attribute distribution significance characteristic of the attribute, ku represents the kurtosis value, IQR represents the interquartile range, and p represents the test p value.
[0065] It should be noted that the greater the kurtosis, the smaller the interquartile range, the greater the difference between the data distribution and the uniform distribution, the more significant the corresponding attribute characteristics, that is, the greater the significance characteristics of the attribute distribution, and the test p-value represents the difference between the empirical distribution function and the uniform distribution function. The larger the value, the greater the possibility that the iterative analysis data of the attribute comes from the uniform distribution, and the smaller the corresponding significance. Thus, the significance characteristics of the attribute distribution are calculated.
[0066] Repeat the above operation to obtain the attribute distribution significance characteristics of different attributes of the application software development data under different risk levels. Then, each attribute can be specifically analyzed based on the attribute distribution significance characteristics.
[0067] First, attribute screening is performed. Furthermore, in some embodiments of the present invention, target attributes are screened from iterative analysis data under the same iteration based on the significance characteristics of attribute distribution, including: determining the importance index of the attribute based on the distribution discreteness of the significance characteristics of the attribute distribution of the same attribute under different risk level classifications; and taking the attributes whose importance index values are less than a preset importance threshold as the target attributes.
[0068] In the embodiment of the present invention, the importance of an attribute specifically represents whether the attribute has the ability to distinguish the risk level of development data. The greater the ability of the attribute to distinguish the risk level, the greater the importance of the attribute in risk level judgment.
[0069] Therefore, in an embodiment of the present invention, the importance index of the attribute is determined according to the distribution discreteness of the attribute distribution significance characteristics of the same attribute under different risk level classifications, including: under the same iteration, determining the standard deviation of the attribute distribution significance characteristics of the same attribute under different risk level classifications, and using the standard deviation as the significance distribution discrete index; normalizing the significance distribution discrete index to obtain the importance index of each attribute.
[0070] In the embodiment of the present invention, the larger the standard deviation of the attribute distribution significance feature, the greater the distribution discreteness of the attribute distribution significance feature of the corresponding attribute under different risk level classifications, that is, the greater the distinguishing ability for risk level judgment, which means that the attribute is more important in risk level judgment.
[0071] Among them, the preset importance threshold is the threshold value of the importance index. Specifically, the preset importance threshold can be 0.75. The attribute with a value of the importance index less than 0.75 is taken as the target attribute. If it is less than the preset importance threshold, it means that there is a large difference in the classification status of the attribute under different risk levels, indicating that the attribute shows a high correlation with the risk level of the development data during the iteration of SAMME. Therefore, the similarity coefficient between the change of the attribute distribution significance characteristics of any target attribute under different risk levels and the change of risk level is determined.
[0072] In an embodiment of the present invention, the attribute distribution significance features of the target attribute are arranged in the order of risk level into a first sequence, and the risk levels are arranged in order into a second sequence. Thus, the similarity between the first sequence and the second sequence is calculated to obtain a similarity coefficient. Of course, the method for calculating the similarity of the two sequences is a technology well known in the art and is not further limited to this. Specifically, the dtw values of the two sequences can be calculated, and the opposite numbers of the dtw values can be normalized to obtain the similarity coefficient. Alternatively, the correlation between the two sequences can be determined by calculating the correlation coefficient, so that the correlation is used as the similarity coefficient.
[0073] Furthermore, in some embodiments of the present invention, a weight correction coefficient of the iterative analysis data is determined based on the similarity coefficients of all target attributes in the same iterative analysis data, including: taking the mean of the similarity coefficients of all target attributes in the same iterative analysis data as the weight correction coefficient of the iterative analysis data.
[0074] Conduct a specific analysis on the iterative analysis data, that is, calculate the weight correction coefficient as the weight correction indicator of the iterative analysis data to produce weight differentiation with other development data.
[0075] S104: Correct the initial weight of the iterative analysis data of the same iteration in the original SAMME algorithm according to the weight correction coefficient, and determine the correction weight of different iterative analysis data under the same iteration; adjust the correction weight according to the difference between the iterative risk level and the preset risk level to obtain the adjusted weight of the iterative analysis data, perform iterative analysis of the SAMME algorithm according to the adjusted weights of all iterative analysis data, and obtain a security protection model.
[0076] Further, in some embodiments of the present invention, the initial weight of the iterative analysis data of the same iteration in the original SAMME algorithm is corrected according to the weight correction coefficient, and the correction weight of different iterative analysis data under the same iteration is determined, including: taking the ratio of the initial weight to the weight correction coefficient as the correction weight, and the corresponding calculation formula is:
[0077]
[0078] It represents the weight correction coefficient corresponding to the iterative analysis data under the kth iteration, represents the weight before the data obtained by the original SAMME algorithm, It represents the correction weight corresponding to the iterative analysis data under the kth iteration. The stronger the correlation between the application software development data and the risk level, that is, the larger the weight correction coefficient, the better the risk level judgment of the iterative analysis data can be judged by the attributes. Therefore, in the SAMME algorithm, such data should be given less attention and the corresponding weight should be smaller. Therefore, it is used as the denominator for analysis.
[0079] It should be noted that in order to ensure that the calculation results are meaningful, when performing fractional operations in the embodiments of the present invention, when the denominator is 0, it is necessary to add a parameter adjustment factor greater than 0 to the denominator to prevent the denominator from being 0. The value of the parameter adjustment factor is set by the implementer according to actual conditions, such as 0.1 or 1, and this application does not impose any special restrictions.
[0080] Furthermore, in some embodiments of the present invention, the correction weight is adjusted according to the difference between the iterative risk level and the preset risk level to obtain the adjusted weight of the iterative analysis data, including: taking the difference between the preset risk level and the iterative risk level as the adjustment coefficient; and normalizing the product of the adjustment coefficient and the correction weight as the adjustment weight.
[0081] Mistaking high-risk data for low-risk data will result in insufficient risk handling, causing users and corporate data to lack adequate security protection and causing serious damage. Therefore, the harm to security protection is greater at this time. On the contrary, if low-risk data is identified as high-risk, although the classification is wrong at this time, it will not cause data security protection loss.
[0082] Therefore, it is necessary to pay more attention to identifying high-risk development data as low-risk in the SAMME algorithm to avoid this type of classification errors as much as possible; and to pay relatively less attention to identifying low-risk data as high-risk, so that this type of classification error can be allowed in the SAMME algorithm to prevent the SAMME algorithm from excessively pursuing the accuracy of this type of data classification, leading to serious data security protection problems.
[0083] Therefore, in the embodiment of the present invention, the difference between the preset risk level and the iterative risk level is used as the adjustment coefficient. The larger the adjustment coefficient, the greater the possibility that the data with higher risk is mistaken for the data with lower risk level, and the smaller the adjustment coefficient, the greater the possibility that the data with low risk is identified as high risk. The product of the adjustment coefficient and the correction weight is normalized as the adjustment weight. The adjustment weight can effectively characterize the weight value of the corresponding iterative analysis data when the subsequent SAMME algorithm is analyzed under different iterations.
[0084] Furthermore, in some embodiments of the present invention, the SAMME algorithm is iteratively analyzed based on the adjusted weights of all iterative analysis data to obtain a security protection model, including: performing SAMME algorithm processing based on the adjusted weights of the iterative analysis data under all iteration numbers until the rate of change of the weights of all development data before and after the iteration process is within 2%, then stopping the iteration to obtain a security protection model.
[0085] When the weight change rate of all development data is within 2% during the iteration process, it means that the SAMME algorithm is stable and the iteration is stopped. The model is built. The data generated by the user is used as the input of the model, and the risk level corresponding to the output data is output. This risk level is more reliable than the SAMME algorithm analysis directly using the same weight, and the risk level assessment is more accurate.
[0086] The invention obtains development data related to application software run by users, quantifies the development data and labels the development data with preset risk levels, and obtains training data, wherein the training data has different attributes, and the attributes are key-value pairs; a decision tree is used as a base classifier of SAMME, and the training data is input, and the training data is processed by the same initial weight, and the iterative risk levels of different training data under different iteration times are output; based on the similarity between the iterative risk levels under different iteration times and the preset risk levels of the training data, the training data with correct risk level classification under each iteration is determined as iterative analysis data; according to the attribute distribution of the iterative analysis data under the same iteration, the attribute distribution significance feature of each attribute is determined; according to the attribute distribution significance feature, the iterative analysis data under the same iteration is obtained. The target attribute is screened from the iterative analysis data under one iteration; the similarity coefficient between the change of the attribute distribution significance characteristics of any target attribute under different risk levels and the change of the risk level is determined, and the weight correction coefficient of the iterative analysis data is determined according to the similarity coefficient of all target attributes in the same iterative analysis data; the initial weight of the iterative analysis data of the same iteration in the original SAMME algorithm is corrected according to the weight correction coefficient, and the correction weight of different iterative analysis data under the same iteration is determined; according to the difference between the iterative risk level and the preset risk level, the correction weight is adjusted to obtain the adjustment weight of the iterative analysis data, and the SAMME algorithm is iteratively analyzed according to the adjustment weights of all iterative analysis data to obtain the security protection model. Since the adjustment weight is adaptively obtained according to the attribute performance of the development data under different iterations, the overall multi-classification effect of the application software system development data is better and the risk level assessment accuracy is higher.
[0087] It should be noted that the sequence of the above embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0088] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.
Claims
1. A data security protection method for application software system development, characterized in that: The method comprises: Acquire development data related to the application software run by the user, quantify the development data and mark the preset risk level to obtain training data, wherein the training data has different attributes, and the attributes are key-value pairs; The decision tree is used as the base classifier of SAMME, and the training data is input and processed with the same initial weights to output the iterative risk levels of different training data at different iteration times. Based on the similarity between the iterative risk levels at different iteration times and the preset risk levels of the training data, the training data with the correct risk level classification at each iteration is determined as the iterative analysis data. According to the attribute distribution of the iterative analysis data under the same iteration, the attribute distribution significance feature of each attribute is determined; according to the attribute distribution significance feature, the target attribute is screened from the iterative analysis data under the same iteration; the similarity coefficient between the change of the attribute distribution significance feature of any target attribute under different risk levels and the change of the risk level is determined, and the weight correction coefficient of the iterative analysis data is determined according to the similarity coefficient of all target attributes in the same iterative analysis data; The initial weight of the iterative analysis data of the same iteration in the original SAMME algorithm is corrected according to the weight correction coefficient, and the correction weight of different iterative analysis data under the same iteration is determined; according to the difference between the iterative risk level and the preset risk level, the correction weight is adjusted to obtain the adjusted weight of the iterative analysis data, and the SAMME algorithm is iteratively analyzed according to the adjusted weights of all iterative analysis data to obtain a security protection model.
2. A method for protecting data security in application software system development as claimed in claim 1, characterized in that: The similarity between the iterative risk level at different iteration times and the preset risk level of the training data is determined as the iterative analysis data with the correct risk level classification at each iteration, including: Determine whether the iteration risk level of the training data in any iteration is the same as the preset risk level. If they are the same, use the corresponding training data as the iteration analysis data in the corresponding iteration.
3. A method for protecting data security in application software system development as claimed in claim 1, characterized in that: The step of iteratively analyzing the attribute distribution of the data under the same iteration and determining the attribute distribution significance feature of each attribute includes: Iteratively analyze the data under the same iteration and construct an empirical distribution function and a uniform distribution function for the attribute of the same attribute, use the KS test to determine the difference between the empirical distribution function and the uniform distribution function, and determine the test p value according to the difference; The kurtosis value and the interquartile range of the empirical distribution function are determined, and the attribute distribution significance characteristics of the attribute are determined by combining the test p value, the kurtosis value and the interquartile range.
4. A method for protecting data security in application software system development as claimed in claim 3, characterized in that: The attribute distribution significance feature is a numerical representation. The attribute distribution significance feature of the attribute is determined by combining the test p value, the kurtosis value and the interquartile range. The corresponding calculation formula is: In the formula, f represents the attribute distribution significance characteristic of the attribute, ku represents the kurtosis value, IQR represents the interquartile range, and p represents the test p value.
5. The data security protection method for application software system development according to claim 1, characterized in that: The step of selecting and obtaining a target attribute from iterative analysis data under the same iteration according to the significant characteristics of the attribute distribution includes: According to the distribution discreteness of the attribute distribution significance characteristics of the same attribute under different risk level classifications, the importance index of the attribute is determined; The attribute whose value of the importance index is less than a preset importance threshold is taken as the target attribute.
6. A method for protecting data security in application software system development as claimed in claim 5, characterized in that: The importance index of the attribute is determined according to the distribution discreteness of the attribute distribution significance characteristics of the same attribute under different risk level classifications, including: In the same iteration, the standard deviation of the attribute distribution significance characteristics of the same attribute under different risk level classifications is determined, and the standard deviation is used as a significance distribution dispersion indicator; The significance distribution discrete index is normalized to obtain the importance index of each attribute.
7. The data security protection method for application software system development as claimed in claim 1, characterized in that: The step of determining the weight correction coefficient of the iterative analysis data according to the similarity coefficients of all target attributes in the same iterative analysis data includes: The mean of the similarity coefficients of all target attributes in the same iterative analysis data is used as the weight correction coefficient of the iterative analysis data.
8. The method for protecting data security during development of an application software system according to claim 1, characterized in that: The method of correcting the initial weight of the iterative analysis data of the same iteration in the original SAMME algorithm according to the weight correction coefficient to determine the correction weights of different iterative analysis data under the same iteration includes: The ratio of the initial weight to the weight correction coefficient is used as the correction weight.
9. The method for protecting data security during development of an application software system according to claim 1, wherein: The step of adjusting the correction weight according to the difference between the iterative risk level and the preset risk level to obtain the adjustment weight of the iterative analysis data includes: Using the difference between the preset risk level and the iterative risk level as an adjustment coefficient; The product of the adjustment coefficient and the correction weight is normalized to obtain the adjustment weight.
10. The data security protection method for application software system development according to claim 1, characterized in that: The SAMME algorithm is iteratively analyzed according to the adjusted weights of all iterative analysis data to obtain a security protection model, including: The SAMME algorithm is used to process the adjusted weights of the iterative analysis data under all iteration numbers until the change rate of the weights of all development data before and after the iteration is within 2%. The iteration is stopped and the security protection model is obtained.
Citation Information
Patent Citations
Hardware resource configuration for processing system
US20220035679A1
Methods of predicting phenome-wide polygenic risk scores using multi-task learning
US20240379240A1