Data isolation method and platform for an integrated testing service platform
By extracting the historical user data of the integrated test service platform, the current user data account for the stock and divide the independent storage space, the problems of multi-user data confusion and waste of storage space are solved, and the isolation of user data and efficient utilization of storage space are achieved.
Patent Information
- Application Number
- CN202510253144.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The integrated testing service platform may be confusing due to the coexistence of multiple users, and there is a risk of exposing important data information of users. At the same time, repeated storing of the same data can lead to wasting storage space.
By extracting historical user data features, estimate the current user's data to account for stock, divide independent storage space, and establish identity contacts to ensure data isolation. At the same time, optimize the storage of the same data and improve the utilization of database storage space.
It realizes effective isolation of user data, avoids data confusion and interaction, improves the utilization rate of database storage space, and ensures the security of user data.
Smart Images

Figure CN119760784B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a data isolation method and platform for an integrated test service platform. Background Art
[0002] With the development of science and technology, the methods for training AI models have gradually matured, and the main factors affecting the training models have gradually shifted from methods to the use and processing of computing power and training data volume. Currently, in order to apply to model training in actual engineering situations, some enterprises provide an integrated test service platform that can provide model training services for multiple tenants.
[0003] Of course, for an integrated test platform, although it can efficiently and quickly integrate algorithms and data information, due to its multi-user characteristics, it will inevitably lead to the situation where the data of different users are stored together in the platform's database, which may cause data confusion between different users, is not conducive to training processing, and also has the risk of exposing important data information of users.
[0004] Therefore, designing a data isolation method and platform for an integrated test service platform, which realizes the reasonable storage space division for the current user by extracting features from the historical data of the platform, improves the utilization rate of the database storage space, and effectively ensures the effective data isolation between different users, is an urgent problem to be solved at present. Summary of the Invention
[0005] The purpose of the present invention is to provide a data isolation method for an integrated test service platform, which provides a relatively independent storage space for different users in the database and stores all the corresponding data of the users in this relatively independent storage space. To achieve this method, first, it is necessary to determine the possible data storage volume of the users to ensure that the divided independent storage space fully meets the data storage requirements of the users. On the other hand, it is necessary to ensure that the data in the storage space can be uniquely and stably corresponded to the users to ensure that there is no mutual extraction and confusion of data in different storage spaces. Considering these two aspects, this application estimates the data storage volume of the current user by extracting features from the data storage situation of historical users on the platform to determine the size of the storage space division and establish an identity connection for a close association with the storage space. In addition, considering that in the case of similar training models, there are some identical data information among different users, the repeated storage of these data wastes the storage space of the database. Therefore, the optimized storage of the same data can be carried out to improve the utilization rate of the database storage space and avoid the confusion of data use.
[0006] The object of the present invention also lies in providing a data isolation platform for an integrated test service platform. This platform obtains the data of historical users through a data acquisition unit, and while extracting feature information through an analysis and processing unit, it realizes the estimation of the data occupancy of the current user, thereby ensuring that a reasonable-sized independent storage space is provided for the user, making the user data confidential and avoiding confusion and interaction with the data of other users. Different units are closely linked, effectively realizing the mutual isolation of data usage for multiple users, effectively ensuring the security of user data usage, and at the same time improving the utilization efficiency of the platform database space, which is an important material basis for realizing effective data isolation.
[0007] In a first aspect, the present invention provides a data isolation method for an integrated test service platform, including: obtaining historical user data, performing data storage volume analysis based on application types to form database space usage feature data; obtaining current user usage information, and combining the database space usage feature data to perform space division to form multi-user database division information; according to the multi-user database division information, performing data optimization based on type similarity to form multi-user shared application isolation database data.
[0008] In the present invention, this method provides a relatively independent storage space for different users in the database and stores all the corresponding data of the users in this relatively independent storage space. To achieve this, first, it is necessary to determine the possible data storage volume of the users to ensure that the divided independent storage space fully meets the data occupancy requirements of the users. On the other hand, it is necessary to ensure that the data in the storage space can be uniquely and stably corresponded to the users to ensure that there is no mutual extraction and confusion of data in different storage spaces. Considering these two aspects, this application estimates the data occupancy of the current user by extracting features from the data occupancy of historical users on the platform to determine the size of the storage space division and establish an identity connection for a close association with the storage space. Additionally, considering that there is a situation where the data information of different users is partially the same when the training models are similar, the repeated storage of this data wastes the storage space of the database. Therefore, the optimized storage of the same data can be carried out to improve the utilization efficiency of the database storage space and also avoid confusion in data usage.
[0009] As a possible implementation, historical user data is obtained, data storage volume analysis based on application types is performed, and database space usage characteristic data is formed, including: extracting the storage usage information corresponding to different historical users with the same application type in the historical user data to form the storage usage information of historical users of the same type; performing occupancy characteristic analysis based on the influx volume of single training data on the storage usage information of all historical users of the same type with the same application type to determine the type envelope occupancy corresponding to the application type; aggregating the type envelope occupancies corresponding to different application types to form database space usage characteristic data.
[0010] In the present invention, to achieve data isolation between different users on the integrated test service platform, first, reasonable storage space needs to be provided for different users in the database to ensure the data storage requirements of the users. Therefore, an accurate and reasonable estimation of the data storage volume of the users is required to reasonably divide the storage space in the database. Here, the corresponding storage space estimation reference data is established by extracting the characteristics of the storage space size of the usage data of the platform's historical users to ensure the storage space of the current user. It can be understood that the training types required for different users are different, so the amount of training data provided may be different. Especially, there is a large difference in the influx volume of the training data provided each time. Therefore, when extracting the data storage characteristic information of historical users, clustering by the application type of the model can more reasonably extract the characteristics of the data storage volume of the users and ensure that a reasonable data reference is made for providing a storage space size that fully meets the requirements of the current user in the future. In addition, considering that users do not provide all the training data at one time when using data storage, but provide it continuously in batches, the size of the storage space used has a certain dynamic stability. Therefore, when extracting characteristic information, it is necessary to fully consider the characteristics of this multi-term data influx to provide a more reasonable estimation of the storage space size.
[0011] As a possible implementation, performing occupancy characteristic analysis based on the influx volume of single training data on the storage usage information of all historical users of the same type with the same application type to determine the type envelope occupancy corresponding to the application type, including: extracting occupancy characteristics based on the intersection stability of the influx terms of training data for the storage usage information of different historical users of the same type with the same application type to determine the stable occupancy of the same type of users corresponding to the historical users. ; According to the stable occupancy of the same type of users corresponding to different historical users with the same application type , perform average occupancy analysis based on data clustering to determine the type envelope occupancy corresponding to the application type.
[0012] In the present invention, to determine a reasonable data occupancy corresponding to an application type, it is first necessary to determine the amount of data occupied by historical users in the long term under this type, and perform an average occupancy analysis based on big data in combination with these data amounts, so as to obtain a representative data occupancy that can reflect this type.
[0013] As a possible implementation, occupancy feature extraction based on the stability of the intersection of the influx times of training data is performed on the storage usage information of different historical users of the same application type to determine the stable occupancy of users of the same type corresponding to the historical users. This includes: for the storage usage information of different historical users of the same type, extracting the average amount of training data influx per user per single training during the entire usage period, the corresponding duration of the continuous occupancy period of the training data, and the basic storage amount of the training model. Based on the duration of the continuous occupancy period of all training data for each single training and the corresponding average amount of training data influx per user per single training, determine the maximum amount of training data influx of users during the period when the training data influxed during the usage period does not overlap with the training data influxed in the adjacent next item in terms of time. And the amount of overlapping training data influx of users corresponding to each overlapping period during the period when the training data influxed during the usage period overlaps with the training data influxed in the adjacent next item in terms of time. Here, n represents the number of different application types, m represents the number of different historical users under the application type numbered n, and k represents the sequence number of the multiple occurrences of overlapping training data corresponding to the historical user numbered m under the application type numbered n; for the different amounts of overlapping training data influx of users corresponding to the historical users. Perform the following effective stability analysis: Aggregate all the amounts of overlapping training data influx of users. To form a set of overlapping training data influx amounts of users. For the set of overlapping training data influx amounts of users. If , and the duration of the overlapping occurrence of the largest amount of overlapping training data influx in the set reaches the stable overlapping duration threshold, then determine the largest amount of overlapping training data influx as the effective stable overlapping influx amount of users. Otherwise, determine the effective stable overlapping influx amount of users =0. represents the variance of the set of the determined amounts of overlapping training data influx of users. represents the stable variance limit value of the overlapping influx data; based on the basic storage amount of the training model , the maximum amount of training data influx of users and the effective stable overlapping influx amount of users , determine the stable occupancy of users of the same type corresponding to historical users , where .
[0014] In the present invention, considering that historical users commonly provide training data in multiple batches during training data provision, it is necessary to fully consider this dynamic stability when extracting the feature quantity of the occupied storage space size. For different historical users in the same application type, it is necessary to obtain information on the amount of training data provided each time to determine the maximum single influx of training data. However, considering that after each amount of training data is provided, it is possible that the amount of training data provided last time has not been fully used and consumed, and the amount of training data for the next time will influx again. In this way, the size of the storage space occupied by the data is not just the size of the single influx of data. Because while obtaining the maximum single influx of data of the user, it is also necessary to confirm the size of the data volume during the storage period where adjacent data influx sub-items overlap. Of course, it can be understood that the independent storage space divided is not exactly the same as the maximum occupancy that has occurred for historical users, but a certain redundancy will be provided. Therefore, it is necessary to confirm whether it is necessary to consider the size of the data volume during the overlapping period of adjacent sub-items by checking whether there is often a certain amount of data volume overlap. After all, the redundancy of the storage space divided by occasional or small data volume regular overlaps can be temporarily compensated, while the long-term repeated occupation of a large space needs to consider the situation where the input of other data may cause the storage space size to be insufficient. Whether to consider the data volume occupancy situation during this long-term large repetition period is determined by judging the data volume stability and the duration of the repetition period. The stability is reflected in the form of variance, and the length of the occupancy period is judged by a reasonable time threshold. The stable variance limit and the stable overlap duration threshold for the overlapping influx data can be set according to the actual situation or determined based on big data analysis.
[0015] As a possible implementation, based on the stable occupancy of users of the same type corresponding to different historical users under the same application type, perform an average occupancy analysis based on data clustering to determine the type envelope occupancy corresponding to the application type, including: for different historical users under the same application type, according to the corresponding stable occupancy of users of the same type , perform the following clustering average occupancy analysis: for each stable occupancy of users of the same type , determine the total difference from the stable occupancy of other users of the same type , sort the total differences from small to large, and take the average value of the stable occupancy of users of the same type corresponding to the first half of the total differences , and calibrate it as the type envelope occupancy corresponding to the application type.
[0016] In the present invention, for the stable occupancy amounts of different historical users under the same application type in historical data, there will be a certain degree of difference in the data volume of the trained model. The features to be extracted are big data features, mainly reflecting the average level of the occupancy amount used by users under this type. And the size of the gap between this average level and the stable occupancy amounts of different users is relevant. Therefore, in this application, by obtaining the gap between the stable occupancy amounts of different users, the stable occupancy amount data of the part with a smaller gap is extracted to determine the occupancy amount space required for the corresponding application type. Here, the quantity extracted according to the total difference can be half of the quantity of the stable occupancy amounts of users, or it can be reasonably extracted according to the size gap between the total differences. It can be understood that the part with a smaller total difference gap indicates a higher degree of aggregation, and the obtained data can fully reflect the characteristics of the space occupancy usage under this application type.
[0017] As a possible implementation, obtain the usage information of the current user, and combine it with the database space usage feature data for space division to form multi-user database division information, including: according to the usage information of the current user of different current users, determine the corresponding current application type of different current users; according to the current application type corresponding to different current users, and combine the database space usage feature data, determine the type envelope occupancy amount corresponding to the application type that is the same as the current application type, and determine the type envelope occupancy amount as the user preset occupancy amount corresponding to the current user; obtain the total occupancy amount of the platform database, and perform division processing on the user occupancy space according to the user preset occupancy amounts corresponding to different current users to form multi-user database division information.
[0018] In the present invention, with the feature data information as a reference for the space size of the independent storage space divided by users, after obtaining the types of training models of different current users, the corresponding type envelope occupancy amount can be determined based on the application type. Dividing the database into storage spaces for different current users based on the type envelope occupancy amount can provide a certain amount of space redundancy to handle the space usage amount not covered by big data. Of course, more importantly, after completing the division and storing the usage data of the current user in the divided storage space, it is also necessary to label the storage space with the identity information of the user to ensure that the data in the storage space can only be used through correct user identity information authentication, achieving a better data isolation effect.
[0019] As a possible implementation, obtain the total occupancy amount of the platform database, and perform division processing on the user occupancy space according to the user preset occupancy amounts corresponding to different current users, including: according to the total occupancy amount of the platform database and the user preset occupancy amounts corresponding to different current users , the storage space of users is divided in the following way. i represents the numbers of different current users: If , then the user preset storage amounts corresponding to different current users are used as the storage space division amounts to divide independent storage spaces corresponding to each current user in the platform database, which are marked as user independent storage spaces, and the current usage data of different current users are stored in the corresponding user independent storage spaces. α is the space division limit factor; if , then the user preset storage amounts corresponding to different current users are compared with each other to perform proportional division on the spaces of sizes in the platform database to determine the user independent storage spaces corresponding to different current users, and the current usage data of different current users are stored in the corresponding user independent storage spaces; the user independent storage spaces are marked with the identity information of the corresponding current users, and the identity information of all user independent storage spaces and the corresponding current users is extracted to form an isolation space identity connection form; combining the user independent storage spaces corresponding to different current users and the isolation space identity connection form forms the multi-user database division information.
[0020] In the present invention, of course, since the size of the database of the integrated test service platform is also limited, when dividing the storage space sizes of different users, it is necessary to consider the total space size of the database. After all, the database not only stores the usage data of different users, but also provides storage space for its own data information analysis and processing and necessary data such as different model data. Therefore, when dividing the independent storage spaces of different users, it is also necessary to fully consider. In this application, when the space is sufficient, the division is performed according to the envelope storage space size provided by the characteristic data, and when the database space is insufficient, the space division for different users is performed according to the proportional relationship of the envelope storage space. Of course, in this case, it is necessary to provide a limit on the increase in the number of users provided by the platform and a limit on the size of the influx of current user data. The space division limit factor can be set according to the size of the necessary data actually stored in the platform.
[0021] As a possible implementation manner, based on the multi-user database division information, data optimization based on type similarity is performed to form multi-user shared application isolation database data, including: determining the current users with the same current application type according to the current application types corresponding to different current users, and aggregating them into a set of users with the same type of usage; for different sets of users with the same type of usage, obtaining the current usage data in the user independent storage spaces corresponding to all current users in the set and performing extraction and sharing processing of the same data to form multi-user shared application isolation database data.
[0022] In the present invention, since the models trained by current users of the same application type are similar or identical, there may be the same data inputs during model training. Although different users independently store such the same data to ensure data isolation, it will reduce the utilization rate of storage space, especially when the amount of the same data is large. Based on reasonably and effectively isolating the data of different users, this application establishes a shared storage space in the database for unified storage of the same data, avoiding waste of space caused by multiple storages.
[0023] As a possible implementation manner, for different sets of users of the same type, obtain the current usage data in the user-independent occupied storage spaces corresponding to all current users in the set and perform extraction and sharing processing of the same data to form multi-user shared application isolation database data, including: for all current usage data corresponding to the set of users of the same type, if the same data information exists at the same time, extract the same data information and store it in the shared storage space of the platform database, and associate the identity information of the corresponding current user with the same data information; extract the identity association information of all the same data in the shared storage space to form a shared data form; combine all the same data and the shared data form in the shared storage space to form multi-user shared application isolation database data.
[0024] In the present invention, considering that the same data is for different users, in order to avoid errors in the sharing of the same data, it is necessary to mark the identity information of the corresponding users for the same data. Therefore, all data in the entire shared storage space has identity identification. When users use it, they can obtain it through the identity information, ensuring effective data isolation while also improving the utilization rate of the storage space.
[0025] In a second aspect, the present invention provides a data isolation platform for an integrated test service platform, including: a data acquisition unit for obtaining historical user data and current user usage information; an analysis and processing unit for performing data storage amount analysis based on the application type on the historical user data obtained by the data acquisition unit to form database space usage characteristic data, performing space division on the current user usage information in combination with the database space usage characteristic data to form multi-user database division information, and performing shared optimization processing on the same data according to the current user usage information of the current user to form multi-user shared isolation database data; a database storage unit for independently storing the current user usage information according to the multi-user database division information and performing shared storage on the same data according to the multi-user shared isolation database data.
[0026] In the present invention, the platform acquires the data of historical users through the data acquisition unit, and while extracting feature information through the analysis and processing unit, it estimates the data occupancy of the current user, thereby ensuring that a reasonable-sized independent storage space is provided for the user, making the user data confidential and avoiding confusion and interaction with the data of other users. Different units are closely connected, effectively realizing the mutual isolation of data usage for multiple users, effectively ensuring the security of user data usage, and at the same time improving the utilization efficiency of the platform database space, which is an important material basis for realizing effective data isolation.
[0027] The beneficial effects of the data isolation method and platform of an integrated test service platform provided by the present invention are as follows:
[0028] This method provides a relatively independent storage space for different users in the database and stores all the corresponding data of the users in this relatively independent storage space. To achieve this method, first, it is necessary to determine the possible data storage capacity of the users to ensure that the divided independent storage space fully meets the data occupancy requirements of the users. On the other hand, it is necessary to ensure that the data in the storage space can be uniquely and stably corresponded to the users to ensure that there is no mutual extraction and confusion of data in different storage spaces. Considering these two aspects, this application estimates the data occupancy of the current user by extracting the feature of the data occupancy of historical users on the platform to determine the size of the storage space division and establish an identity connection to closely associate with the storage space. Additionally, considering that there are some identical data information among different users in the case of similar training models, the repeated storage of these data wastes the storage space of the database. Therefore, the optimized storage of the same data can be carried out to improve the utilization rate of the database storage space and avoid the confusion of data usage at the same time.
[0029] The platform acquires the data of historical users through the data acquisition unit, and while extracting feature information through the analysis and processing unit, it estimates the data occupancy of the current user, thereby ensuring that a reasonable-sized independent storage space is provided for the user, making the user data confidential and avoiding confusion and interaction with the data of other users. Different units are closely connected, effectively realizing the mutual isolation of data usage for multiple users, effectively ensuring the security of user data usage, and at the same time improving the utilization efficiency of the platform database space, which is an important material basis for realizing effective data isolation. Brief Description of the Drawings
[0030] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for use in the embodiments of the present invention. It should be understood that the following drawings only show certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0031] Figure 1 It is a step diagram of the data isolation method for the integrated test service platform provided by the embodiments of the present invention;
[0032] Figure 2 It is a schematic structural diagram of the data isolation platform of the integrated test service platform provided by the embodiments of the present invention. Specific embodiments
[0033] The following will describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention.
[0034] With the development of science and technology, the methods for AI model training have gradually matured, and the main factors affecting the training model have gradually shifted from methods to the use and processing of computing power and training data volume. Currently, in order to apply to model training in actual engineering situations, some enterprises have provided an integrated test service platform that can provide model training services for multiple tenants.
[0035] Of course, for the integrated test platform, although it can efficiently and quickly integrate algorithms and data information, due to its multi-user characteristics, it will inevitably lead to the situation where data of different users is jointly stored in the platform's database, which may cause data confusion among different users, is not conducive to training processing, and also has the risk of exposing important data information of users.
[0036] Refer to Figure 1 - Figure 2, an embodiment of the present invention provides a data isolation method for an integrated test service platform. This method provides a relatively independent storage space for different users in the database and stores all the data corresponding to the users in this relatively independent storage space. To achieve this method, first, it is necessary to determine the possible data storage volume of the users to ensure that the divided independent storage space fully meets the data storage requirements of the users. On the other hand, it is necessary to ensure that the data in the storage space can be uniquely and stably corresponding to the users to ensure that there is no mutual extraction and confusion of data in different storage spaces. Considering these two aspects, this application estimates the data storage volume of the current user by extracting the characteristics of the data storage situation of historical users on the platform to determine the size of the storage space division and establish an identity connection for a close association with the storage space. In addition, considering that there is a situation where some data information of different users is the same when the training models are similar, the repeated storage of these data wastes the storage space of the database. Therefore, the optimized storage of the same data can be carried out to improve the utilization rate of the database storage space and avoid the confusion of data use at the same time.
[0037] The data isolation method for the integrated test service platform specifically includes the following steps:
[0038] S1: Obtain historical user data, perform data storage volume analysis based on the application type, and form database space usage characteristic data.
[0039] Obtain historical user data, perform data storage volume analysis based on the application type, and form database space usage characteristic data, including: extracting the storage volume usage information corresponding to different historical users with the same application type in the historical user data to form the storage usage volume information of the same type of historical users; performing occupancy characteristic analysis based on the influx volume of single training data on the storage usage volume information of all the same type of historical users with the same application type to determine the type envelope occupancy corresponding to the application type; aggregating the type envelope occupancies corresponding to different application types to form database space usage characteristic data.
[0040] To achieve data isolation between different users on the integrated test service platform, it is first necessary to provide reasonable storage space for different users in the database to ensure the data storage requirements of users. Therefore, the data storage volume of users needs to be accurately and reasonably estimated to reasonably divide the storage space in the database. Here, by extracting the characteristics of the storage space size of the usage data of historical users on the platform, corresponding reference data for storage space estimation is established to ensure the storage space of current users. It can be understood that the training types required for different users are different, so the amount of training data provided may be different. In particular, there are significant differences in the influx volume of training data provided each time. Therefore, when extracting the data storage characteristic information of historical users, clustering conditions can be based on the application type of the model, which can more reasonably extract the characteristics of the data storage volume of users and ensure that subsequent reasonable data references are made for providing sufficient storage space to meet the needs of current users. In addition, considering that users do not provide all training data at once when using data storage, but provide it continuously in batches, the size of the storage space used has a certain dynamic stability. Therefore, when extracting characteristic information, it is necessary to fully consider the characteristics of this multi-term data influx to provide a more reasonable estimate of the storage space size.
[0041] Conduct occupancy characteristic analysis based on the influx volume of single training data for the storage usage information of all historical users of the same application type, and determine the type envelope occupancy corresponding to the application type, including: extracting occupancy characteristics based on the intersection stability of training data influx terms for the storage usage information of different historical users of the same application type, and determining the stable occupancy of the same type of users corresponding to the historical users. ; Based on the stable occupancy of the same type of users corresponding to different historical users under the same application type Conduct average occupancy analysis based on data clustering to determine the type envelope occupancy corresponding to the application type.
[0042] To determine the reasonable data occupancy corresponding to the application type, first, the size of the data volume that historical users in this type have occupied for a long time should be determined, and based on these data volumes, average occupancy analysis based on big data should be conducted to obtain the representative data occupancy that can reflect this type.
[0043] Extract occupancy characteristics based on the intersection stability of training data influx terms for the storage usage information of different historical users of the same application type, and determine the stable occupancy of the same type of users corresponding to the historical users. , including: storing the usage information of different historical users of the same type, extracting the average amount of training data influx per single user for each training during the entire usage period, the corresponding duration of the training data's continuous storage period, and the basic storage amount of the training model ; according to the duration of the continuous storage period of all training data for a single training and the corresponding average amount of training data influx per single user, determining the maximum amount of training data influx of users during the period when the training data influxed during the usage period does not overlap with the training data influxed in the adjacent item in terms of time and the amount of overlapping user training data influx corresponding to each overlapping period during the period when the training data influxed during the usage period overlaps with the training data influxed in the adjacent item in terms of time , n represents the number of different application types, m represents the number of different historical users under the application type numbered n, and k represents the sequential number of the overlapping occurrences of multiple training data corresponding to the historical user numbered m under the application type numbered n; for the different amounts of overlapping user training data influx corresponding to historical users perform effective and stable analysis in the following way: gather all the amounts of overlapping user training data influx , forming a set of overlapping amounts of user training data influx ; for the set of overlapping amounts of user training data influx , if , and the duration of the overlapping occurrence of the largest amount of overlapping user training data influx in the set reaches the stable overlapping duration threshold, then determine the largest amount of overlapping user training data influx as the effective and stable overlapping influx amount of users , otherwise determine the effective and stable overlapping influx amount of users = 0, represents the variance of determining the set of overlapping amounts of user training data influx , represents the stable variance limit value of overlapping influx data; according to the basic storage amount of the training model , the maximum amount of training data influx of users and the effective and stable overlapping influx amount of users , determine the stable storage amount of users of the same type corresponding to historical users , where .
[0044] Considering the common way that historical users provide training data in batches multiple times when providing training data, it is necessary to fully consider this dynamic stability when extracting the feature quantity of the occupied storage space size. For different historical users in the same application type, it is necessary to obtain information about the amount of training data provided each time to determine the maximum single influx of training data. However, considering that after each amount of training data is provided, it is possible that the amount of training data provided last time has not been fully used and consumed, and the next amount of training data will flood in again. In this way, the size of the occupied storage space is not just the size of the single influx of data. Because when obtaining the maximum single influx of data of the user, it is also necessary to confirm the size of the data volume during the storage period where adjacent data influx items overlap. Of course, it can be understood that the independently divided storage space is not exactly the same as the maximum occupied storage amount that has occurred for historical users, but a certain redundancy will be provided. Therefore, it is necessary to confirm whether it is necessary to consider the size of the data volume during the overlapping period of adjacent items by checking whether there is often a certain amount of data volume overlap. After all, occasional or small-data-volume frequent overlaps can be temporarily compensated by the redundancy of the divided storage space, while long-term repeated occupation of a large space needs to consider the situation where the input of other data may cause the storage space to be insufficient. Whether to consider the data volume occupation situation during such a long-term large repetition period is determined by judging the data volume stability and the duration of the repetition period. The stability is reflected in the form of variance, and the length of the occupation period is judged by a reasonable duration threshold. The stable variance limit and stable overlap duration threshold for the overlapping influx data can be set according to the actual situation or determined based on big data analysis.
[0045] Based on the stable occupied storage amounts of users of the same type corresponding to different historical users under the same application type, perform an analysis of the average occupied storage amount based on data clustering to determine the type envelope occupied storage amount corresponding to the application type, including: for different historical users under the same application type, according to the stable occupied storage amounts of users of the same type corresponding to them , perform the following clustering average occupied storage amount analysis: for each stable occupied storage amount of users of the same type , determine the total difference from the stable occupied storage amounts of other users of the same type . Sort the total differences from small to large, and take the average value of the stable occupied storage amounts of users of the same type corresponding to the first half of the total differences as the type envelope occupied storage amount corresponding to the application type.
[0046] For the stable stockpiles of different historical users under the same application type in historical data, there will be a certain degree of difference in the amount of data in the trained model. The features to be extracted are big data features, mainly reflecting the average level of the stockpiles used by users under this type. And the size of the gap between this average level and the stable stockpiles of different users is relevant. Therefore, in this application, by obtaining the gap between the stable stockpiles of different users, the stable stockpile data of the part with a smaller gap is extracted to determine the stockpile space required for the corresponding application type. Here, the quantity extracted according to the total difference can be half of the quantity of the stable stockpiles of users, or it can be reasonably extracted according to the size gap between the total differences. It can be understood that the part with a smaller total difference gap indicates a higher degree of aggregation, and the obtained data can fully reflect the characteristics of the space occupancy usage under this application type.
[0047] S2: Obtain the usage information of the current user, and perform space division in combination with the database space usage characteristic data to form multi-user database division information.
[0048] Obtain the usage information of the current user, and perform space division in combination with the database space usage characteristic data to form multi-user database division information, including: determining the corresponding current application type of different current users according to the usage information of different current users; determining the type envelope stockpile corresponding to the application type that is the same as the current application type according to the corresponding current application type of different current users and in combination with the database space usage characteristic data, and determining the type envelope stockpile as the user preset stockpile corresponding to the current user; obtaining the total stockpile of the platform database, and performing division processing on the user storage space according to the user preset stockpiles corresponding to different current users to form multi-user database division information.
[0049] With the feature data information as a reference for the size of the independent storage space divided by users, after obtaining the types of training models of different current users, the corresponding type envelope stockpile can be determined based on the application type. Dividing the database into storage spaces for different current users based on the type envelope stockpile can provide a certain amount of space redundancy to handle the space usage not covered by big data. Of course, more importantly, after completing the division and storing the usage data of the current user in the divided storage space, it is also necessary to label the storage space with the user's identity information to ensure that the data in the storage space can only be used through correct user identity information authentication, achieving a better data isolation effect.
[0050] Obtain the total stockpile of the platform database, and perform division processing on the user storage space according to the user preset stockpiles corresponding to different current users to form multi-user database division information, including: according to the total stockpile of the platform database User preset occupied storage corresponding to different current users , perform user occupied space division in the following manner, where i represents the numbers of different current users: If , then use the user preset occupied storage corresponding to different current users as the occupied space division quantity to divide the independent occupied space corresponding to each current user in the platform database, mark it as the user independent occupied space, and store the current usage data of different current users into the corresponding user independent occupied space, where α is the space division limit factor; If , then use the user preset occupied storage corresponding to different current users to perform proportional division on the space whose size is compared in the platform database , determine the user independent occupied space corresponding to different current users, and store the current usage data of different current users into the corresponding user independent occupied space; Mark the user independent occupied space with the identity information of the corresponding current user, and extract all user independent occupied spaces and the identity information of the corresponding current users to form an isolation space identity connection form; Combine the user independent occupied space corresponding to different current users and the isolation space identity connection form to form multi-user database division information.
[0051] Of course, since the size of the database of the integrated test service platform is also limited, when dividing the storage space sizes of different users, it is necessary to consider the total space size of the database. After all, the database not only needs to store the usage data of different users, but also provide storage space for its own data information analysis and processing and necessary data such as different model data. Therefore, when dividing the independent storage spaces of different users, it is also necessary to fully consider. In this application, when the space is sufficient, it is divided according to the envelope occupied space size provided by the characteristic data, and when the database space is insufficient, the space division for different users is performed according to the proportional relationship of the envelope occupied space. Of course, in this case, it is necessary to provide a limit on the increase in the number of users provided by the platform and a limit on the size of the influx of current user data. For the space division limit factor, it can be set according to the size of the necessary data actually stored in the platform.
[0052] S3: According to the multi-user database division information, perform data optimization based on type similarity to form multi-user shared application isolation database data.
[0053] Based on the multi-user database partitioning information, data optimization is performed based on type similarity to form multi-user shared application isolation database data, including: determining the current users with the same current application type according to the current application types corresponding to different current users, and aggregating them into a set of users with the same type of usage; for different sets of users with the same type of usage, obtaining the current usage data in the user-independent storage spaces corresponding to all current users in the set and performing extraction and sharing processing of the same data to form multi-user shared application isolation database data.
[0054] Current users with the same application type may have similar or identical data inputs during model training because the trained models are similar or identical. Although different users store such identical data independently to ensure data isolation, it will reduce the utilization rate of storage space, especially when the amount of identical data is large. In this application, on the basis of reasonably and effectively isolating the data of different users, a shared storage space is established in the database to uniformly store the identical data, avoiding waste of space caused by multiple storage.
[0055] For different sets of users with the same type of usage, obtaining the current usage data in the user-independent storage spaces corresponding to all current users in the set and performing extraction and sharing processing of the same data to form multi-user shared application isolation database data, including: for all the current usage data corresponding to the set of users with the same type of usage, if there are identical data information at the same time, extracting the identical data information and storing it in the shared storage space of the platform database, and associating the identity information of the corresponding current users with the identical data information; extracting the identity association information of all the identical data in the shared storage space to form a shared data form; combining all the identical data and the shared data form in the shared storage space to form multi-user shared application isolation database data.
[0056] Considering that the identical data is for different users, in order to avoid errors in the sharing of identical data, it is necessary to mark the identity information of the corresponding users for the identical data. Therefore, all the data in the entire shared storage space has identity identification, and can be obtained through the identity information when the user uses it, ensuring effective data isolation while also improving the utilization rate of the storage space.
[0057] The present invention also provides a data isolation platform for an integrated test service platform, which platform comprises: a data acquisition unit for obtaining historical user data and current user usage information; an analysis and processing unit for performing data storage volume analysis based on application types on the historical user data obtained by the data acquisition unit to form database space usage characteristic data, performing space division on the current user usage information in combination with the database space usage characteristic data to form multi-user database division information, and performing shared optimization processing on the same data according to the current user usage information of the current user to form multi-user shared isolation database data; and a database storage unit for independently storing the current user usage information according to the multi-user database division information and performing shared storage on the same data according to the multi-user shared isolation database data.
[0058] Through the data acquisition unit, the platform obtains the data of historical users, and while extracting characteristic information through the analysis and processing unit, it realizes the estimation of the data occupancy of the current user, thereby ensuring that a reasonable-sized independent storage space is provided for the user, making the user data confidential and avoiding confusion and interaction with the data of other users. Different units are closely linked, effectively realizing the mutual isolation of data usage by multiple users, effectively ensuring the security of user data usage, and at the same time improving the utilization efficiency of the platform database space, which is an important material basis for realizing effective data isolation.
[0059] In summary, the beneficial effects of the data isolation method and platform provided by the embodiments of the present invention are as follows:
[0060] This method provides a relatively independent storage space for different users in the database and stores all the data corresponding to the users in this relatively independent storage space. To achieve this method, first, it is necessary to determine the possible data storage volume of the users to ensure that the divided independent storage space fully meets the data occupancy requirements of the users. On the other hand, it is necessary to ensure that the data in the storage space can be uniquely and stably corresponding to the users to ensure that there will be no mutual extraction and confusion of data in different storage spaces. Considering these two aspects, this application extracts the characteristic of the data occupancy situation of historical users on the platform to estimate the data occupancy of the current user to determine the division size of the storage space and establish an identity connection to be closely associated with the storage space. Additionally, considering that in the case of similar training models, there are some identical data information among different users, the repeated storage of these data wastes the storage space of the database. Therefore, the optimized storage of the same data can be performed to improve the utilization rate of the database storage space and avoid confusion in data usage.
[0061] In the embodiments of the present application, "indication" may include direct indication and indirect indication, and may also include explicit indication and implicit indication. If the information indicated by a certain piece of information is called the information to be indicated, then in the specific implementation process, there are many ways to indicate the information to be indicated. For example, but not limited to, the information to be indicated can be directly indicated, such as the information to be indicated itself or the index of the information to be indicated, etc. It is also possible to indirectly indicate the information to be indicated by indicating other information, where there is an association relationship between the other information and the information to be indicated. It is also possible to only indicate a part of the information to be indicated, while the other parts of the information to be indicated are known or pre-agreed. For example, it is also possible to use the arrangement order of each piece of information pre-agreed (such as stipulated in the protocol) to achieve the indication of specific information, thereby reducing the indication overhead to a certain extent. At the same time, it is also possible to identify the common parts of each piece of information and indicate them uniformly to reduce the indication overhead caused by separately indicating the same information.
[0062] In addition, the specific indication method can also be various existing indication methods, such as, but not limited to, the above-mentioned indication methods and their various combinations, etc. The specific details of various indication methods can refer to the prior art and will not be elaborated herein. As can be seen from the above, for example, when it is necessary to indicate multiple pieces of information of the same type, there may be a situation where the indication methods of different pieces of information are different. In the specific implementation process, the required indication method can be selected according to specific needs. The embodiments of the present application do not limit the selected indication method. In this way, the indication methods involved in the embodiments of the present application should be understood to cover various methods that can enable the party to be indicated to obtain the information to be indicated.
[0063] It should be understood that the information to be indicated can be sent as a whole or divided into multiple sub-information and sent separately, and the sending periods and / or sending times of these sub-information can be the same or different. The specific sending method is not limited in the embodiments of the present application. Among them, the sending periods and / or sending times of these sub-information can be predefined, such as predefined according to the protocol, or can be configured by the sending device by sending configuration information to the receiving device.
[0064] "Predefined" or "pre-configured" can be achieved by pre-saving the corresponding code, table or other ways that can be used to indicate relevant information in the device. The embodiments of the present application do not limit its specific implementation method. Among them, "saving" can mean saving in one or more memories. The one or more memories can be separately set, or can be integrated in the encoder or decoder, processor, or communication device. The one or more memories can also be partly separately set and partly integrated in the decoder, processor, or communication device. The type of the memory can be any form of storage medium, which is not limited in the embodiments of the present application.
[0065] In the embodiments of the present application, the "protocol" may refer to a protocol family in the communication field, a standard protocol with a frame structure similar to that of a protocol family, or a related protocol applied to future communication systems. The embodiments of the present application do not make specific limitations on this.
[0066] In the embodiments of the present application, descriptions such as "when...", "in the case of...", "if", and "if" all mean that the device will perform corresponding processing under a certain objective situation, which does not limit the time, and it is not required that the device must have a judgment action during implementation, nor does it mean that there are other limitations.
[0067] In the description of the embodiments of the present application, unless otherwise specified, " / " means that the objects associated before and after are in an "or" relationship. For example, A / B may represent A or B. The "and / or" in the embodiments of the present application is only a description of the association relationship of the associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B may be singular or plural. Also, in the description of the embodiments of the present application, unless otherwise specified, "a plurality of" means two or more than two. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c may be single or multiple. Additionally, for the convenience of clearly describing the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and roles. Those skilled in the art can understand that the terms "first" and "second" do not limit the quantity and execution order, and the terms "first" and "second" do not necessarily mean different. At the same time, in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way for easy understanding.
[0068] It should be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0069] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0070] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0071] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.
[0072] In the present application, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0073] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0074] Those of ordinary skill in the art will realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0075] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0076] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0077] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0078] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0079] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0080] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A data isolation method for an integrated testing service platform, characterized in that: include: Obtain historical user data, perform data storage volume analysis based on application types, and form database space usage feature data; The acquisition of historical user data, analysis of data storage volume based on application type, and formation of database space usage feature data include: Extracting storage usage information corresponding to different historical users with the same application type from the historical user data to form storage usage information of historical users of the same type; Performing a storage feature analysis based on a single training data influx on all the same type of historical user storage usage information of the same application type to determine the type envelope storage corresponding to the application type; Gather the type envelope occupancy corresponding to different application types to form the database space usage characteristic data; Acquire current user usage information, and perform space division in combination with the database space usage characteristic data to form multi-user database division information; According to the multi-user database division information, data optimization based on type homogeneity is performed to form multi-user shared application isolation database data; Among them, the storage usage information of all the same type of historical users of the same application type is analyzed based on the storage characteristics of the single training data influx, and the type envelope storage corresponding to the application type is determined, including: The storage usage information of different historical users of the same type under the same application type is extracted based on the intersection stability of the training data influx items to determine the stable storage usage of the same type of users corresponding to the historical users. The influx of training data of different overlapping users corresponding to the historical users Perform effective stability analysis in the following ways: Collect all the overlapping user training data influx Forming a user training data overlap influx set The user training data overlaps the influx set like And the largest influx of overlapping user training data in the set If the overlap duration reaches the stable overlap duration threshold, the largest overlap user training data influx Determine the effective stable overlap influx for users Otherwise, determine the effective stable overlap influx of the user Indicates determining the user training data overlap influx set The variance of , D0 represents the stable variance limit of the coincident influx data; Based on the basic storage capacity of the training model Maximum user training data influx And the effective stable overlap influx of the user Determine the stable share of the same type of users corresponding to the historical users in According to the stable share of users of the same type corresponding to different historical users under the same application type Perform average inventory analysis based on data clustering to determine the type envelope inventory corresponding to the application type: For different historical users of the same application type, the corresponding stable share of users of the same type is calculated. Perform cluster average inventory analysis in the following ways: Stable inventory for each user of the same type It is indeed stable with other similar users. The total difference is sorted from small to large, and the stable stock of the same type of users corresponding to the first half of the total difference is taken. The average value is calibrated as the type envelope share corresponding to the application type.
2. The data isolation method of the integrated testing service platform according to claim 1, characterized in that: The storage usage information of different historical users of the same type under the same application type is extracted based on the intersection stability of the training data influx items to determine the stable storage usage of the same type of users corresponding to the historical users include: For different historical users of the same type, storage usage information is extracted to extract the average training data influx of a single user for each training in the entire usage cycle, the corresponding training data continuous storage cycle duration, and the basic storage capacity of the training model. According to the duration of the continuous storage period of all training data of a single training and the corresponding average training data influx of a single user, the maximum user training data influx during the period when the training data influx in the usage cycle does not overlap with the training data influx in the adjacent items in time is determined. and the amount of overlapping user training data inflow corresponding to each overlapping period during which the training data inflow in the usage cycle overlaps in time with the training data inflow in the adjacent sub-item n represents the number of different application types, m represents the number of the historical user not used in the application type numbered n, and k represents the sequential number of the overlapping training data corresponding to the historical user numbered m in the application type numbered n.
3. The data isolation method of the integrated testing service platform according to claim 2, characterized in that: The obtaining of the current user usage information and combining the database space usage characteristic data to perform space division to form multi-user database division information includes: Determining the current application types corresponding to different current users according to the current user usage information of different current users; According to the current application types corresponding to different current users and in combination with the database space usage characteristic data, determine the type envelope occupancy corresponding to the application type that is the same as the current application type, and determine the type envelope occupancy as the user preset occupancy corresponding to the current user; The total storage space occupied by the platform database is obtained, and user storage space is divided according to the preset storage spaces corresponding to different current users to form the multi-user database division information.
4. The data isolation method of the integrated testing service platform according to claim 3 is characterized in that: The acquiring of the total storage space occupied by the platform database and the partitioning of user storage space according to the preset storage spaces occupied by different current users to form the multi-user database partitioning information includes: According to the total occupied capacity C0 of the platform database and the preset occupied capacity C0 of different users corresponding to the current users i , the user storage space is divided in the following manner, where i represents the number of different current users: like Then the user preset storage capacity C corresponding to different current users is used. i For the occupied space division amount, an independent occupied space corresponding to each current user is divided in the platform database, marked as user independent occupied space, and the current usage data of different current users is stored in the corresponding user independent occupied space, α is the space division restriction factor; like Then the user preset storage capacity C corresponding to different current users is used. i The space of αC0 in the platform database is divided proportionally by comparing them with each other, the independent user storage space corresponding to different current users is determined, and the current usage data of different current users is stored in the corresponding independent user storage space; Mark the user's independent occupied space with the corresponding identity information of the current user, and extract all the user's independent occupied spaces and the corresponding identity information of the current user to form an isolated space identity contact form; The multi-user database partition information is formed by combining the user independent storage spaces corresponding to different current users and the isolation space identity contact forms.
5. The data isolation method of the integrated testing service platform according to claim 4 is characterized in that: The step of performing data optimization based on type homogeneity according to the multi-user database partition information to form multi-user shared application isolation database data includes: According to the current application types corresponding to different current users, the current users with the same current application type are determined and grouped into a user set of the same type; For different sets of users of the same type, current usage data in the user independent storage space corresponding to all current users in the set is obtained, and the same data is extracted and shared to form the multi-user shared application isolation database data.
6. The data isolation method of the integrated testing service platform according to claim 5, characterized in that: The method of obtaining the current usage data in the independent user storage space corresponding to all current users in the set for different sets of users of the same type and performing extraction and sharing processing of the same data to form the multi-user shared application isolation database data includes: For all the current usage data corresponding to the same type of user set, if the same data information exists at the same time, the same data information is extracted and stored in the shared storage space of the platform database, and the same data information is associated with the identity information of the corresponding current user; Extracting identity association information of all identical data in the shared storage space to form a shared data form; All the same data in the shared storage space and the shared data table are combined to form the multi-user shared application isolation database data.
7. A data isolation platform for an integrated testing service platform, using the data isolation method for an integrated testing service platform according to any one of claims 1 to 6, characterized in that: include: Data collection unit, used to obtain historical user data and current user usage information; an analysis and processing unit, configured to analyze the data storage volume of the historical user data acquired by the data acquisition unit based on the application type to form database space usage characteristic data, perform spatial division on the current user usage information in combination with the database space usage characteristic data to form multi-user database division information, and perform sharing optimization processing on the same data according to the current user usage information of the current user to form multi-user shared isolated database data; The performing of data storage volume analysis based on application type on the historical user data acquired by the data acquisition unit to form database space usage feature data includes: Extracting storage usage information corresponding to different historical users with the same application type from the historical user data to form storage usage information of historical users of the same type; Performing a storage feature analysis based on a single training data influx on all the same type of historical user storage usage information of the same application type to determine the type envelope storage corresponding to the application type; Gather the type envelope occupancy corresponding to different application types to form the database space usage characteristic data; The database storage unit is used to independently store the current user's usage information according to the multi-user database division information, and to share the same data according to the multi-user shared isolation database data.
Citation Information
Patent Citations
Data isolation method and device, computer equipment and storage medium
CN110851853A
Generating and utilizing pre-allocated storage space
US20220019574A1