Database service platform based on AI large model and use method
By dividing the database service platform into a central computing layer, a dynamic topology node layer and a user layer, and using the dynamic topology node layer to process the data of non-autonomous nodes, the problem of integration difficulties and model distortion among small and medium-sized participants in federated learning is solved, and efficient and secure model training is achieved.
Patent Information
- Application Number
- CN202510726129.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The existing federated learning technology has problems such as difficulty in integrating small and micro participants, complex model synchronization, data imbalance and model distortion in data storage and model training, which is difficult to effectively ensure data privacy and training efficiency.
The database service platform is divided into a central computing layer, a dynamic topology node layer and a user layer. The data of non-autonomous nodes are mapped, standardized and obfuscated through the dynamic topology node layer, and weighted average calculations are performed at the central computing layer to realize data privacy protection and model parameters fusion.
It realizes the improvement of model training efficiency and accuracy while ensuring user privacy and security, solves the problems of high cost and low efficiency of small and micro participants, and ensures load balancing and output accuracy of model training.
Smart Images

Figure CN120256525A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of AI large models, and more specifically, to a database service platform based on an AI large model and a usage method thereof. Background Art
[0002] In recent years, technologies related to the artificial intelligence (AI) industry have developed rapidly and are being applied more and more widely and deeply in various industries with completely different technical principles, such as healthcare, finance, industrial manufacturing, autonomous driving, and supply chain. The development of AI technology is based on data, which is the core resource supporting the development of AI technology. Generally, AI large model data has various characteristics such as large quantity, multi-modal, fluidity, complexity, etc. Importantly, data samples imply users' privacy behaviors. For example, medical data contains users' medical privacy, code data contains the company's future decision-making characteristics, and user behavior data contains users' daily life information, etc. After these highly sensitive big data are exposed in the AI large model, it will introduce huge risks. In federated learning technology, data is stored in an autonomous database, and users independently manage data assets. Multiple participants (such as devices, enterprises, or institutions) collaborate to train and share models on the premise of keeping the data local, and the original data does not need to leave the local. Independently managing data assets can basically solve the problem of data privacy exposure, but this distributed architecture actually has many disadvantages. First, not every participant has the software and hardware environment basis for model deployment and training, and small and micro participants (compared with large participants) are difficult to integrate and have low utilization rates. Second, the maintenance and alignment of models and data require all participants to synchronize, which will introduce complex problems other than technology such as management and decision-making. Third, when synthesizing the training parameters obtained by the participating parties' training models with asynchronous shared models or samples, it will lead to problems such as invalid / imbalanced parameters and inaccurate models. Finally, it is difficult to prevent model distortion problems, such as any participant fabricating or forging big data samples, resulting in overall model distortion. To solve the above problems, the existing federated learning technology has conducted extensive research and improvement in ensuring the deployment environment, model synchronization, compensating for the differences in data imbalance, data calibration, participation permissions, etc.; including adopting a distributed / hierarchical storage architecture, and adopting complex database logics (such as distributed databases (HDFS architecture, Spark computing framework), time series databases (InfluxDB / TimescaleDB), vector databases (Pinecone embedded data representation / PostageSQL+pgvector architecture), multi-modal databases (MongoDB / Neo4j), blockchain databases, etc.). However, these AI data management methods have not changed the characteristics of autonomous data management, so the problem of federated learning model distortion caused by the above situation cannot be completely avoided.
[0003] The information disclosed in the background section of the present invention is only intended to deepen the understanding of the general background of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art. Summary of the Invention
[0004] The present invention proposes a database service platform based on an AI large model and its usage method, which can fully train user samples, ensure user privacy and security, and improve the efficiency and accuracy of model training.
[0005] In a first aspect, an embodiment of the present disclosure provides a usage method for a database service based on an AI large model, including: S100, dividing the database service platform into a central computing layer, a dynamic topology node layer, and a user layer; the database service platform receives user node registration information from the user layer and returns a user token and node set information of the dynamic topology node layer to the user node; S110, the non-autonomous user nodes in the user layer generate a data upload mapping table based on the node set information and a random algorithm, the data upload mapping table includes random digest information of the user token, and the data upload mapping table is used to map the nodes in the node set of the dynamic topology node layer where the user layer data is stored; the user layer uploads user data slices to the mapped nodes based on the data upload mapping table; S120, the nodes in the dynamic topology node layer receive the user data slices, perform node standardization processing on the user data slices in the data format corresponding to this node, perform confusion processing on the data slices after node standardization processing, and store them in the database sub-nodes; the training sub-nodes train the sample data in the database sub-nodes based on the configured AI model and send the trained model parameters to the central computing layer; the autonomous nodes in the user layer train their sample data based on the configured AI model and send the trained model parameters to the central computing layer; S130, the central computing layer receives the model parameters of each node in the dynamic topology node layer and the model parameters of the autonomous nodes in the user layer, and performs weighted average calculation on the model parameters to obtain the aggregated total model parameters.
[0006] Preferably, the user slices the user data to be uploaded by unit, and the slicing algorithm is a random generation algorithm, which ensures that the minimum slicing unit is the data unit that can be processed by training; the non-autonomous nodes generate a data upload mapping table based on the received node set information of the dynamic topology node layer and the random algorithm, the data upload mapping table is used to map the nodes in the node set of the dynamic topology node layer where the user layer data is stored, and adapt the user data slices and the corresponding upload paths, so as to upload the data slices according to the upload paths of the nodes in the mapped node set of the dynamic topology node layer.
[0007] Preferably, the obfuscation processing implementation method is as follows: (1) temporarily store the received data slices in a non-decrypted state; (2) within a preset time, determine whether the data source of the data slices is less than a predetermined quantity threshold. If so, decrypt the data slices and add controllable noise data, and store them in the database sub-node; if not, perform obfuscation processing on the data in the data slices from different data sources, and store them in the database sub-node.
[0008] Preferably, the dynamic topology node layer includes a specific number of distributed physical nodes. Within a preset time, determine whether the number of nodes whose data source of the data slices is less than the predetermined quantity threshold exceeds a second predetermined quantity threshold. If it exceeds, select several of them for shutdown processing. Before shutdown, dynamically migrate their data to other nodes and release the physical node; at the same time, within a preset time, determine whether the number of nodes whose data source of the data slices is greater than a third predetermined quantity threshold exceeds a fourth predetermined quantity threshold. If it exceeds, add new physical nodes to the topology network and update the node set information of the dynamic topology node layer.
[0009] Preferably, randomly eliminate the sample data of several training nodes, determine the quality of the sample data of the several training nodes based on the degree of change in the difference between the model output and the true value, and based on the determined result, uniformly reduce or increase the data weights of the parties involved in the training nodes.
[0010] In a second aspect, the embodiments of the present disclosure further provide a database service platform based on an AI large model, including: A platform construction function module, which is used to divide the central computing layer, the dynamic topology node layer, and the user layer; receive the user node registration information of the user layer, and return a user token and the node set information of the dynamic topology node layer to the user node; A data upload function module, which is used for non-autonomous user nodes in the user layer to generate a data upload mapping table based on the node set information and a random algorithm. The data upload mapping table includes random digest information of the user token, and the data upload mapping table is used to map the nodes in the node set of the dynamic topology node layer where the user layer data is stored; the user layer uploads user data slices to the mapped nodes based on the data upload mapping table; A data training function module. The nodes in the dynamic topology node layer receive the user data slices, perform node standardization processing on the user data slices in the data format corresponding to this node, perform obfuscation processing on the data slices after the node standardization processing, and store them in the database sub-node; the training sub-node trains the sample data in the database sub-node based on the configured AI model, and sends the trained model parameters to the central computing layer; the autonomous nodes in the user layer train their sample data based on the AI model configured by them, and send the trained model parameters to the central computing layer; The parameter fusion functional module, where the central computing layer receives the model parameters of the nodes in each dynamic topology node layer and the model parameters of the autonomous nodes in the user layer, and performs weighted average calculation on the model parameters to obtain the aggregated total model parameters.
[0011] Preferably, the user slices the user data to be uploaded by unit. The slicing algorithm is a random generation algorithm, which ensures that the minimum slicing unit is the data unit that can be processed by training. The non-autonomous node generates a data upload mapping table based on the received information of the node set in the dynamic topology node layer and the random algorithm. The data upload mapping table is used to map the nodes in the node set of the dynamic topology node layer where the user layer data is stored, and adapt the user data slices and the corresponding upload paths, so as to upload the data slices according to the upload paths of the nodes in the mapped dynamic topology node layer node set.
[0012] Preferably, the implementation method of the confusion processing is as follows: (1) temporarily store the received data slices in a non-decrypted state; (2) within a preset time, judge whether the data source of the data slices is less than a predetermined quantity threshold. If so, decrypt the data slices and add controllable noise data, and store them in the database sub-node; if not, perform confusion processing on the data in the data slices from different data sources and store them in the database sub-node.
[0013] Preferably, the dynamic topology node layer includes a specific number of distributed physical nodes. Within a preset time, judge whether the number of nodes whose data source of the data slices is less than the predetermined quantity threshold exceeds the second predetermined quantity threshold. If it exceeds, select several of them for shutdown processing. Before shutdown, dynamically migrate its data to other nodes and release the physical node; at the same time, within a preset time, judge whether the number of nodes whose data source of the data slices is greater than the third predetermined quantity threshold exceeds the fourth predetermined quantity threshold. If it exceeds, add a new physical node to the topology network and update the information of the node set in the dynamic topology node layer.
[0014] Preferably, randomly eliminate the sample data of several training nodes, determine the quality of the sample data of the several training nodes based on the change degree of the difference between the model output and the true value, and based on the determined result, uniformly reduce or increase the data weights of the parties involved in the training nodes.
[0015] The present invention provides a database service platform based on an AI large model and a usage method, achieving at least the following technical effects: (1) User nodes are divided into autonomous nodes and non-autonomous nodes. The data processing of non-autonomous nodes is proxied by nodes in the dynamic topology node layer. This creative distributed training meets the comprehensive requirements of all training participants for the training environment, efficiency, and cost; (2) Data privacy processing is jointly performed from two aspects: the data uploader and the data receiver. At the same time, during the privacy processing, sample data with a small sample size is fused. Compared with the training of samples with a small sample size in the prior art, it not only maximally ensures the privacy and security of the data but also improves the model training efficiency; (3) An innovative training architecture for data processing proxy based on nodes in the dynamic topology node layer is proposed, which centralizes the data of non-autonomous node participants, improves the training efficiency, facilitates ensuring the load balance of training nodes, and solves a series of problems of high cost and low efficiency for small and micro participants in the prior art; (4) A data sample weight adjustment method is set to ensure the output accuracy of the model and the privacy and security of users. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] By describing the exemplary embodiments of the present invention in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present invention will become more apparent. Among them, in the exemplary embodiments of the present invention, the same reference numerals generally represent the same components.
[0017] Figure 1 The flowchart showing the steps of a usage method of a database service platform based on an AI large model according to an embodiment of the present invention is shown.
[0018] Figure 2 The block diagram of a database service platform based on an AI large model according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.
[0020] In the description and claims of this application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0021] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0022] Embodiment 1: As Figure 1 shown, the present invention provides a method for using a database service based on an AI large model, and the method includes the steps of: S100, dividing the database service platform into a central computing layer, a dynamic topology node layer and a user layer; the database service platform receives user node registration information from the user layer and returns a user token and dynamic topology node layer node set information to the user node; The user layer is composed of at least one user node; the user node includes an autonomous node and a non-autonomous node. The dynamic topology nodes, in the form of a set, serve as an intermediate layer between the central computing layer and the user layer. On the one hand, they receive the data uploaded by the non-autonomous nodes in the user layer and perform training-related processing. On the other hand, they send the results obtained from the training-related processing to the central computing layer for further processing. Among them, an autonomous node refers to a node where the model participant manages user data and model training by itself, and a non-autonomous node refers to a node that uploads user data to the dynamic topology node layer and uses it as an agent node for training. Among users who wish to participate in model training, not all users have the software and hardware environment for user data management and model training. At the same time, for some small and micro enterprises, small data volume units or private type users, the cost of fully configuring the training environment is high, and the model training results for small data volumes are also unreliable. Therefore, for the sake of effectiveness, security and cost, the user nodes are divided into autonomous nodes and non-autonomous nodes. As an important asset of data resources, any part of the data makes a significant contribution to model algorithm training. The sum of the data possessed by autonomous nodes and non-autonomous nodes constitutes the big data foundation of the model.
[0023] The user requests registration from the database service platform to join the AI large model training. The database service platform receives the user node registration information at the user layer. After registration verification, it returns the user token and the node set information of the dynamic topology node layer to the user, and maintains the user information database in the system to record the full user information. Among them, the returned user token is a sequence identifier for identifying the user's unique identity information, and the node set of the dynamic topology node layer is the dynamic topology nodes that the user has the right to use. The user node implements related functional operations such as data upload, privacy processing, and data obfuscation based on the node set information. As can be seen from the subsequent content, the topology architecture of the dynamic topology node layer is dynamically changed, and this change information can be synchronized to the user layer nodes based on various well-known methods. The data in the user information database is used for user management, including basic information management, membership management, permission management, contribution management, etc.
[0024] S110, the non-autonomous user node at the user layer generates a data upload mapping table based on the node set information and a random algorithm. The data upload mapping table contains the random digest information of the user token, and the data upload mapping table is used to map the nodes in the node set of the dynamic topology node layer stored in the user layer data; the user layer uploads the user data slices to the mapped nodes based on the data upload mapping table.
[0025] Privacy is a basic condition for the AI large model during the joint training process. The data uploaded by the user should ensure its privacy. During the process of uploading data from the non-autonomous node to the dynamic topology node layer, since the dynamic topology node layer is not managed locally by the user, there will inevitably be a risk of data exposure. Therefore, in this step, the user side dominates the data transmission process. Specifically, in a preferred embodiment, the user slices the user data to be uploaded by unit. The slicing algorithm can be a random generation algorithm, and it should be ensured that the minimum slicing unit is the data unit that can be processed by training; correspondingly, the non-autonomous node generates a data upload mapping table based on the node set information of the dynamic topology node layer received in step S100 and a random algorithm. The data upload mapping table contains the random digest information of the user token, which is used to verify the permission and legality of the user's data upload behavior; the data upload mapping table is used to map the nodes in the node set of the dynamic topology node layer stored in the user layer data, adapt the user data slices and the corresponding upload paths, and thus upload the data slices according to the upload paths of the nodes in the mapped node set of the dynamic topology node layer. In the above method, randomly slicing the data and determining the nodes of the dynamic topology node layer as proxy nodes based on a random algorithm hides the integrity of the data and the data processing party, ensuring the privacy of the training sample data.
[0026] In a preferred embodiment, the generation algorithm of the data upload mapping table further includes: (1) user-defined configuration, in which the fuzzy proxy node conditions are determined, including the geographical location of the storage node, the load condition, the security level, the permission level, etc.; (2) randomly generating mapping nodes with the user-defined configuration as a constraint.
[0027] S120, the nodes in the dynamic topology node layer receive the user data slices, perform node standardization processing on the user data slices in the data format corresponding to this node, perform fusion and confusion processing on the data slices after the node standardization processing, and store them in the database sub-node; the training sub-node trains the sample data in the database sub-node based on the configured AI model, and sends the trained model parameters to the central computing layer; the autonomous nodes in the user layer train their sample data based on the configured AI model, and send the trained model parameters to the central computing layer; Each node in the dynamic topology node layer includes a database sub-node and a training sub-node. The database sub-node stores the sample big data for training, and the training sub-node deploys an AI model for model training on the sample big data; the dynamic topology node layer contains a specific number of distributed physical nodes. When the average node load exceeds a predetermined value, a new physical node can be added to the topology network; when the model training result of a specific node does not meet the convergence condition, its data is dynamically migrated to other nodes to release the physical node.
[0028] There are differences in the training content of different nodes, and this difference is determined by the database service platform according to the training algorithm of the deployed model at the initial stage of environment construction. According to the difference in the training content, the nodes in the dynamic topology node layer perform node standardization processing on the user data slices in the data format corresponding to this node after receiving the user data slices.
[0029] In a preferred embodiment, before storing the received user data slices in the database sub-node, the nodes in the dynamic topology node layer perform fusion and confusion processing on the data slices after the node standardization processing. The principle of the confusion processing is to obscure the source of the user data. Specifically, its implementation method is: (1) temporarily store the received data slices in a non-decrypted state; (2) within a preset time, determine whether the data source of the data slice is less than a predetermined quantity threshold. If so, decrypt the data slice and add controllable noise data, and store it in the database sub-node; if not, perform confusion processing on the data in the data slices from different data sources, and store it in the database sub-node. Each node in the dynamic topology node layer performs confusion processing on the user-uploaded data slices before training with the application samples, eliminating the possibility of tracing the data source, and ensuring data privacy and security from the perspective of the recipient.
[0030] In a preferred embodiment, the obfuscation process includes: randomly scrambling data in units of data slices, and extracting and fusing the data content in the data slices. The non-decryption state refers to temporarily delaying the decryption process after receiving encrypted data, or performing encryption processing after receiving data, so as to ensure the security during data storage. The controllable noise data includes: data containing specific fields for noise filtering, adding fields that are meaningless for model training, changing field mapping, and encrypting fields.
[0031] Further, in a preferred embodiment, the dynamic topology node layer includes a specific number of distributed physical nodes. Within a preset time, it is determined whether the number of nodes with the data source of the data slice less than a predetermined quantity threshold exceeds a second predetermined quantity threshold. If it exceeds, several of these nodes are selected for shutdown processing. Before shutdown, their data is dynamically migrated to other nodes to release the physical node. At the same time, within a preset time, it is determined whether the number of nodes with the data source of the data slice greater than a third predetermined quantity threshold exceeds a fourth predetermined quantity threshold. If it exceeds, new physical nodes are added to the topology network to update the node set information of the dynamic topology node layer. Even for controllable noise, the data after noise addition processing will still be distorted to a certain extent and introduce unnecessary data processing loads; the number of data sources of data slices is an indication of node load; moreover, if the data volume of the dynamic topology node layer is too small, the training results are not easily convergent or inaccurate, which will in turn hinder the training process. Therefore, for the above reasons, the dynamic topology node layer is dynamically adjusted, improving the training efficiency of the system and ensuring data privacy and security.
[0032] The training sub-nodes train the sample data in the database sub-nodes based on the configured AI model, and send the trained model parameters to the central computing layer. These training results reflect the training data contribution of the non-autonomous nodes. The autonomous nodes in the user layer train their sample data based on the configured AI model and send the trained model parameters to the central computing layer, reflecting the training data contribution of the autonomous nodes.
[0033] S130, the central computing layer receives the model parameters of each node in the dynamic topology node layer and the model parameters of the autonomous nodes in the user layer, and performs weighted average calculation on the model parameters to obtain the aggregated total model parameters.
[0034] The non-autonomous nodes use the nodes in the dynamic topology node layer as proxy nodes to achieve model training under the premise of data privacy and security. However, due to the uncontrollability of the proxy process of data transmission and processing, this privacy and security is relative. Therefore, the data processing method of autonomous nodes is still an option that cannot be discarded. The database service platform of the present invention provides database services to users. It differentiates user nodes into autonomous nodes and non-autonomous nodes, constructs a distributed machine learning framework formed by nodes in the dynamic topology node layer and autonomous nodes, so that the autonomous nodes and non-autonomous nodes can form a distributed machine learning framework, allowing multiple participants (such as mobile devices, enterprises, or data centers) to collaboratively train a global model without sharing the original data. In a preferred embodiment, when obtaining the aggregated model parameters, the model parameters are weighted and averaged using the ratio of the data load of the node to the total data of each node in the dynamic topology node layer and the autonomous nodes in the user layer as the weight.
[0035] In a preferred embodiment, the accuracy of the real-time evaluation model is evaluated. Before parameter fusion, sample data of several nodes are randomly excluded, and the quality of the sample data is determined based on the change in the difference between the model output and the true value. Based on the determined results, the data weights of the involved parties are reduced or increased. By controlling the node data as described above, the probability of model distortion is reduced. In a preferred embodiment, reducing or increasing the data weights of the involved parties specifically means: determining the data source of the sample data of the training nodes and setting training weights for the data of the data source nodes. Since the data of the training nodes comes from non-autonomous user nodes, the quality of the sample data of the training nodes is essentially determined by the data slices of the non-autonomous user nodes. Therefore, the optimal way to process the training sample weights is to directly process the data source. In a preferred embodiment, to further ensure user privacy, setting training weights for the data of the data source nodes means: performing the same normalization process of reducing or increasing the weights on all the data sources of the sample data involved in the determined training nodes. From the perspective of a single training node, the normalization process may seem unfair to nodes that do not need to adjust the sample weights. However, from an overall perspective, since each training node involves user nodes from multiple data sources, the weight adjustment for all training nodes will ultimately accurately affect the user nodes of each data source. For example, the data of training node A comes from non-autonomous nodes a, b, and c, the data of training node B comes from non-autonomous nodes b, c, and d, and the data of training node C comes from non-autonomous nodes a, c, and d. The original weights are a = b = c = d = 100%. Periodically adjust the weights of A, that is, adjust the weight numbers of the non-autonomous nodes a, b, and c involved in it to -10%, then the weight values are a = 90%, b = 90%, c = 90%, and d = 100%; periodically adjust the weights of B, that is, adjust the weight numbers of the non-autonomous nodes b, c, and d involved in it to +5%, then the weight values are a = 90%, b = 95%, c = 95%, and d = 105%; periodically adjust the weights of C, that is, adjust the weight numbers of the non-autonomous nodes a, c, and d involved in it to -10%, then the weight values are a = 80%, b = 95%, c = 85%, and d = 95%. It can be seen that through the overall weight adjustment of all training nodes, the quality of the data of each non-autonomous user node of the actual data source is reflected. This weight adjustment method first ensures the superiority and inferiority of the training sample data, so that the model training result will not be distorted and improves the model accuracy; secondly, the normalization process is performed on all the data sources of the sample data involved in the same training node, which is simple and efficient. Since this method does not distinguish the data sources, it greatly ensures the privacy security of users.
[0036] Embodiment 2: As Figure 2As shown in the figure, the present invention also provides a database service platform based on an AI large model, including: A platform construction function module for dividing the central computing layer, the dynamic topology node layer, and the user layer; receiving user node registration information from the user layer and returning a user token and node set information of the dynamic topology node layer to the user node; A data upload function module for non-autonomous user nodes in the user layer to generate a data upload mapping table based on the node set information and a random algorithm, where the data upload mapping table contains random digest information of the user token, and the data upload mapping table is used to map nodes in the node set of the dynamic topology node layer where the user layer data is stored; the user layer uploads user data slices to the mapped nodes based on the data upload mapping table; A data training function module, where nodes in the dynamic topology node layer receive user data slices, perform node standardization processing on the user data slices in the data format corresponding to the node, perform confusion processing on the data slices after node standardization processing, and store them in the database sub-nodes; the training sub-nodes train the sample data in the database sub-nodes based on the configured AI model and send the trained model parameters to the central computing layer; autonomous nodes in the user layer train their sample data based on the configured AI model and send the trained model parameters to the central computing layer; A parameter fusion function module, where the central computing layer receives the model parameters of each node in the dynamic topology node layer and the model parameters of autonomous nodes in the user layer, performs weighted average calculation on the model parameters, and obtains the aggregated total model parameters.
[0037] In a preferred embodiment, the user slices the user data to be uploaded by unit, and the slicing algorithm is a random generation algorithm that ensures that the minimum slicing unit is a data unit that can be processed by training; the non-autonomous node generates a data upload mapping table based on the received node set information of the dynamic topology node layer and the random algorithm, and the data upload mapping table is used to map nodes in the node set of the dynamic topology node layer where the user layer data is stored, adapts the user data slices and the corresponding upload paths, and thus uploads the data slices according to the upload paths of the nodes in the mapped node set of the dynamic topology node layer.
[0038] In a preferred embodiment, the implementation manner of the confusion processing is: (1) temporarily storing the received data slices in a non-decrypted state; (2) within a preset time, determining whether the data source of the data slices is less than a predetermined quantity threshold. If so, decrypt the data slices and add controllable noise data, and store them in the database sub-nodes; if not, perform confusion processing on the data in the data slices from different data sources and store them in the database sub-nodes.
[0039] In a preferred embodiment, the dynamic topology node layer includes a specific number of distributed physical nodes. Within a preset time, it is determined whether the number of nodes with the data source of data slices less than a predetermined quantity threshold exceeds a second predetermined quantity threshold. If it exceeds, several of these nodes are selected for shutdown processing. Before shutdown, their data is dynamically migrated to other nodes, and the physical node is released. At the same time, within the preset time, it is determined whether the number of nodes with the data source of data slices greater than a third predetermined quantity threshold exceeds a fourth predetermined quantity threshold. If it exceeds, new physical nodes are added to the topology network, and the node set information of the dynamic topology node layer is updated.
[0040] In a preferred embodiment, the sample data of several training nodes is randomly excluded, and the quality of the sample data of the several training nodes is determined based on the degree of change in the difference between the model output and the true value. Based on the determined result, the data weights of the parties involved in the training nodes are uniformly reduced or increased.
[0041] The present invention provides a database service platform and a usage method based on an AI large model, achieving at least the following technical effects: (1) User nodes are divided into autonomous nodes and non-autonomous nodes, and the data processing of non-autonomous nodes is agented by the nodes in the dynamic topology node layer. This creative distributed training meets the comprehensive requirements of each training participant for the training environment, efficiency, and cost; (2) Data privacy processing is jointly carried out from two aspects: the data uploader and the data receiver. At the same time, during the privacy processing process, sample data with a small sample size is fused. Compared with the training of samples with a small sample size in the prior art, it not only maximally ensures the privacy and security of data but also improves the model training efficiency; (3) An innovative training architecture for data processing agency based on the nodes in the dynamic topology node layer is proposed, which centralizes the data of non-autonomous node participants, improves the training efficiency, facilitates ensuring the load balance of training nodes, and solves a series of problems of high cost and low efficiency of small and micro participants in the prior art; (4) A data sample weight adjustment method is set to ensure the output accuracy of the model and the user privacy security.
[0042] According to one embodiment, a program product such as a machine-readable medium is provided. The machine-readable medium may have instructions (i.e., the elements implemented in software as described above), which when executed by the machine, cause the machine to perform the various operations and functions described above in connection with Figure 1 the various embodiments of this specification. Specifically, a system or device equipped with a readable storage medium may be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer or processor of the system or device reads and executes the instructions stored in the readable storage medium.
[0043] In this case, the program code read from the readable medium itself can implement the functions of any one of the above-described embodiments. Therefore, the machine-readable code and the readable storage medium storing the machine-readable code constitute a part of this specification.
[0044] Examples of the readable storage medium include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD-RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program code can be downloaded from a server computer or a cloud via a communication network.
[0045] Those skilled in the art should understand that various modifications and variations can be made to the above-disclosed embodiments without departing from the essence of the invention. Therefore, the protection scope of this specification should be defined by the appended claims.
[0046] It should be noted that not all steps and units in the above-mentioned processes and system structure diagrams are necessary, and some steps or units can be omitted according to actual needs. The execution order of each step is not fixed and can be determined as needed. The device structures described in the above embodiments can be physical structures or logical structures. That is, some units may be implemented by the same physical entity, or some units may be implemented by multiple physical entities respectively, or some components in multiple independent devices may be jointly implemented.
[0047] In the above embodiments, the hardware units or modules can be implemented mechanically or electrically. For example, a hardware unit, module, or processor can include permanent dedicated circuits or logics (such as dedicated processors, FPGAs, or ASICs) to complete corresponding operations. The hardware unit or processor can also include programmable logics or circuits (such as general-purpose processors or other programmable processors), which can be temporarily set by software to complete corresponding operations. The specific implementation method (mechanical method, or dedicated permanent circuit, or temporarily set circuit) can be determined based on cost and time considerations.
[0048] The specific embodiments described above in conjunction with the accompanying drawings describe exemplary embodiments, but do not represent all embodiments that can be implemented or fall within the protection scope of the claims. The term "exemplary" used throughout this specification means "serving as an example, instance, or illustration", and does not mean "preferred" or "superior" to other embodiments. For the purpose of providing an understanding of the described technology, the specific embodiments include specific details. However, these technologies can be implemented without these specific details. In some instances, well-known structures and devices are shown in block diagram form to avoid obscuring the concepts of the described embodiments.
[0049] The foregoing description of the present disclosure is provided to enable any person of ordinary skill in the art to make or use the present disclosure. Various modifications to the present disclosure will be apparent to those of ordinary skill in the art, and the general principles corresponding to the present disclosure herein can also be applied to other variations without departing from the scope of protection of the present disclosure. Therefore, the present disclosure is not limited to the examples and designs described herein, but is consistent with the broadest scope that conforms to the principles and novel features disclosed herein.
Claims
1. A method for using a database service platform based on a large AI model, characterized in that, The method includes the steps of: S100, dividing the database service platform into a central computing layer, a dynamic topology node layer, and a user layer; the database service platform receives user node registration information from the user layer and returns a user token and node set information of the dynamic topology node layer to the user node; S110, the non-autonomous user nodes in the user layer generate a data upload mapping table based on the node set information and a random algorithm, the data upload mapping table includes user token random digest information, and the data upload mapping table is used to map the nodes in the node set of the dynamic topology node layer where the user layer data is stored; the user layer uploads user data slices to the mapped nodes based on the data upload mapping table; S120, the nodes in the dynamic topology node layer receive the user data slices, perform node standardization processing on the user data slices in the data format corresponding to the current node, perform obfuscation processing on the data slices after node standardization processing, and store them in the database sub-nodes; the training sub-nodes train the sample data in the database sub-nodes based on the configured AI model and send the trained model parameters to the central computing layer; the autonomous nodes in the user layer train their sample data based on the AI model they configure and send the trained model parameters to the central computing layer; S130, the central computing layer receives the model parameters of each node in the dynamic topology node layer and the model parameters of the autonomous nodes in the user layer, and performs weighted average calculation on the model parameters to obtain the aggregated total model parameters.
2. The method according to claim 1, wherein: The user slices the user data to be uploaded using a slicing algorithm by unit, and the slicing algorithm is a random generation algorithm to ensure that the minimum slicing unit is a data unit that can be processed by training; The non-autonomous nodes generate a data upload mapping table based on the received node set information of the dynamic topology node layer and a random algorithm, the data upload mapping table is used to map the nodes in the node set of the dynamic topology node layer where the user layer data is stored, and adapt the user data slices and the corresponding upload paths, so as to upload the data slices according to the upload paths of the nodes in the mapped node set of the dynamic topology node layer.
3. The method according to claim 1, characterized in that: The obfuscation processing method includes: Temporarily storing the received data slices in a non-decrypted state; Within a preset time, judge whether the data source of the data slices is less than a predetermined quantity threshold. If so, decrypt the data slices and add controllable noise data, and store them in the database sub-nodes; if not, perform obfuscation processing on the data in the data slices from different data sources and store them in the database sub-nodes.
4. The method according to claim 1, wherein: The dynamic topology node layer includes a specific number of distributed physical nodes. Within a preset time, judge whether the number of nodes whose data source of the data slices is less than the predetermined quantity threshold exceeds a second predetermined quantity threshold. If it exceeds, select several of them for shutdown processing, and dynamically migrate their data to other nodes before shutdown to release the physical node; And within a preset time, judge whether the number of nodes whose data source of the data slices is greater than a third predetermined quantity threshold exceeds a fourth predetermined quantity threshold. If it exceeds, add new physical nodes to the topology network and update the node set information of the dynamic topology node layer.
5. The method according to claim 1, characterized in that: Randomly eliminate the sample data of several training nodes, determine the quality of the sample data of the several training nodes based on the degree of change in the difference between the model output and the true value, and uniformly reduce or increase the data weights of the parties involved in the training nodes based on the determined results.
6. A database service platform based on an AI large model, characterized in that, The platform includes: A platform construction function module for dividing the central computing layer, the dynamic topology node layer, and the user layer; receiving the user node registration information of the user layer and returning the user token and the node set information of the dynamic topology node layer to the user node; A data upload function module for non-autonomous user nodes in the user layer to generate a data upload mapping table based on the node set information and a random algorithm, where the data upload mapping table includes the random digest information of the user token, and the data upload mapping table is used to map the nodes in the node set of the dynamic topology node layer stored by the user layer data; the user layer uploads user data slices to the mapped nodes based on the data upload mapping table; A data training function module, where the nodes in the dynamic topology node layer receive the user data slices, perform node standardization processing on the user data slices in the data format corresponding to the current node, perform confusion processing on the data slices after the node standardization processing, and store them in the database sub-node; the training sub-node trains the sample data in the database sub-node based on the configured AI model and sends the trained model parameters to the central computing layer; the autonomous nodes in the user layer train their sample data based on the configured AI model and send the trained model parameters to the central computing layer; A parameter fusion function module, where the central computing layer receives the model parameters of each node in the dynamic topology node layer and the model parameters of the autonomous nodes in the user layer, and performs weighted average calculation on the model parameters to obtain the aggregated total model parameters.
7. The database service platform based on the AI large model according to claim 6, wherein: The user slices the user data to be uploaded using a slicing algorithm, and the slicing algorithm is a random generation algorithm to ensure that the minimum slice unit is the data unit that can be processed by training; The non-autonomous node generates a data upload mapping table based on the received node set information of the dynamic topology node layer and a random algorithm, where the data upload mapping table is used to map the nodes in the node set of the dynamic topology node layer stored by the user layer data, and adapts the user data slices and the corresponding upload paths, so as to upload the data slices according to the upload paths of the nodes in the mapped node set of the dynamic topology node layer.
8. The database service platform based on the AI large model according to claim 6, wherein: The confusion processing method includes: temporarily storing the received data slices in a non-decrypted state; within a preset time, determining whether the data source of the data slices is less than a predetermined quantity threshold, if so, decrypting the data slices and adding controllable noise data, and storing them in the database sub-node; if not, performing confusion processing on the data in the data slices from different data sources and storing them in the database sub-node.
9. For the database service platform based on the AI large model as described in claim 6, the dynamic topology node layer includes a specific number of distributed physical nodes. Within a preset time, it is determined whether the number of nodes with the data source of the data slice less than a predetermined quantity threshold exceeds a second predetermined quantity threshold. If it exceeds, several of these nodes are selected for shutdown processing. Before shutdown, their data is dynamically migrated to other nodes to release the physical node. And within a preset time, it is determined whether the number of nodes with the data source of the data slice greater than a third predetermined quantity threshold exceeds a fourth predetermined quantity threshold. If it exceeds, new physical nodes are added to the topology network to update the node set information of the dynamic topology node layer.
10. The database service platform based on the AI large model according to claim 6, characterized in that: Randomly eliminate the sample data of several training nodes, determine the quality of the sample data of the several training nodes based on the degree of change in the difference between the model output and the true value, and based on the determined result, uniformly reduce or increase the data weights of the parties involved in the training nodes.
Citation Information
Patent Citations
Private data protection method and device, equipment and storage medium
CN118133351A
Off-site distributed security data transaction method for privacy protection
CN119885257A
Federated learning with partitioned and dynamically-shuffled model updates
US20220374763A1